Head-related transfer function applied to sound localization in emergency situations

Through the combined processing circuit of the head-mounted device image sensor and audio sensor, the problem of obscuring the visibility of the signal source in an emergency environment is solved, and accurate positioning and guidance of the signal source is achieved.

CN120379728APending Publication Date: 2025-07-253M INNOVATIVE PROPERTIES CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380081916.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-30
Filing Date
2023-11-21
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In an emergency environment, the visibility of the signal source is obscured, resulting in the first responder being unable to accurately locate and identify the source of the sound or radio beacon.

Method used

The head-mounted device is used to combine an image sensor and an audio sensor to determine the location of the signal source through a processing circuit, and to guide the user to the signal source using the user interface of the head-mounted device.

Benefits of technology

Under limited visibility conditions, accurately identifying the signal source and guiding the user to the signal source improves positioning accuracy and efficiency in emergencies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120379728A_ABST
    Figure CN120379728A_ABST
Patent Text Reader

Abstract

A head-mounted device configured to be worn by a user is provided. The head-mounted device includes at least one microphone, at least one image sensor, and processing circuitry configured to receive an audio signal detected by the at least one microphone, the audio signal originating from a signal source. The processing circuit is configured to determine a location of the signal source based on the received audio signal. The processing circuitry is further configured to receive image data from the at least one image sensor, the image data associated with at least one of: the face of the user and at least one of the eyes of the user. The processing circuitry is further configured to determine a gaze direction of the user based on the received image data; and determining a user instruction based on the determined location and the determined gaze direction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to location detection and, more particularly, to devices and systems for guiding first responders to a signal source in an emergency environment with limited visibility and related methods of use thereof. Background Art

[0002] In emergency situations and environments, the visibility of the target item or individual being sought may be obscured. Such obscuration may be caused by smoke, suspended matter, debris piles, low or no light, etc., or the target may be covered with dust, dirt, soot, etc., making rapid identification impossible. Similarly, in industrial environments, individuals may be partially hidden by machinery, wires and equipment, pipes and pipe racks, etc. In such limited visibility environments, first responders may not be able to identify the source / location of a signal source (such as a sound generating source and / or a wireless beacon / signal). Summary of the Invention

[0003] Head-mounted devices such as face masks, goggles, and / or self-contained breathing apparatus (SCBA) may use image sensors and / or audio sensors to accurately identify the source of a sound (such as a sound emitted by a person-down alarm in an emergency environment) and determine the gaze of the user of the head-mounted device. The head-mounted device may include a user interface configured to indicate to the user of the head-mounted device the signal source location and guide the user to that location based on the determined gaze.

[0004] Some embodiments advantageously provide methods and systems for a head-mounted device configured to be worn by a user. In some embodiments, the head-mounted device includes at least one microphone, at least one image sensor, and processing circuitry configured to receive an audio signal detected by the at least one microphone, the audio signal originating from a signal source. The processing circuitry is configured to determine the location of the signal source based on the received audio signal. The processing circuitry is further configured to receive image data from the at least one image sensor, the image data being associated with at least one of the following: the face of the user and at least one of the eyes of the user. The processing circuitry is further configured to determine the gaze direction of the user based on the received image data; and determine a user instruction based on the determined location and the determined gaze direction. Brief Description of the Drawings

[0005] A more complete understanding of the embodiments described herein and their attendant advantages and features will be more readily understood by reference to the following detailed description when considered in conjunction with the accompanying drawings, in which:

[0006] Figure 1 is a schematic diagram of various devices and components according to some embodiments of the present invention;

[0007] Figure 2Block diagram of an exemplary head-mounted device according to some embodiments of the present invention;

[0008] Figure 3 Block diagram of an exemplary hand-held device according to some embodiments of the present invention;

[0009] Figure 4 Illustration of a technique for sound localization using a head-mounted device according to some embodiments of the present invention;

[0010] Figure 5 Illustration of a technique for gaze tracking using a head-mounted device according to some embodiments of the present invention;

[0011] Figure 6 Illustration of another technique for gaze tracking using a head-mounted device according to some embodiments of the present invention;

[0012] Figure 7 Illustration of another technique for gaze tracking using a head-mounted device according to some embodiments of the present invention;

[0013] Figure 8 Illustration of another technique for gaze tracking using a head-mounted device according to some embodiments of the present invention;

[0014] Figure 9 Illustration of another technique for sound localization using a head-mounted device according to some embodiments of the present invention;

[0015] Figure 10 Illustration of another technique for sound localization using a head-mounted device according to some embodiments of the present invention;

[0016] Figure 11 Illustration of another technique for sound localization using a head-mounted device according to some embodiments of the present invention; and

[0017] Figure 12 Flowchart of an exemplary method of a head-mounted device according to some embodiments of the present invention. Detailed Description

[0018] Before describing the exemplary embodiments in detail, it should be noted that the embodiments mainly lie in the combination of device components and processing steps related to signal source localization and gaze tracking for first responders. Therefore, system and method components have been represented by conventional symbols in the drawings, showing only those specific details relevant to understanding the embodiments of the present disclosure, so as not to obscure the present disclosure with details that are obvious to those of ordinary skill in the art who benefit from the present disclosure herein.

[0019] As used herein, relational terms such as "first" and "second", "top" and "bottom", etc. may be used solely to distinguish one entity or element from another entity or element, and do not necessarily require or imply any physical or logical relationship or order between these entities or elements. The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the concepts described herein. As used herein, unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" are intended to also include the plural forms. It should also be understood that the terms "comprises", "comprising", "includes", and / or "including", when used herein, specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0020] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. It should also be understood that the terms used herein should be interpreted as having a meaning consistent with their meaning in the context of this specification and the relevant art, and should not be construed in an idealized or overly formal sense unless expressly so defined herein.

[0021] In the embodiments described herein, the conjunctive term "communicating with" and the like may be used to indicate electrical or data communication, which may be accomplished, for example, by physical contact, induction, electromagnetic radiation, wireless telecommunication signaling, infrared signaling, or optical signaling. One of ordinary skill in the art will appreciate that multiple components may interoperate and that modifications and variations may achieve electrical and data communication.

[0022] In the embodiments described herein, the term "signal source" may be any detectable signal that may, for example, indicate the distress status of a person (such as a first responder) and / or a device (such as a handheld alert device) and / or be associated therewith. For example, the signal source may include an audible sound, such as an alert sound emitted by a personal alert device, which may include a shrill (e.g., high-frequency) sound and / or a low-frequency sound. The signal source may include sounds associated with an emergency / life-threatening situation, such as sounds generated by a person in distress, by equipment / machinery, by running water, by an explosion, etc. The signal source may also / alternatively include a radio signal / beacon, such as a distress radio signal emitted by a personal alert device. Other types of detectable signals may be employed without departing from the scope of the present disclosure.

[0023] Now referring to the drawings, where like reference numerals refer to like elements, Figure 1An embodiment of a signal source localization detection system 10 utilizing a head-mounted device 12 worn by a user 14 is shown. The head-mounted device may include a mask body 15, lenses 16, top microphones 18a-b (collectively "microphones 18"), and side microphones 20a-b (collectively "microphones 20"). The top microphones may be molded into the head-mounted device 12, for example, around the perimeter of the top portion of the lens 16, and the side microphones may be molded into the head-mounted device 12, for example, around the side portion of the perimeter of the lens 16. The head-mounted device 12 may include, for example, a display 22 integrated into the lens 16 for providing visual indications / messages to the user 14 of the head-mounted device 12. The head-mounted device 12 may include, for example, a speaker 26 integrated into the head-mounted device 12 for providing audio indications / messages to the user 14, and may include, for example, a user microphone 28 integrated into the head-mounted device 12 for receiving verbal commands from the user 14. The microphones 18 and / or the microphones 20 may be configured to detect audio signals, such as sounds originating from a signal source 30, which may be a sound generating object / entity / event (e.g., wood breaking, a person shouting / breathing, a pressurized gas release (e.g., a jet release), etc.) in the environment / vicinity of the user 14 of the head-mounted device 12. The signal source localization detection system 10 may include a handheld device 31, which may communicate with the head-mounted device 12, for example, via a wired / wireless connection. In some embodiments, the head-mounted device 12 may be a mask, such as a mask that is part of a respirator.

[0024] Now referring to Figure 2 , the head-mounted device 12 may include hardware 32 that includes the microphones 18, the microphones 20, the display 22, the speaker 26, the microphone 28, an accelerometer 34, a light emitter 36, an image sensor 38, a communication interface 40, and processing circuitry 42. The processing circuitry 42 may include a processor 44 and a memory 46. In addition to or instead of a processor (such as a central processing unit) and a memory, the processing circuitry 42 may include an integrated circuit system for processing and / or control, such as one or more processors and / or processor cores and / or an FPGA (field programmable gate array) and / or an ASIC (application specific integrated circuit) adapted to execute instructions. The processor 44 may be configured to access (e.g., write to and / or read from) the memory 46, which may include any kind of volatile and / or non-volatile memory, such as a cache and / or a buffer memory and / or a RAM (random access memory) and / or a ROM (read only memory) and / or an optical memory and / or an EPROM (erasable programmable read only memory). The hardware 32 may be removable from the mask body 15 to allow for replacement, upgrade, etc., or may be integrated as part of the head-mounted device 12.

[0025] The head - mounted device 12 may also include software 48 stored, for example, inside a memory 46 or in an external memory (e.g., a database, a storage array, a network storage device, etc.) accessible by the head - mounted device 12 via an external connection. The software 48 may be executable by the processing circuitry 42. The processing circuitry 42 may be configured to control any of the methods and / or processes described herein and / or cause such methods and / or processes to be executed, for example, by the head - mounted device 12. The processor 44 corresponds to one or more processors 44 for performing the functions of the head - mounted device 12 described herein. The memory 46 is configured to store data, programming software code, and / or other information described herein. In some embodiments, the software 48 may include instructions that, when executed by the processor 44 and / or the processing circuitry 42, cause the processor 44 and / or the processing circuitry 42 to perform the processes described herein with respect to the head - mounted device 12. For example, the head - mounted device 12 may include a gaze tracker 50 configured to perform one or more functions of the head - mounted device 12 as described herein, such as detecting a gaze point (i.e., where the user 14 is looking), tracking the 3 - D line of sight of the user 14, tracking the movement of the user 14's eyes, the user 14's face, etc., as described herein. The processing circuitry 42 of the head - mounted device 12 may include a sound locator 52 configured to perform one or more functions of the head - mounted device 12 as described herein, such as determining the location / origin of the signal source 30, as described herein. The processing circuitry 42 of the head - mounted device 12 may include a sound classifier 54 configured to perform one or more functions of the head - mounted device 12 as described herein, such as classifying, tagging, and / or identifying the cause / type of the signal source 30, as described herein. The processing circuitry 42 of the head - mounted device 12 may include a user interface 56 configured to perform one or more functions of the head - mounted device 12 as described herein, such as displaying to the user 14 (e.g., using the display 22) or announcing (e.g., using the speaker 26) an indication / message, such as an indication of the location of the signal source 30 and / or an indication of the distance / direction of the signal source 30 relative to the user 14; and / or receiving from the user 14 an oral command and / or other commands from the user 14 (e.g., the user 14 presses a button communicating with the processing circuitry 42, the user 14 interacts with a separate device such as a smartphone or a handheld device 31, which conveys the user's interaction to the processing circuitry 42 via the communication interface 40, etc.) or receiving other commands from other users (e.g., via a remote server that communicates with the processing circuitry 42 via the communication interface 40), as described herein.

[0026] Although Figure 1Two of each of microphones 18 and 20 are shown, but it should be understood that the embodiments are not limited to two sets of two microphones, and there may be different numbers of sets of microphones, each set having a different number of individual microphones.

[0027] The display 22 can be implemented by any device (a stand-alone device or part of the head-mounted device 12) that can be configured to display an indication / message to the user 14 (e.g., an indication regarding the location of the signal source 30 and / or an indication regarding the distance / direction of the signal source 30 relative to the user 14). In some embodiments, the display 22 can be configured to display an icon (e.g., an arrow) that indicates in which direction the user should adjust his gaze, e.g., as determined by the sound locator 52 and / or the gaze tracker 50. In some embodiments, the display 22 can be configured to display the relative distance separating the user from the signal source 30 (e.g., "5 meters"), e.g., as determined by the sound locator 52 and / or the gaze tracker 50. In some embodiments, the display 22 can be configured to display the predicted classification / type / tag / etc. of the signal source 30, e.g., as determined by the sound classifier 54. In some embodiments, the display 22 can be configured to display an indication of the location of the signal source 30, e.g., as an AR overlay on the lens 16, such as by drawing a circle as an augmented reality (AR) overlay on the area of the lens 16 and / or the display 22 corresponding to the location of the signal source 30 within the field of view of the user 14. In some embodiments, if the location of the signal source 30 is outside the field of view of the user 14, the display 22 can be configured to instruct the user 14 to change direction (e.g., turn left, turn right, look up, look down, turn around, etc.).

[0028] The speaker 26 can be implemented by any device (a stand-alone device or part of the head-mounted device 12) that can be configured to produce sound that the user 14 can hear while wearing the head-mounted device 12, and can be configured to announce an indication / message to the user 14 (e.g., using the speaker 26), such as an indication regarding the location of the signal source 30 and / or an indication regarding the distance / direction of the signal source 30 relative to the user 14. In some embodiments, the speaker 26 is configured to provide an audio message corresponding to the indications described above with respect to the display 22.

[0029] The microphone 28 can be implemented by any device (a stand-alone device or part of the head-mounted device 12 and / or the user interface 56) that can be configured to detect the spoken commands of the user 14 while the user 14 is wearing the head-mounted device 12.

[0030] The accelerometer 34 can be implemented by any device (a stand-alone device or part of the head-mounted device 12) that can be configured to detect the acceleration of the head-mounted device 12.

[0031] The light emitter 36 can be implemented by any device (a stand-alone device or a part of the head-mounted device 12), which can be configured to generate light such as infrared radiation and direct the generated light into the eyes of the user 14 for detecting the positioning of the iris / cornea of the user 14 and / or detecting the gaze direction of the user 14. The direction, phase, amplitude, frequency, etc. of the light emitted by the light emitter 36 can be controlled by the processing circuit 42 and / or the gaze tracker 50. The light emitter 36 can include multiple light emitters (e.g., co-located and / or mounted at different positions on the head-mounted device 12), or can include a single light emitter.

[0032] The image sensor 38 can be implemented by any device (a stand-alone device or a part of the head-mounted device 12), which can be configured to detect images such as images of the eyes, face of the user 14 and / or images of the surrounding environment of the user 14 such as an image of the signal source 30. The image sensor 38 can include multiple image sensors (e.g., co-located and / or mounted at different positions on the head-mounted device 12 and / or on other devices / equipment communicating with the head-mounted device 12 via the communication interface 40), or can include a single image sensor.

[0033] The communication interface 40 can include a radio interface configured to establish and maintain a wireless connection (e.g., via a public land mobile network with a remote server, with a handheld device such as a smart phone, etc.). The radio interface can be formed as or can include, for example, one or more radio frequency (RF) transmitters, one or more RF receivers, and / or one or more RF transceivers. The communication interface 40 can include a wired interface configured to establish and maintain a wired connection (e.g., an Ethernet connection, a universal serial bus connection, etc.). In some embodiments, the head-mounted device 12 can send sensor readings and / or data (e.g., image data, orientation data, etc.) from one or more of the microphone 18, microphone 20, display 22, speaker 26, microphone 28, accelerometer 34, light emitter 36, image sensor 38, communication interface 40, and processing circuit 42 to another head-mounted device 12 (not shown), a handheld device 31, and / or a remote server (e.g., an incident command server, not shown) via the communication interface 40.

[0034] In some embodiments, microphones 18 and 20 may be mounted / arranged to optimize the reception of sound in a direction (e.g., of sound from signal source 30) and / or to optimize the strength of the shaping of the head-mounted device 12. In other embodiments, microphones 18 and 20 may be incorporated into various parts of the head-mounted device 12, such as the front, back, sides, top, bottom, etc. of the head-mounted device 12, to optimize sound detection from multiple directions. In some embodiments, microphones 18 and 20 may be omnidirectional / non-directional to detect sound in all directions. In other embodiments, microphones 18 and 20 may be directional to detect sound in a specific direction relative to the head-mounted device 12.

[0035] In some embodiments, the user interface 56 and / or the display 22 may be a superimposed / enhanced reality (AR) overlay that may be configured such that a user 14 of the head-mounted device 12 can view through the transparent lens 16, and the images / icons displayed on the display 22 are presented to the user 14 of the head-mounted device 12 as being superimposed on the transparent / translucent field of view (FOV) through the lens 16. In some embodiments, the display 22 may be separate from the lens 16. The display 22 may be implemented using a variety of techniques known in the art, such as a liquid crystal display built into the lens 16, an optical head-mounted display built into the head-mounted device 12, a retinal scan display built into the head-mounted device 12, etc.

[0036] In some embodiments, the gaze tracker 50 may use a variety of techniques known in the art to track the gaze and / or eye movements of the user 14. In some embodiments, the gaze tracker 50 may direct light (e.g., infrared / near-infrared light emitted by the light emitter 36) into the eyes of the user 14 (e.g., into the iris and / or cornea). The emitted light may be reflected on each eye corneal surface and produce a "flash", and the position of each flash may be detected, for example, using an image sensor 38, which may be configured to filter the detected light (e.g., such that only infrared / near-infrared light is detected). The gaze tracker 50 (e.g., using the image sensor 38) may detect points on each eye of the user 14 corresponding to the pupil center in each eye. The gaze tracker 50 may calculate the relative movement / distance between the pupil center of each eye and the flash position. For example, the gaze tracker 50 may calculate the optical axis, which is a vector connecting the pupil center, the corneal center, and the center of the eyeball. The gaze tracker 50 may calculate the visual axis, which is a vector connecting the fovea and the corneal center. The visual axis and the optical axis may intersect at the corneal center (also known as the nodal point of the eye). The gaze tracker 50 may utilize pre-configured / estimated physiological data (e.g., stored in the memory 46) regarding eye dimensions (e.g., corneal curvature, eye diameter, distance between the pupil center and the corneal center, etc.) to estimate the direction and angle of the optical axis, and this data may be based on the demographic information of the user 14 (e.g., male users and female users may have different average / estimated eye dimensions). The crossing angle between the flash vector and the pupil center vector may be used to estimate the angle between the optical axis and the visual axis. Using the estimated optical axis, the estimated crossing angle, and / or the pre-configured / estimated physiological information, the gaze tracker 50 may estimate the visual axis corresponding to the estimated gaze of the user.

[0037] In some embodiments, the gaze tracker 50 may utilize regression and / or machine learning models to estimate the gaze direction of the user 14. For example, the gaze tracker 50 (e.g., using the image sensor 38) may detect the physical / geometric features of the face / head / eyes of the user 14, and may use a machine learning model to determine the gaze direction based on the detected features.

[0038] In some embodiments, the gaze tracker 50 may be configured to perform a calibration procedure. For example, the user 14 of the head-mounted device 12 may initiate the calibration procedure, for example, when using the device for the first time. The calibration procedure may include, for example, displaying reference points on the display 22, instructing (e.g., using visual and / or audio commands via the user interface 56) the user 14 to direct his gaze to the reference points, and adjusting one or more parameters utilized by the gaze tracker 50 based thereon. Without departing from the scope of the present disclosure, other calibration procedures may be used to improve the accuracy of the gaze tracker 50, such as using machine learning (e.g., based on a dataset of multiple users of the head-mounted device 12).

[0039] Without departing from the scope of the present invention, the gaze tracker 50 can use any technique known in the art for determining / estimating the gaze of the user 14.

[0040] The sound locator 52 can use a variety of techniques known in the art to determine the location, relative direction, and / or relative distance of the signal source 30 to the user 14. In some embodiments, the sound locator 52 can apply a head-related transfer function (HRTF) to the signals received by the microphone 18, and this HRTF can be used to determine the left / right / horizontal orientation of the signal source relative to the user 14. The sound locator 52 applies the HRTF to the signals received by the microphone 20, compares the HRTF results of the microphone 20 with the HRTF results of the microphone 18, and determines the up / down / vertical orientation of the signal source 30 relative to the user 14 based on this comparison. The sound locator 52 can be configured to determine a vector passing through the center of the plane formed by four points corresponding to the locations of the microphone 18 and the microphone 20 from an appropriate point (such as the bridge of the nose of the user 14); this vector can point to the source of the signal source 30.

[0041] In some embodiments, the sound locator 52 can additionally or alternatively utilize image data corresponding to the visual environment of the user 14, such as from the image sensor 38 and / or from the handheld device 31, to locate the signal source 30, for example, using edge detection, boundary tracking, machine learning techniques, etc. When estimating the location of the signal source 30, as an alternative to or in addition to using sound data, the sound locator 52 can also utilize radio signal data, such as from the antenna array 60 of the handheld device 31, as described herein.

[0042] Without departing from the scope of the present invention, the sound locator 52 can use any technique known in the art for determining / estimating the location of the signal source 30.

[0043] In some embodiments, the sound classifier 54 can estimate / predict / determine the type / tag / category / cause of the sound originating from the signal source 30 based on the sounds detected by the microphone 18 and / or 20. For example, it can be determined that the sound has the characteristics of events such as wood breaking, a person shouting / breathing, a person falling to the ground alarm, a pressurized gas release (e.g., a jet release), etc. For example, in some embodiments, the sound classifier 54 can utilize a sound / sound tag library to classify the signal source, for example, by applying regression / machine learning techniques to a pre-configured data set / library of labeled sounds / sound tags (e.g., stored in the memory 46), generating a model for predicting / classifying the detected sounds, and using this model to classify specific detected sounds. Without departing from the scope of the present disclosure, other noise classification techniques known in the art can be used.

[0044] Now refer to Figure 3 ,the handheld device 31 may include hardware 58, which includes an antenna array 60, an image sensor 62, a communication interface 64, and processing circuitry 66. The processing circuitry 66 may include a processor 68 and a memory 70. In addition to or instead of a processor (such as a central processing unit) and a memory, the processing circuitry 66 may include an integrated circuit system for processing and / or control, such as one or more processors and / or processor cores and / or FPGA (field programmable gate array) and / or ASIC (application specific integrated circuit system) adapted to execute instructions. The processor 68 may be configured to access (e.g., write to and / or read from) the memory 70, which may include any kind of volatile and / or non-volatile memory, such as cache and / or buffer memory and / or RAM (random access memory) and / or ROM (read only memory) and / or optical memory and / or EPROM (erasable programmable read only memory).

[0045] The handheld device 31 may also include software 72 stored, for example, inside the memory 70 or in an external memory (such as a database, a storage array, a network storage device, etc.) accessible by the handheld device 31 via an external connection. The software 72 may be executable by the processing circuitry 66. The processing circuitry 66 may be configured to control any of the methods and / or processes described herein and / or cause such methods and / or processes to be executed, for example, by the handheld device 31. The processor 68 corresponds to one or more processors 68 for performing the functions of the handheld device 31 described herein. The memory 70 is configured to store data, programming software code, and / or other information described herein. In some embodiments, the software 72 may include instructions that, when executed by the processor 68 and / or the processing circuitry 66, cause the processor 68 and / or the processing circuitry 66 to perform the processes described herein with respect to the handheld device 31. For example, the handheld device 31 may include a locator 74 configured to perform one or more functions of the handheld device 31 as described herein, such as determining the location / origin of the signal source 30, as described herein.

[0046] The antenna array 60 may be implemented by any device (a stand-alone device or part of the handheld device 31) that may be configured to detect a beacon signal from a wireless beacon, such as a wireless beacon transmitted by an alarm device attached to the equipment of a fallen first responder. The antenna array 60 may include one or more directional antennas for following the wireless beacon.

[0047] The image sensor 62 may be implemented by any device (a stand-alone device or part of the handheld device 31) that may be configured to detect an image (such as a thermal image) and / or may be configured to detect light within and / or outside the visible spectrum.

[0048] The communication interface 64 may include a radio interface configured to establish and maintain a wireless connection (e.g., with a remote server via a public land mobile network, with the head-mounted device 12 via a Bluetooth connection, etc.). The radio interface may be formed as or may include, for example, one or more radio frequency (RF) transmitters, one or more RF receivers, and / or one or more RF transceivers. The communication interface 64 may include a wired interface configured to establish and maintain a wired connection (e.g., an Ethernet connection, a universal serial bus connection, etc.). In some embodiments, the handheld device 31 may send sensor readings and / or data (e.g., image data, orientation data, wireless beacon data, etc.) to the head-mounted device 12 and / or to a remote server (e.g., an incident command server, not shown) via the communication interface 64 and / or receive the sensor readings and / or data from one or more of the antenna array 60 and the image sensor 62. In some embodiments, the sound locator 52 may utilize such sensor / image data to determine the location of the signal source 30.

[0049] In some embodiments, the head-mounted device 12 (e.g., using the gaze tracker 50) may monitor the functionality of the iris of the user 14's eyes as a measure of focus. The head-mounted device 12 may include one or more microphones (e.g., microphone 18 and microphone 20). In some embodiments, microphone 18 and / or 20 may be mounted on either side, and / or on the front and / or back of the head-mounted device 12. In some embodiments, microphone 18 and / or 20 may be directional, such that sound is detected from the front of the head-mounted device 12. In other embodiments, microphone 18 and / or 20 may be placed on the body of the user 14 and / or at other locations on the head of the user 14. In some embodiments, microphone 18 and / or 20 may be placed on and / or pointed at the sides and back of the user 14, e.g., to provide directionality of sound. In some embodiments, microphone 18 and / or 20 may be fixed to the head-mounted device 12 by a mounting member (not shown) molded into the mask body 15 of the head-mounted device 12 (e.g., around the top portion of the lens 16).

[0050] In some embodiments, the sound classifier 54 may compare the detected sound (e.g., from microphone 18 and / or 20) with a library of sounds / sound signatures specific to events such as wood breaking, a person shouting or breathing, a pressurized gas release (such as a jet release), a person falling to the ground alarm, etc.

[0051] In some embodiments, the head-mounted device 12 may be a breathing mask, goggles, a face shield, and / or glasses, and / or may be part of a self-contained breathing apparatus (SCBA).

[0052] In some embodiments, microphones 18 and / or 20 may provide an audio cue (e.g., by detecting an alarm, voice, other audio, etc.). The head-mounted device 12 may apply the HRTF algorithm, for example using a sound locator 52, and determine an approximate location of the signal source 30 in a three-dimensional (x-y-z) space. In some embodiments, the determined approximate location may be iconically represented on the display 22 as an area in front of the user 14 (e.g., within the field of view of the user 14). In the case where the signal source 30 is outside the display 22 (e.g., outside the field of view of the user 14), an arrow or other icon may appear in the visual space (e.g., within the field of view of the user 14) to show the user 14 where to look. In some embodiments, when the user 14 moves throughout the environment, the user interface 56 may utilize the accelerometer 34 to adjust the indication. For example, the accelerometer 34 may be configured to sense movement of the user 14's body and / or head, such as sensing that the user 14 has moved his head to face the signal source 30.

[0053] In some embodiments, once the user 14 is looking / gazing generally in the vicinity of the signal source 30, the head-mounted device 12 (e.g., using the gaze tracker 50) may monitor the user 14's eyes for visual cues as to whether the user 14 is looking in the correct location (e.g., in the direction of the determined approximate location of the signal source 30). In some embodiments, the display 22 may display a series of virtual concentric boxes with location icons that may indicate the location (e.g., as determined by the sound locator 52 using the HRTF function) and / or may indicate an icon representing the focus location / gaze direction of the user 14's eyes. As the user 14 progresses towards the location of the signal source 30, the display 22 may be continuously / periodically refreshed to provide the user 14 with updated cues as to the location of the signal source 30 until the location is reached and / or until the user 14 terminates the program (e.g., via a voice command or a toggle button).

[0054] In some embodiments, the user 14 may activate / deactivate the search program via a voice command (e.g., via the microphone 28 and / or the user interface 56) and / or via a toggle switch / button (e.g., in communication with the user interface 56). In some embodiments, one or more components of the head-mounted device 12 such as the gaze tracker 50, the sound locator 52, the sound classifier 54, and / or the user interface 56 may be located in / executed by a separate circuitry (elsewhere on the user 14's body), such as in a handheld device 31 or may communicate remotely with the head-mounted device 12 via a wired or wireless connection to the communication interface 40, for example.

[0055] In some embodiments, the sound locator 52 may be configured to perform a calibration procedure, for example, when the user 14 first uses the head-mounted device 12. In some embodiments, the calibration procedure is configured to compensate for any head and / or hearing protection worn by the user 14, which may affect the directionality of sound.

[0056] In some embodiments, the sound locator 52 may be configured to utilize thermal imaging and / or other visual data (e.g., received from the image sensor 38 and / or from other image sensors, such as the image sensor 62 in a separate handheld device 31 that communicates with the head-mounted device 12 via the communication interface 40) when determining the location of the signal source 30. Such visual data may include images of light outside the visible spectrum.

[0057] In some embodiments, the handheld device 31 that communicates with the head-mounted device 12 is configured to follow a radio beacon, for example, using the antenna array 60. The handheld device 31 is configured to be scanned by the user 14 back and forth, up and down, etc., to attempt to identify the maximum beacon intensity, for example, by comparing measurements of the signal strength detected by the antenna array 60. In some embodiments, the handheld device 31 may determine the direction of the radio beacon by making multiple directional measurements (e.g., using the antenna array 60), which may form a virtual cone structure as the user 14 approaches the beacon source. The locator 74 may be configured to identify the detected segment of the cone and back-calculate it to identify the vertex of the cone, which represents the location of the signal source 30. For example, the signal source 30 may be a personal alert / distress alert device worn by a fallen first responder (e.g., a Scott Pak-Alert Personal Alert Safety System (PASS) device), which may emit an audible sound (e.g., a piercing sound) and / or may emit a radio signal / beacon when activated. Detecting the radio signal may be advantageous in addition to / as an alternative to detecting the sound signal, for example, in scenarios where detecting the audible sound signal is impractical, such as when the fallen first responder is at least partially submerged underwater, or in scenarios where the sound chamber of the personal alert device has been blocked by debris due to environmental conditions, etc. Thus, detecting the radio signal in addition to the sound signal may improve the accuracy of estimating the location / direction of the signal source 30.

[0058] In some embodiments of the present disclosure, the handheld device 31 and / or the head-mounted device 12 may be configured to determine the location of the personal alert device based on the characteristics of the audible sound signal and / or the transmitted radio signal, and the head-mounted device 12 may be configured to determine whether the user 14 is looking at and / or facing the direction of the location of the personal alert device, and / or may be configured to guide the user 14 to the location of the signal source 30, even in scenarios where the audible sound signal cannot be detected and / or the visibility is at least partially blocked.

[0059] Locator 74 can detect signals (such as sound signals / waves, radio signals / waves, etc.) at multiple different points in the entire environment (i.e., 3D space), and record the characteristics of those signals, such as signal strength, power, amplitude, frequency, noise, etc. Locator 74 can construct / utilize a 3D model based on the detected signals to determine / estimate the source of signal source 30 and / or guide user 14 to signal source 30. As a non-limiting example, the signal can be modeled as a 3D cone, and Locator 74 can utilize one or more formulas known in the art (such as the surface equation for a straight cone) to determine one or more characteristics of the signal. The handheld device 31 can detect signals as user 14 moves through the environment, and / or user 14 can intentionally move the handheld device (e.g., in a sweeping motion) to collect detected signal data points at various positions relative to user 14.

[0060] For example, the personal alert device of a fallen first responder (i.e., signal source 30) can emit a sound signal and / or a radio signal / beacon, which can be represented as a spherical field in 3D space. Inside the field, there can be layers of constant field strength, which can be represented as concentric spheres (i.e., spheres of radio energy). The handheld device 31 can include a large directional antenna (e.g., as part of antenna array 60) and can also include a display / indicator to indicate the signal strength to user 14, and / or can provide such information to user 14 via the head-mounted device 12 (e.g., display 22 and / or speaker 26). The handheld device 31 can be configured to sample one or more points in 3D space as user 14 moves through the environment to generate / estimate the shape of the signal (e.g., a virtual cone) and / or predict the source of the signal (i.e., signal source 30). The handheld device 31 can be configured to capture the detected signal data (e.g., signal strength) and position / location data, and determine the shape (e.g., a cone) and / or the source of the sound and / or radio signal (signal source 30), and can provide the data to the user, e.g., as a translucent conical shape overlaid on the display 22 of the head-mounted device 12 to assist in guiding user 14 towards signal source 30.

[0061] Figure 4Depicts an exemplary scenario in accordance with some embodiments of the present disclosure. In such a scenario, user 14 is looking directly forward, and at this time, the sound detected from signal source 30 is detected at a downward angle and to the left of user 14. For example, the head-mounted device 12 and / or the sound locator 52 may apply the head-related transfer function (HRTF) based on the characteristics of the audio signals received from microphones 18 and / or 20 to determine the HRTF focus 76 and the vector 78 from the focus 76 to the signal source 30. In some embodiments, such directional measurements may be approximations for guiding user 14 closer to the target (e.g., signal source 30). As user 14 approaches the target, the quality of the measurement may improve. For example, the gaze tracker 50 may determine that user 14 can direct his gaze straight ahead, e.g., as represented by the vector 80 in the gaze direction from the eye 82a (and / or eye 82b). For example, the vector 80 may be projected from the cornea of the eye 82a to the back of the eye 82a, and the average diameter of the eye 82a may be a known / predetermined value, e.g., based on population averages and / or demographic information (e.g., of user 14). The back of the eyes 82a-b is fixed to allow the optic nerve to pass through the skull. The vector 80 may be determined based on, for example, detecting the positioning of the iris and / or cornea of the eyes 82a and / or 82b using infrared radiation emitted from the light emitter 36, which may be used by the gaze tracker 50 to determine the vector 80. Using the lens 16 (which may be part of the face mask of the head-mounted device 12) as a plane and adjusting the depth and / or curvature difference using known geometric relationships, two vectors may be determined on a common basis, which may allow the application of Euclid's parallel postulate. In particular, if the sum of the interior angles of two lines is 180 degrees or less, then these lines are parallel or convergent. The specific geometric formulas to be applied are known in the art and are outside the scope of the present disclosure. Indicators within the face mask (e.g., the user interface 56 and / or the display 22) may display arrows, concentric circle displays, or other directions / icons to the user to indicate that user 14 should adjust his gaze.

[0062] For example, the gaze direction of user 14 determined by the gaze tracker 50 may include a direction pointing to user 14's face and / or a direction pointing to user 14's eyes. The gaze tracker 50 may employ various gaze tracking techniques known in the art. Eye tracking involves locating a fixed point on the surface of user 14's eyes and monitoring the movement of that fixed point. The gaze tracker 50 may employ various gaze tracking techniques known in the art.

[0063] As Figure 5As shown, gaze tracking or gaze direction determination techniques can utilize the commonality of the eye physiology, muscle tissue, orbits, geometric structures, etc. of a human population to fix one point and then track the pupil of the eye to locate a second point. These two points can define a line / vector that can be defined by an equation. There are various techniques known in the art for determining such vectors. For example, the first point can be the center of the eye pupil, and the second point can be the light reflection on the cornea, such as the reflection of light emitted by the light emitter 36 and captured by the image sensor 38. The visual axis corresponding to the gaze of the user 14 can be estimated by determining, for example, the keratometric angle (i.e., the angle between the optical axis and the visual axis with the corneal surface as the vertex) based on the calibration of the eyes 82a-b of the user 14.

[0064] In some embodiments, the gaze tracker 50 can perform eye detection and / or gaze mapping. Feature-based mapping can be employed, for example, using a 2D model (support vector regression, neural network, etc.) or a 3D model. Landmark-based methods are shown in Figure 6 which set coordinates and / or points, for example, in a.dat file. These features can utilize datasets and / or machine learning techniques known in the art to map facial features to the predicted gaze direction. In some embodiments, a dedicated dataset can be generated when the user 14 wears the head-mounted device 12. The eye / face / image / location data associated with the user 14 can be stored, for example, in the head-mounted device 12 and can be used to generate / improve the machine learning model, thereby improving the prediction accuracy over time as more data is collected. Additionally or alternatively, the head-mounted device 12 can use, for example, a synthetic dataset with a large source of participants. Such datasets can include infrared image samples of the faces and / or eyes of the participants. Using the dataset / machine learning enables the gaze tracker 50 to adapt / adjust / calibrate to various users 14. In some non-limiting embodiments, the gaze tracker 50 can utilize various computer vision libraries, such as Python OpenCV. Although the gaze tracker 50 can be implemented using any suitable hardware and / or software arrangement, some embodiments can utilize / execute software code written in C++ as well as Python, making it more suitable for deployment on various microcontrollers.

[0065] In some embodiments, the gaze tracker 50 is configured to estimate the gaze of the user 14, as Figure 7As shown. In these embodiments, the diameter of the eye 82a (or 82b) is a known value and / or an estimated value. The estimation can be based on population averages, the demographic information of the user 14, and / or machine learning techniques. For example, an image sensor 38 can be used to measure the distance between the line representing the eye diameter and the line representing the light reflection. Using known geometric formulas and relationships (the specific details of which are beyond the scope of the present disclosure), the angles between the eye center and the flash and between the characteristic light reflection line and the optical axis can be determined. For example, when the light rays from the light reflection (e.g., emitted by the light emitter 36) approach the light rays from the pupil center, the right-angled triangle formed becomes a single triangle instead of two back-to-back triangles. Given an approximate eye diameter (or the base of the triangle), the height of the right-angled side of the triangle can be measured (e.g., using the image sensor 38). For example, the height of the triangle can be taken as the distance from the line representing the eye diameter from the corneal surface to the back of the eyeball, and the hypotenuse of the triangle can be taken as the distance from the corneal surface at the flash to the back of the eyeball. The third right-angled side of the triangle is formed by the straight line from the flash on the cornea to the line representing the eye diameter.

[0066] These techniques can be improved using machine / neural network learning. For example, the neural network can be initially trained using publicly available datasets and can be further trained / boosted for a specific group of users 14 (e.g., employees of a specific fire department) by collecting data from actual use and / or from artificial training / calibration scenarios. For example, by setting sound targets to simulate the signal source 30 and instructing the user 14 to complete movement patterns such as a mask seal sequence. During actual use and / or training / calibration use, the head-mounted device 12 can collect data (e.g., images of the user 14's eyes, face, environment, signal / location data of the signal source 30, etc.) to train the neural network.

[0067] For example, as Figure 7 shown, when the light reflection rays and the rays from the pupil center are close enough, a right-angled triangle 84 is formed, and there is no smaller right-angled triangle on the right side of the triangle 84. The tangent á is equal to the ratio of the opposite side to the adjacent side and is equal to the slope b of the light reflection rays. The equation of the line of the light reflection is in the form y = mx + b and thus y = (tan á)x. This analysis can be iteratively repeated in the circle to refine the equation in three-dimensional space. Multiple lines can be tested to select the best fit and avoid polar coordinates. By using multiple image sensors 38, the head-mounted device 12 can determine the flash / optical axis from multiple perspectives, which can improve the accuracy of the estimation.

[0068] For further illustration, as Figure 8As shown, as the light from the light reflection approaches the light from the optical axis of the pupil, the base of the right triangle becomes smaller and the angle á becomes smaller. If user 14 does not look directly at the source of the flash, there will be a measurable length to the base of the triangle. If user 14 looks directly at the source of the flash, the length of the base will approach zero. Known geometric formulas / relationships can be used to determine the right triangle, the details of which are outside the scope of this disclosure. The line defined as the optical axis can be compared with the line calculated according to HRTF (i.e., the direction of the signal source 30). When user 14 looks in the direction of the signal source 30, the slopes of the two lines should be equal. The base of the triangle can be fixed and represents the distance from the optical axis and the HRTF line exactly inside the lens 16. To adjust the slope, the head-mounted device 12 can instruct the user to look in the direction that minimizes the slope difference and / or brings the difference closer to zero. Trial and error can be used to determine how far user 14 must move his gaze and in which direction to minimize the slope difference. This process can be iteratively repeated, for example, using multiple image sensors 38 and multiple microphones 18 and 20 to improve the accuracy of the model, select the best fit, etc.

[0069] In some embodiments, the gaze tracker 50 determines the gaze direction based on two components: the direction the eyes are pointing and the direction the face is pointing. In some embodiments, the gaze tracker 50 employs / considers / compensates for the Wollaston Effect, the Mona Lisa Effect, and / or the mirror effect. The specific techniques for compensating for these effects are known in the art and are outside the scope of this disclosure.

[0070] In some embodiments, the sound locator 52 utilizes HRTF. HRTF is a measure of the difference in hearing between the right ear and the left ear of a listener (e.g., user 14). By placing the microphones 18 and / or 20 on either side of the mask, the sound locator 52 can simulate the auditory system in the form of a simplified head, without the need to consider the pinna structure, inefficiencies inside the ear, etc. HRTF can consider two signal collection points separated in space and use this information to determine / estimate the signal source location. Utilizing two pairs of microphones 18 and / or 20 can further improve the accuracy and / or provide additional information, such as how the head of user 14 is tilted and / or pointed relative to the signal source 30. The specific formulas for HRTF are known in the art and are outside the scope of this disclosure.

[0071] Because the head-mounted device 12 with microphones is not actually a human head with ears, source estimation can be simplified for the head-mounted device 12 with microphones 18 and 20. In particular, based on the HRTF formula, the head-mounted device 12 can determine that the user 14 is facing the source. In some embodiments, a simplified transfer function can be used to calibrate the face of the user 14. For example, a tightly fitting head-mounted device 12 can have a different transfer function compared to a loosely fitting head-mounted device 12.

[0072] In some embodiments, the head-mounted device 12 can include two or more sets of microphones, such as microphones 18 and 20. The sound locator 52 can apply a transfer function to the top two microphones 18 to derive a left-right orientation. The sound locator 52 can apply a transfer function to the bottom two microphones 20, compare it with the top two microphones 18, and derive an up-down orientation therefrom. For example, if two microphones 18a and 20a on the left side of the user 14's head detect a comparable sound intensity higher than two microphones 18b and 20b on the right side of the user 14's head, the source 30 can be predicted to be on the left side of the user 14. Similarly, if the top two microphones 18 have a comparable sound intensity higher than the two lower microphones 20, it is inferred that the source is above the user 14. Once the microphones 18 and 20 are balanced (i.e., the detected sound signals have similar amplitudes, intensities, etc.), the sound locator 52 can determine the direction of the source 30 relative to the user 14. In some embodiments, as Figure 9 shown, the microphones 18 and 20 can be arranged on the head-mounted device 12 such that the center of the plane 86 formed by the positioning of the microphones 18 and 20 is in the general positioning of the user 14's nose bridge. The sound locator 52 can calculate a line passing through the nose bridge and the focus 76 of the plane 86 (e.g., in the form of y = mx + b), and this line can be determined as the sound source 30. As Figure 7 shown geometric relationships / formulas (the details of which are well known in the art and outside the scope of this disclosure) can be used to estimate the line. The geometric position of the plane 86 can be defined by the microphones 18 and 20. The positioning of the bridge of the nose cup can be other points that define the line, which points in the direction of the source 30.

[0073] In some embodiments, the first line is formed by the direction of the user 14's gaze (e.g., as determined by the gaze tracker 50), and the second line is formed by the direction the user is facing (e.g., as determined by the gaze tracker 50). If the two lines coincide (i.e., y1 = mx1 and y2 = mx2 such that x1 = x2), then the gaze tracker 50 can determine that the user 14 is looking at the signal source 30. If the two lines do not meet this condition, using a comparison algorithm, the user interface 56 can notify the user 14 to adjust the direction of the user 14's face. The degree of adjustment can follow a trial-and-error algorithm. For example, after multiple trials, once the iteration of the slope change is less than a threshold (e.g., 10%), the process stops. Alternatively, a 2D or 3D least squares algorithm or various other geometric calculations known in the art (the specific details of which are beyond the scope of this disclosure) can be used to iteratively improve the accuracy. The ability of the user 14 to accurately adjust can be assisted by providing a display 22 attached to the accelerometer 34 and a representation of the two lines in that distance in the display 22 (e.g., as an AR overlay). Using this display 22, the user 14 can adjust his gaze to the direction of the signal source 30. In some cases, the two lines may not coincide or may be at least parallel. For example, the user 14 may point his face in the correct direction of the signal source 30, but may look up, down, right, or left at the signal source 30. As another example, the user 14 may be confused about the sound source and may be facing completely the wrong direction. As another example, the head-mounted device 12 can detect echoes or sound reflections. The sound locator 52 can be configured to compensate for sound reflections. For example, the sound locator 52 can assume that the amplitude of the sound will change after reflection, but the frequency will not. The sound locator 52 can utilize additional microphones (e.g., mounted on the side / back of the head-mounted device 12), compare the amplitude and / or frequency of the sound received from the side / back microphones with the amplitude and / or frequency of the sound received from the front microphones (18 and 20), and determine whether the front microphones 18 and 20 are detecting the sound or an echo of the sound. If the lines share a common slope and direction, then the sound locator 52 can determine that the user 14 is looking approximately near the sound source. The sound locator 52 can further analyze these lines using vertical angles and auxiliary lines according to geometric formulas known in the art that are beyond the scope of this disclosure.

[0074] Figure 10 and Figure 11 Another example of using HRTF to determine sound direction and gaze detection is shown. The sound direction (from the signal source 30) can be modeled as a straight line. The gaze direction can also be modeled as a straight line. The head-mounted device 12 uses HRTF to describe / determine the sound direction / source and the direction the user 14 is looking, and instructs the user 14 to change the gaze direction so that the two lines are parallel / converge, and this situation indicates that the user 14 is gazing in the direction of the signal source 30.

[0075] In some cases, there may be multiple estimates of the location of the signal source 30, such as a point behind the user 14 and a point in front of the user 14. The head-mounted device 12 may provide (e.g., via the display 22) multiple estimates of the location of the signal source 30 to the user 14, e.g., by displaying multiple vectors overlaid on the display 22, and the user 14 may determine which vector to follow. For example, the user 14 may have just walked through a room and not found the signal source 30 in the room, and the user 14 may use this information to decide to ignore the vector on the display 22 that points the user 14 back to the room and instead follow the vector that points to a new room that the user 14 has not previously entered.

[0076] In some embodiments, multiple head-mounted devices 12 associated with multiple users 14 may cooperate (e.g., by wirelessly transmitting data directly or indirectly to each other, by communicating with a remote server, etc.) to improve the accuracy of the estimated location of the signal source 30. For example, if multiple users 14 in a first responder team are present in an emergency scenario, each equipped with a corresponding head-mounted device 12, the detected signals (e.g., sound waves) from each head-mounted device 12 may be distributed to the other head-mounted devices 12 in the first responder team, where each head-mounted device may utilize the additional data to improve the accuracy of the signal source 30 location detection. Similarly, one or more users 14 in the first responder team may be equipped with handheld devices 31 that can detect radio signals transmitted by the signal source 30, as described herein, and the head-mounted devices 12 may utilize the location information (e.g., as determined by the locator 74) generated by one or more of the multiple handheld devices 31. As another example, if one head-mounted device 12 locates the signal source 30, the location data (e.g., geographical coordinates, distance / angle information, etc.) may be shared with the other head-mounted devices 12 in the first responder team.

[0077] Figure 12FIG. 0 is a flowchart of an exemplary method of a head-mounted device 12 according to some embodiments of the present invention. One or more of the boxes described herein may be performed by one or more elements of the head-mounted device 12, such as by one or more of the processing circuitry 42, microphone 18, microphone 20, display 22, speaker 26, microphone 28, accelerometer 34, light emitter 36, image sensor 38, communication interface 40, processing circuitry 42, processor 44, memory 46, software 48, gaze tracker 50, sound locator 52, sound classifier 54, and / or user interface 56. The head-mounted device 12 is configured to receive (block S100) an audio signal detected by at least one microphone (e.g., microphone 18 and microphone 20), the audio signal originating from a signal source 30. The head-mounted device 12 is configured to determine (block S102) the location of the signal source 30 based on the received audio signal. The head-mounted device 12 is configured to receive (block S104) image data from at least one image sensor 38, the image data being associated with at least one of the following: the face of the user 14 and at least one of the eyes of the user 14. The head-mounted device 12 is configured to determine (block S106) the gaze direction of the user 14 based on the received image data. The head-mounted device is configured to determine (block S108) a user instruction based on the determined location and the determined gaze direction.

[0078] In some embodiments, the user instruction is determined based on at least one of the relative distance and relative direction from the user 14 to the signal source 30. In some embodiments, the user instruction indicates the location and / or direction that the user is looking at.

[0079] In some embodiments, the head-mounted device 12 includes at least one light emitter 36 communicatively coupled to the processing circuitry 42. In some embodiments, the at least one light emitter 36 is configured to emit light into at least one of the eyes of the user 14 to cause at least one reflected flash for determining the gaze direction.

[0080] In some embodiments, determining the gaze direction of the user 14 includes using a machine learning model to determine a plurality of features based on the image data to predict the gaze direction based on the determined plurality of features.

[0081] In some embodiments, the processing circuitry is further configured to receive a radio signal from the signal source, and determining the location of the signal source is further based on the received radio signal.

[0082] In some embodiments, the processing circuitry 42 is further configured to determine a plurality of features based on the received audio signal. In some embodiments, the processing circuitry is configured to use a machine learning model to determine a sound classification based on the determined plurality of features to predict the sound classification of the signal source 30 based on the determined plurality of features.

[0083] In some embodiments, the head-mounted device 12 includes a user interface 56. In some embodiments, the user interface 56 is configured to display user instructions as augmented reality (AR) indications.

[0084] Those skilled in the art will appreciate that the embodiments of the present invention are not limited to what has been specifically shown and described above. Additionally, unless the contrary has been mentioned above, it should be noted that all the drawings are not drawn to scale. There can be various modifications and variations based on the above teachings.

Claims

1. A head-mounted device configured to be worn by a user, the head-mounted device comprising at least one microphone, at least one image sensor, and processing circuitry in communication with the at least one microphone and the at least one image sensor, the processing circuitry being configured to: Receive an audio signal detected by the at least one microphone, the audio signal originating from a signal source; Determine the localization of the signal source based on the received audio signal; Receive image data from the at least one image sensor, the image data being associated with at least one of the following: the face of the user and at least one of the eyes of the user; Determine the gaze direction of the user based on the received image data; Determine a user instruction based on the determined localization and the determined gaze direction.

2. The head-mounted device according to claim 1, wherein the user instruction is determined based on at least one of a relative distance and a relative direction from the user to the signal source.

3. The head-mounted device according to any one of claims 1 and / or 2, wherein the head-mounted device comprises at least one light emitter in communication with the processing circuitry, the at least one light emitter being configured to emit light into at least one of the eyes of the user to cause at least one reflected flash for determining the gaze direction.

4. The head-mounted device according to any one of claims 1, 2, and / or 3, wherein determining the gaze direction of the user comprises using a machine learning model to determine a plurality of features based on the image data to predict the gaze direction based on the determined plurality of features.

5. The head-mounted device according to any one of claims 1, 2, and / or 3, wherein the processing circuitry is further configured to receive a radio signal from the signal source, and the determination of the localization of the signal source is further based on the received radio signal.

6. The head-mounted device according to any one of claims 1, 2, 3, 4, and / or 5, wherein the processing circuitry is further configured to: Determine a plurality of features based on the received audio signal; Use a machine learning model to determine a sound classification based on the determined plurality of features to predict the sound classification of the signal source based on the determined plurality of features.

7. The head-mounted device according to any one of claims 1, 2, 3, 4, 5, and / or 6, wherein the head-mounted device comprises a user interface configured to display the user instruction as an augmented reality (AR) indication.