Head-mounted computing device using microphone beam steering
Adaptive beamforming in head-mounted devices adjusts sensitivity based on user head movements, enhancing audio quality and privacy by maintaining focus on the target speaker with reduced computational demands.
Patent Information
- Application Number
- JP2023544348
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-01-28
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2041-01-28
AI Technical Summary
Existing head-mounted computing devices face challenges in maintaining beamformed audio sensitivity alignment with user head movements, leading to potential privacy violations and reduced audio quality.
Adaptive beamforming techniques that utilize sensors like IMUs and cameras to track head movements, adjusting beamformed sensitivity dynamically to maintain focus on a target speaker while reducing computational demands.
Improves audio quality and privacy by maintaining beamformed sensitivity on the intended speaker, even with head movements, while minimizing processing requirements.
Smart Images

Figure 0007717170000001 
Figure 0007717170000002 
Figure 0007717170000003
Abstract
Description
Technical Field
[0001] Field of Disclosure The present disclosure relates to acoustic beam steering, and more particularly, to steering the beam of a microphone array of a head-mounted computing device.
Background Art
[0002] Background A head-mounted computing device can be configured to capture information from the environment and from the user. The captured information can be processed to determine the relative orientation and position of objects and the user in the environment so that a virtual scene is generated and displayed. As a result, the user can become aware that the environment includes both a real scene and a virtual scene that change as the user interacts with or moves within the environment. Thus, a head-mounted computing device can include numerous subsystems for capturing and displaying sensory information (e.g., auditory, visual) and for determining orientation and position (e.g., head pose). Thus, there may be an opportunity for a head-mounted computing device to assist the user with conversations. However, this assistance can provide an opportunity to violate the privacy of others.
Summary of the Invention
[0003] Summary In at least one aspect, the present disclosure generally describes a head-mounted computing device. The head-mounted computing device includes a microphone array that includes a plurality of microphones. The microphone array is configured to generate a beamformed audio signal according to the beamformed sensitivity of the microphone array based on the sound received by the plurality of microphones. The head-mounted computing device further includes a plurality of loudspeakers configured to transmit sound. The head-mounted computing device further includes a plurality of sensors configured to measure the orientation of the head-mounted computing device relative to a fixed reference system. The head-mounted computing device further includes a processor coupled to the plurality of microphones, the plurality of loudspeakers, and the plurality of sensors. The processor of the head-mounted computing device is configured by software instructions that direct the performance of a method. The method includes identifying the orientation of the reference system of the microphone array based on the orientation of the head-mounted computing device. The method further includes calculating a target direction relative to the orientation of the reference system. The method further includes directing the beamformed sensitivity of the microphone array in the target direction. The method further includes detecting a change in the orientation of the head-mounted computing device and, based on the detected change in the orientation of the head-mounted computing device, detecting a change in the orientation of the reference system to obtain an updated target direction relative to the reference system of the microphone array. The method further includes redirecting the beamformed sensitivity of the microphone array in the updated target direction.
[0004] According to possible realizations of the head-mounted computing device, the head-mounted computing device may include one or more (e.g., all) (or any combination thereof) of the following features.
[0005] In one possible implementation example of a head-mounted computing device, redirecting the beamformed sensitivity of a microphone array to an updated target direction includes reducing the sensitivity of the microphone array in directions other than the target direction.
[0006] In one possible implementation example of a head-mounted computing device, the processor is further configured to delay the channels of the audio from the microphone array relative to each other in order to redirect the beamformed sensitivity of the microphone array.
[0007] In another possible implementation example of a head-mounted computing device, the plurality of microphones includes omnidirectional microphones, and the omnidirectional microphones are configured to generate unfocused audio based on the sound received according to the omnidirectional sensitivity of the omnidirectional microphones. In this possible implementation example, the processor may further be configured to detect a speaker based on the sound received by the omnidirectional microphones and redirect the beamformed sensitivity of the microphone array towards the speaker.
[0008] In another possible implementation example of a head-mounted computing device, the plurality of sensors includes a camera configured to capture an image from the perspective of a user wearing the head-mounted computing device, and the processor is further configured to identify a conversation between the user and a participant and calculate the target direction as being towards the participant.
[0009] In another possible implementation example of a head-mounted computing device, the plurality of sensors includes an inertial measurement unit (IMU) configured to measure the orientation of the head-mounted computing device. The inertial measurement unit may be configured to track changes in the orientation of the microphone array. The processor may be configured to detect changes in the orientation of the reference system based on the tracked changes and obtain an updated target direction.
[0010] In another possible implementation example, the head-mounted computing device generates beamformed audio based on the beamformed sensitivity of the microphone array and is configured to transmit the beamformed audio to an augmented reality application running on the head-mounted computing device.
[0011] In another possible implementation example of the head-mounted computing device, the processor generates beamformed audio based on the beamformed sensitivity of the microphone array and is configured to transmit the beamformed audio to a plurality of loudspeakers. In this possible implementation example, the plurality of loudspeakers includes one or more hearing devices configured to be worn on one or both ears of the user. For example, one or more hearing devices configured to be worn on one or both ears of the user may be configured to communicate wirelessly with the processor.
[0012] In another aspect, the present disclosure generally describes a method for generating beamformed audio based on a conversation layout. The method includes detecting a conversation between a user and a participant based on an image or video captured by a camera of a head-mounted computing device worn by the user. The method further includes determining a head pose of the user based on measurements captured by sensors of the head-mounted computing device worn by the user. The method further includes calculating a conversation layout based on the relative positions of the participant and the head pose. The method further includes receiving audio channels from a microphone array of the head-mounted computing device and processing the audio channels to generate beamformed audio based on the conversation layout. Alternatively, or in addition, a method for generating beamformed audio based on a conversation layout may include identifying an orientation of a reference system of the microphone array based on an orientation of the head-mounted computing device, calculating a target direction relative to the orientation of the reference system, directing a beamformed sensitivity of the microphone array in the target direction, updating the reference system to obtain an updated target direction when a change in the orientation of the head-mounted computing device is detected, and redirecting the beamformed sensitivity of the microphone array in the updated target direction.
[0013] According to possible implementations of the method, the method may include one or more (e.g., all) (or any combination thereof) of the following features.
[0014] In one possible implementation of the method, the beamformed voice corresponds to the received voice according to the beamformed sensitivity directed to the participant. In this implementation, the method further includes, when detecting a change in the orientation of the head-mounted computing device using an inertial measurement unit, obtaining an updated conversation layout, and processing the audio channels so that the beamformed sensitivity is redirected to the participants in the updated conversation layout.
[0015] In another possible implementation of the method, the method further includes presenting the beamformed voice to the user.
[0016] In another possible implementation of the method, the method further includes reducing the sensitivity of the beamformed voice in the direction towards the bystander.
[0017] In another possible implementation of the method, the method further includes presenting an augmented reality visual on the display of the head-mounted computing device, where the augmented reality visual corresponds to the beamformed voice. For example, the augmented reality visual can be subtitles of the conversation.
[0018] In another aspect, the present disclosure generally describes a computer program product tangibly embodied on a non-transitory computer-readable medium and comprising instructions configured to cause at least one processor of a head-mounted computing device to perform a method. The method includes identifying a reference system of a microphone array based on an orientation of the head-mounted computing device. The method further includes calculating a target direction relative to the reference system. The method further includes directing a beamformed sensitivity of the microphone array in a direction toward the target direction. The method further includes updating the reference system to obtain an updated target direction when a change in the orientation of the head-mounted computing device is detected. The method further includes redirecting the beamformed sensitivity of the microphone array toward the updated target direction to provide privacy to the person present.
[0019] The foregoing exemplary summary, other exemplary objects and / or advantages of this disclosure, and the manner in which they are achieved are further described in the following detailed description and its accompanying drawings.
Brief Description of the Drawings
[0020]
Figure 1A
Figure 1B
Figure 2
Figure 3A
Figure 3B
Figure 3C
Figure 4
Figure 5
Figure 6
Best Mode for Carrying Out the Invention
[0021] The components in the drawings are not necessarily to scale relative to each other. Throughout some of the drawings, the same reference numerals refer to corresponding parts.
[0022] Detailed Description Beamforming is a technique for enhancing the reception sensitivity of a microphone array in a specific direction (s) compared to other directions. Beamforming can be used to steer the sensitivity of a head-mounted microphone array towards a sound source in order to improve the quality of the audio from the sound source. However, when the head position / orientation (i.e., head pose) of the user wearing the head-mounted microphone array changes, there is a risk of a problem of misalignment of the steered sensitivity. Therefore, the disclosed devices and methods provide an adaptive beamforming technique for a head-mounted microphone array that can adapt to (i.e., be tolerant of) changes in the user's head position / orientation. The disclosed solution can have a technical effect of improving the quality of the audio captured by the head-mounted microphone array while providing the user with more freedom of movement. Adaptive beamforming can also have a technical effect of providing a layer of privacy. For example, beamforming can maintain the focus of the microphone array on a specific person during a conversation with the user and prevent amplification of the audio received from the people present. The problem associated with adaptive beamforming is its processing requirements. The disclosed devices and methods provide means for reducing the processing requirements of adaptive beamforming.
[0023] FIG. 1A is an exemplary polar chart of the sensitivity of the omnidirectional microphone 100 in an acoustic environment. The omnidirectional microphone 100 has a sensitivity pattern (i.e., sensitivity 101) that does not vary with angle (i.e., is isotropic). Thus, the omnidirectional microphone 100 will receive the speech voice 104 from the speech source 103 (e.g., a person) along the speech direction 105 with a sensitivity 101 that substantially matches the sensitivity 101 with which the omnidirectional microphone 100 receives the noise voice 109 from the noise source 108 (e.g., machinery) along the noise direction 110. In some applications (e.g., hearing devices such as headphones, earphones, or hearing aids), it may be desirable to reduce the sensitivity of the microphone in the noise direction 110 and / or increase the sensitivity of the microphone in the speech direction 105 so that the speech voice 104 can be preferentially amplified over the noise voice 109 for the user of the microphone.
[0024] Beamforming (i.e., beam steering) is signal processing in which multiple channels of audio can be processed (e.g., filtered, delayed, phase shifted) to generate a beamformed audio signal in which audio from different directions can be increased or decreased. For example, a first microphone and a second microphone can be spatially separated by a distance along the array direction. This spatial separation distance and the direction of the sound (relative to the array direction) may cause an interaural delay between a first audio stream at the first microphone and a second audio stream at the second microphone. Beamforming may include further delaying one of the audio streams by a beamforming delay. Thus, after beamforming, the first audio stream and the second audio stream are phase shifted by the interaural delay and the beamforming delay. The phase-shifted audio streams are then combined (e.g., summed) to generate the beamformed audio. By adjusting the beamforming delay with respect to the interaural delay, audio from a particular direction can be adjusted (e.g., canceled, attenuated, increased) by the summing process. For example, a pure sine wave received by the first microphone and the second microphone can be completely canceled for a particular direction if the phase shift between versions of the sine wave at the combiner after the interaural delay and the beamforming delay is 180 degrees. Alternatively, if the versions of the sine wave at the combiner after the interaural delay and the beamforming delay are in phase (i.e., the phase shift is 0 degrees), the versions of the sine wave at the combiner can be increased.
[0025] Multiple channels of audio can be captured (i.e., collected) by an array of microphones (i.e., a microphone array). Each microphone in the microphone array may be of the same type, or the types of different microphones in the array may be different. The microphone array may include a plurality of microphones spaced in one dimension, two dimensions, or three dimensions (e.g., equally spaced). For example, each microphone in the microphone array may be omnidirectional. However, due to beamforming, the microphone array may have a beamformed sensitivity that is directional (i.e., has a beam for reception). Thus, steering the beamformed sensitivity can be understood as steering (i.e., repositioning) the beam of the priority sensitivity of the microphone array.
[0026] Figure 1B is an exemplary polar plot of the beamformed sensitivity of microphone array 120. In particular, each microphone in microphone array 120 may generate an audio channel. Different audio channels may be processed (e.g., phase shifted relative to each other and summed) to generate a beamformed audio channel having a non-isotropic beamformed sensitivity. In other words, microphone array 120 may focus beam 121 in beam direction 122 that may be steered to align with speech direction 105 by beamforming processing. The number and spacing of microphones in microphone array 120 may correspond to the directivity (i.e., focus, angular range) of beam 121. As shown in FIG. 1B, the increased directivity created by the microphone array may result in beamformed audio that includes speech audio 104 having a higher amplitude than ambient audio 109. Thus, beamforming may assist the user in identifying speech audio 104 (e.g., in a noisy environment). In addition (or alternatively), beamforming may improve the accuracy of other computer-aided speech applications (e.g., speech recognition, voice-to-text (VTT), language translation, etc.). Additionally, beamforming may enhance privacy. This is because other audio 133 received from directions other than the speech direction (e.g., conversations of people present) can be amplified much less than speech audio 104.
[0027] A head-mounted computing device can include various sensing and computing resources to enable various technologies. For example, a head-mounted computing device can be configured to provide augmented reality (AR). In AR, sensors in the head-mounted computing device can be configured to capture sensory data from the environment and from the user wearing the head-mounted computing device. Based on this sensory data, virtual elements can be generated to enhance (i.e., augment) the user's perceptual experience. For example, when virtual elements are merged (e.g., overlaid) with the real environment, the generation of sound (e.g., tones, music, speech, etc.) and / or the display of visual (e.g., graphics, text, colors, etc.) can add information to the environment perceived by the user.
[0028] This disclosure describes a head-mounted computing device configured to enhance the natural perception of a user's real environment. This enhancement may or may not include the virtual aspect of AR. For example, the head-mounted computing device can be configured to beamform the captured audio to assist the user in listening in a direction associated with a conversation or to record sound from a direction associated with a conversation. In addition (or alternatively), the head-mounted computing device can further be configured to beamform the captured audio to assist (e.g., improve the accuracy of) AR applications such as adding subtitles to a conversation in real time.
[0029] The head-mounted computing device may further be configured to beamform the captured audio to help prevent the user from eavesdropping on conversations (e.g., listening in, recording) in directions other than the direction associated with the conversation. To achieve this privacy, the head-mounted computing device may be configured to detect conversations to determine the participant(s) in conversation with the user. This detection may require a computationally expensive process, including, for example, running a computer vision algorithm on an image captured from a camera(s) of the head-mounted computing device. This computationally expensive process may exceed the processing and / or power budget of the head-mounted computing device if they are continuously run at a speed sufficient to respond to the movement of the user's head. Thus, the disclosed devices and methods can hand off beam steering to a process that is not as computationally expensive after a conversation(s) has been determined. The process that is not as computationally expensive may include using a position / orientation sensor(s) to determine a change in head movement from an initial position and then updating the position of the participant(s) relative to the change in head movement. Beamforming may then adjust the beam steering in response to head movement to maintain focus on the participant(s). This approach can be run fast enough to respond to the movement of the user's head because it requires less processing and / or power.
[0030] Figure 2 is a perspective view of a head-mounted computing device configured to generate beamformed audio, according to one possible implementation of the present disclosure. As shown, the head-mounted computing device may be implemented as smart glasses. In this specification, "smart glasses" will be described and referred to in the same sense as the term "head-mounted computing device" for the purposes of explaining this disclosure. However, the techniques presented herein are more generally applicable to any head-mounted computing device that includes a microphone array (s) that can be focused (i.e., steered) in accordance with head movement (e.g., to allow for some modifications to the functionality of the head-mounted computing device). For example, it is contemplated that this disclosure may be implemented as a virtual-reality (VR) headset or smart earbuds.
[0031] The head-mounted computing device 200 shown in FIG. 2 is configured to be worn on a user's head / face. The head-mounted computing device 200 may be configured to have various sensors and various interfaces. In addition, the head-mounted computing device may include a power source (e.g., a battery) to enable portable operation, a memory for storing data and computer-readable instructions, one or more cameras 201 (e.g., cameras) for capturing video / image / depth information, and a projector / display for presenting visuals to the user in a display area 220 of the lens(es). Thus, the head-mounted computing device 200 can be configured for AR as described above and can present augmented reality visuals to the user in the display area 220. In one implementation example, the augmented reality visuals may include subtitles of a conversation. In addition, the head-mounted computing device 200 may include a subsystem and circuitry that can preferentially capture audio from a certain direction(s) that can change (e.g., automatically) with the movement of the head. This preference in the capture direction(s) can help improve what the user hears, enhance the functionality of the application, and / or provide a layer of privacy for bystanders not engaged in conversation with the user.
[0032] The head-mounted computing device 200 may further include a plurality of microphones 211A - F that can operate together as a microphone array 210. The microphones in the microphone array 210 can be configured to capture sound from the user's environment. For example, when the user is wearing the head-mounted computing device, the microphones can be directed towards the user's field of view. The microphones in the microphone array 210 can be spaced apart in various ways. In one possible implementation, one or more microphones (211A, 211B) can form a right channel, while one or more microphones (211C, 211D) can form a left channel. The microphones in the right channel and the microphones in the left channel can be spaced apart to simulate a natural inter-aural distance. For example, the microphones in the left channel (211C, 211D) can be positioned near the left lens 242 of the head-mounted computing device, while the microphones in the right channel (211A, 211B) can be positioned near the right lens 241 of the head-mounted computing device. In another possible implementation, the microphone layout can help with beamforming in a certain direction (e.g., the basic direction) of the reference system 250 of the head-mounted computing device 200. The reference system 250 of the head-mounted computing device 200 is not fixed in space, but rather has an orientation that tracks the orientation of the head-mounted computing device 200 as it moves in space. The reference system 250 can be defined by three orthogonal axes: a first axis 251 that can be parallel to, for example, the horizontal direction of the device, a second axis 252 that can be parallel to, for example, the vertical direction of the device, and a third axis 253 that is orthogonal to the first and second axes. As shown in FIG. 2, the first array of microphones (211A, 211B, 211C, 211D) can be aligned parallel to the first axis 251 of the reference system 250 of the head-mounted computing device 200. The beamforming algorithm applied to the first array of microphones can steer the beam in response to the left / right (i.e., yaw) movement of the user's head.The second array of microphones (211E, 211F) can be aligned parallel to a second axis 252 of a reference system 250 of the head-mounted computing device 200. The beamforming algorithm applied to the second array of microphones can steer the beam in response to up / down (i.e., pitch) movement of the user's head. In general, there can be an array of microphones aligned in any number of directions. Beamforming can include combining different microphone arrays to handle beam steering in multiple directions.
[0033] Figures 3A and 3B illustrate beamforming as described above. In Figure 3A, the microphone array of a head-mounted computing device worn on the head of user 301 can be beamformed (i.e., focused) towards a first speaker 302 (e.g., during conversation with user 301). Here, the user's head is in a first position, and the beam 320 of the microphone array of the head-mounted computing device is aligned with the user's line of sight direction 310. In Figure 3B, the user's head is rotated by an angle 330 from the first position. The microphone array of the head-mounted computing device is configured to adjust beamforming so as to remain focused on the first speaker 302 despite movement (i.e., yaw) of the user's head. In other words, if the user's head is rotated by an angle 330 in a first direction (e.g., to the right), the head-mounted computing device can be triggered to adjust beamforming such that the beam 320 is rotated by an angle 330 in a second direction opposite to the first direction (e.g., to the left).
[0034] The head-mounted computing device 200 may further include ambient microphones 212 configured to capture various sounds that may be excluded by the microphone array 210. For example, the directional sensitivity of the ambient microphones 212 may be similar to the sensitivity shown in FIG. 1A, while the sensitivity of the microphone array 210 may be similar to the sensitivity shown in FIG. 1B. The ambient microphones 212 may be useful for beamforming in response to sound rather than head movement. For example, beamforming may be applied in response to a change in the speaker during a conversation. For example, at a first time, based on the audio from the ambient microphones 212, it may be recognized that a first speaker is speaking. Accordingly, the microphone array 210 may be focused (i.e., beamformed) towards the first speaker. Next, at a second time, based on the audio from the ambient microphones 212, it may be recognized that a second speaker is speaking. The second speaker may be recognized based on the quality (e.g., tone, pitch) of the sound from the second speaker. For this recognition, this quality may be compared to the quality stored in a participant list associated with the participants in the conversation. After the second speaker is recognized, the microphone array 210 may be focused (i.e., beamformed) towards the second speaker. Focusing may include focusing towards the direction of the second speaker stored in the conversation layout. Alternatively, focusing may include adjusting the relative phase of the microphones in the microphone array so as to scan the beam of the microphone array until the received sound having a quality that matches the second speaker is maximized. Thus, beamforming may be applied without head movement.
[0035] Figures 3A and 3C illustrate beamforming as described above. In Figure 3A, the microphone array of the worn head-mounted computing device can be beamformed (i.e., focused) towards the first speaker 302 to capture the voice 304 of the first speaker (e.g., while in conversation with user 301) with high directional sensitivity. As shown in Figure 3C, the ambient microphones of the head-mounted computing device have an omnidirectional sensitivity 340. When the ambient microphone 212 captures the voice 305 of the second speaker from the second speaker 303, the head-mounted computing device can be triggered to adjust the beamforming such that the beam 320 of the microphone array is rotated (i.e., focused) towards the second speaker 303 to capture the second speaker voice 305 with high sensitivity, as described above. Generally, the direction of the beam can be updated when a change in the target (e.g., a change in speaker) is detected based on the sound (e.g., speech) received by the omnidirectional microphone. For example, each speaker in a conversation can have a corresponding direction stored in the conversation layout. To determine the speaker who is speaking, the voice from the omnidirectional microphone (i.e., unfocused voice) can be processed (e.g., using speech recognition). The conversation layout can then be addressed for the speaker who is speaking to determine the direction for focusing the beam of the microphone array. The examples shown in Figures 3A - 3C are not limiting. For example, the number of speakers is not limited to two, and in some implementations, the microphone array may be configured to focus multiple beams on multiple speakers.
[0036] The head-mounted computing device 200 may further include a plurality of loudspeakers. In one possible implementation, the plurality of loudspeakers includes a left loudspeaker(s) configured to transmit sound to the user's left ear and a right loudspeaker(s) configured to transmit sound to the user's right ear. These loudspeakers may be integrated within the frame of the head-mounted computing device 200. For example, the left loudspeaker 231 may be integrated into the left arm of the head-mounted computing device, and the right loudspeaker 230 may be integrated into the right arm of the head-mounted computing device. In a possible implementation, the head-mounted computing device may include a left earbud 234 and a right earbud 233. The left earbud 234 and the right earbud 233 may be communicatively coupled to processing in the head-mounted computing device via a wired or wireless communication link 232 (e.g., Bluetooth®, WiFi, etc.). The earbuds may be worn on the user's respective ears. The earbuds may be configured to reproduce sound received by the microphone array 210. The sound reproduced by the earbuds may be beamformed sound resulting from a beamforming (i.e., beam steering, focusing) process as described in connection with FIG. 1B.
[0037] The sensitivity of the microphone array 210 may be focused towards a target (e.g., a person) in a first direction. When the orientation of the microphone array (i.e., the head-mounted computing device 200) changes, the sensitivity of the microphone array 210 may be focused towards the target in a second direction relative to the reference system of the head-mounted display that changes with the orientation of the microphone array. The second direction corresponds to the change in orientation. In other words, when the user is wearing the head-mounted computing device, the focus of the microphone array may be maintained on the target even as the user's head moves. This movement may include a change in the orientation of the user's head and / or a translation of the position.
[0038] The first direction and the second direction can be determined using various sensors on the head-mounted computing device. The first direction is established using a first sensor, while the second direction can be determined using a second sensor. For example, a camera can capture an image / video that is analyzed to determine a target (e.g., a person) and a first direction towards the target. After the first direction is established, the second direction can be determined as a change from the first direction. Calculating this change from the first direction can be achieved using sensors with reduced processing requirements. For example, the head-mounted computing device can include an IMU for measuring changes in orientation to determine the second direction relative to the first direction. Since the IMU can quickly respond to changes with fewer processing requirements needed to perform computer vision techniques, it can provide a faster tracking speed, which can be useful for some beamforming applications (e.g., the privacy of a person in a kata). For applications that require a slower tracking speed, the change in position can also be obtained using images, depth data, and / or position data (e.g., GPS data) captured by sensors in the head-mounted computing device.
[0039] The IMU can be configured to determine (and track) the orientation of the microphone array 210. For example, data from the camera and / or the IMU can help define an initial orientation of the reference system 250 of the head-mounted computing device 200 relative to a reference frame fixed in space. Thereafter, data from the IMU can help detect a change in the orientation of the reference system 250 of the head-mounted computing device 200 from the initial orientation and quantify the change or orientation. In other words, the IMU can help detect the movement of the head wearing the head-mounted computing device 200, quantify the head movement, and establish a new head orientation after the movement.
[0040] The IMU of the head-mounted computing device 200 may include a multi-axis accelerometer, a gyroscope, and / or a magnetometer. The IMU may be preferable for applications that require high tracking speeds and / or lower processing requirements. For example, a head-mounted computing device operating via a battery may have limited power resources. The IMU can continuously measure orientation without consuming a large amount of power. Additionally, reading data from the IMU can be achieved using a relatively simple controller. Thus, the IMU may be able to help track the orientation of the microphone array 210 very quickly and continuously without overtaxing the limited processing / power resources of the head-mounted computing device. This power / processing-efficient continuous tracking can be useful for quickly responding to head movements.
[0041] Determination of orientation may use alternative or additional sensors of the head-mounted computing device 200. For example, the head-mounted computing device may include a depth camera (e.g., using structured light, LIDAR) for optically sensing the movement of the head-mounted computing device 200 corresponding to the movement of the microphone array 210.
[0042] FIG. 4 is a flowchart of a method for focusing a microphone array of a head-mounted computing device onto a target according to one possible implementation example of the present disclosure. Method 400 includes a step 405 of determining an orientation of a reference system for a microphone array mounted on a user's head. The microphone array may be included as part of a head-mounted computing device as shown in FIG. 2, although it may also be part of a system. For example, the microphone array may be part of a system in which components for focusing the microphone array are communicatively coupled (e.g., wirelessly) even though they are physically separate. The orientation of the reference system of the microphone array can be determined by one or more sensors such as an IMU. The reference system of the microphone array may or may not be aligned with the user's perspective.
[0043] Method 400 further includes a step 410 of identifying a target. The target can be a sound source such as a person or an object (e.g., a TV, a radio, a speaker, etc.). The step of identifying the target can include a step of indicating the position or direction of the target relative to a reference system. The identifying step can be automatic or manual. In one possible implementation, to identify (e.g., acquire) the target, the user's viewpoint can be positioned on the target. For example, the user positions the target within the user's field of view and then triggers the head-mounted computing device to identify the target using keywords (e.g., "lock on (continuously track) the speaker", "switch the speaker") spoken by the user and recognized by the head-mounted computing device. Alternatively, the user positions the target within the user's field of view and then triggers the head-mounted computing device to identify the target by physically interacting with the head-mounted computing device (e.g., pressing a button, tapping the device, etc.). In automatic target recognition, the sound from the microphone of the head-mounted computing device and / or the image from the camera of the head-mounted computing device can be monitored to identify the target. For example, these sounds and images can be processed using computer recognition algorithms to identify the speech patterns (pauses, speaker changes, etc.) and visual cues (e.g., eye contact) indicating a conversation with the target. In another possible implementation, a specific sound can be recognized to identify the target.
[0044] Once the target is identified, the method may include step 415 of determining the target direction relative to a reference system. In one possible implementation, the target direction may be determined using light. For example, imaging sensing and computer vision algorithms may be configured to determine the direction of the target (e.g., relative to the user's perspective) based on still images and / or videos captured by one or more cameras 201 of a head-mounted computing device. In another possible implementation, the target direction may be determined using sound. For example, sound sensing and computing hearing algorithms may be configured to determine the direction of the target based on the sound emitted by the target, as described above.
[0045] Once the initial reference system of the head-mounted computing device 200 and the target are acquired and spatially determined, the method may include step 420 of focusing the microphone array in the target direction (e.g., onto the target). The focusing step may be performed by processing the audio signals received by each microphone in the microphone array such that the sensitivity of the microphone array is higher (e.g., highest) in the target direction of the reference system of the head-mounted computing device 200 than in other directions. In other words, the focusing step may result from signal processing rather than a physical change to the system.
[0046] The microphone array may remain focused on the target direction until movement of the microphone array is detected (e.g., using an IMU). When movement of the microphone array is detected (425), the method may include step 430 of determining a change in the orientation of a reference system with respect to a fixed reference frame in space and / or with respect to an initial reference system of the head-mounted computing device. For example, the angle(s) of the new orientation of the reference system with respect to the initially determined orientation of the reference system may be determined. The change in the orientation of the reference system may be used to update the target direction (440). After the target direction is updated, the beamformed sensitivity of the microphone array may be redirected towards the updated target direction and remain there until another movement is detected. Redirecting the beamformed sensitivity of the microphone array towards the updated target direction may include reducing the sensitivity of the microphone array in directions other than the updated target direction. This process may be repeated to maintain the focus of the microphone array of the head-mounted computing device on the target. Step 430 of determining a change in the reference system and step 440 of updating the target direction may be repeated at a first rate that is fast enough to accommodate natural movement of the user's head so that the user's movement is not inhibited. Step 410 of identifying the target and step 415 of calculating the target direction may also be repeated at a second rate that may be slower than the first rate because the target may be added or removed on a longer time scale compared to head movement. Thus, an approach of determining an initial reference system orientation with respect to the target and then tracking changes in orientation from the initial orientation may use fewer resources and be conveniently used in applications with limited processing / power, such as battery-powered head-mounted computing devices, compared to an approach of continuously identifying the target and calculating the target direction.For example, the processor may first perform a computer vision algorithm to detect a target, and once the target direction is determined with respect to the initial orientation of the reference system, it may then be configured to continuously update the target direction based on changes to the reference system. The method shown in FIG. 4 can be applied to various head-mounted microphone arrays, various targets, and various means for identifying the target. A more specific implementation example is shown in FIG. 5.
[0047] FIG. 5 is a flowchart of a method for generating beamformed audio based on conversation. The method can be implemented to provide the security level of the person present to any conversation enhancement provided by the head-mounted computing device. For example, beamforming can prevent the user from eavesdropping on and / or accidentally hearing the conversation of the person present. The method can also be implemented to assist the user in understanding the conversation. The method can also be used to assist an application (e.g., running on the head-mounted computing device) in accurately recognizing words within the conversation. Thus, the beamformed audio can also be used in an AR application running on the head-mounted computing device of FIG. 2 (e.g., smart glasses) configured for AR. The method of FIG. 5 is described in this context.
[0048] Method 500 includes monitoring a head-mounted sensor for the user. In particular, the IMU in the head-mounted computing device 200 is monitored (525) and can be used to determine (530) the orientation / position of the head of the user wearing the head-mounted computing device 200 (i.e., the head pose 535). For example, the user's head pose 535 may include a reference system 250 aligned with the microphone array 210. When the user wears the head-mounted computing device 200, the head pose 535 can be repeatedly updated such that changes in the head pose can trigger changes in beamforming.
[0049] Method 500 also includes the step of monitoring a head-mounted sensor for a conversation. For example, video / images that can be applied to a computer vision algorithm configured to detect a conversation based on visual features (such as eye contact, lip movement, etc.) associated with a conversation between a conversation participant (i.e., a participant) and a user can be monitored (i.e., captured) by one or more cameras 201 (505). Similarly, audio that can be applied to an audio / speech recognition algorithm configured to recognize a conversation based on audio features (such as speech-to-text conversion, pauses, transcription, etc.) associated with a conversation between a participant and a user can be monitored (i.e., captured) by ambient microphones 212 (505). In one possible implementation, both visual and audio features can be monitored (505) to detect a conversation. After a conversation is detected, the head-mounted sensor can be monitored (510) to determine the status (i.e., active, inactive) of the conversation between the participant and the user. For example, if no audio or visual features are detected over a period of time, the conversation can be determined to be inactive (i.e., ended, finished).
[0050] The conversation may include two or more participants. Thus, method 500 may further include step (515) of adding, excluding, or otherwise updating a participant list 520 corresponding to the detected conversation. For example, at a first time, a conversation may be detected between a user and a first participant. At the first time, the participant list includes one conversation (i.e., the first participant). At a second time, a second participant may join the conversation or start a new conversation. At the second time, the second participant may be added to the participant list so that the participant list includes the first participant and the second participant. At a third time after the conversation by the second participant becomes inactive for a period, the second participant is excluded (i.e., removed, deleted) from the participant list 520 so that only the first participant remains. At a fourth time after the conversation by the first participant becomes inactive for a period, the first participant may be excluded from the list so that the participant list becomes empty. The participant list may change automatically based on the user's conversation with people. The participant list may include various information associated with a participant. For example, the identifier of the participant and the status of the conversation between the participant and the user may be included in the participant list 520. In addition, the participant list 520 may include the (e.g., visual, auditory) characteristics of each participant. In this way, when a previous participant reappears, the conversation can be recognized more easily.
[0051] Based on the participant list, the method further includes step 540 of monitoring the head-mounted sensors (e.g., cameras, microphones) of the head-mounted computing device 200 to determine the position of each participant on the participant list 520 relative to the user's head pose 535. Based on the relative positions of the user and the participants, a conversation layout 560 can be calculated (or updated) (545). The conversation layout 560 can include the participants and the directions to the user, as shown in FIG. 3C. The microphone array 210 of the head-mounted computing device 200 then monitors (i.e., captures) the audio from the microphone array 210 (555) to generate beamformed audio 575 and can process it according to a beamforming algorithm. For example, the captured audio can be processed (filtered, delayed, phase-shifted) according to the conversation layout to beamform the microphone array to each participant simultaneously or sequentially (570).
[0052] As the conversation layout 560 based on the participant list 520 and the user's head pose 535 changes, the beamformed audio 575 can be automatically updated in real time. Regardless of how the conversation layout (e.g., head pose) changes, the beamformed audio 575 can be provided to the user 585 via a loudspeaker (e.g., in the user's ear) to help the user hear the audio from the participants during the conversation. The beamformed audio 575 can also be provided to the AR application 580. The AR application can modify or transform the beamformed audio 575 into an output that can be experienced by the user. For example, the AR application can process the beamformed audio to generate subtitles (e.g., text-to-speech, translation) that can be displayed to the user in the display area 220. The beamformed audio can have the technical effect of providing privacy by preventing or interfering with the user from receiving audio from people present in the detected conversation.
[0053] FIG. 6 is a block diagram of a head-mounted computing device configured to generate beamed audio based on a conversation layout, according to one possible implementation of the present disclosure. The head-mounted computing device 600 may include a plurality of sensors 610. The sensors 610 may include one or more image sensors 611 (e.g., cameras) configured to capture an image / video of the field of view. The sensors 610 may also include an IMU 612 configured to measure the orientation and / or movement of the head-mounted computing device. The sensors 610 further include a microphone array 615, which may include a plurality of microphones 613A, 613B, 613C.
[0054] The head-mounted computing device 600 may further include a plurality of interfaces 640. The interface 640 may include a communication interface 641 configured to transmit / receive data with the head-mounted computing device 600. For example, the communication interface 641 may include a short-range wireless communication transceiver (e.g., Bluetooth). In one possible implementation, the communication interface 641 is coupled to a hearing device (e.g., a hearing aid, earbuds, etc.) worn by the user. The interface 640 may include a display 642 configured to present images, graphics, and / or text to the user. The interface may further include one or more loudspeakers 643A, 643B, 643C. The one or more loudspeakers may include a left loudspeaker and a right loudspeaker. In one possible implementation, the loudspeakers may be included in a loudspeaker array 645.
[0055] The head-mounted computing device 600 may further include a non-transitory computer-readable medium (i.e., memory 630). The memory 630 may store data and / or computer programs. These computer programs (also known as modules, programs, software, software applications, or code) may include machine instructions for a programmable processor and may be implemented in high-level procedural and / or object-oriented programming languages and / or in assembly / machine language. For example, the memory may include a computer program product tangibly embodied on a non-transitory computer-readable medium. The computer program product, when executed, may include computer-executable instructions (i.e., software instructions) that cause at least one processor 620 to perform a method for generating beamformed speech based on a conversation, as shown in FIG. 5. Thus, the memory 630 may further be configured to store the (latest) participant list 520, the head pose 535, and the conversation layout 560.
[0056] The head-mounted computing device 600 may further include at least one processor 620. The at least one processor 620 may execute one or more modules to perform various aspects of a method for generating beamformed audio. The one or more modules may include a conversation detector 621 configured to receive measurements from the sensor 610 and detect a conversation between the user and a participant based on the measurements. The results of the conversation detector may be stored in the participant list 520 within the memory 630. The one or more modules may further include a head pose calculator 622 configured to receive measurements from the sensor 610 and calculate the orientation of the user's head (i.e., the orientation of the microphone array 615) based on the measurements. The results of the head pose calculator 622 may be stored in the memory 630 as the head pose 535. The one or more modules may further include a conversation layout generator 623 configured to generate a layout (i.e., a map) of the conversation based on the participant list 520 and the head pose 535. The results of the head pose calculator 622 may be stored in the memory 630 as the conversation layout 560. The one or more modules may further include a beamformer 624 configured to receive audio signals (i.e., channels) from each microphone (or a portion of a microphone) in the microphone array 615 and process the audio signals to generate beamformed audio (e.g., a beamformed audio signal). The beamformed audio may be communicated to the interface 640 or, in some implementations, to the user.
[0057] In the specification and / or the drawings, exemplary embodiments have been disclosed. The present disclosure is not limited to such exemplary embodiments. The use of the term "and / or" includes any and all combinations of one or more of the associated listed items. The drawings are schematic representations and, therefore, are not necessarily drawn to scale. Unless otherwise specified, specific terms have been used in an inclusive and descriptive sense rather than for purposes of limitation.
[0058] Unless otherwise noted, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. Methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure. As used in the specification and the appended claims, the singular forms also include the plural unless the context clearly dictates otherwise. The term "comprising" and variations thereof as used herein are used synonymously with the term "including" and variations thereof and are open-ended and non-limiting terms. The term "optional" or "optionally" as used herein means that the subsequent described feature, event or circumstance may or may not occur, and that the description includes instances where the said feature, event or circumstance occurs and instances where it does not. In this specification, ranges may be expressed as "about (a particular value)" to, and / or "about (another particular value)". When such a range is expressed, aspects include from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations by use of the preceding "about", it will be understood that the particular value forms another aspect. Further, it will be understood that each endpoint of a range is significant both in relation to the other endpoint and independently of the other endpoint.
[0059] Although certain features of the described embodiments have been shown as described herein, many modifications, substitutions, changes, and equivalents will now occur to those of ordinary skill in the art. Accordingly, it should be understood that the appended claims are intended to cover all modifications and changes that fall within the scope of the embodiments. It is to be understood that they have been presented by way of example only and not by way of limitation, and that various changes in form and detail may be made. Any part of the apparatus and / or method described herein may be combined in any combination except combinations that are mutually inconsistent. The embodiments described herein may include various combinations and / or sub-combinations of the functions, components, and / or features of the different embodiments described.
[0060] In the foregoing description, when an element is referred to as being on, connected to, electrically connected to, coupled to, or electrically coupled to another element, it will be understood that it may be directly on, connected to, or coupled to the other element, or that one or more intervening elements may be present. In contrast, when an element is referred to as being directly on, directly connected to, or directly coupled to another element, there are no intervening elements present. The terms "directly on", "directly connected", or "directly coupled" may not be used throughout the detailed description, but an element shown as being directly on, directly connected, or directly coupled may be referred to as such. The claims of the present application may be amended, if necessary, to recite exemplary relationships described in the specification or shown in the drawings.
[0061] As used herein, the singular forms include the plural forms unless the context clearly dictates otherwise for a particular case. Spatially relative terms (such as "on", "above", "upper", "under", "below", "lower", etc.) are intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the drawings. In some implementations, the relative terms "above" and "below" may include "vertically above" and "vertically below", respectively. In some implementations, the term "adjacent" may include "adjacent laterally" or "adjacent horizontally".
Claims
1. A head-mounted computing device comprising: A microphone array including a plurality of microphones configured to generate beamformed audio according to beamformed sensitivity; An omnidirectional microphone configured to generate unfocused audio; A plurality of sensors configured to detect a first speaker and a second speaker; The microphone array, the omnidirectional microphone, and a processor coupled to the plurality of sensors, the processor being configured by software instructions to: Direct the beamformed sensitivity of the microphone array in a first direction toward the first speaker; Detect a change in the speaker based on the unfocused audio; Redirect the beamformed sensitivity of the microphone array in a second direction toward the second speaker based on the change in the speaker. A head-mounted computing device configured by software instructions.
2. The head-mounted computing device according to claim 1, wherein redirecting the beamformed sensitivity of the microphone array in the second direction includes reducing the beamformed sensitivity of the microphone array in the first direction.
3. The head-mounted computing device according to claim 1 or 2, wherein the processor is further configured to delay channels of audio from the microphone array relative to each other in order to redirect the beamformed sensitivity of the microphone array.
4. The head-mounted computing device according to any one of claims 1 to 3, wherein the plurality of sensors includes a camera configured to capture an image from the perspective of a user wearing the head-mounted computing device, and the processor is further configured to identify the first direction toward the first speaker and the second direction toward the second speaker.
5. The head-mounted computing device according to any one of claims 1 to 4, wherein the plurality of sensors includes an inertial measurement unit configured to measure the orientation of the head-mounted computing device.
6. The inertial measurement unit is configured to detect a change in the orientation, and the processor is configured to obtain an updated second direction toward the second speaker based on the change, the head-mounted computing device according to claim 5.
7. The processor is further configured to generate the beamformed audio based on the beamformed sensitivity of the microphone array and transmit the beamformed audio to an augmented reality application running on the head-mounted computing device, the head-mounted computing device according to any one of claims 1 to 6.
8. The processor is configured to generate the beamformed audio based on the beamformed sensitivity of the microphone array and transmit the beamformed audio to a plurality of loudspeakers, the head-mounted computing device according to any one of claims 1 to 7.
9. The plurality of loudspeakers includes one or more hearing devices configured to be worn on one or both ears of a user, the head-mounted computing device according to claim 8.
10. The one or more hearing devices are configured to communicate wirelessly with the processor, the head-mounted computing device according to claim 9.
11. A method for generating beamformed audio based on a conversation layout, comprising: detecting a conversation between a user, a first speaker, and a second speaker based on an image or video captured by a camera of a head-mounted computing device worn by the user; determining a head pose of the user based on measurements captured by a sensor of the head-mounted computing device worn by the user; calculating the conversation layout based on the relative positions of the first speaker, the second speaker, and the head pose, the conversation layout including a first direction toward the first speaker and a second direction toward the second speaker, the method further comprising: receiving unfocused audio from an omnidirectional microphone of the head-mounted computing device; Detecting a change in speakers between the first speaker and the second speaker based on the unfocused voice; Receiving a voice channel from a microphone array of the head-mounted computing device; Processing the voice channel so as to shift the beamformed voice between the first direction and the second direction based on the change in speakers and the conversation layout, a method for generating beamformed voice based on a conversation layout.
12. Shifting the beamformed voice reduces the beamformed sensitivity directed to the first speaker, a method for generating beamformed voice based on the conversation layout according to claim 11.
13. When detecting a change in the orientation of the head-mounted computing device using an inertial measurement unit, obtaining an updated conversation layout; Further comprising processing the voice channel so as to redirect the beamformed sensitivity to the second speaker in the updated conversation layout, a method for generating beamformed voice based on the conversation layout according to claim 12.
14. Further comprising presenting the beamformed voice to the user, a method for generating beamformed voice based on the conversation layout according to any one of claims 11 to 13.
15. Further comprising reducing the sensitivity of the beamformed voice in the direction towards the person present, a method for generating beamformed voice based on the conversation layout according to any one of claims 11 to 14.
16. Further comprising presenting an augmented reality visual on a display of the head-mounted computing device, the augmented reality visual corresponding to the beamformed voice, a method for generating beamformed voice based on the conversation layout according to claim 15.
17. The augmented reality visual is a subtitle of the conversation, a method for generating beamformed voice based on the conversation layout according to claim 16. A computer program product tangibly embodied on a non-transitory computer-readable medium and comprising instructions configured to cause at least one processor of a head-mounted computing device to perform the following when executed: Receive unfocused audio from an omnidirectional microphone of the head-mounted computing device; Identify a first speaker and a second speaker in the unfocused audio; Detect a first direction toward the first speaker and a second direction toward the second speaker based on data from a plurality of sensors of the head-mounted computing device; Direct the beamformed sensitivity of a microphone array of the head-mounted computing device in the first direction toward the first speaker; Detect a change in speakers based on the unfocused audio; Redirect the beamformed sensitivity of the microphone array in the second direction toward the second speaker based on the change in speakers. A computer program product. Claim 19. The head-mounted computing device according to claim 1, wherein the processor is further configured to detect the change in speakers based on a comparison of the quality of the unfocused audio. Claim 20. Detecting the change in speakers between the first speaker and the second speaker includes comparing the quality of the unfocused audio. A method for generating beamformed audio based on the conversation layout according to claim 11. Claim 21. Detecting the change in speakers includes comparing the quality of the unfocused audio. The computer program product according to claim 18.
Citation Information
Patent Citations
Acoustic input device
JP2009135594A
Sound analyzer, sound acquisition device, sound analysis system and program
JP2014192546A
Circuit device system and associated computer executable code for acquired acoustic signals
JP2017521902A
Conversation assistance audio device control
US20200128322A1