Wearable System Speech Processing

The wearable system with integrated sensors improves speech processing accuracy by using sensor inputs to address environmental and positional challenges, enhancing reliability in mobile and outdoor applications.

JP7745603B2Active Publication Date: 2025-09-29MAGIC LEAP INC
View PDF 13 Cites 0 Cited by

Patent Information

Application Number
JP2023142856
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-06-21
Filing Date
2023-09-04
Publication Date
2025-09-29
Estimated Expiration
2039-06-21

AI Technical Summary

Technical Problem

Conventional speech processing systems struggle with accuracy due to unpredictable real-world conditions, such as multiple speech sources, environmental noise, speaker movement, and varying distances from the microphone, especially in mobile or outdoor applications.

Method used

A wearable system equipped with sensors, like cameras and IMUs, processes acoustic signals by determining control parameters based on sensor inputs to improve the fidelity and reliability of speech processing, including echo cancellation and noise reduction.

Benefits of technology

Enhances the accuracy and reliability of speech processing by compensating for environmental and positional variables, particularly in mobile and outdoor settings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007745603000001
    Figure 0007745603000001
  • Figure 0007745603000002
    Figure 0007745603000002
  • Figure 0007745603000003
    Figure 0007745603000003
Patent Text Reader

Abstract

To provide suitable wearable system speech processing.SOLUTION: A method of processing an acoustic signal is disclosed. According to one or more embodiments, a first acoustic signal is received via a first microphone. The first acoustic signal is associated with a first speech of a user of a wearable headgear unit. A first sensor input is received via the sensor. A control parameter is determined based on the sensor input. The control parameters are applied to one or more of the first acoustic signal, the wearable headgear unit, and the first microphone. A step of determining the control parameter includes the step of determining, based on the first sensor input, a relation between the first speech and the first acoustic signal.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of U.S. Provisional Patent Application No. 62 / 687,987, filed June 21, 2018, which is incorporated herein by reference in its entirety.

[0002] The present disclosure relates generally to systems and methods for processing acoustic speech signals, and more particularly to systems and methods for processing acoustic speech signals generated by a user of a wearable device. [Background technology]

[0003] Tools for speech processing are typically tasked with receiving audio input representing human speech via a microphone, processing the audio input, and determining words, logical structures, or other output that correspond to the audio input. For example, an automatic speech recognition (ASR) tool may generate text output based on the human speech corresponding to the audio input signal, and a natural language processing (NLP) tool may generate logical structures or computer data that correspond to the meaning of the human speech. It is desirable for such processes to occur accurately and quickly, and some applications require results in real time.

[0004] Computer speech processing systems have a history of producing inaccurate results. In general, the accuracy of a speech processing system can depend heavily on the quality of the input audio signal, with the highest accuracy obtained from input provided under controlled conditions. For example, a speech processing system can perform reliably when the audio input is clearly articulated speech captured by a microphone at a direct angle and close distance, with no ambient noise, a high signal-to-noise ratio, and a constant volume level. However, speech processing systems can struggle to adapt to the many variables that may be introduced into the input audio signal by real-world conditions. For example, speech processing signals may demonstrate limited accuracy when multiple speech sources are present (e.g., multiple people speaking at once in the same space), when environmental noise (e.g., wind, rain, electrical interference, ambient noise) mixes with the source speech signal, when a human speaker does not articulate or speaks with unique or inconsistent tones, accents, or inflections, when a speaker moves or rotates relative to the microphone, when a speaker is in an acoustically reflective environment (e.g., a tiled bathroom or a cathedral), when a speaker is far from the microphone, when a speaker faces away from the microphone, or when any number of other variables are present and compromise the fidelity of the input audio signal. These problems may be magnified in mobile or outdoor applications, where unpredictable noise sources may be present and attempts to control or understand speaker proximity may be difficult or impossible.

[0005] It would be desirable to use sensor-equipped wearable systems, such as those that incorporate head-mounted units to compensate for the effects of such variables on the input audio signal for a speech processing system. By presenting more predictable and higher fidelity input to speech processing systems, the output of those systems can produce more accurate and more reliable results. In addition, wearable systems are well suited to mobile outdoor applications, i.e., precisely the types of applications where many conventional speech processing systems can perform particularly poorly. Summary of the Invention [Means for solving the problem]

[0006] Examples of the present disclosure describe systems and methods for processing acoustic signals. According to one or more embodiments, a first acoustic signal is received via a first microphone. The first acoustic signal is associated with a first speech of a user of a wearable headgear unit. A first sensor input is received via a sensor. A control parameter is determined based on the sensor input. The control parameter is applied to one or more of the first acoustic signal, the wearable headgear unit, and the first microphone. Determining the control parameter includes determining a relationship between the first speech and the first acoustic signal based on the first sensor input. The present specification also provides, for example, the following items: (Item 1) 1. A method for processing an acoustic signal, the method comprising: receiving, via a first microphone, a first acoustic signal associated with a first utterance of a user of the wearable headgear unit; receiving a first sensor input via the sensor; determining a control parameter based on the sensor input; applying the control parameters to one or more of the first acoustic signal, the wearable headgear unit, and the first microphone; Including, The method, wherein determining the control parameter includes determining a relationship between the first speech utterance and the first acoustic signal based on the first sensor input. (Item 2) Item 10. The method of item 1, wherein the control parameters are applied to the first acoustic signal to generate a second acoustic signal, the method further comprising providing the second acoustic signal to a speech recognition engine to generate a text output corresponding to the first utterance. (Item 3) 2. The method of claim 1, wherein the control parameters are applied to the first acoustic signal to generate a second acoustic signal, and the method further includes providing the second acoustic signal to a natural language processing engine to generate natural language data corresponding to the first utterance. (Item 4) Item 10. The method of claim 1, wherein the wearable headgear unit comprises the first microphone. (Item 5) Determining a control parameter based on the sensor input comprises: detecting a surface based on the sensor input; determining an effect of the surface on a relationship between the first speech utterance and the first acoustic signal; determining control parameters that, when applied to one or more of the first acoustic signal, the wearable headgear unit, and the first microphone, reduce the effect of the surface on a relationship between the first speech utterance and the first acoustic signal; The method according to item 1, comprising: (Item 6) 6. The method of claim 5, further comprising determining an acoustic property of the surface, wherein an effect of the surface on a relationship between the first utterance and the first acoustic signal is determined based on the acoustic property. (Item 7) Determining a control parameter based on the sensor input comprises: detecting a person different from the user based on the sensor input; determining an effect of the person's speech on a relationship between the first speech and the first acoustic signal; determining control parameters that, when applied to one or more of the first acoustic signal, the wearable headgear unit, and the first microphone, reduce an effect of the first speech on a relationship between the first speech and the first acoustic signal; The method according to item 1, comprising: (Item 8) Item 10. The method of item 1, wherein determining a control parameter based on the sensor input comprises applying the sensor input to an input of an artificial neural network. (Item 9) the control parameters are control parameters for an echo cancellation module; Determining the control parameter based on the sensor input comprises: detecting a surface based on the sensor input; determining a time of flight between the surface and the first microphone; The method according to item 1, comprising: (Item 10) 2. The method of claim 1, wherein the control parameter is a control parameter for a beamforming module, and determining the control parameter based on the sensor input includes determining a time of flight between the user and the first microphone. (Item 11) 2. The method of claim 1, wherein the control parameter is a control parameter for a noise reduction module, and determining the control parameter based on the sensor input includes determining frequencies to be attenuated within the first acoustic signal. (Item 12) the wearable headgear unit includes a second microphone; the sensor input includes a second acoustic signal detected via the second microphone; the control parameter is determined based on a difference between the first acoustic signal and the second acoustic signal. The method according to item 1. (Item 13) the wearable headgear unit includes a plurality of microphones excluding the first microphone; The method further includes receiving, via the plurality of microphones, a plurality of acoustic signals associated with the first utterance; the control parameter is determined based on a difference between the first acoustic signal and the plurality of acoustic signals. The method according to item 1. (Item 14) Item 10. The method of claim 1, wherein the sensor is coupled to the wearable headgear unit. (Item 15) Item 10. The method of item 1, wherein the sensor is positioned within the user's environment. (Item 16) 1. A system comprising: Wearable Head Gear Unit Equipped with The wearable head gear unit includes: a display for displaying the mixed reality environment to a user; A speaker and One or more processors, the one or more processors configured to perform a method, the method comprising: receiving, via a first microphone, a first acoustic signal associated with a first speech utterance of the user; receiving a first sensor input via the sensor; determining a control parameter based on the sensor input; applying the control parameters to one or more of the first acoustic signal, the wearable headgear unit, and the first microphone; Including, one or more processors, wherein determining the control parameter includes determining a relationship between the first speech utterance and the first acoustic signal based on the first sensor input; Including, the system. (Item 17) 17. The system of claim 16, wherein the control parameters are applied to the first acoustic signal to generate a second acoustic signal, and the method further includes providing the second acoustic signal to a speech recognition engine to generate a text output corresponding to the first utterance. (Item 18) 17. The system of claim 16, wherein the control parameters are applied to the first acoustic signal to generate a second acoustic signal, and the method further includes providing the second acoustic signal to a natural language processing engine to generate natural language data corresponding to the first utterance. (Item 19) Item 17. The system of item 16, wherein the wearable headgear unit further includes the first microphone. (Item 20) Determining a control parameter based on the sensor input comprises: detecting a surface based on the sensor input; determining an effect of the surface on a relationship between the first speech utterance and the first acoustic signal; determining control parameters that, when applied to one or more of the first acoustic signal, the wearable headgear unit, and the first microphone, reduce the effect of the surface on a relationship between the first speech utterance and the first acoustic signal; Item 17. The system according to item 16, comprising: (Item 21) 21. The system of claim 20, wherein the method further includes determining an acoustic property of the surface, and an effect of the surface on a relationship between the first speech utterance and the first acoustic signal is determined based on the acoustic property. (Item 22) Determining a control parameter based on the sensor input comprises: detecting a person different from the user based on the sensor input; determining an effect of the person's speech on a relationship between the first speech and the first acoustic signal; determining control parameters that, when applied to one or more of the first acoustic signal, the wearable headgear unit, and the first microphone, reduce an effect of the first speech on a relationship between the first speech and the first acoustic signal; Item 17. The system according to item 16, comprising: (Item 23) Item 17. The system of item 16, wherein determining a control parameter based on the sensor input comprises applying the sensor input to an input of an artificial neural network. (Item 24) the control parameter is a control parameter for an echo cancellation module, and determining the control parameter based on the sensor input includes: detecting a surface based on the sensor input; determining a time of flight between the surface and the first microphone; Item 17. The system according to item 16, comprising: (Item 25) 17. The system of claim 16, wherein the control parameter is a control parameter for a beamforming module, and determining the control parameter based on the sensor input includes determining a time of flight between the user and the first microphone. (Item 26) Item 17. The system of item 16, wherein the control parameters are control parameters for a noise reduction module, and determining the control parameters based on the sensor input includes determining frequencies to be attenuated within the first acoustic signal. (Item 27) the wearable headgear unit includes a second microphone; the sensor input includes a second acoustic signal detected via the second microphone; the control parameter is determined based on a difference between the first acoustic signal and the second acoustic signal. Item 17. The system according to item 16. (Item 28) the wearable headgear unit includes a plurality of microphones excluding the first microphone; The method further includes receiving, via the plurality of microphones, a plurality of acoustic signals associated with the first utterance; the control parameter is determined based on a difference between the first acoustic signal and the plurality of acoustic signals. Item 17. The system according to item 16. (Item 29) Item 17. The system of item 16, wherein the sensor is coupled to the wearable headgear unit. (Item 30) Item 17. The system of item 16, wherein the sensor is positioned within the user's environment. [Brief explanation of the drawings]

[0007] [Figure 1] FIG. 1 illustrates an exemplary wearable head device that may be used as part of a wearable system, according to some embodiments.

[0008] [Figure 2] FIG. 2 illustrates an exemplary handheld controller that may be used as part of a wearable system, according to some embodiments.

[0009] [Figure 3] FIG. 3 illustrates an exemplary auxiliary unit that may be used as part of a wearable system, according to some embodiments.

[0010] [Figure 4] FIG. 4 illustrates an example functional block diagram for an example wearable system, according to some embodiments.

[0011] [Figure 5] FIG. 5 illustrates a flow chart of an exemplary speech processing system, according to some embodiments.

[0012] [Figure 6] FIG. 6 illustrates a flowchart of an exemplary system for processing an acoustic speech signal, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0013] In the following description of the embodiments, reference is made to the accompanying drawings which form a part hereof, and in which is shown, by way of illustration, specific embodiments which may be practiced. It is to be understood that other embodiments may be used and structural changes may be made without departing from the scope of the disclosed embodiments.

[0014] Exemplary Wearable System

[0015] 1 illustrates an exemplary wearable head device 100 configured to be worn on a user's head. Wearable head device 100 may be part of a broader wearable system that includes one or more components, such as a head device (e.g., wearable head device 100), a handheld controller (e.g., handheld controller 200 described below), and / or an auxiliary unit (e.g., auxiliary unit 300 described below). In some examples, wearable head device 100 can be used for virtual reality, augmented reality, or mixed reality systems or applications. The wearable head device 100 includes one or more displays, such as displays 110A and 110B (which may comprise left and right transmissive displays and associated components for coupling light from the displays to the user's eyes, such as orthogonal pupil expansion (OPE) grating sets 112A / 112B and exit pupil expansion (EPE) grating sets 114A / 114B), left and right acoustic structures, such as speakers 120A and 120B (which may be mounted on temple arms 122A and 122B, respectively, and positioned adjacent the user's left and right ears), and an infrared sensor. The wearable head device 100 may include one or more sensors, such as a microphone, an accelerometer, a GPS unit, an inertial measurement unit (IMU) (e.g., IMU 126), an acoustic sensor (e.g., microphone 150), a quadrature coil electromagnetic receiver (e.g., receiver 127 shown mounted on left temple arm 122A), left and right cameras (e.g., depth (time-of-flight) cameras 130A and 130B) oriented away from the user, and left and right eye cameras (e.g., for detecting the user's eye movements) (e.g., eye cameras 128 and 128B) oriented toward the user. However, the wearable head device 100 may incorporate any suitable display technology and any suitable number, type, or combination of sensors or other components without departing from the scope of the invention.In some examples, wearable head device 100 may incorporate one or more microphones 150 configured to detect audio signals generated by the user's voice, and such microphones may be positioned within the wearable head device adjacent to the user's mouth. In some examples, wearable head device 100 may incorporate networking features (e.g., Wi-Fi capabilities) for communicating with other devices and systems, including other wearable systems. Wearable head device 100 may further include components such as a battery, a processor, memory, a storage unit, or various input devices (e.g., buttons, touchpad), or may be coupled to a handheld controller (e.g., handheld controller 200) or auxiliary unit (e.g., auxiliary unit 300) that comprises one or more such components. In some examples, sensors may be configured to output a set of coordinates of the head-mounted unit relative to the user's environment and may provide input to a processor to implement a simultaneous localization and mapping (SLAM) procedure and / or a visual odometry algorithm. In some embodiments, the wearable head device 100 may be coupled to a handheld controller 200 and / or an auxiliary unit 300, as described further below.

[0016] 2 illustrates an exemplary mobile handheld controller component 200 of an exemplary wearable system. In some examples, handheld controller 200 may communicate wired or wirelessly with wearable head device 100 and / or auxiliary unit 300, described below. In some examples, handheld controller 200 includes a handle portion 220 to be held by a user and one or more buttons 240 disposed along a top surface 210. In some examples, handheld controller 200 may be configured for use as an optical tracking target; for example, a sensor (e.g., a camera or other optical sensor) of wearable head device 100 can be configured to detect the position and / or orientation of handheld controller 200, which in turn may indicate the position and / or orientation of a user's hand holding handheld controller 200. In some examples, handheld controller 200 may include a processor, memory, a storage unit, a display, or one or more input devices, such as those described above. In some examples, the handheld controller 200 includes one or more sensors (e.g., any of the sensors or tracking components described above with respect to the wearable head device 100). In some examples, the sensors can detect the position or orientation of the handheld controller 200 relative to the wearable head device 100 or relative to another component of the wearable system. In some examples, the sensors may be positioned within the handle portion 220 of the handheld controller 200 and / or may be mechanically coupled to the handheld controller. The handheld controller 200 can be configured to provide one or more output signals corresponding, for example, to the press state of the button 240 or the position, orientation, and / or movement of the handheld controller 200 (e.g., via an IMU). Such output signals may be used as inputs to a processor of the wearable head device 100, to the auxiliary unit 300, or to another component of the wearable system.In some embodiments, the handheld controller 200 may include one or more microphones to detect sounds (e.g., a user's speech, environmental sounds) and, in some cases, provide signals corresponding to the detected sounds to a processor (e.g., a processor of the wearable head device 100).

[0017] 3 illustrates an exemplary auxiliary unit 300 of an exemplary wearable system. In some examples, the auxiliary unit 300 may communicate wired or wirelessly with the wearable head device 100 and / or the handheld controller 200. The auxiliary unit 300 may include a battery to provide energy for operating one or more components of the wearable system, such as the wearable head device 100 and / or the handheld controller 200 (including a display, sensors, an acoustic structure, a processor, a microphone, and / or other components of the wearable head device 100 or the handheld controller 200). In some examples, the auxiliary unit 300 may include a processor, memory, a storage unit, a display, one or more input devices, and / or one or more sensors, such as those described above. In some examples, the auxiliary unit 300 includes a clip 310 for attaching the auxiliary unit to a user (e.g., to a belt worn by the user). An advantage of using auxiliary unit 300 to store one or more components of a wearable system is that doing so may allow large or heavy components to be carried on the user's waist, chest, or back, which are relatively better suited to supporting large, heavy objects, rather than being mounted on the user's head (e.g., when stored in wearable head device 100) or carried by the user's hands (e.g., when stored in handheld controller 200). This may be particularly advantageous with respect to relatively heavy or bulky components, such as batteries.

[0018] 4 shows an example functional block diagram that may correspond to an example wearable system 400, such as may include the example wearable head device 100, handheld controller 200, and auxiliary unit 300 described above. In some examples, the wearable system 400 may be used for virtual reality, augmented reality, or mixed reality applications. As shown in FIG. 4 , the wearable system 400 may include an example handheld controller 400B, referred to herein as a “totem” (and which may correspond to the handheld controller 200 described above), which may include a totem / headgear six-degree-of-freedom (6DOF) totem subsystem 404A. The wearable system 400 may also include an example wearable head device 400A (which may correspond to the wearable headgear device 100 described above), which includes a totem / headgear 6DOF headgear subsystem 404B. In some embodiments, the 6DOF totem subsystem 404A and the 6DOF headgear subsystem 404B cooperate to determine six coordinates (e.g., offsets in three translational directions and rotations along three axes) of the handheld controller 400B relative to the wearable head device 400A. The six degrees of freedom may be expressed relative to the coordinate system of the wearable head device 400A. The three translational offsets may be expressed as X, Y, and Z offsets within such a coordinate system, a translation matrix, or some other representation. The rotational degrees of freedom may be expressed as a sequence of yaw, pitch, and roll rotations, a vector, a rotation matrix, a quaternion, or some other representation. In some embodiments, one or more depth cameras 444 (and / or one or more non-depth cameras) and / or one or more optical targets (e.g., buttons 240 of the handheld controller 200 as described above or dedicated optical targets included in the handheld controller) included within the wearable head device 400A can be used for 6DOF tracking.In some embodiments, the handheld controller 400B can include a camera as described above, and the headgear 400A can include optical targets for optical tracking in conjunction with the camera. In some embodiments, the wearable head device 400A and the handheld controller 400B each include a set of three orthogonally oriented solenoids used to wirelessly transmit and receive three distinguishable signals. By measuring the relative magnitudes of the three distinguishable signals received in each of the coils used for receiving, the 6DOF of the handheld controller 400B relative to the wearable head device 400A can be determined. In some embodiments, the 6DOF totem subsystem 404A can include an inertial measurement unit (IMU), which is useful for providing improved accuracy and / or more timely information regarding high-speed movement of the handheld controller 400B.

[0019] In some examples involving augmented reality or mixed reality applications, it may be desirable to transform coordinates from a local coordinate space (e.g., a coordinate space that is fixed relative to the wearable head device 400A) to an inertial coordinate space or to an environmental coordinate space. For example, such a transformation may be necessary for the display of the wearable head device 400A to present virtual objects in an expected position and orientation relative to the real environment (e.g., a virtual person sitting in a real chair facing forward, regardless of the position and orientation of the wearable head device 400A), rather than in a fixed position and orientation on the display (e.g., at the same position on the display of the wearable head device 400A). This can maintain the illusion that the virtual objects exist in the real environment (and do not appear unnaturally positioned in the real environment, e.g., as the wearable head device 400A shifts and rotates). In some examples, a compensatory transformation between coordinate spaces can be determined by processing images from the depth camera 444 (e.g., using simultaneous localization and mapping (SLAM) and / or visual odometry procedures) to determine a transformation of the wearable head device 400A relative to an inertial or environmental coordinate system. In the example shown in FIG. 4 , the depth camera 444 can be coupled to the SLAM / visual odometry block 406 and can provide images to the block 406. The SLAM / visual odometry block 406 implementation can include a processor configured to process the images and then determine the position and orientation of the user's head, which can be used to identify a transformation between the head coordinate space and the real coordinate space. Similarly, in some examples, an additional source of information regarding the user's head pose and location is obtained from the IMU 409 of the wearable head device 400A. Information from the IMU 409 can be integrated with information from the SLAM / visual odometry block 406 to provide improved accuracy and / or more timely information regarding rapid adjustments of the user's head pose and position.

[0020] In some examples, depth camera 444 can provide 3D images to hand gesture tracker 411, which can be implemented within a processor of wearable head device 400A. Hand gesture tracker 411 can identify the user's hand gestures, for example, by matching the 3D images received from depth camera 444 to stored patterns representing hand gestures. Other suitable techniques for identifying the user's hand gestures will also be apparent.

[0021] In some embodiments, one or more processors 416 may be configured to receive data from the headgear subsystem 404B, the IMU 409, the SLAM / visual odometry block 406, the depth camera 444, the microphone 450, and / or the hand gesture tracker 411. The processor 416 may also send and receive control signals to and from the 6DOF totem system 404A. The processor 416 may be wirelessly coupled to the 6DOF totem system 404A, such as in embodiments in which the handheld controller 400B is untethered. The processor 416 may further communicate with additional components, such as an audiovisual content memory 418, a graphical processing unit (GPU) 420, and / or a digital signal processor (DSP) audio spatializer 422. The DSP audio spatializer 422 may be coupled to a head-related transfer function (HRTF) memory 425. The GPU 420 may include a left channel output coupled to a left source of imagewise modulated light 424 and a right channel output coupled to a right source of imagewise modulated light 426. The GPU 420 may output stereoscopic image data to the imagewise modulated light sources 424, 426. The DSP audio spatializer 422 may output audio to the left speaker 412 and / or the right speaker 414. The DSP audio spatializer 422 may receive an input from the processor 419 indicating a direction vector from the user to a virtual sound source (which may be moved by the user, e.g., via the handheld controller 400B). Based on the direction vector, the DSP audio spatializer 422 may determine a corresponding HRTF (e.g., by accessing an HRTF or by interpolating multiple HRTFs). The DSP audio spatializer 422 may then apply the determined HRTF to an audio signal, such as an audio signal corresponding to a virtual sound generated by a virtual object.This can improve the believability and realism of virtual sounds by incorporating the user's relative position and orientation to the virtual sounds in the mixed reality environment, i.e., by presenting virtual sounds that match the user's expectations of what they would hear if the virtual sounds were real sounds in a real environment.

[0022] 4 , one or more of the processor 416, GPU 420, DSP audio spatializer 422, HRTF memory 425, and audio / visual content memory 418 may be included within auxiliary unit 400C (which may correspond to auxiliary unit 300 described above). Auxiliary unit 400C may include battery 427 to power its components and / or provide power to wearable head device 400A and / or handheld controller 400B. Including such components within an auxiliary unit, which may be mounted on the user's waist, can limit the size and weight of wearable head device 400A, which in turn can reduce fatigue in the user's head and neck.

[0023] While FIG. 4 presents elements corresponding to various components of exemplary wearable system 400, various other suitable arrangements of these components will be apparent to those skilled in the art. For example, elements shown in FIG. 4 as associated with auxiliary unit 400C may instead be associated with wearable head device 400A or handheld controller 400B. Furthermore, some wearable systems may dispense with handheld controller 400B or auxiliary unit 400C entirely. Such variations and modifications are understood to be within the scope of the disclosed embodiments.

[0024] Speech Processing Engine

[0025] A speech processing system generally includes a system that receives an input audio signal corresponding to human speech (a source signal), processes and analyzes the input audio signal, and produces, as a result of the analysis, an output corresponding to the human speech. A process or module for performing these tasks may be considered a speech processing engine. In the case of an automatic speech recognition system, the output of the speech processing engine may be a text transcript of the human speech. In the case of a natural language processing system, the output may be one or more commands or instructions indicated by the human speech, or a semantically meaningful representation (e.g., a logical representation or data structure) of the human speech. Other types of speech processing systems (e.g., automatic translation systems) are also contemplated and within the scope of this disclosure.

[0026] Speech processing systems are found in a wide variety of products and applications, including traditional telephone systems, automated voice messaging systems, voice assistants (including stand-alone and smartphone-based voice assistants), vehicles and aircraft, desktop and document processing software, data entry, home appliances, medical devices, language translation software, closed captioning systems, and others. An advantage of speech processing systems is that they can enable users to provide input to computer systems using naturally spoken language as presented to a microphone instead of traditional computer input devices such as keyboards or touchscreens. Thus, speech processing systems can be particularly useful in environments where traditional input devices (e.g., keyboards) may be unavailable or impractical. Furthermore, by allowing users to provide intuitive voice-based input, speech recognition systems can enhance the sense of immersion. Thus, speech recognition may be a natural fit for wearable systems, particularly for virtual reality, augmented reality, and / or mixed reality applications of wearable systems, where user immersion is a primary goal and it may be desirable to limit the use of traditional computer input devices, the presence of which may detract from the sense of immersion.

[0027] FIG. 5 illustrates an automatic speech recognition engine 500 according to some embodiments. Engine 500 is intended to be illustrative of automatic speech recognition systems generally; other specific systems are also possible and within the scope of this disclosure. Engine 500 may be implemented using one or more processors (e.g., CPU, GPU, and / or DSP), memory, input devices (e.g., microphones), output devices (e.g., displays, speakers), networks, databases, and / or other suitable components. In engine 500, an audio signal 510 corresponding to a source human speech signal is presented to a signal pre-processing stage 520. In some examples, signal pre-processing stage 520 can apply one or more signal processing functions to audio signal 510. For example, pre-processing functions can include audio processing functions such as peak compression, noise reduction, band-limiting, equalization, signal attenuation, or other suitable functions. These pre-processing functions can simplify the task of later processing and analyzing audio signal 510. For example, a feature extraction algorithm may be calibrated to perform best on an input signal having certain audio characteristics, such as gain and frequency response, where the signal-to-noise ratio of the input signal is maximized. In some embodiments, the pre-processing audio signal 510 may condition the signal so that it can be more reliably analyzed elsewhere within the engine 500. For example, the signal pre-processing stage 520 may re-encode the audio signal 510 (e.g., re-encode at a specific bit rate) or convert the audio signal 510 from a first form (e.g., a time-domain signal) to a second form (e.g., a frequency-domain signal or a parameter representation) that may simplify subsequent processing of the audio signal 510. In some embodiments, one or more functions of the pre-processing stage 520 may be implemented by a DSP.

[0028] In step 530, a feature extraction process can be applied to the audio signal 510 (as pre-processed in step 520). The goal of feature extraction is to identify individual speech features of the audio signal 510 and reduce or eliminate variation in these features so that the features can be processed effectively and consistently (e.g., compared to patterns stored in a database). For example, feature extraction can reduce, eliminate, or control for variation in speaker pitch, gender, accent, pronunciation, and pace. Feature extraction can also reduce, eliminate, or control for variation in recording equipment (e.g., microphone type), signal transmission (e.g., over land-based telephone lines or cellular telephone networks), or recording environment (e.g., room acoustics, background noise level, speaker distance from microphone, speaker angle relative to microphone). Various suitable techniques for feature extraction are known in the art.

[0029] The speech features extracted from the audio signal 510 in stage 530 may be presented to a decoder stage 540. The goal of the decoder stage 540 is to determine a text output 570 that corresponds to the original human speech from which the audio signal 510 was generated. In some embodiments, the text output 570 need not be text, but may be another data representation of the original speech. Various techniques exist for decoding the speech features into text, such as hidden Markov models, Viterbi decoding, beam search, dynamic search, multi-pass search, weighted finite state transducers (WFSTs), or any suitable combination of the above. Other suitable techniques will be familiar to those skilled in the art.

[0030] In some embodiments, decoder 540 may utilize an acoustic modeling stage 550 to facilitate the generation of text output 570. Acoustic modeling stage 550 may identify one or more linguistic units from audio signal 510 (including one or more features extracted in stage 530) using a model of the relationship between the speech signal and linguistic units (e.g., phonemes). Those skilled in the art will be familiar with a variety of suitable acoustic modeling techniques that may be applied to acoustic modeling stage 550.

[0031] In some embodiments, decoder 540 may utilize a language modeling stage 560 to facilitate generation of text output 570. Linguistic modeling stage 560 may use a model of a language's grammar, vocabulary, and other characteristics to determine linguistic units (e.g., phonemes) that likely best correspond to features of audio signal 510. For example, the linguistic model applied by stage 560 may conclude that a particular extracted feature is more likely to correspond to a high-frequency word in a speaker's language than a low-frequency word in that language. Those skilled in the art will be familiar with various suitable linguistic modeling techniques that may be applied to linguistic modeling stage 560. Additionally, decoder 540 may utilize other suitable techniques or models to facilitate generation of text output 570. This disclosure is not limited to any particular technique or group of techniques.

[0032] Typically, the text output 570 does not correspond with perfect certainty to the source human speech. Instead, the likelihood that the text output 570 correctly corresponds to the source human speech can be expressed as a probability or confidence interval. Due to the many variables that can affect the audio signal 510, even advanced speech recognition systems do not consistently produce perfect text output for all speakers. For example, the reliability of a speech recognition system such as engine 500 can depend heavily on the quality of the input audio signal 510. If the audio signal 510 is recorded under ideal conditions, for example, in an acoustically controlled environment, with a human speaker speaking clearly and directly into a microphone from a close distance, the source speech can be more easily determined from the audio signal. For example, features can be more reliably extracted from the audio signal 510 in stage 530, and the decoder 540 can more effectively determine the text output 570 corresponding to those features (e.g., acoustic modeling can be more reliably applied to the features in stage 550 and / or linguistic modeling can be more reliably applied to the features in stage 560).

[0033] However, in real-world applications, the audio signal 510 may deviate from ideal conditions to the extent that determining the source human speech may become more difficult. For example, the audio signal 510 may incorporate environmental noise, such as that introduced by an outdoor environment or substantial distance between the human speaker and the microphone; electrical noise, such as from electrical interference (e.g., a battery charger for a smartphone); natural reverberations, such as from nearby surfaces (e.g., concrete, bathroom tiles) or acoustic spaces (e.g., caves, cathedrals); or other undesirable effects. In addition, the audio signal 510 may suffer from attenuation of certain frequencies, such as may occur when the human speaker turns away from the microphone. This is particularly problematic when the attenuated frequencies carry significant speech-related information (e.g., formant frequencies that can be used to distinguish vowel sounds). Similarly, the audio signal 510 may suffer from an overall low amplitude or low signal-to-noise ratio, which may occur when there is a large distance between the human speaker and the microphone. Additionally, if a human speaker moves and re-orients while speaking, the audio signal 510 may change characteristics over the course of the signal, further complicating the effort to determine the underlying speech.

[0034] Although exemplary system 500 illustrates an exemplary speech recognition engine, other types of speech processing engines may follow a similar structure. For example, a natural language processing engine, in response to receiving an input audio signal corresponding to human speech, may perform a signal pre-processing stage, extract components from the signal (e.g., via a segmentation and / or tokenization stage), and, in some cases, perform component detection / analysis with the assistance of one or more linguistic modeling subsystems. Furthermore, in some embodiments, the output of an automatic speech recognition engine such as that shown in exemplary system 500 may be used as input to a further language processing engine. For example, a natural language processing engine may receive the text output 570 of exemplary system 500 as input. Such systems may suffer challenges similar to those faced by exemplary system 500. For example, variations in the input audio signal, which may make it more difficult to recover the underlying source speech signal of the audio signal as described above, may also make it more difficult to provide other forms of output (e.g., a logical representation or data structure, in the case of a natural language processing engine). Therefore, such systems will also benefit from the present invention as described below.

[0035] Improving speech processing using wearable systems

[0036] The present disclosure is directed to systems and methods for improving the accuracy of speech processing systems by using input from sensors, such as those associated with wearable devices (e.g., head-mounted devices such as those described above with respect to FIG. 1 ), to reduce, eliminate, or control variations in input audio signals, such as those described above with respect to audio signal 510. Such variations may be particularly noticeable in mobile applications of speech processing or applications of speech processing in uncontrolled environments, such as outdoor environments. Wearable systems are often intended for use in such applications and may be subject to such variations. A wearable system may generally refer to any combination of a head device (e.g., wearable head device 100), a handheld controller (e.g., handheld controller 200), an auxiliary unit (e.g., auxiliary unit 300), and / or the head device's environment. In some embodiments, sensors of a wearable system may be on the head device, handheld controller, auxiliary unit, and / or in the head device's environment. For example, because wearable systems may be designed to be mobile, audio signals recorded at a single stationary microphone from a user of the wearable system (e.g., a stand-alone voice-assistance device) may suffer from low signal-to-noise ratios if the user is far from the single stationary microphone, or from "acoustic shadowing" or undesirable frequency responses if the user faces away from the microphone. Furthermore, the audio signal may change characteristics over time as the user moves and turns relative to the single stationary microphone, as might be expected with a mobile user of a wearable system. In addition, because some wearable systems are intended for use in uncontrolled environments, there is a high potential for environmental noise (or other human speech) to be recorded along with the target human's speech. Similarly, such uncontrolled environments may introduce undesirable echoes and reverberations into the audio signal 510 that may obscure the underlying speech.

[0037] As described above with respect to the exemplary wearable head device 100 in FIG. 1 , a wearable system can include one or more sensors that can provide input about a user and / or the environment of the wearable system. For example, the wearable head device 100 can include a camera (e.g., camera 444 illustrated in FIG. 4 ) and output a visual signal corresponding to the environment. In some examples, the camera can be a forward-facing camera on a head-mounted unit that shows what is currently in front of the user of the wearable system. In some examples, the wearable head device 100 can include a LIDAR unit, a radar unit, and / or an acoustic sensor, which can output a signal corresponding to the physical geometry of the user's environment (e.g., walls, physical objects). In some examples, the wearable head device 100 can include a GPS unit, which can indicate geographic coordinates corresponding to the current location of the wearable system. In some examples, the wearable head device 100 can include an accelerometer, a gyroscope, and / or an inertial measurement unit (IMU) and can indicate the orientation of the wearable head device 100. In some examples, wearable head device 100 can include an environmental sensor, such as a temperature or pressure sensor. In some examples, wearable head device 100 can include a biometric sensor, such as an iris camera, a fingerprint sensor, an eye-tracking sensor, or a sensor for measuring the user's vital signs. In examples where wearable head device 100 includes a head-mounted unit, such orientation may correspond to the orientation of the user's head (and, for that matter, the direction of the user's mouth and the user's speech). Other suitable sensors can also be included. In some embodiments, handheld controller 200, auxiliary unit 300, and / or the environment of wearable head device 100 can include any suitable one or more of the sensors described above for wearable head device 100. Additionally, in some cases, one or more sensors may be installed in the environment with which the wearable system interacts.For example, a wearable system may be designed to be worn by a car driver, and appropriate sensors (e.g., depth cameras, accelerometers, etc.) may be installed inside the car. One advantage of this approach is that the sensors may occupy known locations in the environment. Compared to sensors that may be attached to wearable devices that move around in the environment, this configuration may simplify the interpretation of the data provided by those sensors.

[0038] Signals provided by such sensors in the wearable system (e.g., wearable head device 100, handheld controller 200, auxiliary unit 300, and / or the environment of wearable head device 100) can be used to provide information about the characteristics of the audio signal recorded by the wearable system and / or about the relationship between the audio signal and the underlying source speech signal. This information can then be used to more effectively determine the underlying source speech of that audio signal.

[0039] To illustrate, FIG. 6 shows an example speech recognition system 600 that incorporates a wearable system to improve speech recognition of audio signals recorded by one or more microphones. FIG. 6 shows a user of wearable system 601, which may correspond to wearable system 400 described above, which may include one or more of example wearable head device 100, handheld controller 200, and auxiliary unit 300. The user of wearable system 601 provides an oral utterance 602 (“source speech”), which is detected at one or more microphones 604, which output a corresponding audio signal 606. Wearable system 601 may include one or more sensors described above, including one or more of a camera, a LIDAR unit, a radar unit, an acoustic sensor, a GPS unit, an accelerometer, a gyroscope, an IMU, a microphone (which may be one of microphones 604), a temperature sensor, a biometric sensor, or any other suitable sensor or combination of sensors. The wearable system 601 may also include one or more processors (e.g., CPU, GPU, and / or DSP), memory, input devices (e.g., microphones), output devices (e.g., displays, speakers), networks, and / or databases. These components may be implemented using a combination of the wearable head device 100, handheld controller 200, and auxiliary unit 300. In some examples, sensors in the wearable system 601 provide sensor data 608 (which may be multi-channel sensor data, such as in examples where one or more sensors are presented in parallel, including sensors from two or more sensor types). The sensor data 608 may be provided in parallel with the microphone 604 that detects the source speech. That is, the sensors in the wearable system 601 may provide sensor data 608 that corresponds to the conditions at the time the source speech is provided.In some embodiments, one or more of the microphones 604 may be included in a wearable head device, a handheld controller, and / or an auxiliary unit of the wearable system 601 and / or its environment.

[0040] In example system 600, subsystem 610 may receive sensor data 608 and audio signal 606 as inputs, determine and apply control parameters to process audio signal 606, and provide the processed audio signal as an input to a signal processing engine (e.g., speech recognition engine 650 and / or natural language processing (NLP) engine 670), which may generate an output (e.g., text output 660 and / or natural language processing output 680, respectively). Subsystem 610 includes one or more processes, stages, and / or modules described below and illustrated in FIG. 6. Subsystem 610 may be implemented using any combination of one or more processors (e.g., CPUs, GPUs, and / or DSPs), memory, networks, databases, and / or other suitable components. In some examples, some or all of subsystems 610 can be implemented on wearable system 601, such as on one or more of a head device (e.g., wearable head device 100), a handheld controller (e.g., handheld controller 200), and / or an auxiliary unit (e.g., auxiliary unit 300). In some examples, some or all of subsystems 610 can be implemented on a device containing microphone 604 (e.g., a smartphone or a standalone voice assistance device). In some examples, some or all of subsystems 610 can be implemented on a cloud server or another network-enabled computing device. For example, a local device may perform latency-sensitive functions of subsystem 610, such as those related to processing audio signals and / or sensor data, while a cloud server or other network-enabled device may perform functions of subsystem 610 that require large computational or memory resources (e.g., training or applying complex artificial neural networks) and transmit the output to the local device. Other implementations of subsystem 610 will be apparent to those skilled in the art and are within the scope of this disclosure.

[0041] In the exemplary system 600, the subsystem 610 includes a sensor data analysis stage 620. The sensor data analysis stage 620 can, for example, process and analyze the sensor data 608 to determine information about the environment of the wearable system 601. In examples where the sensor data 608 includes sensor data from disparate sources (e.g., camera data and GPS data), the sensor data analysis stage 620 can combine the sensor data (“sensor fusion”) according to techniques known to those skilled in the art (e.g., a Kalman filter). In some examples, the sensor data analysis stage 620 can incorporate data from other sources in addition to the sensor data 608. For example, the sensor data analysis stage 620 can combine the sensor data 608 from the GPS unit with map data and / or satellite data (e.g., from a memory or database that stores such data) to determine location information based on the output of the GPS unit. As an example, the GPS unit may output GPS coordinates corresponding to the latitude and longitude of the wearable system 601 as part of the sensor data 608. The sensor data analysis stage 620 may use latitude and longitude, along with map data, to identify the county, town, street, unit (e.g., commercial or residential unit), or room in which the wearable system 601 is located, or to identify nearby businesses or points of interest. Similarly, architectural data (e.g., from public building records) can be combined with the sensor data 608 to identify the building in which the wearable system 601 is located, or weather data (e.g., from a real-time feed of satellite data) can be combined with the sensor data 608 to identify the current weather conditions at the location. Other exemplary applications will be apparent and are within the scope of this disclosure.

[0042] On a smaller scale, a sensor data analysis stage 620 can analyze the sensor data 608 to generate information related to objects and geometry in the immediate vicinity of the wearable system 601 or associated with the user of the wearable system 601. For example, using techniques known in the art, data from a LIDAR sensor or radar unit of the wearable system 601 may indicate that the wearable system 601 is facing a wall located 8 feet away at an angle θ relative to the normal to that wall, and image data from the camera of the wearable system 601 can identify that the wall is likely made of ceramic tile (an acoustically reflective material). In some examples, an acoustic sensor of the wearable system 601 can be used to measure the acoustic effect that a surface may have on an acoustic signal (e.g., by comparing a signal reflected from the surface with a source signal transmitted to the surface). In some examples, the sensor data analysis stage 620 can use the sensor data 608 to determine the position and / or orientation of the user of the wearable system 601, for example, using an accelerometer, gyroscope, or IMU associated with the wearable system 601. In some examples, such as augmented or mixed reality applications, the stage 620 can incorporate a map or other representation of the user's current environment. For example, if the sensors of the wearable system 601 are used to build a 3D representation of the geometry of a room, that 3D representation data can be used in conjunction with the sensor data 608. Similarly, the stage 620 can incorporate information such as the materials of nearby surfaces and the acoustic properties of those surfaces, information related to other users in the environment (e.g., their location and orientation and / or the acoustic characteristics of their voices), and / or information about the user of the wearable system 601 (e.g., the user's age group, gender, native language, and / or voice characteristics).

[0043] In some examples, the sensor data analysis stage 620 can analyze the sensor data 608 and generate information related to the microphones 604. For example, the sensor data 608 may provide the position and / or orientation of one or more microphones 604, such as their position and / or orientation relative to the wearable system 601. In examples in which the wearable system 601 includes one or more microphones 604, the position and orientation of the one or more microphones 604 may be directly linked to the position and orientation of the wearable system 601. In some examples, the wearable system 601 can include one or more additional microphones other than the one or more microphones 604. Such additional one or more microphones can be used to provide a baseline audio signal, for example, corresponding to the user's speech as detected from a known position and orientation and from a short distance. For example, the additional one or more microphones can be at a known position with a known orientation in the user's environment. The amplitude, phase, and frequency characteristics of this baseline audio signal can be compared to the audio signal 606 to identify a relationship between the source speech 602 and the audio signal 606. For example, if a first audio signal detected at a first time has half the amplitude of the baseline audio signal and a second audio signal detected at a second time has a quarter the amplitude of the baseline audio signal, it can be inferred that the user moved away from microphone 604 during the interval between the first and second times. This can be extended to any suitable number of microphones (e.g., an initial microphone and two or more additional microphones).

[0044] Based on the sensor data 608 and / or other data as described above, the information output by the sensor data analysis stage 620 can identify a relationship between the source speech 602 and the corresponding audio signal 606. Information describing this relationship can be used in stage 630 to calculate one or more control parameters that can be applied to the audio signal 606, the one or more microphones 604, and / or the wearable system 601. Application of these control parameters can improve the accuracy with which the system 600 can recover the underlying source speech 602 from the audio signal 606. In some embodiments, the control parameters calculated in stage 630 can include digital signal processing (DSP) parameters that can be applied to process the audio signal 606. For example, such control parameters may include parameters for a digital signal processing (DSP) noise reduction process (e.g., a signal threshold below which gated noise reduction will be applied to the audio signal 606, or a noise frequency in the audio signal 606 to be attenuated), parameters for a DSP echo cancellation or reverberation process (e.g., a time value corresponding to the delay between the audio signal 606 and an echo of that signal), or parameters for other audio DSP processes (e.g., phase correction, limiting, pitch correction).

[0045] In some embodiments, the control parameters may define a DSP filter to be applied to the audio signal 606. For example, the sensor data 608 (e.g., from a microphone in a head-mounted unit of the wearable system 601) may indicate a characteristic frequency curve corresponding to the voice of a user of the wearable system 601 (i.e., the user generating the source speech 602). This frequency curve can be used to determine control parameters that define a digital band-pass filter to apply to the audio signal 606. This band-pass filter can isolate frequencies that more closely correspond to the source speech 602 to make the source speech 602 more prominent in the audio signal 606. In some embodiments, the sensor data 608 (e.g., from a microphone in a head-mounted unit of the wearable system 601) may indicate a characteristic frequency curve corresponding to the voice of a different user (other than the user of the wearable system 601) in the vicinity of the wearable system 601. This frequency curve can be used to determine control parameters that define a digital notch filter to apply to the audio signal 606. This notch filter can remove undesired sounds from the audio signal 606 to render the source speech 602 more prominent in the audio signal 606. Similarly, sensor data 608 (e.g., the camera of the wearable system 601) can identify specific other individuals in the vicinity and their location relative to the wearable system 601. This information can determine the level of the notch filter (e.g., the closer the individual, the more likely their voice is to be louder in the audio signal 606, and the greater the level of attenuation that may need to be applied). As another example, the presence of certain surfaces and / or materials in the user's vicinity can affect the frequency characteristics of the user's voice as detected by the microphone 604. For example, if the user is standing in a corner of a room, certain low frequencies of the user's voice may become prominent in the audio signal 606. This information can be used to generate parameters (e.g., cutoff frequency) of a high-pass filter to be applied to the audio signal 606.These control parameters can be applied to the audio signal 606 in stage 640 or as part of the speech recognition engine 650 and / or the natural language processing engine 670 (e.g., in a stage corresponding to the signal pre-processing stage 520 or the feature extraction stage 530 described above).

[0046] In some examples, the control parameters calculated in stage 630 can be used to configure the microphone 604. For example, such control parameters can include hardware configuration parameters such as gain levels for hardware amplifiers coupled to the microphone 604, beamforming parameters for adjusting the directivity of the microphone 604 (e.g., the vector to which the microphone 604 should be pointed), parameters for determining which microphones 604 should be enabled or disabled, or parameters for controlling where the microphone 604 should be positioned or oriented (e.g., in examples where the microphone 604 is mounted on a mobile platform). In some examples, the microphone 604 may be a component of a smartphone or another mobile device, and the control parameters calculated in stage 630 can be used to control the mobile device (e.g., to enable various components of the mobile device or to configure or operate software on the mobile device).

[0047] In some examples, the control parameters calculated in stage 630 can be used to control the wearable system 601 itself. For example, such control parameters may include parameters for presenting a message to a user of the wearable system 601 via display 110A / 110B or speaker 120A / 120B, etc. (e.g., an audio or video message that the user should move away from a nearby wall to improve speech recognition accuracy), or parameters for enabling, disabling, or reconfiguring one or more sensors of the wearable system 601 (e.g., to reorient a servo-mounted camera to obtain more useful camera data). In examples in which the wearable system 601 includes a microphone 604, the control parameters can be sent to the wearable system 601 to control the microphone 604 as described above.

[0048] In some embodiments, the control parameters calculated in stage 630 can be used to influence a decoding process (e.g., decoding process 540 described above) of a speech processing system (e.g., speech recognition engine 650 and / or natural language processing engine). For example, sensor data 608 may indicate characteristics of a user's environment, behavior, or mental state, which may affect the user's language use. For example, sensor data 608 (e.g., from a camera and / or GPS unit) may indicate that a user is watching a football game. Because the user's utterance (i.e., source utterance 602) may be much more likely than usual to include football-related words (e.g., "coach," "quarterback," "touchdown") while the user is watching a football game, the control parameters of the speech processing system (e.g., language modeling stage 560) can be temporarily set to reflect a higher probability that the audio signal corresponds to football-related words.

[0049] 6 , individual update modules 632, 634, and 636 may determine the control parameters calculated in stage 630 or may apply the control parameters calculated in stage 630 and / or the sensor data 608 to the audio signal 606, the microphone 604, the wearable system 601, or any hardware or software subsystems described above. For example, the beamformer update module 632 may determine, based on the sensor data 608 or one or more control parameters calculated in stage 630, how a beamforming module (e.g., the beamforming module of the microphone 604) may be updated to improve recognition or natural language processing of the audio signal 606. In some examples, the beamforming update module 632 may control the directionality of the sensor array (e.g., the array of microphones 604) to maximize the signal-to-noise ratio of the signal detected by the sensor array. For example, the beamforming update module 632 may adjust the directionality of the microphone 604 so that the source speech 602 is detected with minimal noise and distortion. For example, in a room with multiple voices, the adaptive beamforming module software can direct the microphone 604 to maximize the signal power corresponding to the voice of interest (e.g., the voice corresponding to the source speech 602). For example, a sensor in the wearable system 601 can output data indicating that the wearable system 601 is located a certain distance from the microphone 604, from which a time-of-flight value of the source speech 602 to the microphone 604 can be determined. This time-of-flight value can be used to calibrate the beamforming, using techniques familiar to those skilled in the art, to maximize the ability of the speech recognition engine 650 to distinguish the source speech 602 from the audio signal 606.

[0050] In some examples, the noise reduction update module 634 can determine, based on the sensor data 608 or one or more control parameters calculated in stage 630, how the noise reduction process can be updated to improve recognition or natural language processing of the audio signal 606. In some examples, the noise reduction update module 634 can control parameters of the noise reduction process applied to the audio signal 606 to maximize the signal-to-noise ratio of the audio signal 606. This, in turn, can facilitate automatic speech recognition performed by the speech recognition engine 650. For example, the noise reduction update module 634 can selectively apply signal attenuation to frequencies of the audio signal 606 where noise is likely to be present, while increasing (or decreasing to attenuate) frequencies of the audio signal 606 that carry information of the source speech 602. Sensors of the wearable system 601 can provide data that helps the noise reduction update module 634 identify frequencies of the audio signal 606 that are likely to correspond to noise and frequencies that are likely to carry information about the source speech 602. For example, sensors (e.g., GPS, LIDAR, etc.) of the wearable system 601 may identify that the wearable system 601 is located on an aircraft. The aircraft may be associated with background noise having certain characteristic frequencies. For example, aircraft engine noise may be centered around a known frequency f. Based on this information from the sensors, the noise reduction update module 634 may attenuate the frequency f of the audio signal 606. Similarly, sensors (e.g., microphones mounted on the wearable system 601) of the wearable system 601 may identify a frequency signature corresponding to the voice of a user of the wearable system 601. The noise reduction update module 634 may apply a bandpass filter to the frequency range corresponding to that frequency signature, or may ensure that noise reduction is not applied to that frequency range.

[0051] In some embodiments, the echo cancellation (or reverberation) update module 636 may determine, based on the sensor data 608 or one or more control parameters calculated in stage 630, how the echo cancellation unit may be updated to improve recognition or natural language processing of the audio signal 606. In some embodiments, the echo cancellation update module 636 may control parameters of the echo cancellation unit applied to the audio signal 606 to maximize the ability of the speech recognition engine 650 to determine the source speech 602 from the audio signal 606. For example, the echo cancellation update module 636 may instruct the echo cancellation unit to detect and correct (e.g., via a comb filter) echoes in the audio signal 606 that last 100 milliseconds after the source speech. Because such echoes can interfere with the ability of a speech processing system (e.g., speech recognition engine 650, natural language processing engine 670) to determine the source speech (e.g., by affecting the ability to extract features from audio signal 606, such as those described above with respect to FIG. 5 in stage 530), removing these echoes can result in greater accuracy in speech recognition, natural language processing, and other speech processing tasks. In some examples, sensors in wearable system 601 can provide sensor data 608 that can be used to determine control parameters for echo cancellation. For example, such sensors (e.g., cameras, LIDAR, radar, acoustic sensors) can determine that a user of wearable system 601 is located 10 feet from a surface and facing the surface at an angle θ1 relative to the normal to the surface, and further that microphone 604 is located 20 feet from the surface and facing the surface at an angle θ2. From this sensor data, it can be calculated that the surface is likely to produce an echo that reaches the microphone 604 some time after the source signal 602 (i.e., the flight time from the user to the surface + the flight time from the surface to the microphone 604).Similarly, it can be determined from the sensor data that the surface corresponds to a bathroom tile surface or another surface with known acoustically reflective properties, which can be used to calculate control parameters for an echo cancellation unit that will attenuate the resulting acoustic reflections in the audio signal 606.

[0052] Similarly, in some embodiments, signal conditioning may be applied to the speech audio signal 606 to account for equalization applied to the speech audio signal 606 depending on the acoustic environment. For example, a room may boost or attenuate certain frequencies of the speech audio signal 606 due to, for example, the room's geometry (e.g., dimensions, cubic volume), materials (e.g., concrete, bathroom tile), or other characteristics that may affect the signal as detected by a microphone (e.g., microphone 604) in the acoustic environment. These effects may complicate the ability of the speech recognition engine 650 to perform consistently across different acoustic environments. Sensors in the wearable system 601 may provide sensor data 608 that may be used to counter such effects. For example, the sensor data 608 may indicate the size or shape of the room or the presence of acoustically significant materials, as described above, from which one or more filters may be determined and applied to the audio signal to counteract the room's effects. In some embodiments, sensor data 608 can be provided by a sensor (e.g., camera, LIDAR, radar, acoustic sensor) that indicates the cubic volume of the room in which the user is present, and the acoustic effects of the room can be modeled as a filter, and the inverse of that filter can be applied to the audio signal to compensate for the acoustic effects.

[0053] In some examples, modules 632, 634, or 636, or other suitable elements of exemplary system 600 (e.g., stage 620 or stage 640), can determine control parameters based on sensor data 608 using a predetermined mapping. In some examples described above, the control parameters are calculated directly based on sensor data 608. For example, as described above, echo cancellation update module 636 can determine control parameters for an echo cancellation unit that can be applied to attenuate echoes present in audio signal 606. As described above, such control parameters can be calculated by geometrically determining the distance between wearable system 601 and microphone 604, calculating the time of flight that an audio signal traveling at the speed of sound in air would require to travel from wearable system 601 to the microphone, and setting the echo period of the echo cancellation unit to correspond to that time of flight. However, in some cases, the control parameters can be determined by comparing sensor data 608 with a mapping of sensor data for controlling parameters. Such a mapping can be stored in a database, such as on a cloud server or another networked device. In some embodiments, one or more elements of the exemplary system 601 (e.g., the beamformer update module 632, the noise reduction update module 634, the echo cancellation update module 636, the stage 630, the stage 640, the speech recognition engine 650, and / or the natural language processing engine 670) can query a database to retrieve one or more control parameters that correspond to the sensor data 608. This process can occur instead of, or in addition to, directly calculating the control parameters as described above. In such a process, the sensor data 608 can be provided to a predetermined mapping, and one or more control parameters that most closely correspond to the sensor data 608 within the predetermined mapping can be returned. Using a predetermined mapping of sensor data to parameters can have several advantages.For example, performing a lookup from a predetermined mapping may be computationally cheaper than processing sensor data in real time, especially when the calculation may involve complex geometric data or if the sensors of the wearable system 601 may suffer from significant latency or bandwidth limitations. Furthermore, the predetermined mapping can capture relationships between sensor data and control parameters that may be difficult to calculate from the sensor data alone (e.g., by mathematical modeling).

[0054] In some examples, machine learning techniques can be used to generate or refine the mapping between sensor data and control parameters or otherwise determine the control parameters from the sensor data 608. For example, a neural network (or other suitable machine learning technique) can be trained to identify desired control parameters based on sensor data input, according to techniques familiar to those skilled in the art. The desired control parameters can be identified, and the neural network further refined through user feedback. For example, a user of the system 600 can be prompted to rate the quality of the speech recognition output (e.g., text output 660) and / or the natural language processing output (e.g., natural language processing output 680). Such user ratings can be used to adjust the likelihood that a given set of sensor data will result in a particular set of control parameters. For example, if a user reports a high rating of the text output 660 for a particular set of control parameters and a particular set of sensor data, a mapping between that sensor data and those control parameters can be created (or the link between them strengthened). Conversely, a low rating can weaken or eliminate the mapping between sensor data and control parameters.

[0055] Similarly, machine learning techniques can be utilized to improve the ability of a speech recognition engine (e.g., 650) to distinguish between speech belonging to the user and speech belonging to other entities, such as speech emanating from a television or stereo system. As explained above, a neural network (or other suitable machine learning technique) can be trained to perform this discrimination (i.e., determine whether input audio belongs to the user or some other source). In some cases, a calibration routine can be used to train the neural network, in which a user provides a set of input audio known to belong to that user. Other suitable machine learning techniques can also be used for the same purpose.

[0056] Although update modules 632, 634, and 636 are described as comprising a beamforming update module, a noise reduction update module, and an echo cancellation module, respectively, other suitable modules may also be included in any combination. For example, in some embodiments, the EQ update module may determine how a filtering process (as described above) may be updated based on the sensor data 608 and / or one or more control parameters calculated in stage 630. Furthermore, the functionality described above with respect to modules 632, 634, and 636 may be implemented in other elements of exemplary system 600, such as stage 630 and / or stage 640, or as part of the speech recognition engine 650 or natural language processing engine 670.

[0057] Although the disclosed embodiments have been fully described with reference to the accompanying drawings, it should be noted that various changes and modifications will be apparent to those skilled in the art. For example, elements of one or more implementations may be combined, deleted, modified, or supplemented to form further implementations. Such changes and modifications are to be understood as being included within the scope of the disclosed embodiments as defined by the appended claims.

Claims

1. 1. A method of processing a first acoustic signal, the method comprising: receiving, via a microphone, the first acoustic signal associated with a user of a wearable device; receiving one or more sensor inputs via one or more sensors, the one or more sensor inputs indicating a location of the user; determining the location of the user based on the one or more sensor inputs; determining a cutoff frequency of a filter based on the location of the user; applying the filter to the first acoustic signal; determining a three-dimensional representation of an environment of the wearable device; determining a control parameter based on the one or more sensor inputs and further based on the three-dimensional representation of an environment of the wearable device; applying the control parameters to the first acoustic signal; A method comprising:

2. 10. The method of claim 1 , wherein the filter is applied to the first acoustic signal to generate a second acoustic signal, the method further comprising providing the second acoustic signal to a speech recognition engine to generate a text output associated with the first acoustic signal.

3. 10. The method of claim 1, wherein the filter is applied to the first acoustic signal to generate a second acoustic signal, the method further comprising providing the second acoustic signal to a natural language processing engine to generate natural language data associated with the first acoustic signal.

4. A method for processing a first acoustic signal, the method comprising: receiving, via a microphone, the first acoustic signal associated with a user of a wearable device; receiving one or more sensor inputs via one or more sensors, the one or more sensor inputs indicating a location of the user; determining the location of the user based on the one or more sensor inputs; determining a cutoff frequency of a filter based on the location of the user; applying the filter to the first acoustic signal; detecting a surface based on the one or more sensor inputs; determining an effect of the surface on the first acoustic signal; applying a control parameter to the first acoustic signal; Including, The method, wherein the control parameter causes a reduction in the effect of the surface on the first acoustic signal.

5. The method of claim 4 , further comprising determining an acoustic property of the surface, wherein an effect of the surface on the first acoustic signal is determined based on the acoustic property.

6. A method for processing a first acoustic signal, the method comprising: receiving, via a microphone, the first acoustic signal associated with a user of a wearable device; receiving one or more sensor inputs via one or more sensors, the one or more sensor inputs indicating a location of the user; determining the location of the user based on the one or more sensor inputs; determining a cutoff frequency of a filter based on the location of the user; applying the filter to the first acoustic signal; detecting a person different from the user based on the one or more sensor inputs; determining an effect of the person on the first acoustic signal; applying a control parameter to the first acoustic signal; Including, The method, wherein the control parameter results in a reduction of the influence of people on the first acoustic signal.

7. A method for processing a first acoustic signal, the method comprising: receiving, via a microphone, the first acoustic signal associated with a user of a wearable device; receiving one or more sensor inputs via one or more sensors, the one or more sensor inputs indicating a location of the user; determining the location of the user based on the one or more sensor inputs; determining a cutoff frequency of a filter based on the location of the user; applying the filter to the first acoustic signal; detecting a surface based on the one or more sensor inputs; determining control parameters for echo cancellation, including determining a time of flight between the surface and the microphone; applying the control parameters for echo cancellation to the first acoustic signal; A method comprising:

8. A method for processing a first acoustic signal, the method comprising: receiving, via a microphone, the first acoustic signal associated with a user of a wearable device; receiving one or more sensor inputs via one or more sensors, the one or more sensor inputs indicating a location of the user; determining the location of the user based on the one or more sensor inputs; determining a cutoff frequency of a filter based on the location of the user; applying the filter to the first acoustic signal; determining control parameters for beamforming, including determining a time of flight between the user and the microphone; applying the control parameters for beamforming to the first acoustic signal; A method comprising:

9. determining a quality of the first acoustic signal; pursuant to determining that the quality of the first acoustic signal is below a quality threshold, presenting a message to the user to move from the location to a new location to improve the quality of the first acoustic signal; The method of claim 1 further comprising:

10. The method of claim 1 , wherein the wearable device comprises the one or more sensors.

11. The method of claim 1 , wherein the filter comprises a high-pass filter.

12. Wearable devices and A microphone and one or more sensors; one or more processors configured to perform the method; A system comprising: The method comprises: receiving, via the microphone, a first acoustic signal associated with a user of the wearable device; receiving one or more sensor inputs via the one or more sensors, the one or more sensor inputs indicating a location of the user; determining the location of the user based on the one or more sensor inputs; determining a cutoff frequency of a filter based on the location of the user; applying the filter to the first acoustic signal; determining a three-dimensional representation of an environment of the wearable device; determining a control parameter based on the one or more sensor inputs and further based on the three-dimensional representation of an environment of the wearable device; applying the control parameters to the first acoustic signal; Including, the system.

13. A wearable device, A microphone and one or more sensors; one or more processors configured to perform the method; A system comprising: The method comprises: receiving, via the microphone, a first acoustic signal associated with a user of the wearable device; receiving one or more sensor inputs via the one or more sensors, the one or more sensor inputs indicating a location of the user; determining the location of the user based on the one or more sensor inputs; determining a cutoff frequency of a filter based on the location of the user; applying the filter to the first acoustic signal; detecting a surface based on the one or more sensor inputs; determining an effect of the surface on the first acoustic signal; applying a control parameter to the first acoustic signal; Including, The control parameter causes a reduction in the effect of the surface on the first acoustic signal.

14. 13. The system of claim 12, wherein the filter is applied to the first acoustic signal to generate a second acoustic signal, the method further comprising providing the second acoustic signal to a speech recognition engine to generate a text output associated with the first acoustic signal.

15. 13. The system of claim 12, wherein the filter is applied to the first acoustic signal to generate a second acoustic signal, the method further comprising providing the second acoustic signal to a natural language processing engine to generate natural language data associated with the first acoustic signal.

16. The system of claim 12 , wherein the wearable device comprises the one or more sensors.

17. The system of claim 12 , wherein the filter comprises a high-pass filter.

18. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform a method, the method comprising: receiving, via a microphone, an acoustic signal associated with a user of the wearable device; receiving one or more sensor inputs via one or more sensors, the one or more sensor inputs indicating a location of the user; determining the location of the user based on the one or more sensor inputs; determining a cutoff frequency of a filter based on the location of the user; applying the filter to the acoustic signal; determining a three-dimensional representation of an environment of the wearable device; determining a control parameter based on the one or more sensor inputs and further based on the three-dimensional representation of an environment of the wearable device; applying said control parameters to said acoustic signal; 1. A non-transitory computer-readable medium comprising:

19. The one or more sensor inputs are further indicative of a nature of the location of the user, and the method further comprises:

10. The method of claim 1, comprising determining the nature of the location of the user based on the one or more sensor inputs, wherein the cutoff frequency of the filter is determined further based on the nature of the location of the user.

Citation Information

Patent Citations

  • Speech recognition device

    JP1994075588A

  • Speech recognizing device

    JP2000148184A

  • Handset

    JP2000261534A

  • Voice recognition method and voice recognition device using the method

    JP2001296887A

  • Telephone call and hands-free call for cordless terminals with echo compensation

    JP2002135173A