Far-field noise reduction using spatial filtering with a microphone array

The spatial filtering technique using a microphone array effectively isolates user speech from ambient noise in XR systems by reducing far-field noise, enhancing communication and recognition performance.

JP2026510853APending Publication Date: 2026-04-10MAGIC LEAP INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
MAGIC LEAP INC
Filing Date
2024-03-13
Publication Date
2026-04-10

Smart Images

  • Figure 2026510853000001_ABST
    Figure 2026510853000001_ABST
Patent Text Reader

Abstract

A system and method for reducing far-field noise (background noise) from near-field audio signals generated by a microphone array. The system and method disclosed herein uses an innovative spatial filtering technique to filter far-field ambient noise from near-field audio, reduce far-field noise within the near-field audio signal, thereby improving the near-field audio and isolating it from far-field interference. While not limited to this use, the system and method are useful for effectively filtering speech (i.e., near-field audio) from ambient noise (i.e., far-field noise), as described herein.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to Related Applications This application claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 453,022, entitled "FAR - FIELD NOISE REDUCTION VIA SPATIAL FILTERING USING A MICROPHONE ARRAY", filed on March 17, 2023. The content of the foregoing application is hereby expressly incorporated by reference herein for all purposes.

[0002] Field of the Invention The present invention generally relates to systems and methods for reducing far - field noise by spatial filtering using a microphone array that can be utilized in spatial audio systems within virtual reality, augmented reality, and / or mixed reality systems.

Background Art

[0003] Background Modern computing and display technologies have facilitated the development of mixed reality systems ("MR") for so - called "virtual reality" or "augmented reality" experiences, where digitally reproduced images or portions thereof are presented to a user in a way that they appear or are recognized as if they were real. Virtual reality, i.e., "VR" scenarios, typically involve the presentation of digital or virtual image information that is opaque to visual input from the actual real - world. Augmented reality, i.e., "AR" scenarios, typically involve the presentation of digital or virtual image information as an augmentation (i.e., transparency to other visual inputs from the actual real - world) to the visualization of the actual world around the user. Thus, AR scenarios involve the presentation of digital or virtual image information that is transparent to other visual inputs from the actual real - world. As used herein, the terms "extended reality" and "XR" are used to collectively refer to any of VR, AR, and / or MR. Further, the term "AR" means either or both of AR and MR.

[0004] VR and AR systems typically use head-mounted displays (or helmet-mounted displays, or smart glasses), which are at least loosely attached to the user's head and therefore move with the end user's head. When the end user's head movement is detected by the display system, the displayed data can be updated to take into account the change in head posture (i.e., the orientation and / or position of the user's head).

[0005] For example, if a user wearing a head-mounted display device views a virtual representation of a virtual object on the display device and walks around the area where the virtual object appears, the virtual object can be rendered for each viewpoint (corresponding to the position and / or orientation of the head-mounted display device), giving the user the perception of walking around an object that occupies real space. If the head-mounted display device is used to present multiple virtual objects at different depths, head pose measurements can be used to render a scene that matches the user's dynamically changing head pose, providing an enhanced sense of immersion. However, there is an unavoidable delay between the rendering of the scene and the display / projection of the rendered scene.

[0006] Head-mounted displays that enable AR (i.e., simultaneous viewing of virtual and real objects) can have several different types of configurations. In one such configuration, often called a “video see-through” display, a camera captures elements of a real-world scene, a computing system overlays virtual elements onto the captured real-world scene, and an opaque display presents the composite image to the eye. Another configuration, often called an “optical see-through” display, allows the end user to see directly the light from real objects in the environment by looking through a transparent (or semi-transparent) element within the display system. The transparent element, often called a “combiner,” overlays the light from the display onto the end user’s view of the real world. A camera may be attached to the head-mounted display device to capture images or videos of the scene the user is viewing.

[0007] XR systems also typically include a microphone array containing one or more microphones for sensing user utterances (i.e., voice) and audio (i.e., sound), such as ambient / environmental noises in the user's real-world surroundings, and generating audio signals corresponding to that audio. For example, users often want to communicate using various speech-based transmission protocols (e.g., group chat, IP-based speech communication platforms, etc.) in real-world environments where ambient noise levels can be relatively high. For example, a user might be in a factory floor, another commercial environment, near children playing, or near media such as a television or music playing in the background. Relatively high levels of ambient noise can hinder effective speech communication and / or negatively impact speech recognition performance by speech recognition systems, such as systems configured to recognize and process speech commands, because those communicating with the user may not be able to clearly hear the user's utterances through the XR system.

[0008] Therefore, there is still a need for improved means to effectively distinguish between user speech and ambient noise in audio signals generated by microphone arrays, for example, which can improve speech communication and speech recognition. [Overview of the project] [Means for solving the problem]

[0009] overview This disclosure relates to a system and method for reducing unwanted noise from audio signals generated by a microphone array. The system and method disclosed herein uses an innovative spatial filtering technique to filter far-field ambient noise from near-field audio, reduce far-field noise within the near-field audio signal, thereby improving the near-field audio and isolating it from far-field interference. As described herein, but not limited to this use, the system and method is useful for effectively filtering speech (i.e., near-field audio) from ambient noise (i.e., far-field noise). As described herein, such a system and method may be implemented on an XR system to improve and isolate user speech using a microphone array on a headset of an XR system. It should be understood that the system and method disclosed herein may be used to filter unwanted far-field audio noise from a target near-field audio signal in any suitable scenario, and may be used not only for isolating near-field speech signals from far-field noise, or only within an XR system. For example, systems and methods can be used to improve and isolate the near-field signal of a target audio (e.g., the sound produced by sliding machine parts) from ambient audio noise (e.g., noise from a factory floor). It may then be desirable to use the improved and isolated target audio signal for purposes such as machine testing, defect detection, or other applications.

[0010] The systems and methods disclosed herein may be implemented and performed on any suitable hardware system having a microphone array, a computer processor, and software (which may include firmware) configured to program the system to perform processes for reducing far-field noise in near-field audio signals.

[0011] Accordingly, one embodiment disclosed herein relates to a computer implementation method for reducing far-field noise in a near-field audio signal using a microphone array. The method includes acquiring a near-field audio signal from at least one primary microphone of the microphone array and acquiring one or more reference audio signals from one or more reference microphones of the microphone array. The far-field signal is determined from one or more reference audio signals. For example, at least one primary microphone may include a speech / voice microphone positioned near the user's mouth, while one or more reference microphones are positioned further away from the user's mouth, such as on the side of the user's head. The far-field signal is determined from one or more reference audio signals. For example, the far-field signal may be calculated by combining the reference audio signals from one or more reference microphones, such as by calculating the difference between one or more reference microphones.

[0012] Next, the near-field audio signal and the far-field signal are divided into multiple audio frequency bands. For example, in one embodiment, the near-field audio signal and the far-field audio signal may be divided using weighted overlap addition (WOLA) analysis. The number of audio frequency bands may be any appropriate number, such as 20 to 60 bands, 40 to 50 bands, more than 20 bands, more than 40 bands, or more than 50 bands.

[0013] Next, the near-field energy of the near-field audio signal and the far-field energy of the far-field audio signal within each audio frequency band are calculated. The resulting near-field and far-field energies within each audio frequency band are used to calculate the near-field to far-field energy ratio for each respective audio frequency band. In other words, the ratio of the near-field energy to the far-field energy within the same first audio frequency band is calculated, the ratio of the near-field energy to the far-field energy within the same second audio frequency band is calculated, and so on, for each audio frequency band.

[0014] The method then determines whether the energy ratio of each audio frequency band falls below a predetermined threshold for each audio frequency band. For example, in another embodiment, the predetermined threshold may be frequency band-dependent such that each audio frequency band has its own predetermined threshold, which may differ for each audio frequency band. Alternatively, the predetermined threshold may be the same for all audio frequency bands.

[0015] For each audio frequency band having an energy ratio below a predetermined threshold, the respective time-varying masking gain is calculated based on the energy ratio of that audio frequency band. For example, in one embodiment, the time-varying masking gain for each audio frequency band may be proportional to the amount by which its respective energy ratio falls below its respective predetermined threshold. In such a case, if the energy ratio is only slightly below (i.e., relatively close to) the predetermined threshold, the time-varying masking gain is small compared to the time-varying masking gain of frequency bands having energy ratios far below their respective thresholds and having larger time-varying masking gains (i.e., greater reductions). In another embodiment, the time-varying masking gains of all audio frequency bands can be used to generate a spatial filtering mask.

[0016] The time-varying masking gain for each audio frequency band is applied to the near-field audio signal of each respective audio frequency band to generate the filtered near-field audio signal for each audio frequency band. As a result, the filtered near-field audio signal is obtained in which the audio signals within the audio frequency band determined to contain significant far-field audio signals are reduced by their respective time-varying masking gains.

[0017] Finally, the filtered near-field audio signals from each audio frequency band are combined (i.e., combined) to produce a synthesized near-field audio signal. This synthesized near-field audio signal corresponds to filtered near-field audio in which the near-field audio is separated from far-field noise. In other words, background noise and interference are reduced, but the near-field sound remains prominent, thereby improving the perception of the near-field sound.

[0018] In another aspect of this method, for each audio frequency band having an energy ratio equal to or greater than a predetermined threshold, the near-field audio signal of such audio frequency band may be passed through without modification, thereby forming a filtered near-field audio signal of such audio frequency band. In other words, the audio signals in an audio frequency band determined to contain lower levels of far-field audio signals are the near-field audio signals that are passed through without modification as the filtered near-field audio signals of each such audio frequency band.

[0019] In yet another embodiment of this method, the step of synthesizing filtered near-field audio signals for each audio frequency band may include the step of performing weighted overlap sum (WOLA) synthesis of the filtered near-field audio signals for each audio frequency band to produce a synthesized near-field audio signal.

[0020] In yet another embodiment of this method, the step of splitting the near-field audio signal and the far-field audio signal may include performing weighted overlap summation (WOLA) analysis and subband splitting of the near-field audio signal and the far-field audio signal, respectively, to split the near-field audio signal and the far-field audio signal into multiple audio frequency bands.

[0021] In another embodiment, the method may further include the step of performing adaptive frequency smoothing of the time-varying gain for each audio frequency band based on an estimate of the signal-to-noise ratio of the near-field audio signal, wherein, for low signal-to-noise ratio signals, the weighting of adjacent subbands is heavier than that of adjacent subbands for high signal-to-noise ratio signals. This step may be performed after the step of calculating the time-varying masking gain for each audio frequency band using the energy ratio of each audio frequency band.

[0022] In another aspect of the method, the far-field audio signal may be generated by combining multiple audio signals from multiple reference microphones. For example, the microphone array may include two symmetrical reference microphones. The far-field audio signal may then be generated by taking the difference between the respective audio signals from the two or more reference microphones.

[0023] In yet another embodiment of this method, the time-varying masking gain for each audio frequency band may be linearly related to the amount by which their respective energy ratios fall below a predetermined threshold, thereby increasing as they fall below the predetermined threshold. In yet another embodiment, the relationship may be nonlinear rather than linear.

[0024] In yet another embodiment, the method may be specifically implemented to filter user voice utterances from background noise. Thus, the near-field audio signal is the user's near-field utterance signal, and the main microphone is positioned close to the user's mouth. The far-field signal is the background signal, and one or more reference microphones are positioned further from the user's mouth than the main microphone. Thus, the method improves the isolation of user utterances from background noise by isolating the user's utterances, attenuating ambient noise, and increasing the signal-to-noise ratio of the user's utterances.

[0025] Another embodiment disclosed herein relates to an XR system that implements one of the methods disclosed herein for reducing far-field noise in a near-field audio signal using a microphone array. In particular, noise reduction methods are implemented on the XR system to improve user speech and isolate it from background noise. In one embodiment, the XR system comprises an XR computer system having a computer processor, memory, a storage device, and executable software stored in the storage device and for programming the computer to perform operations that enable the XR system. The XR system also has a wearable support structure, such as a headset configured to be worn on the subject's head, and a display system for displaying a 3D virtual image (i.e., an XR image) within the user's XR field of view. In one embodiment, the display system is held by a support structure, such as an eyepiece on the headset. For example, the display may include a pair of optical projectors, a panel display, etc., and optical elements for projecting a 3D virtual image in the XR field of view into the user's eyes. The XR system is configured to present the user with a 3D virtual image in the XR field of view that simulates the precise location of virtual objects in a world coordinate system. If the XR system provides an AR and / or MR experience, the headset may also allow for some degree of transparency to the real world surrounding the user, so that the XR images enhance the visualization of the real world. The 3D virtual images may simulate the precise location of virtual objects in a world coordinate system.

[0026] The XR system also includes a microphone array. The microphone array includes a main microphone configured to be positioned proximate to the user's mouth to sense near-field audio and generate a near-field audio signal corresponding to the near-field audio sound. The microphone array also includes one or more reference microphones configured to be positioned further from the user's mouth than the main microphone to sense far-field audio sound and generate a far-field audio signal corresponding to the far-field audio sound. For example, in one aspect, the main microphone may be held on the front side of the XR headset (support structure) of the headset to position the main microphone in front of the user's face, and the one or more reference microphones may be held on the XR headset on both sides of the headset to position the one or more reference microphones towards the sides of the user's head.

[0027] The XR system also has an audio processor operably coupled to the main microphone and the one or more reference microphones to obtain the near-field audio signal and the far-field audio signal. The audio processor includes a microprocessor and software configured to program the audio processor to execute a process for reducing far-field noise from the microphone array, including any of the methods disclosed herein. For example, the process may include

[0028] a) obtaining a near-field audio signal from the main microphone of the microphone array;

[0029] b) obtaining one or more reference audio signals from each of the one or more reference microphones of the microphone array and determining an estimate of the far-field audio signal from the one or more reference audio signals;

[0030] c) dividing the near-field audio signal and the far-field signal into a plurality of audio frequency bands;

[0031] d) Calculate the near-field energy of the near-field audio signal within each audio frequency band and the far-field energy of the far-field audio signal within each audio frequency band,

[0032] e) Calculate the energy ratio of near-field energy to far-field energy for each audio frequency band,

[0033] f) Determining whether the energy ratio of each audio frequency band falls below a predetermined threshold,

[0034] g) For each audio frequency band having an energy ratio below a predetermined threshold, a time-varying masking gain for each audio frequency band is generated using the energy ratio of each audio frequency band.

[0035] h) Applying the time-varying masking gain of each audio frequency band to the near-field audio signal of each respective audio frequency band to generate a filtered near-field audio signal for each audio frequency band,

[0036] i) This includes synthesizing filtered near-field audio signals from each audio frequency band to produce a synthesized near-field audio signal.

[0037] In another embodiment of the XR system, the audio processor may be housed in a body pack configured to be detachably attached to the user's body. For example, the body pack may be a belt pack for attachment to the user's buttocks, or any other suitable structure.

[0038] In a further embodiment of the XR system, the audio processor may be integrated with the XR computer system, or it may be a separate system or module.

[0039] In a further embodiment, the XR system may be configured such that the process for reducing far-field noise includes any combination of one or more embodiments of method embodiments for reducing far-field noise in a near-field audio signal using a microphone array.

[0040] Another embodiment disclosed herein relates to a non-temporary computer-readable medium storing a set of instructions stored in memory and, when executed by the processor, to program the processor to perform a process for reducing far-field noise in a near-field audio signal using a microphone array. Thus, in one embodiment, the process is

[0041] a) Acquiring a near-field audio signal from the main microphone of the microphone array,

[0042] b) Obtaining one or more reference audio signals from one or more reference microphones of a microphone array, and determining a far-field audio signal (or an estimate thereof) from one or more reference audio signals,

[0043] c) Dividing near-field audio signals and far-field signals into multiple audio frequency bands,

[0044] d) Calculate the near-field energy of the near-field audio signal within each audio frequency band and the far-field energy of the far-field audio signal within each audio frequency band,

[0045] e) Calculate the energy ratio of near-field energy to far-field energy for each audio frequency band,

[0046] f) Determining whether the energy ratio of each audio frequency band falls below a predetermined threshold,

[0047] g) For each audio frequency band having an energy ratio below a predetermined threshold, a time-varying masking gain for each audio frequency band is generated using the energy ratio of each audio frequency band.

[0048] h) Applying the time-varying masking gain of each audio frequency band to the near-field audio signal of each respective audio frequency band to generate a filtered near-field audio signal for each audio frequency band,

[0049] i) This includes synthesizing filtered near-field audio signals from each audio frequency band to generate a synthesized near-field audio signal.

[0050] Additional and other purposes, features, and advantages of the technologies described herein are described in the detailed description, drawings, and claims. [Brief explanation of the drawing]

[0051] Brief explanation of the drawing U.S. Provisional Patent Application No. 63 / 453,022 included several drawings created in color. These color drawings have been converted to grayscale drawings for the purposes of this application and are intended to convey all the same information as the color drawings. The color drawings are expressly incorporated herein by reference for all purposes. Copies of the color drawings are available from the U.S. Patent and Trademark Office upon request and payment of the necessary fees.

[0052] The drawings illustrate the design and utility of preferred embodiments of the present invention, and similar elements are referenced by common reference numerals. To better understand how the above and other advantages and objectives of the present invention are obtained, a more specific description of the present invention, as briefly described above, is made by reference to specific embodiments of the present invention shown in the accompanying drawings. Understanding that these drawings only show typical embodiments of the present invention and should not be considered limiting its scope, the present invention is described and explained with additional specificities and details by using the accompanying drawings.

[0053] [Figure 1] Figure 1 shows photographs of three-dimensional augmented reality scenes that can be displayed to end users by an augmented reality system, according to several embodiments.

[0054] [Figure 2] Figure 2 shows a perspective view and a block diagram of an augmented reality system constructed according to one embodiment of the present invention.

[0055] [Figure 3] Figure 3 shows the augmented reality system of Figure 2, which has a body pack configured to be detachably attached to the user's body.

[0056] [Figure 4] Figure 4 is a flowchart showing the process for reducing far-field noise in a speech audio signal using the microphone array of the augmented reality system shown in Figure 2.

[0057] [Figure 5] Figure 5 shows an example of the frequency band energy ratio from a test case of speech captured by one or more primary microphones and ambient sound captured by two reference microphones.

[0058] [Figure 6]Figure 6 is a graph representing the spatial filter mask from Figure 5 after thresholding and warping.

[0059] [Figure 7] Figure 7 is a graph of the estimated signal-to-noise ratio of the near-field audio signal in the test case shown in Figure 5.

[0060] [Figure 8A] Figure 8A shows an example of an adaptive frequency smoothing curve graph for a low-signal-to-noise near-field audio signal.

[0061] [Figure 8B] Figure 8B shows an example of an adaptive frequency smoothing curve graph for a high-signal-to-noise near-field audio signal.

[0062] [Figure 9] Figure 9 shows the near-field audio signal before noise reduction.

[0063] [Figure 10] Figure 10 shows the near-field audio signal after noise reduction. [Modes for carrying out the invention]

[0064] Detailed explanation Next, various embodiments will be described in detail with reference to the drawings provided as exemplary examples of the disclosure, so that those skilled in the art can implement the disclosure. In particular, the following drawings and embodiments are not intended to limit the scope of the disclosure. Where certain elements of the disclosure can be partially or completely implemented using known components (or methods or processes), only the portion of such known components (or methods or processes) necessary for understanding the disclosure will be described, and detailed descriptions of other portions of such known components (or methods or processes) will be omitted so as not to obscure the disclosure. Furthermore, the various embodiments will include current and future known equivalents to the components mentioned herein as examples.

[0065] The following description discloses a technique for reducing far-field noise from near-field audio using a microphone array, such as that implemented in an exemplary XR system 200 (see Figures 2-3). As described herein, the disclosed system and method for reducing far-field noise from near-field audio using a microphone array utilizes a spatial filtering technique that filters far-field ambient noise from the near-field audio and reduces far-field noise within the near-field audio. This filtering technique improves the near-field audio by separating it from far-field interference. In non-limiting embodiments in an XR system, this filtering technique is used to effectively filter speech (i.e., near-field audio) from ambient noise (i.e., far-field noise). However, embodiments may also be useful for applications in other types of display systems (including other types of VR, AR, and / or MR systems), and therefore, embodiments should not be limited to the exemplary systems disclosed herein. Furthermore, systems and methods for reducing far-field noise from near-field audio using microphone arrays are not limited to use in XR systems, but may be used in any suitable application, device, or system for improving a target near-field audio signal and isolating it from far-field audio interference.

[0066] Referring to Figure 1, an AR scenario typically involves the presentation of virtual content (e.g., images and speech) corresponding to virtual objects relating to real-world objects. For example, Figure 1 illustrates an XR scenario (specifically, an AR scenario) having specific virtual reality objects and specific physical real-world objects seen by a user on the 3D display system of an XR system 200 (see Figure 2). As shown in Figure 1, an XR scene 100 is shown where a user of the XR system 200 sees a setting 102 that resembles a real-world physical park, featuring people, trees, buildings, and a real-world physical concrete platform 104 in the background. In addition to these items, the user 252 of the XR system 200 also perceives "seeing" a virtual robot figure 106 standing on the physical concrete platform 104, and a virtual cartoon-like avatar character 108 flying nearby that appears to be a personification of a bumblebee, although these virtual objects 106 and 108 do not even exist in the real world.

[0067] Figure 2 shows an XR system 200 according to several embodiments disclosed herein. The XR system 200 is a wearable system comprising a display-mounted headset 205 that is worn on the head 250 of a user 252. The XR headset 205 includes a wearable support structure comprising a frame structure 206 configured to be worn on the head 250 of the user 252, similar to eyeglass frames. The XR system 200 does not have to be a wearable system and may instead include a separate display, which may be a portable monitor, desktop monitor, tablet computer, smartphone, etc. However, a wearable system has the advantage of allowing the user to keep their hands free while using the XR system 200 and, in the case of a headset, provides an immersive XR experience.

[0068] Referring to Figure 2, in the illustrated embodiment, the display screen 204 is a partially transparent display screen that allows the end user 252 to see real objects in the surrounding environment and may display images of virtual objects. The frame structure 206 holds the partially transparent display screen 204 so that it is positioned in front of the end user 50's eyes 248, and in particular within the end user 252's field of view between the end user 252's eyes 248 and the surrounding environment.

[0069] The display subsystem 208 is designed to present two-dimensional content to the eyes 248 of an end user 252, presenting a light-based radiation pattern that can be comfortably perceived as an extension of physical reality with high levels of image quality and three-dimensional perception. The display subsystem 208 presents a series of frames at high frequency, providing the perception of a single coherent scene.

[0070] In an alternative embodiment, the XR system 200 may use one or more imagers (e.g., cameras) to capture images of the surrounding environment and convert them into video data, the video data can then be mixed with video data representing virtual objects, in which case the XR system 200 can display the mixed video data to the end user 252 on an opaque display surface.

[0071] Further details describing the display subsystem are provided in U.S. Provisional Patent Application No. 14 / 212,961, entitled "Display Subsystem and Method," and U.S. Provisional Patent Application No. 14 / 331,216, entitled "Planar Waveguide Apparatus With Diffraction Element(s) and Subsystem Employing Same," which are expressly incorporated herein by reference.

[0072] The XR system 200 further comprises one or more speakers 210 for presenting only sounds from virtual objects to the end user 252, while allowing the end user 252 to hear sounds directly from real objects. The speaker(s) 210 are held by a frame structure 206 so that the speaker(s) 210 are positioned adjacent to (inside or around) the end user 252's ear canal, e.g., earphones or headphones. The speaker(s) 210 may provide stereo / shapeable sound control. Although the speaker(s) 210 have been described as being positioned adjacent to the ear canal, other types of speakers not positioned adjacent to the ear canal can be used to deliver sound to the end user 252. For example, the speaker(s) 210 may be placed away from the ear canal, for example, using bone conduction technology. In an optional embodiment, a plurality of spatialized speakers 210 (e.g., four speakers) may be positioned around the head 250 of the end user 252 and configured to direct sound from the left, right, front, and rear of the head 250 to the left and right ears 254 of the end user 252. Further details regarding spatialized speakers that can be used in an augmented reality system are described in U.S. Provisional Patent Application No. 62 / 369,561, entitled "Mixed Reality System with Spatialized Audio," which is expressly incorporated herein by reference.

[0073] The XR system 200 further comprises a microphone array 220 configured to capture actual sounds, including the speech of user 252 and sounds arising from the surrounding environment of user 252, and convert them into audio signals output to an audio processor 246. The microphone array 220 includes a plurality of microphones 222, 224 held on the frame structure 206 of the headset 205. In the illustrated embodiment of Figure 2, the microphone array 220 includes four microphones, including two front / primary microphones 222a-222b and two side / reference microphones 224a-224b. The microphone array 220 may include any suitable number of microphones 222, 224 for effectively capturing sound and processing the sound for use by the XR system 200, as disclosed herein. The sound captured by microphone 222 can be used for communication by the user and / or can be cross-mixed with audio data from a virtual sound, in which case speaker 210 can transmit a sound representing the cross-mixed audio data to the end user 252. The front microphone 222 (also called the main microphone 222) is configured and positioned to capture the user's speech / voice and ambient sounds in front of the user 252. The side microphone 224 (also called the reference microphone 224) is configured and positioned to capture ambient sounds originating from the surrounding environment. Although called a “side” microphone, microphone 224 may be positioned anywhere relative to the user, including in front of and behind the user, but is positioned and configured to capture the smallest speech / voice sounds, or at least significantly quieter than the front microphone 222. Thus, the side microphone 224 is typically positioned further from the user’s mouth than the front microphone 222 and / or directed to capture sounds originating from the surrounding environment rather than from the user’s mouth.

[0074] The microphone array 220 includes a first front microphone 222a (also referred to as the first main microphone 222a) located on the lower right side (from the user's perspective) of the front side of the headset 205 to effectively capture the user's utterances / voice by positioning the first main microphone 222a close to the user's mouth. The microphone array 220 also includes a second front microphone 222b located on the left front side of the headset 205. Thus, the second microphone 222b is also located relatively close to the user's mouth so that it can be used as the second main microphone 222b in the far-field noise reduction process described herein. The microphone array 220 also includes a right-side microphone 224a (also referred to as the first reference microphone 224a), which is located on the right temple piece of the frame structure 206 to position the right-side microphone 224a on the right side of the user's head. A left microphone 224b (also referred to as a second reference microphone 224b) is positioned on the left temple piece of the frame structure 206 so as to position the left microphone 224b on the left side of the user's head. The XR system 200 may also include an additional front microphone 222 (i.e., a primary microphone 222) and / or additional side microphones 224 (i.e., reference microphones 224 as appropriate, configured and positioned).

[0075] Each of the microphones 222 and 224 is configured to sense sound and output an audio signal corresponding to the sensed sound. Microphones 222 and 224 may be digital microphones that output a digital audio signal, or analog microphones that output an analog audio signal. Thus, if additional primary microphones 222 and / or reference microphones 224 are present, the first primary microphone 222a generates and outputs a first near-field audio signal (which may also be called the "first primary audio signal"), the second primary microphone 222b generates and outputs a second near-field audio signal (which may also be called the "second first primary audio signal"), the first reference microphone 224a generates and outputs a first reference audio signal, the second reference microphone 224b generates and outputs a second reference audio signal, and so on.

[0076] In the illustrated embodiment, the XR system 200 may, if necessary, render and present spatialized audio corresponding to virtual objects having known virtual positions and orientations in real physical three-dimensional (3D) space, thereby influencing the clarity or presence of sound using a spatialized audio system that makes sound appear to the end user 252 as if it originates from the virtual positions of real objects. The XR system 200 tracks the position of the end user 252 to render the spatialized audio more accurately so that the audio associated with various virtual objects appears as if it originates from their virtual positions. Furthermore, the XR system 200 may track the head pose of the end user 252 to render the spatialized audio more accurately so that the directional audio associated with various virtual objects appears as if it propagates in the appropriate virtual direction to each virtual object (for example, coming out of the mouth of a virtual character rather than coming out from behind the head of the virtual character). Furthermore, the XR system 200 may take into account other real physical and virtual objects when rendering spatialized audio, so that audio associated with various virtual objects appears to be appropriately reflected from, or blocked from, or obstructed by, real physical and virtual objects.

[0077] For this purpose, the XR system 200 may further include a head / object tracking subsystem 240 for tracking the position and orientation of the end user 252's head 250 relative to a virtual three-dimensional scene, as well as the position and orientation of real objects relative to the end user 252's head 250, if necessary. For example, the head / object tracking subsystem 240 may include one or more sensors configured to collect head pose data (position and orientation) of the end user 252, and a processor (not shown) configured to determine the head pose of the end user 252 in a known coordinate system based on the head pose data collected by the sensors. The sensors may include one or more image capture devices (such as visible and infrared cameras), inertial measurement units (including accelerometers and gyroscopes), compasses, microphones, GPS units, or wireless devices. In the illustrated embodiment, the sensors include a forward-facing camera 230. When mounted on the head in this manner, the forward-facing camera 230 is particularly well-suited for capturing information indicating the distance and angular position of the end user 252's head 250 relative to the environment in which the end user 250 is located (i.e., the direction the head is facing). The head orientation may be detected in any direction (e.g., up / down, left, right relative to the end user 252's reference frame). The forward-facing camera 230 is also configured to acquire video data of real objects in the surrounding environment in order to facilitate the video recording capabilities of the XR system 200. Cameras may also be provided to track real objects in the surrounding environment. The frame structure 206 may be designed so that cameras can be mounted on the front and back of the frame structure 106. In this way, the array of cameras may surround the end user 252's head 250 to cover all directions of the relevant objects.

[0078] The XR system 200 may also include, as necessary, one or more rear-facing cameras 232 and corresponding processors that track the end user 252's eyes 248, particularly the direction and / or distance to which the end user 252 is focused. The rear-facing cameras 232 may track the angular position (direction in which one or both eyes are looking), blinking, and depth of focus (by detecting eye convergence) of the end user 252's eyes 248. Further details discussing the eye-tracking device are provided in U.S. Provisional Patent Application No. 14 / 212,961 entitled "Display Subsystem and Method," U.S. Patent Application No. 14 / 726,429 entitled "Methods and Subsystem for Creating Focal Planes in Virtual and Augmented Reality," and U.S. Patent Application No. 14 / 205,126 entitled "Subsystem and Method for Augmented and Virtual Reality," which are expressly incorporated herein by reference.

[0079] The XR system 200 further comprises a three-dimensional database 242 configured to store virtual three-dimensional scenes, including virtual objects (both the content data of the virtual objects and the absolute metadata associated with these virtual objects, e.g., the absolute position and orientation of these virtual objects in a 3D scene) and virtual objects (both the content data of the virtual objects and the absolute metadata associated with these virtual objects, e.g., the volume and absolute position and orientation of these virtual objects in a 3D scene, as well as spatial acoustic effects surrounding each virtual object, including any virtual or real objects in the vicinity of the virtual source, room dimensions, wall / floor materials, etc.).

[0080] The augmented reality system 200 further comprises a control subsystem 202, which further records video data arising from virtual and real objects appearing in the field of view, as well as audio data captured by the microphone array 220. The XR system 200 may also record metadata associated with the video and audio data so that the synchronized video and audio can be accurately re-rendered during playback.

[0081] For this purpose, the control subsystem 202 is configured to include a video processor 244, which retrieves absolute metadata associated with video content and virtual objects from a three-dimensional database 242, retrieves head pose data of the end user 252 from the head / object tracking subsystem 240 (which can be used to localize the absolute metadata of the video to the head 250 of the end user 252), renders the video from there, and then transmits it to the display subsystem 208 to convert it into an image that is intermixed with images arising from real objects in the surrounding environment within the end user 252's field of view. The video processor 244 is also configured to retrieve video data from a forward-facing camera 230 arising from real objects in the surrounding environment, which is then recorded along with video data arising from virtual objects, as will be further described below.

[0082] The audio processor 246 is configured to retrieve audio content and metadata associated with virtual objects from the three-dimensional database 242, retrieve head pose data of the end user 252 from the head / object tracking subsystem 240 (which can be used to localize the absolute metadata of the audio to the head 250 of the end user 252), and render spatialized audio from there, which is then transmitted to the speaker 210 to be converted into spatialized sound that is intermixed with sounds originating from real objects in the surrounding environment.

[0083] The audio processor 246 is also configured to acquire audio data captured by microphones 222, 224 of the microphone array 220, including sounds from the user's mouth (speech / voice) and the surrounding environment. This audio data, along with spatialized audio data from selected virtual objects, along with any resulting metadata localized to the end user 252's head 250 (e.g., position, orientation, and volume data) for each virtual object, as well as global metadata (e.g., volume data set globally by the XR system 200 or the end user 252), may be recorded as described further below. The audio processor 246 and the XR system 200 are also configured to enable voice communication between the user 252 and other parties, and / or to use voice sounds for speech activation functions and commands.

[0084] As described herein, the audio processor 246 is further configured to process the speech audio signal (i.e., near-field audio signal) captured by the main microphone 222 and to filter far-field noise from the speech audio signal using the microphone array 220 to produce a filtered speech audio signal (i.e., filtered near-field audio signal) having improved separation of speech sounds with reduced background noise. The audio processor 246 includes a noise filtering software program (which may be in the form of software and / or firmware) that programs the microprocessor to perform a process for reducing far-field noise in the speech audio signal (i.e., near-field audio signal) captured by the main microphone 222 using the microphone array 220.

[0085] Referring here to Figure 4, one embodiment of process 400 for reducing far-field noise in speech audio signals (i.e., near-field audio signals) captured by the main microphones 222 is shown using a microphone array 220. In step 402, the audio processor acquires the respective near-field audio signals from one or more of the main microphones 222. Process 400 may utilize only the first near-field audio signal from only the first main microphone 222a, or the process may utilize both the first and second near-field audio signals from both the first and second main microphones 222a and the second main microphone 222b, and combine them to form a near-field audio signal. By using near-field audio signals from multiple main microphones 222, the near-field audio signals can be combined to provide some directivity to the near-field audio signal, such as a dipole pattern.

[0086] In step 404, the audio processor 246 acquires each reference audio signal from one or more of the reference microphones 224. In process 400, the audio processor 246 may acquire and use only one of the reference audio signals from either the first reference microphone 224a or the second reference microphone 224b, or both the first reference audio signal and the second reference audio signal from both the first reference microphone 224a and the second reference microphone 224b.

[0087] In step 406, the audio processor 246 determines the far-field audio signal using one or more reference audio signals. Typically, the far-field is defined as being at least 1 meter away from the near-field, so process 400 determines the far-field audio signal by making an estimation. When only a single reference audio signal is used, the far-field audio signal is simply the single reference audio signal. When multiple reference audio signals are used, the audio processor 246 combines the reference audio signals. As an example, in the embodiment of Figure 2, the far-field audio signal may be the difference between a first reference audio signal and a second reference audio signal, where this difference assumes a dipole pattern in the array of the first reference microphone 224a and the second reference microphone 224b. Again, by using reference audio signals from multiple reference microphones 224, the reference audio signals can be combined to provide some directivity to the far-field audio signal, such as a dipole audio signal. In embodiments utilizing three or more reference microphones 224, the reference microphones 224 may be configured to utilize beamforming techniques such as delay and sum, frost, or minimum dispersion-free response (MVDR) to determine the far-field audio signal.

[0088] In step 408, the audio processor 246 divides the near-field audio signal and the far-field signal into multiple audio frequency bands. This may be achieved by any suitable process, including the use of WOLA analysis. In the example described below with reference to Figures 5-10, the near-field audio signal and the far-field signal are divided into 42 frequency bands in the range of 0-25000 Hz. The number of audio frequency bands may be any suitable number, such as 20-60 bands, 40-50 bands, more than 20 bands, more than 40 bands, or more than 50 bands.

[0089] In step 410, the audio processor 246 calculates the near-field energy of the near-field audio signal and the far-field energy of the far-field audio signal within each audio frequency band. This provides the respective near-field energy value and far-field energy value for each audio frequency band.

[0090] In step 412, the resulting near-field and far-field energies within each audio frequency band are used to calculate the near-field to far-field energy ratio for each respective audio frequency band. In other words, the ratio of near-field energy to far-field energy within the first audio frequency band is calculated, the ratio of near-field energy to far-field energy within the second audio frequency band is calculated, and so on for each audio frequency band.

[0091] In step 414, the audio processor 246 determines whether the energy ratio of each audio frequency band falls below a predetermined threshold for each audio frequency band. The predetermined threshold sets a limit on the ratio of near-field energy to far-field energy for each audio frequency band. This step of the process determines audio frequency bands with a relatively high energy ratio of near-field energy to far-field energy such that the near-field audio signal in that frequency band is considered to be a prominent speech sound (i.e., a target, desired sound), and audio frequency bands with a relatively low energy ratio of near-field energy such that the near-field audio signal in that frequency band is considered to be mostly background noise or interference and not speech. As an example, the predetermined threshold may be 1 or close to 1. The predetermined threshold may be frequency band-dependent such that each audio frequency band has its own predetermined threshold which may differ for each audio frequency band. In an alternative embodiment, the predetermined threshold may be the same for all audio frequency bands. The predetermined threshold may be tuned by empirical experimentation to determine a predetermined threshold that provides the desired or best far-field filtering effect. Figure 5 shows examples of energy ratios across 42 frequency bands from test cases of speech captured by one or more primary microphones and ambient sound captured by two reference microphones. The graphs in Figure 5 represent the spatial filtering masks before thresholding and warping, as will be discussed later.

[0092] In step 416, for each audio frequency band having an energy ratio below a predetermined threshold, the audio processor 246 calculates a time-varying masking gain for each audio frequency band based on its energy ratio. This is also known as warping of the spatial filtering mask. In one embodiment, the time-varying masking gain for each audio frequency band may be proportional to the amount by which its energy ratio falls below its respective predetermined threshold. Thus, if the energy ratio is only slightly below (i.e., relatively close to) the predetermined threshold, the time-varying masking gain is small compared to the time-varying masking gain for frequency bands with energy ratios far below their respective thresholds and with larger time-varying masking gains (i.e., greater reductions). In another embodiment, the time-varying masking gains for all audio frequency bands can be used to generate a spatial filtering mask.

[0093] The following equation represents one embodiment of the thresholding step 414 and the warping step 416.

[0094]

number

[0095]

number

[0096]

number

[0097] In step 418, for each audio frequency band having an energy ratio equal to or greater than a predetermined threshold, the near-field audio signal of such audio frequency band passes through unchanged as the filtered near-field audio signal of that respective audio frequency band.

[0098] The graph in Figure 6 shows the spatial filter mask from Figure 5 after thresholding and warping. In Figure 6, it can be seen that the audio frequency band with low energy ratios in Figure 5 has been significantly reduced.

[0099] Step 420 is an optional step and is included in this embodiment of Method 400. In step 420, the audio processor 246 performs adaptive frequency smoothing of time-varying masking gain for each audio frequency band based on an estimate of the signal-to-noise ratio of the near-field audio signal. Figure 7 shows a graph of the estimated signal-to-noise ratio of the near-field audio signal in the test case of Figure 5. In this embodiment, for low signal-to-noise ratio signals, the weighting of adjacent subbands is greater than that of adjacent subbands for high signal-to-noise ratio signals. Adaptive frequency smoothing smooths the frequencies of the spatial masking gain, thereby reducing musical noise artifacts. During periods of low signal-to-noise ratio, higher degrees of smoothing are utilized, as shown in Figures 8A and 8B. Figure 8A shows an example graph of adaptive frequency smoothing curves for a low signal-to-noise near-field audio signal. Each curve in the graph represents the smoothing curve for its respective frequency band. Figure 8B shows an example graph of adaptive frequency smoothing curves for a high signal-to-noise near-field audio signal. Here again, each curve in the graph represents the smoothing curve for its respective frequency band. Comparing the adaptive frequency curves in Figures 8A and 8B, it can be seen that for near-field audio signals with a low signal-to-noise ratio, adjacent frequency bands are weighted more strongly than for near-field audio signals with a high signal-to-noise ratio.

[0100] In step 422, the audio processor 246 applies a time-varying masking gain for each audio frequency band to the near-field audio signal of each respective audio frequency band in order to generate filtered near-field audio signals for each audio frequency band. As a result, filtered near-field audio signals are produced in which audio signals within the audio frequency bands determined to contain significant far-field audio signals are reduced by their respective time-varying masking gains. Figures 9 and 10 illustrate the filtering of near-field audio signals in the test case of Figure 5.

[0101] In step 424, the audio processor 246 synthesizes (i.e., combines) the filtered near-field audio signals for each audio frequency band, thereby producing a synthesized near-field audio signal. The synthesized near-field audio signal corresponds to filtered near-field audio in which the near-field audio is separated from far-field noise. As shown in the examples in Figures 9 and 10, background noise and interference are reduced, while the near-field sound remains prominent, thereby improving the perception of near-field speech. Figure 9 shows the near-field audio signal before noise reduction. Figure 10 shows the near-field audio signal after noise reduction using Method 400. In Figure 9, it can be seen that broadband interference or background noise is present at approximately 75 dBA, while Figure 10 shows that this broadband interference is significantly reduced at 75 dBA, and the user speech signal is extracted with minimal spectral artifacts.

[0102] Returning to Figure 2, the XR system 200 further comprises a memory 260 and a recorder 262 configured to store video and audio that can be accessed for playback in the memory 260.

[0103] The control subsystem that performs the functions of the video processor 1244, audio processor 246, recorder 262, and audio / video player (not shown) can take any of many different forms and may include several controllers, for example, one or more microcontrollers, microprocessors or central processing units (CPUs), digital signal processors, graphics processing units (GPUs), other integrated circuit controllers such as application-specific integrated circuits (ASICs), programmable gate arrays (PGAs), such as field PGAs (FPGAs), and / or programmable logic controllers (PLUs).

[0104] The functions of the video processor 244, the audio processor 246, and the recorder 262 may each be performed by a single integrated device, at least some of the functions of the video processor 244, the audio processor 246, and the recorder 262 may be combined into a single integrated device, or the functions of each of the video processor 244, the audio processor 246, and the recorder 262 may be distributed among several devices. For example, the video processor 244 may comprise a graphics processing unit (GPU) that retrieves video data of virtual objects from a three-dimensional database 242 and renders composite video frames therefrom, and a central processing unit (CPU) that retrieves video frames of real objects from a forward-facing camera 230. Similarly, the audio processor 246 may comprise a digital signal processor (DSP) that processes audio data acquired from a microphone subsystem and a microphone array 220, and a CPU that processes the audio data. The recording function of the recorder 262 may also be performed by the CPU.

[0105] Furthermore, various processing components of the XR system 200 may be physically contained within a distributed subsystem. The XR system 200 may also further include remote processing modules 300 and remote data repositories 302 operably coupled to the local processing and data modules 308 by wired leads or wireless connections 304, 306, etc., so that these remote modules 300, 302 are operably coupled to each other and available as resources for the local processing and data modules 308. For example, as shown in Figure 3, the XR system 200 may include local processing and data modules 308 operably coupled by wired leads or wireless connections, etc., to components held by a headset 205 worn on the head 250 of an end user 252 (e.g., projection subsystem of display subsystem 208, microphone array 220, audio processor 246, speaker 210, and cameras 230, 232). The local processing and data module 308 may include any one or more of the components and subsystems included in the control subsystem 201 in Figure 2, such as the display subsystem 208 and the audio processor 246. The GPU of the video processor 244, and any one or more of the CPUs of the video processor 244 and / or the audio processor 246 may be included in the remote processing module 308 to improve portability and limit the power consumption of the local processing and data module 308, which may be battery-powered; however, in alternative embodiments, these components or parts thereof may be included in the local processing and data module 308. As shown in Figure 3, the local processing and data module 308 may be mounted in a body pack configured to be detachably attached to the user's body, such as a belt pack having a belt coupling for attachment to the user's buttocks 310.Alternatively, the local processing and data module 308 may be fixedly attached to the frame structure 206, fixedly attached to a helmet or other headpiece worn by the user 252, embedded in headphones, or detachably attached to the user's torso.

[0106] The local processing and data module 308 may comprise a power-efficient processor or controller, as well as digital memory such as flash memory, both of which may be used to assist in processing, caching, and storing data captured from sensors, and / or retrieved and / or processed using the remote processing module 300 and / or remote data repository 302, and optionally proceed to the display subsystem 208 after such processing or retrieval. The remote processing module 300 may comprise one or more relatively powerful processors or controllers configured to analyze and process data and / or image information. The remote data repository 302 may comprise a relatively large digital data storage facility that may be available via the internet or other networking configuration in a “cloud” resource configuration. In an alternative embodiment, all data is stored and all calculations are performed within the local processing and data module 308, allowing for fully autonomous use from any remote module.

[0107] The connections 304, 306 between the various components described above may include one or more wired interfaces or ports for providing wired or optical communication, or one or more wireless interfaces or ports via RF, microwave, and IR, etc., for providing wireless communication. In some implementations, all communication may be wired, while in other implementations, all communication may be wireless, except for the optical fiber used in the display subsystem 208. In yet another embodiment, the choice between wired and wireless communication may differ from that shown in Figure 3. Therefore, any particular choice between wired or wireless communication should not be considered limiting.

[0108] The present invention has been described in the above specification with reference to specific embodiments thereof. However, it will be apparent that various modifications and changes may be made without departing from the broader spirit and scope of the invention. For example, the process flow described above is described with reference to a specific order of process actions. However, many of the described order of process actions may be changed without affecting the scope or operation of the invention. Therefore, this specification and the drawings should be considered illustrative rather than restrictive.

Claims

1. A method for reducing far-field noise in a near-field audio signal using a microphone array, a) Acquiring a near-field audio signal from the main microphone of the microphone array, b) Obtaining one or more reference audio signals from one or more reference microphones of the microphone array, and determining a far-field audio signal from the one or more reference audio signals, c) Dividing the near-field audio signal and the far-field signal into multiple audio frequency bands, d) Calculate the near-field energy of the near-field audio signal within each audio frequency band and the far-field energy of the far-field audio signal within each audio frequency band, e) Calculate the energy ratio of the near-field energy to the far-field energy for each audio frequency band, f) Determining whether the energy ratio of each audio frequency band falls below a predetermined threshold, g) For each audio frequency band having an energy ratio below the predetermined threshold, a time-varying masking gain for each audio frequency band is generated using the energy ratio of each audio frequency band, h) Applying the time-varying masking gain for each audio frequency band to the respective audio frequency band's near-field audio signal to generate a filtered near-field audio signal for each audio frequency band, i) To synthesize the filtered near-field audio signals of each audio frequency band to produce a synthesized near-field audio signal. Methods that include...

2. The method according to claim 1, further comprising, after step f), allowing each audio frequency band having an energy ratio greater than or equal to a predetermined threshold to pass through as the filtered near-field audio signal of such audio frequency band without modification.

3. The method according to any one of claims 1 to 2, step i) comprising generating the synthesized near-field audio signal by weighted overlap sum (WOLA) synthesis of the filtered near-field audio signals for each audio frequency band.

4. Step c) is The method according to any one of claims 1 to 3, comprising performing weighted overlap sum (WOLA) analysis and subband partitioning on the near-field audio signal and the far-field audio signal, respectively, to divide the near-field audio signal and the far-field audio signal into the plurality of audio frequency bands.

5. The method according to any one of claims 1 to 4, further comprising, after step g), performing adaptive frequency smoothing of the time-varying gain of each audio frequency band based on an estimate of the signal-to-noise ratio of the near-field audio signal, wherein, in the case of a low signal-to-noise ratio signal, the weighting of adjacent subbands is weighted more heavily than the weighting of adjacent subbands of a high signal-to-noise ratio signal.

6. The method according to any one of claims 1 to 5, wherein the far-field audio signal is generated by combining multiple reference audio signals from multiple reference microphones.

7. The method according to claim 6, wherein the far-field audio signal is generated by taking the difference between the respective audio signals of the two or more reference microphones.

8. The method according to any one of claims 1 to 7, wherein the time-varying masking gain for each audio frequency band is linearly related to the amount by which the respective energy ratio falls below a predetermined threshold, and thereby the amount of the time-varying masking gain increases as it falls below the predetermined threshold.

9. The method according to any one of claims 1 to 7, wherein the time-varying masking gain for each audio frequency band is nonlinearly related to the amount by which the respective energy ratio falls below a predetermined threshold, and thereby the amount of the time-varying masking gain increases as it falls below the predetermined threshold.

10. The aforementioned near-field audio signal is the user's near-field speech signal, and the main microphone is positioned close to the user's mouth. The aforementioned far-field signal is a background signal, and the one or more reference microphones are positioned further away from the user's mouth than the main microphone. The method according to any one of claims 1 to 9, wherein the method separates the user's speech by attenuating the far-field ambient noise on each audio frequency band.

11. The method according to any one of claims 1 to 10, wherein the predetermined threshold depends on the frequency band, and thereafter each audio frequency band has a predetermined threshold, and the predetermined threshold differs for at least two of the audio frequency bands.

11. An augmented reality (XR) system, An XR headset configured to be worn on the head of a subject, wherein the XR headset is A support structure configured to be attached to the head of the subject, A display system held by the support structure, wherein the display system is for displaying virtual images generated by the video processor of the XR system, and The XR headset is equipped with, It is a microphone array, A main microphone configured to be positioned close to the user's mouth, the main microphone being configured to sense near-field audio sound and generate a near-field audio signal corresponding to the near-field audio sound, One or more reference microphones configured to be positioned further away from the user's mouth than the main microphone in order to sense far-field audio sound, wherein the reference microphones are configured to generate far-field audio signals corresponding to the far-field audio sound. A microphone array, including An audio processor, wherein the audio processor is operably coupled to the main microphone and one or more reference microphones to acquire the near-field audio signal and the far-field audio signal, and the audio processor is configured to perform a process for reducing far-field noise from the microphone array, the process being: a) Acquiring the near-field audio signal from the main microphone of the microphone array, b) Obtaining one or more reference audio signals from one or more reference microphones of the microphone array, and determining a far-field audio signal from the one or more reference audio signals, c) Dividing the near-field audio signal and the far-field signal into multiple audio frequency bands, d) Calculate the near-field energy of the near-field audio signal within each audio frequency band and the far-field energy of the far-field audio signal within each audio frequency band, e) Calculate the energy ratio of the near-field energy to the far-field energy for each audio frequency band, f) Determining whether the energy ratio of each audio frequency band falls below a predetermined threshold, g) For each audio frequency band having an energy ratio below the predetermined threshold, a time-varying masking gain for each audio frequency band is generated using the energy ratio of each audio frequency band, h) Applying the time-varying masking gain for each audio frequency band to the respective audio frequency band's near-field audio signal to generate a filtered near-field audio signal for each audio frequency band, i) To synthesize the filtered near-field audio signals of each audio frequency band to generate a synthesized near-field audio signal. Includes an audio processor and An augmented reality (XR) system equipped with [features / technology].

12. The aforementioned support structure is an XR headset, The main microphone is held on the front side of the XR headset, positioning the main microphone in front of the user's face. The system according to claim 11, wherein the one or more reference microphones are held on both sides of the headset to position the one or more reference microphones toward the sides of the user's head.

13. The system according to any one of claims 11 to 12, wherein the audio processor is housed in a body pack configured to be detachably attached to the user's body.

14. The aforementioned process, The system according to any one of claims 11 to 13, further comprising, after step f), allowing each audio frequency band having an energy ratio greater than or equal to a predetermined threshold to pass through as the filtered near-field audio signal of such audio frequency band without modification.

15. The system according to any one of claims 11 to 14, step i) comprising generating the synthesized near-field audio signal by WOLA synthesis of the filtered near-field audio signals for each audio frequency band.

16. Step c) is The system according to any one of claims 11 to 15, comprising performing weighted overlap summation analysis and subband partitioning on the near-field audio signal and the far-field audio signal, respectively, to divide the near-field audio signal and the far-field audio signal into the plurality of audio frequency bands.

17. The system according to any one of claims 11 to 16, further comprising, after step g), performing adaptive frequency smoothing of the time-varying gain of each audio frequency band based on an estimate of the signal-to-noise ratio of the near-field audio signal, wherein, in the case of a low signal-to-noise ratio signal, the weighting of adjacent subbands is weighted more heavily than the weighting of adjacent subbands for a high signal-to-noise ratio signal.

18. The system according to any one of claims 11 to 17, wherein the far-field audio signal is generated by combining multiple reference audio signals from multiple reference microphones.

19. The system according to claim 18, wherein the far-field audio signal is generated by taking the difference between the respective audio signals of the two or more reference microphones.

20. The system according to any one of claims 11 to 19, wherein the time-varying masking gain for each audio frequency band is linearly related to the amount by which the respective energy ratio falls below a predetermined threshold, and thereby the amount of the time-varying masking gain increases as it falls below the predetermined threshold.

21. The system according to any one of claims 11 to 20, wherein the time-varying masking gain for each audio frequency band is nonlinearly related to the amount by which the respective energy ratio falls below a predetermined threshold, and thereby the amount of the time-varying masking gain increases as it falls below the predetermined threshold.

22. The aforementioned near-field audio signal is the user's near-field speech signal, and the main microphone is positioned close to the user's mouth. The aforementioned far-field signal is a background signal, and the one or more reference microphones are positioned further away from the user's mouth than the main microphone. The system according to any one of claims 11 to 21, wherein the method separates the user's speech by attenuating the far-field ambient noise on each audio frequency band.

23. A non-temporary computer-readable medium storing software instructions, wherein the software instructions are executable by a computer processor of an audio processor to perform a process for reducing far-field noise using a microphone array, and the process is a) Acquiring a near-field audio signal from the main microphone of the microphone array, b) Obtaining one or more reference audio signals from one or more reference microphones of the microphone array, and determining a far-field audio signal from the one or more reference audio signals, c) Dividing the near-field audio signal and the far-field signal into multiple audio frequency bands, d) Calculate the near-field energy of the near-field audio signal within each audio frequency band and the far-field energy of the far-field audio signal within each audio frequency band, e) Calculate the energy ratio of the near-field energy to the far-field energy for each audio frequency band, f) Determining whether the energy ratio of each audio frequency band falls below a predetermined threshold, g) For each audio frequency band having an energy ratio below the predetermined threshold, a time-varying masking gain for each audio frequency band is generated using the energy ratio of each audio frequency band, h) Applying the time-varying masking gain for each audio frequency band to the respective audio frequency band's near-field audio signal to generate a filtered near-field audio signal for each audio frequency band, i) To synthesize the filtered near-field audio signals of each audio frequency band to produce a synthesized near-field audio signal. Computer-readable media, including [specific examples of computer-readable media].

24. The computer-readable medium according to claim 23, further comprising, after step f), allowing each audio frequency band having an energy ratio greater than or equal to the predetermined threshold to pass through as the filtered near-field audio signal of such audio frequency band without modification.

25. Step i) is to produce the synthesized near-field audio signal by WOLA synthesis of the filtered near-field audio signals for each audio frequency band, the computer-readable medium according to any one of claims 23 to 24.

26. Step c) is A computer-readable medium according to any one of claims 23 to 25, comprising performing weighted overlap summation analysis and subband partitioning on the near-field audio signal and the far-field audio signal, respectively, to divide the near-field audio signal and the far-field audio signal into the plurality of audio frequency bands.

27. The computer-readable medium according to any one of claims 23 to 26, further comprising, after step g), performing adaptive frequency smoothing of the time-varying gain of each audio frequency band based on an estimate of the signal-to-noise ratio of the near-field audio signal, wherein, in the case of a low signal-to-noise ratio signal, the weighting of adjacent subbands is weighted more heavily than the weighting of adjacent subbands of a high signal-to-noise ratio signal.

28. The computer-readable medium according to any one of claims 23 to 27, wherein the far-field audio signal is generated by combining multiple reference audio signals from multiple reference microphones.

29. The computer-readable medium according to claim 28, wherein the far-field audio signal is generated by taking the difference between the respective audio signals of the two or more reference microphones.

30. The computer-readable medium according to any one of claims 23 to 29, wherein the time-varying masking gain for each audio frequency band is linearly related to the amount by which the respective energy ratio falls below a predetermined threshold, and thereby the amount of the time-varying masking gain increases as it falls below the predetermined threshold.

31. The computer-readable medium according to any one of claims 23 to 30, wherein the time-varying masking gain for each audio frequency band is nonlinearly related to the amount by which the respective energy ratio falls below a predetermined threshold, and thereby the amount of the time-varying masking gain increases as it falls below the predetermined threshold.

32. The aforementioned near-field audio signal is the user's near-field speech signal, and the main microphone is positioned close to the user's mouth. The aforementioned far-field signal is a background signal, and the one or more microphones are positioned further away from the user's mouth than the main microphone. The computer-readable medium according to any one of claims 23 to 31, wherein the method separates the user's speech by attenuating the far-field ambient noise on each audio frequency band.