Head-mounted device and method for improving signal-to-noise ratio of signals acquired using a head-mounted device

By setting up a microphone array on a head-mounted device and utilizing dynamic beamforming technology, the problem of environmental noise interference was solved, and the signal-to-noise ratio and voice signal quality were improved.

CN116805998BActive Publication Date: 2026-04-24SNAP INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SNAP INC
Filing Date
2020-06-26
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing head-mounted devices are often affected by environmental noise when collecting voice signals, resulting in a decrease in the signal-to-noise ratio and affecting the quality of voice communication.

Method used

Dynamic beamforming technology is employed, which involves placing microphone arrays on the left and right temples of the head-mounted device and using a beamformer controller to dynamically adjust the direction of the microphone arrays, thereby enhancing the voice signal and attenuating the noise signal.

Benefits of technology

It effectively improves the signal-to-noise ratio, enhances the clarity of voice signals, reduces the impact of environmental noise, and improves the quality of voice communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116805998B_ABST
    Figure CN116805998B_ABST
Patent Text Reader

Abstract

A method of performing dynamic beamforming to reduce signal-to-noise ratio in a signal acquired by a headset begins with a microphone generating an acoustic signal. The microphone is coupled to a first temple of the device and a second temple of the device. First and second beamformers generate first and second beamformer signals, respectively. A noise suppressor attenuates noise content from the first and second beamformer signals. The noise content from the first beamformer signal is the acoustic signal not collocated in the second beamformer signal, and the noise content from the second beamformer signal is the acoustic signal not collocated in the first beamformer signal. A speech enhancer generates a clean signal including speech content from the first and second noise-suppressed signals. The speech content is the acoustic signal collocated in the first and second beamformer signals.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of patent application filed on June 26, 2020, with application number 202080047279.2 and invention title "Dynamic beamforming for improving the signal-to-noise ratio of signals acquired using a head-mounted device".

[0002] Cross-reference to related applications

[0003] This claims priority to U.S. Provisional Patent Application No. 62 / 868,715, filed on June 28, 2019, the contents of which are incorporated herein by reference in their entirety. Background Technology

[0004] Currently, many consumer electronic devices are adapted to receive voice via microphone ports or headphones. While a typical example is portable telecommunications devices (mobile phones), with the advent of Voice over IP (VoIP), desktop computers, laptops, tablets, and wearable devices can also be used to perform voice communications.

[0005] When using these electronic devices, users also have the option to receive their voice using speaker mode or wired or wireless headphones. However, a common complaint about these hands-free operating modes is that the voice picked up by the microphone port or headphones includes ambient noise, such as wind noise, secondary speakers in the background, or other background noise. This ambient noise often makes the user's voice difficult to understand, thus reducing the quality of voice communication. Attached Figure Description

[0006] In the accompanying drawings (which are not necessarily drawn to scale), the same numbers may describe similar components in different figures. Similar numbers with different letter suffixes may represent different instances of similar components. Some embodiments are shown in the drawings by way of example rather than limitation, in which:

[0007] Figure 1 A perspective view of a headset for generating binaural audio according to an example embodiment is shown.

[0008] Figure 2 The following is illustrated according to an example embodiment: Figure 1 Bottom view of the head-mounted device.

[0009] Figure 3 This illustrates the implementation of dynamic beamforming according to an example embodiment to improve the use of... Figure 1 A block diagram of the system for the signal-to-noise ratio of signals acquired by the head-mounted device.

[0010] Figure 4 It is based on various aspects of this disclosure for improving the use of [the technology / method / etc.] Figure 1An exemplary flowchart of the dynamic beamforming process for the signal-to-noise ratio of signals acquired by a head-mounted device.

[0011] Figure 5 This is a block diagram illustrating a representative software architecture that can be used in conjunction with the various hardware architectures described herein.

[0012] Figure 6 This is a block diagram illustrating components of a machine capable of reading instructions from a machine-readable medium (e.g., a machine-readable storage medium) and performing any one or more methods discussed herein, according to some exemplary embodiments.

[0013] Figure 7 This is a high-level functional block diagram of an example head-mounted device that couples mobile devices and server systems via various network communications. Detailed Implementation

[0014] The following description includes systems, methods, techniques, instruction sequences, and computer program products embodying illustrative embodiments of the present disclosure. In the following description, numerous specific details are set forth for purposes of explanation in order to provide an understanding of various embodiments of the subject matter of the invention. However, it will be apparent to those skilled in the art that embodiments of the subject matter of the invention can be practiced without these specific details. Generally, well-known examples of instructions, protocols, structures, and techniques need not be shown in detail.

[0015] To improve the signal-to-noise ratio of signals acquired by current electronic mobile devices, some embodiments of this disclosure relate to head-mounted devices that perform dynamic beamforming and audio processing on beamformer signals to enhance speech content while attenuating noise content. Specifically, the head-mounted device may be a pair of glasses comprising left and right temples coupled to both sides of a frame. Each temple is coupled to a microphone housing comprising two microphones. The microphones on each temple form a microphone array. A beamformer can steer the microphone arrays on each side of the frame toward the user's face or mouth. While a directional beamformer pointing toward the user's mouth will acquire acoustic signals from the user's mouth, it also acquires acoustic content passing through the user's mouth in the same direction. Therefore, some embodiments utilize microphone arrays located on either side of the user's face or mouth to determine what might be speech content in the beamformer signal. For example, when two microphone arrays are pointed toward the user's mouth from opposite directions, the content collapsing between or between the microphone arrays can be considered speech content.

[0016] In one embodiment, the system further includes a beamformer controller that orients the beamformers in different directions. The beamformer controller can dynamically change the orientation of the beamformers relative to each other. Knowing the orientation and configuration of each beamformer, the system can perform audio processing to attenuate acoustic content that is not expected to be received. The system can also attenuate acoustic content that is not between the beamformer beams or that is not juxtaposed.

[0017] In one embodiment, by means of a microphone array on opposite sides of the head-mounted device, the system is able to cycle through various beamforming configurations (e.g., dynamic beamforming) and acquire raw acoustic data as real-time audio processing. This allows the system to maximize the attenuation of noise content (e.g., ambient noise, secondary speakers, etc.), enhance speech content, and thereby reduce the signal-to-noise ratio in the resulting clean signal.

[0018] Figure 1 A perspective view of a head-mounted device 100 according to an example embodiment for performing dynamic beamforming to improve the signal-to-noise ratio of signals acquired using the head-mounted device. Figure 2 The example shown is from Figure 1 Bottom view of the head-mounted device 100. Figure 1 and Figure 2 In this context, the head-mounted device 100 is a pair of glasses. In some embodiments, the head-mounted device 100 may be sunglasses or goggles. Some embodiments may include one or more wearable devices, such as a pendant with an integrated camera that is integrated with, communicates with, or is coupled to the head-mounted device 100 or a client device. Any desired wearable device may be used in conjunction with embodiments of this disclosure, such as a watch, headphones, wristband, earplugs, clothing (such as a hat or jacket with integrated electronics), clip-on electronics, or any other wearable device. It should be understood that, although not shown, one or more parts of the system included in the head-mounted device may be included in a client device (e.g., [missing information]) that can be used in conjunction with the head-mounted device 100. Figure 6 In machine 800). For example, such as Figure 3 One or more of the components shown may be included in the head-mounted device 100 and / or the client device.

[0019] As used herein, the term "client device" can refer to any machine that is connected to a communication network interface to obtain resources from one or more server systems or other client devices. A client device can be, but is not limited to, a mobile phone, desktop computer, laptop computer, portable digital assistant (PDA), smartphone, tablet computer, ultrabook, netbook, laptop computer, multiprocessor system, microprocessor-based or programmable consumer electronics, game console, set-top box, or any other communication device that a user can use to access the network.

[0020] exist Figure 1 and Figure 2 In this design, the head-mounted device 100 is a pair of eyeglasses, including a frame 103. The frame 103 includes eye wires (or frames) coupled to two temples (or eyeglass temples) via hinges and / or end pieces. The eye wires of the frame 103 carry or hold a pair of lenses 104_1, 104_2. The frame 103 includes a first (e.g., right) side coupled to a first temple and a second (e.g., left) side coupled to a second temple. The first side of the frame 103 is opposite to the second side.

[0021] The device 100 further includes a camera module comprising camera lenses 102_1 and 102_2 and at least one image sensor. The camera lenses can be perspective or non-perspective lenses. Non-perspective lenses can be, for example, fisheye lenses, wide-angle lenses, omnidirectional lenses, etc. The image sensor acquires digital video through the camera lenses. The images can also be still image frames or video comprising multiple still image frames. The camera module can be coupled to frame 103. Figure 1 and Figure 2 As shown, frame 103 is coupled to camera lenses 102_1 and 102_2 such that the camera lenses face forward. Camera lenses 102_1 and 102_2 may be perpendicular to lenses 104_1 and 104_2. The camera module may include dual front-facing cameras separated by the width of frame 103 or the width of the user's head of device 100.

[0022] exist Figure 1 and Figure 2 In this device, two temples (or eyeglass temples) are coupled to microphone housings 101_1 and 101_2, respectively. The first temple and the second temple are coupled to opposite sides of the frame 103 of the head-mounted device 100. The first temple is coupled to the first microphone housing 101_1 and the second temple is coupled to the second microphone housing 101_2. The microphone housings 101_1 and 101_2 can be coupled to the temples between the position of the frame 103 and the tip of the eyeglass temple. When the user wears the device 100, the microphone housings 101_1 and 101_2 can be located on either side of the user's eyeglass temple.

[0023] like Figure 2 As shown, microphone housings 101_1 and 101_2 enclose a plurality of microphones 110_1 to 110_N (N>1). Microphones 110_1 to 110_N are air interface pickup devices that convert sound into electrical signals. More specifically, microphones 110_1 to 110_N are transducers that convert sound pressure into electrical signals (e.g., acoustic signals). Microphones 110_1 to 110_N can be digital or analog microelectromechanical systems (MEMS) microphones. The acoustic signals generated by microphones 110_1 to 110_N can be pulse density modulation (PDM) signals.

[0024] exist Figure 2 In the first microphone housing 101_1, microphones 110_3 and 110_4 are enclosed, while microphones 110_1 and 110_2 are enclosed in the second microphone housing 101_2. In the first microphone housing 101_1, the first front microphone 110_3 and the first rear microphone 110_4 are separated by a predetermined distance d1 and can form a first-order differential microphone array. In the second microphone housing 101_2, the second front microphone 110_1 and the second rear microphone 110_2 are also separated by a predetermined distance d2 and can form a first-order differential microphone array. The predetermined distances d1 and d2 can be the same or different. The predetermined distances d1 and d2 can be set based on the Nyquist frequency. Content above the Nyquist frequency of the beamformer is irretrievable, especially speech. The Nyquist frequency is determined by the following equation:

[0025]

[0026] In this equation, c is the speed of sound, and d is the distance between the microphones. Using this equation, in one embodiment, predetermined distances d1 and d2 can be set to any value of d that results in a frequency higher than 6 kHz (the cutoff frequency for wideband speech).

[0027] In one embodiment, the first front microphone 110_3 and the first rear microphone 110_4 form a first microphone array, and the second front microphone 110_1 and the second rear microphone 110_2 form a second microphone array.

[0028] In one embodiment, both the first and second microphone arrays are end-fire arrays. An end-fire array comprises multiple microphones arranged in the desired direction of sound propagation. As described above, this configuration is called a differential array when the first front microphone in the array (e.g., the first microphone to which sound arrives as it propagates along the axis) is added to an inverted and delayed signal from the first rear microphone. Beamformers can be used to steer the first and second microphone arrays to create cardioid or subcardioid pickup patterns. In this embodiment, the sound at the rear of the microphone arrays is significantly attenuated.

[0029] In another embodiment, both the first and second microphone arrays are broadside arrays. A broadside microphone array is an array in which one row of microphones is arranged perpendicular to a preferred direction of sound waves. Broadside microphone arrays attenuate sound originating from the sides of the array. In one embodiment, the first microphone array is a broadside array and the second microphone array is an end-fire array. Alternatively, the first microphone array is an end-fire array, and the second microphone array is a broadside array.

[0030] Although Figure 1 In this system 100, there are four microphones 110_1 to 110_4, but the number of microphones can vary. In some embodiments, microphone housings 101_1 and 101_2 may include at least two microphones and may form a microphone array. Each of the microphone housings 101_1 and 101_2 may also include a battery.

[0031] refer to Figure 2 Each of the microphone housings 101_1 and 101_2 includes a front port and a rear port. The front port of the first microphone housing 101_1 is coupled to microphone 110_3 (e.g., a first front microphone), and the rear port of the first microphone housing 101_1 is coupled to microphone 110_4 (e.g., a first rear microphone). In one embodiment, microphone 110_3 (e.g., the first front microphone) and microphone 110_4 (e.g., the first rear microphone) are located on the same plane (e.g., a first plane). The front port of the second microphone housing 101_2 is coupled to microphone 110_1 (e.g., a second front microphone), and the rear port of the second microphone housing 101_2 is coupled to microphone 110_2 (e.g., a second rear microphone). In one embodiment, microphone 110_1 (e.g., the second front microphone) and microphone 110_2 (e.g., the second rear microphone) are located on the same plane (e.g., a second plane). In one embodiment, microphones 101_1 to 101_4 may be further movable toward the temple tips on the temples of the device 100 (e.g., the back of the device 100).

[0032] Figure 3This illustrates the implementation of dynamic beamforming according to an example embodiment to improve the use of... Figure 1 A block diagram of the system 300 showing the signal-to-noise ratio of the signals acquired by the head-mounted device 100. In some embodiments, one or more portions of the system 300 may be included in the head-mounted device 100 or may be included in a client device (e.g., [missing information]) that can be used in conjunction with the head-mounted device 100. Figure 6 In the machine 800).

[0033] System 300 includes microphones 110_1 to 110_N, beamformers 301_1 and 301_2, a noise suppressor 302, a voice enhancer 303, and a beamformer controller 304. A first front microphone 110_3 and a first rear microphone 110_4, enclosed in a first microphone housing 101_1, form a first microphone array. Similarly, a second front microphone 110_1 and a second rear microphone 110_2, enclosed in a second microphone housing 101_2, form a second microphone array. The first and second microphone arrays can be first-order differential microphone arrays. The first and second microphone arrays can also be, respectively, wide-side arrays, end-fire arrays, or a combination of a wide-side array and an end-fire array. Microphones 110_1 to 110_4 can be analog or digital MEMS microphones. The acoustic signals generated by microphones 110_1 to 110_4 can be pulse density modulation (PDM) signals.

[0034] In one embodiment, the first beamformer 301_1 and the second beamformer 301_2, which have directional steering characteristics, are differential beamformers that allow for a flat frequency response other than the Nyquist frequency. Beamformers 301_1 and 301_2 may use the transfer function of a first-order differential microphone array. In one embodiment, beamformers 301_1 and 301_2 are fixed beamformers that include a subcardioid or cardioid fixed beam pattern.

[0035] like Figure 3 As shown, the first beamformer 301_1 receives acoustic signals from the first front microphone 110_3 and the first rear microphone 110_4 and generates a first beamformer signal based on the received acoustic signals. The second beamformer 301_2 receives acoustic signals from the second front microphone 110_1 and the second rear microphone 110_2 and generates a second beamformer signal based on the received acoustic signals.

[0036] exist Figure 3In this embodiment, beamformer controller 304 steers the first beamformer 301_1 in a first direction and the second beamformer 301_2 in a second direction. When the user wears the head-mounted device, the first and second directions can be aligned with the user's mouth. Since the first beamformer 301_1 and the second beamformer 301_2 receive acoustic signals from opposite sides of the user's head, in this embodiment, the first and second directions are directed from opposite directions toward the user's mouth.

[0037] The beamformer controller 304 can also dynamically change the first and second directions. In one embodiment, the first beamformer 301_1 and the second beamformer 301_2 can be oriented in the first and second directions (which are different directions relative to each other). By dynamically changing the direction, the beamformer controller 304 can cycle between multiple different configurations of the beamformers 301_1 and 301_2. Furthermore, by knowing the configurations of the beamformers 301_1 and 301_2, the location of the voice content can be predicted. For example, the voice content may be between microphone arrays, between beamformer signals, or juxtaposed within beamformer signals.

[0038] Noise suppressor 302 attenuates noise content from the first beamformer signal and the second beamformer signal. Noise suppressor 302 may be a dual-channel noise suppressor, generating a first noise-suppressed signal and a second noise-suppressed signal. In one embodiment, noise suppressor 302 may implement a noise suppression algorithm. The noise content may be, for example, ambient noise, secondary loudspeakers, etc. In one embodiment, system 300 receives acoustic signals from opposite sides of the user's head using first beamformer 301_1 and second beamformer 301_2, such that a first direction (e.g., of the first beamformer 301_1) and a second direction (e.g., of the second beamformer 301_2) point towards the user's mouth from opposite directions. Given that the first and second directions point towards the user from opposite directions, the noise content from the first beamformer signal is an acoustic signal not co-located with the second beamformer signal, and the noise content from the second beamformer signal is an acoustic signal not co-located with the first beamformer signal. Since beamformers 301_1 and 301_2 can point towards the user's mouth from opposite sides or pass through the user's mouth in that direction, the non-overlapping (or non-juxtaposed) areas between the beamformers contain noise content.

[0039] Furthermore, the speech enhancer 303 generates a clean signal that includes speech content from the first noise suppression signal and the second noise suppression signal. For example, when both the first and second beamformer signals are directed from opposite sides of the user's head toward the user's mouth, the overlap (or juxtaposition) between the beamformer beams contains the speech content. In this embodiment, the speech content is an acoustic signal juxtaposed between the first and second beamformer signals. In one embodiment, the speech enhancer 303 may implement a speech enhancement algorithm.

[0040] Figure 4 It is based on various aspects of this disclosure for improving the use of [the technology / method / etc.] Figure 1 An exemplary flowchart of the dynamic beamforming process for the signal-to-noise ratio of signals acquired by a head-mounted device.

[0041] Although flowcharts can describe operations as a sequential process, many operations can be executed in parallel or concurrently. Furthermore, the order of operations can be rearranged. A process terminates when its operations are completed. A process can correspond to a method, procedure, etc. The steps of a method can be executed all or part of the time, can be combined with some or all of the steps in other methods, and can be executed by any number of different systems (such as...). Figure 1 and / or Figure 6 The process 400 can also be executed by the system described herein. Figure 1 The processor in the head-mounted device 100 or the processor included in Figure 6 The processor in the client device 800 executes the commands.

[0042] Process 400 begins with operation 401, where microphones 110_1 to 110_4 generate acoustic signals. Microphones 110_1 to 110_4 may be MEMS microphones that convert sound pressure into electrical signals (e.g., acoustic signals). A first front microphone 110_3 and a first rear microphone 110_4 are housed within a housing of a first microphone 101_1 coupled to a first temple of the head-mounted device 100. In one embodiment, the first front microphone 110_3 and the first rear microphone 110_4 form a first microphone array. The first microphone array may be a first-order differential array.

[0043] The second front microphone 110_1 and the second rear microphone 110_2 are housed within a second microphone housing 101_2 coupled to a second temple of the head-mounted device 100. In one embodiment, the second front microphone 110_1 and the second rear microphone 110_2 form a second microphone array. The second microphone array may be a first-order differential microphone array. The first temple and the second temple are coupled to opposite sides of the frame 103 of the head-mounted device 100.

[0044] At operation 402, the first beamformer 301_1 generates a first beamformer signal based on acoustic signals from the first front microphone 110_3 and the first rear microphone 110_4. At operation 403, the second beamformer 301_2 generates a second beamformer signal based on acoustic signals from the second front microphone 110_1 and the second rear microphone 110_2. In one embodiment, the first beamformer 301_1 and the second beamformer 301_2 are fixed beamformers. Fixed beamformers may include subcardioid or cardioid fixed beam patterns.

[0045] In one embodiment, the beamformer controller 304 steers towards a first beamformer in a first direction and towards a second beamformer in a second direction. When a user wears the head-mounted device, the first and second directions can be aligned with the user's mouth. The beamformer controller can dynamically change the first and second directions.

[0046] At operation 404, noise suppressor 302 attenuates noise content from the first beamformer signal and the second beamformer signal to generate a first noise-suppressed signal and a second noise-suppressed signal. The noise content from the first beamformer signal may be an acoustic signal not co-located in the second beamformer signal, and the noise content from the second beamformer signal may also be an acoustic signal not co-located in the first beamformer signal.

[0047] At operation 405, the speech enhancer 303 generates a clean signal comprising speech content from the first noise suppression signal and the second noise suppression signal. The speech content is an acoustic signal juxtaposed between the first beamformer signal and the second beamformer signal.

[0048] Figure 5 This is a block diagram illustrating an exemplary software architecture 706, which can be used in conjunction with various hardware architectures described herein. Figure 5 This is a non-limiting example of a software architecture, and it should be understood that many other architectures can be implemented to facilitate the functionality described herein. Software architecture 706 can be implemented in, for example... Figure 6 The machine 800 executes on hardware including a processor 804, memory 814, and I / O components 818, etc. A representative hardware layer 752 is shown and can represent, for example... Figure 6 The machine 800. A representative hardware layer 752 includes a processing unit 754 having associated executable instructions 704. The executable instructions 704 represent executable instructions of the software architecture 706, including implementations of the methods, components, etc., described herein. Hardware layer 752 also includes a memory or storage module / memory device 756, which also has executable instructions 704. Hardware layer 752 may also include other hardware 758.

[0049] As used herein, the term "component" means a device, physical entity, or logic having boundaries defined by functional or subroutine calls, branch points, application programming interfaces (APIs), or other technologies that provide specific processing or control functionality, such as partitions or modularization. Components can be combined with other components via their interfaces to perform machine processes. Components can be encapsulated functional hardware units designed for use with other components and as part of a program that typically performs related functions.

[0050] Components can constitute software components (e.g., code embodied on a machine-readable medium) or hardware components. A “hardware component” is a tangible unit capable of performing certain operations and can be configured or arranged in a physical manner. In various example embodiments, one or more computer systems (e.g., standalone computer systems, client computer systems, or server computer systems) or one or more hardware components of a computer system (e.g., a processor or a group of processors) can be configured by software (e.g., an application or an application portion) to operate to perform certain operations as described herein. Hardware components can also be implemented mechanically, electronically, or any suitable combination thereof. For example, a hardware component may include dedicated circuitry or logic permanently configured to perform certain operations.

[0051] Hardware components can be dedicated processors, such as field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs). Hardware components can also include programmable logic or circuitry temporarily configured by software to perform certain operations. For example, a hardware component may include software executed by a general-purpose processor or other programmable processor. After being configured by this software, the hardware component becomes a specific machine (or a specific component of a machine), uniquely tailored to perform the configured functions and no longer a general-purpose processor. It should be understood that cost and time considerations can drive the decision to implement hardware components mechanically in dedicated and permanently configured circuitry or in temporarily configured circuitry (e.g., software-configured).

[0052] A processor can be or includes any circuitry or virtual circuitry (physical circuitry simulated by logic executed on an actual processor) that manipulates data values ​​according to control signals (e.g., "commands", "opcodes", "machine codes", etc.) and generates corresponding output signals applied to operate the machine. For example, a processor can be a Central Processing Unit (CPU), a Reduced Instruction Set Computing (RISC) processor, a Complex Instruction Set Computing (CISC) processor, a Graphics Processing Unit (GPU), a Digital Signal Processor (DSP), an Application-Specific Integrated Circuit (ASIC), a Radio Frequency Integrated Circuit (RFIC), or any combination thereof. A processor can further be a multi-core processor having two or more independent processors (sometimes called "cores") capable of executing instructions simultaneously.

[0053] Therefore, the phrase "hardware component" (or "hardware-implemented component") should be understood to include tangible entities, i.e., entities with physical construction, permanent configuration (e.g., hardwired), or temporary configuration (e.g., programming), that operate or perform certain operations described herein in a certain way. Consider embodiments where hardware components are temporarily configured (e.g., programmed), it is not necessary to configure or instantiate each hardware component at any given time. For example, in cases where hardware components include a general-purpose processor configured by software as a dedicated processor, the general-purpose processor can be configured at different times as correspondingly different dedicated processors (e.g., including different hardware components). The software accordingly configures one or more specific processors, for example, constituting a specific hardware component at one time and different hardware components at different times. Hardware components can provide information to and receive information from other hardware components. Therefore, the described hardware components can be considered communicatively coupled. In cases where multiple hardware components exist simultaneously, communication can be achieved through signal transmission (e.g., via appropriate circuitry and buses) between two or more hardware components. In embodiments where multiple hardware components are configured or instantiated at different times, communication between the hardware components can be achieved, for example, by storing and retrieving information in a memory structure accessible to the multiple hardware components.

[0054] For example, a hardware component may perform an operation and store the output of that operation in a storage device communicatively coupled to it. Another hardware component may then later access the storage device to retrieve and process the stored output. The hardware component may also initiate communication with an input or output device and may operate on a resource (e.g., a collection of information). Various operations of the example methods described herein may be performed, at least in part, by one or more processors configured, either temporarily (e.g., by software) or permanently, to perform the relevant operations. Whether temporarily or permanently configured, the processor may constitute a processor-implemented component for performing one or more operations or functions described herein. As used herein, "processor-implemented component" refers to a hardware component implemented using one or more processors. Similarly, the methods described herein may be implemented, at least in part, by processors, where a particular processor or processors are examples of hardware. For example, at least some operations of the method may be performed by one or more processors or processor-implemented components.

[0055] Furthermore, one or more processors may operate to support the performance of related operations in a “cloud computing” environment or as “Software as a Service” (SaaS). For example, at least some operations in the operation may be performed by a group of computers (as an example of a machine including processors), and these operations may be accessed via a network (e.g., the Internet) and via one or more appropriate interfaces (e.g., application programming interfaces (APIs)). The performance of some operations in the operation may be distributed among processors, residing not only within a single machine but also deployed across multiple machines. In some exemplary embodiments, the processor or processor-implemented components may reside in a single geographic location (e.g., within a home environment, office environment, or server cluster). In other exemplary embodiments, the processor or processor-implemented components may be distributed across multiple geographic locations.

[0056] exist Figure 5 In the exemplary architecture, software architecture 706 can be conceptualized as a stack of layers, where each layer provides specific functionality. For example, software architecture 706 may include layers such as operating system 702, libraries 720, applications 716, and a presentation layer 714. Operationally, applications 716 or other components within a layer can call application programming interface (API) calls 708 through the software stack and receive messages 712 in response to API calls 708. The layers shown are representative in nature, and not all software architectures have all layers. For example, some mobile or dedicated operating systems may not provide a framework / middleware 718, while other operating systems may provide such a layer. Other software architectures may include additional layers or different layers.

[0057] Operating system 702 manages hardware resources and provides public services. Operating system 702 may include, for example, a kernel 722, services 724, and drivers 726. Kernel 722 can act as an abstraction layer between hardware and other software layers. For example, kernel 722 may be responsible for memory management, processor management (e.g., scheduling), component management, networking, security settings, etc. Services 724 can provide other public services to other software layers. Drivers 726 are responsible for controlling or interfacing with the underlying hardware. For example, depending on the hardware configuration, drivers 726 may include display drivers, camera drivers, Bluetooth® drivers, flash memory drivers, serial communication drivers (e.g., Universal Serial Bus (USB) drivers), Wi-Fi® drivers, audio drivers, power management drivers, etc.

[0058] Library 720 provides common infrastructure used by application 916 or other components or layers. Library 720 provides functionality that allows other software components to perform tasks in a way that is easier than directly interfaced with the underlying operating system functions 702 (e.g., kernel 722, services 724, and / or drivers 726). Library 720 may include system libraries 744 (e.g., the C standard library), which provide functions such as memory allocation, string manipulation, and mathematical functions. Additionally, library 720 may include API libraries 746, such as media libraries (e.g., libraries for supporting the rendering and manipulation of various media formats such as MPREG4, H.264, MP3, AAC, AMR, JPG, and PNG), graphics libraries (e.g., OpenGL frameworks for rendering 2D and 3D content on a display), database libraries (e.g., SQLite providing various relational database functions), web libraries (e.g., WebKit providing web browsing functionality), and so on. Library 720 may also include various other libraries 748 to provide numerous other APIs to application 716 and other software components / modules.

[0059] The framework / middleware 718 (sometimes also called middleware) provides a higher level of common infrastructure that can be used by applications 716 or other software components / modules. For example, the framework / middleware 718 can provide various graphical user interface (GUI) functions, advanced resource management, advanced location services, etc. The framework / middleware 718 can provide a wide range of other APIs that can be used by applications 716 or other software components / modules, some of which may be specific to a particular operating system 702 or platform.

[0060] Application 716 includes either built-in application 738 or third-party application 740. Examples of representative built-in applications 738 may include, but are not limited to, contact applications, browser applications, book reader applications, location applications, media applications, messaging applications, or game applications. Third-party applications 740 may include applications developed by entities other than the vendor of a specific platform using a software development kit (SDK), and may be mobile software running on a mobile operating system. Third-party applications 740 may invoke API calls 708 provided by the mobile operating system (such as operating system 702) to facilitate the functions described herein.

[0061] Application 716 can use built-in operating system functions (e.g., kernel 722, service 724, or driver 726), libraries 720, and frameworks / middleware 718 to create user interfaces to interact with the system's users. Alternatively or additionally, in some systems, user interaction may occur through a presentation layer (such as presentation layer 714). In these systems, application / component "logic" can be decoupled from the aspects of the application / component that interact with the user.

[0062] Figure 6 This is a block diagram illustrating components (also referred to herein as "modules") of a machine 800 according to some exemplary embodiments, the machine 800 being capable of reading instructions from a machine-readable medium (e.g., a machine-readable storage medium) and executing any one or more methods discussed herein. Specifically, Figure 6 A graphical representation of a machine 800 in an example form employing a computer system is shown, within which instructions 810 (e.g., software, programs, applications, applets, application software, or other executable code) can be executed to cause the machine 800 to perform any or more of the methods discussed herein. Thus, instructions 810 can be used to implement the modules or components described herein. Instructions 810 transform a general, unprogrammed machine 800 into a specific machine 800 programmed to perform the described and illustrated functions in the described manner. In alternative embodiments, machine 800 operates as a standalone device or can be coupled (e.g., networked) to other machines. In a networked deployment, machine 800 can operate as a server machine or client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. Machine 800 may include, but is not limited to, server computers, client computers, personal computers (PCs), tablet computers, laptop computers, netbooks, set-top boxes (STBs), personal digital assistants (PDAs), entertainment media systems, cellular phones, smartphones, mobile devices, wearable devices (e.g., smartwatches), smart home devices (e.g., smart appliances), other smart devices, web devices, network routers, network switches, network bridges, or any machine capable of sequentially or otherwise executing instructions 810 to be performed by the specified machine 800. Furthermore, although only a single machine 800 is shown, the term "machine" should also be considered as including a collection of machines that individually or jointly execute instructions 1010 to implement any or more of the methods discussed herein.

[0063] Machine 800 may include a processor 804, a memory / storage device 806, and I / O components 818 that can be configured to communicate with each other, such as via bus 802. Memory / storage device 806 may include memory 814, such as main memory or other memory storage devices, and storage cells 816, both accessible by processor 804, such as via bus 802. Storage cells 816 and memory 814 store instructions 810 embodying any one or more methods or functions described herein. During execution of machine 800, instructions 810 may also reside wholly or partially within memory 814, storage cells 816, at least one of processor 804 (e.g., within the processor's cache), or any suitable combination thereof. Thus, the memory of memory 814, storage cells 816, and processor 804 are examples of machine-readable media.

[0064] As used herein, the terms “machine-readable medium,” “computer-readable medium,” etc., refer to any component, device, or other tangible medium capable of temporarily or permanently storing instructions and data. Examples of such media may include, but are not limited to, random access memory (RAM), read-only memory (ROM), buffer memory, flash memory, optical media, magnetic media, cache memory, other types of storage devices (e.g., erasable programmable read-only memory (EEPROM)), and / or any suitable combination thereof. The term “machine-readable medium” should be considered to include a single medium or multiple media capable of storing instructions (e.g., a centralized or distributed database, or associated caches and servers). The term “machine-readable medium” should also be considered to include any medium or combination of media capable of storing instructions (e.g., code) that are executable by a machine, such that when executed by one or more processors of the machine, the instructions cause the machine to perform any or more methods described herein. Thus, “machine-readable medium” can refer to a single storage device or device, as well as a “cloud-based” storage system or storage network that includes multiple storage devices or devices. The term “machine-readable medium” excludes signals themselves.

[0065] I / O component 818 may include a wide variety of components to provide a user interface for receiving input, providing output, generating output, sending information, exchanging information, acquiring measurements, etc. The specific I / O component 818 included in the user interface of a particular machine 800 will depend on the type of machine. For example, a portable machine such as a mobile phone may include a touch input device or other such input mechanism, while a headless server machine may not include such a touch input device. It should be understood that I / O component 818 may include... Figure 6Many other components are not shown. The grouping of I / O components 818 according to function is merely for the sake of simplicity in the following discussion, and the grouping is by no means limiting. In various exemplary embodiments, I / O components 818 may include output components 826 and input components 828. Output components 826 may include visual components (e.g., displays such as plasma display panels (PDPs), light-emitting diode (LED) displays, liquid crystal displays (LCDs), projectors, or cathode ray tubes (CRTs)), acoustic components (e.g., speakers), haptic components (e.g., vibration motors, resistive mechanisms), other signal generators, etc. Input components 828 may include alphanumeric input components (e.g., keyboards, touchscreens configured to receive alphanumeric input, photoelectric keyboards, or other alphanumeric input components), point-based input components (e.g., mice, touchpads, trackballs, joysticks, motion sensors, or other pointing instruments), haptic input components (e.g., physical buttons, touchscreens providing position and / or force for touch or touch gestures, or other haptic input components), audio input components (e.g., microphones), etc. The input component 828 may also include one or more image acquisition devices, such as a digital camera for generating digital images or videos.

[0066] In a further exemplary embodiment, I / O component 818 may include biometric component 830, motion component 834, environmental component 836, or positioning component 838, as well as a host of other components. One or more of these components (or portions thereof) may be collectively referred to herein as “sensor components” or “sensors” for collecting various data relating to machine 800, the environment of machine 800, the user of machine 800, or combinations thereof.

[0067] For example, biometric component 830 may include components for detecting expressions (e.g., hand gestures, facial expressions, vocal expressions, body posture, or eye tracking), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweat, or brain waves), and identifying people (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or EEG-based recognition). Motion component 834 may include accelerometer components (e.g., accelerometers), gravity sensor components, velocity sensor components (e.g., speedometers), rotation sensor components (e.g., gyroscopes), etc. Environmental component 836 may include, for example, illumination sensor components (e.g., photometers), temperature sensor components (e.g., one or more thermometers for detecting ambient temperature), humidity sensor components, pressure sensor components (e.g., barometers), acoustic sensor components (e.g., one or more microphones for detecting background noise), proximity sensor components (e.g., infrared sensors for detecting nearby objects), gas sensor components (e.g., gas detection sensors for detecting hazardous gas concentrations or measuring pollutants in the atmosphere for safety purposes), or other components that may provide indications, measurements, or signals corresponding to the surrounding physical environment. Positioning component 838 may include positioning sensor components (e.g., a Global Positioning System (GPS) receiver component), altitude sensor components (e.g., an altimeter or barometer that can detect from which altitude air pressure can be derived), orientation sensor components (e.g., a magnetometer), etc. For example, the position sensor components can provide location information associated with system 800, such as the GPS coordinates of system 800 or information about the current location of system 1000 (e.g., the name of a restaurant or other business).

[0068] Various technologies can be used to implement communication. I / O component 818 may include communication component 840, operable to couple machine 800 to network 832 or device 820 via coupler 822 and coupler 824, respectively. For example, communication component 840 may include a network interface component or other suitable device to interface with network 832. In further examples, communication component 840 may include wired communication components, wireless communication components, cellular communication components, near field communication (NFC) components, Bluetooth® components (e.g., Bluetooth® Low Energy), Wi-Fi® components, and other communication components to provide communication via other modes. Device 820 may be another machine or any of various peripheral devices (e.g., peripheral devices coupled via Universal Serial Bus (USB)).

[0069] Furthermore, the communication component 840 can detect identifiers or include components operable to detect identifiers. For example, the communication component 840 may include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional barcodes such as Universal Product Code (UPC) barcodes, multi-dimensional barcodes such as Quick Response (QR) codes, Aztec codes, Data Matrix, Dataglyph, MaxiCode, PDF417, SuperCode, UCC RSS-2D barcodes, and other optical codes), or an acoustic detection component (e.g., a microphone for identifying audio signals from the tag). Additionally, various information can be derived via the communication component 840, such as location via Internet Protocol (IP) geolocation, location via Wi-Fi® signal triangulation, location via detection of NFC beacon signals that can indicate a specific location, etc.

[0070] Figure 7 This is a high-level functional block diagram of an example head-mounted device 100 that couples a mobile device 800 and a server system 998 via various network communications.

[0071] Device 100 includes a camera, such as at least one of a visible light camera 950, an infrared emitter 951, and an infrared camera 952. The camera may include [missing information - likely related to camera technology]. Figure 1 and Figure 2 The camera module with lenses 104_1 and 104_2.

[0072] Client device 800 can connect to device 100 using low-power wireless connection 925 and high-speed wireless connection 937. Client device 800 connects to server system 998 and network 995. Network 995 can include any combination of wired and wireless connections.

[0073] Device 100 further includes two image displays 980A-980B of optical components. The two image displays 980A-980B include one image display associated with the left side of device 100 and one image display associated with the right side of device 100. Device 100 also includes an image display driver 942, an image processor 912, low-power circuitry 920, and high-speed circuitry 930. The image displays of optical components 980A-B are used to present images and videos to a user of device 100, and may include images that can include a graphical user interface.

[0074] The image display driver 942 commands and controls the image display of the optical components 980A-B. The image display driver 942 can directly transmit image data to the image display of the optical components 980A-B for presentation, or it may need to convert the image data into a signal or data format suitable for transmission to the image display device. For example, the image data can be video data formatted according to a compression format, such as H.264 (MPEG-4 Part 10), HEVC, Theora, Dirac, RealVideo RV40, VP8, VP9, ​​etc., and still image data can be formatted according to a compression format such as Portable Network Group (PNG), Joint Photo Experts Group (JPEG), Tagged Image File Format (TIFF), or Exchangeable Image File Format (Exif).

[0075] As described above, device 100 includes a frame 103 and temples (or eyeglass temples) extending laterally from the frame 103. Device 100 further includes a user input device 991 (e.g., a touch sensor or button) that includes an input surface on device 100. The user input device 991 (e.g., a touch sensor or button) receives input selections from the user to manipulate a graphical user interface of a presented image.

[0076] Figure 7 The components for device 100 shown are located on one or more circuit boards, such as PCBs or flexible PCBs, within the frame or temple of the lens. Alternatively or additionally, the depicted components may be located in blocks, frames, hinges, or bridges of device 100. The left and right visible light cameras 950 may include digital camera elements, such as complementary metal-oxide-semiconductor (CMOS) image sensors, charge-coupled devices, lenses 104_1, 104_2, or any other corresponding visible light or light-gathering elements that can be used to acquire data, including images of scenes with unknown objects.

[0077] Device 100 includes a memory 934 that stores a subset or all of the functions described herein for generating binaural audio content. The memory 934 may also include a storage device 604. Figure 4 The exemplary process shown in the flowchart can be implemented in instructions stored in memory 934.

[0078] like Figure 7As shown, the high-speed circuit 930 includes a high-speed processor 932, a memory 934, and a high-speed wireless circuit 936. In this example, an image display driver 942 is coupled to the high-speed circuit 930 and operated by the high-speed processor 932 to drive the left and right image displays of the optical components 980A-B. The high-speed processor 932 can be any processor capable of managing the high-speed communication and operation of any general-purpose computing system required by device 100. The high-speed processor 932 includes the processing resources required to manage high-speed data transmission over the high-speed wireless connection 937 to a wireless local area network (WLAN) using the high-speed wireless circuit 936. In some examples, the high-speed processor 932 executes an operating system, such as the LINUX operating system or other such operating system of device 100, and the operating system is stored in the memory 934 for execution. Among other duties, the high-speed processor 932, which executes the software architecture of device 100, is used to manage data transmission with the high-speed wireless circuit 936. In some examples, the high-speed wireless circuit 936 is configured to implement the Institute of Electrical and Electronics Engineers (IEEE) 802.11 communication standard, also referred to herein as Wi-Fi. In other examples, other high-speed communication standards can be implemented by the high-speed wireless circuit 936.

[0079] The low-power wireless circuitry 924 and high-speed wireless circuitry 936 of device 100 may include short-range transceivers (Bluetooth™) and wireless wide area network, local area network, or wide area network transceivers (e.g., cellular or WiFi). Client device 800 (including transceivers communicating via low-power wireless connection 925 and high-speed wireless connection 937) may be implemented using details of the architecture of device 100, as well as other elements of network 995.

[0080] Memory 934 includes any storage device capable of storing various data and applications, including camera data generated by the left and right visible light cameras 950, infrared camera 952, and image processor 912, as well as images generated by image display driver 942 for display on the image display of optical components 980A-B. While memory 934 is shown as integrated with high-speed circuitry 930, in other examples, memory 934 may be a separate, independent element of device 100. In some such examples, electronic wiring may provide a connection from image processor 912 or low-power processor 922 to memory 934 via a chip including high-speed processor 932. In other examples, high-speed processor 932 may manage addressing of memory 934 such that low-power processor 922 will initiate high-speed processor 932 whenever a read or write operation involving memory 934 is required.

[0081] like Figure 7As shown, the processor 932 of device 100 may be coupled to a camera (visible light camera 950; infrared emitter 951, or infrared camera 952), an image display driver 942, a user input device 991 (e.g., a touch sensor or button) and a memory 934.

[0082] Device 100 is connected to a host computer. For example, device 100 is paired with client device 800 via a high-speed wireless connection 937 or connected to server system 998 via network 995. Server system 998 may be one or more computing devices as part of a network computing system or service, and may include, for example, a processor, memory, and a network communication interface for communicating with client device 800 and device 100 via network 995.

[0083] The client device 800 includes a processor and a network communication interface coupled to the processor. The network communication interface allows communication via a network 925 or 937. The client device 800 may further store at least a portion of the instructions for generating binaural audio content in its memory to implement the functions described herein.

[0084] The output components of device 100 include visual components, such as displays, such as liquid crystal displays (LCDs), plasma display panels (PDPs), light-emitting diode (LED) displays, projectors, or waveguides. The image display of the optical components is driven by an image display driver 942. The output components of device 100 further include acoustic components (e.g., speakers), haptic components (e.g., vibration motors), other signal generators, etc. The input components (such as user input devices 991) of device 100, client device 800, and server system 998 may include alphanumeric input components (e.g., keyboards, touchscreens configured to receive alphanumeric input, photoelectric keyboards, or other alphanumeric input components), point-based input components (e.g., mice, touchpads, trackballs, joysticks, motion sensors, or other pointing instruments), haptic input components (e.g., physical buttons, touchscreens that provide position and force for touch or touch gestures, or other haptic input components), audio input components (e.g., microphones), etc.

[0085] Device 100 may optionally include additional peripheral device elements. Such peripheral device elements may include biometric sensors, additional sensors, or display elements integrated with device 100. For example, peripheral device elements may include any I / O components, including output components, motion components, position components, or any other such components described herein.

[0086] For example, biometric components include those for detecting facial expressions (e.g., hand gestures, facial expressions, vocal expressions, body posture, or eye tracking), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweat, or brain waves), and identifying people (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or EEG-based recognition). Motion components include accelerometer components (e.g., accelerometers), gravity sensor components, rotation sensor components (e.g., gyroscopes), etc. Positioning components include position sensor components (e.g., GPS receiver components) for generating location coordinates, WiFi or Bluetooth™ transceivers for generating positioning system coordinates, altitude sensor components (e.g., altimeters or barometers that detect air pressure from which altitude can be derived), orientation sensor components (e.g., magnetometers), etc. Such positioning system coordinates can also be received from client device 800 via low-power wireless circuit 924 or high-speed wireless circuit 936 through wireless connections 925 and 937.

[0087] When phrases such as “at least one of A, B or C”, “at least one of A, B and C”, “one or more of A, B or C” or “one or more of A, B and C” are used, the phrase is intended to be interpreted as indicating that A may exist alone in one embodiment, B may exist alone in one embodiment, C may exist alone in one embodiment, or any combination of elements A, B and C may exist in a single embodiment; for example, A and B, A and C, B and C, or A and B and C.

[0088] Changes and modifications may be made to the disclosed embodiments without departing from the scope of this disclosure. These and other changes or modifications are intended to be included within the scope of this disclosure, as set forth in the following claims.

Claims

1. A head-mounted device, comprising: Audio processor, which includes A noise suppressor attenuates noise content from a first beamformer signal and noise content from a second beamformer signal to generate a first noise suppression signal and a second noise suppression signal, respectively. The attenuation of noise content from the first beamformer signal and noise content from the second beamformer signal includes: The acoustic signals not included in the first beamformer signal are determined, wherein the noise content from the first beamformer signal includes acoustic signals not included in the second beamformer signal, and Determine the acoustic signals in the second beamformer signal that are not included in the first beamformer signal, wherein the noise content from the second beamformer signal includes acoustic signals not included in the first beamformer signal; and A speech enhancer that generates a clean signal comprising speech content from the first noise-suppressed signal and the second noise-suppressed signal, wherein generating the clean signal includes: The acoustic signal included in both the first beamformer signal and the second beamformer signal is determined, wherein the speech content is included in the acoustic signal included in both the first beamformer signal and the second beamformer signal.

2. The head-mounted device according to claim 1, wherein, A first beamformer generates a first beamformer signal, and a second beamformer generates a second beamformer signal, wherein the first beamformer and the second beamformer are fixed beamformers.

3. The head-mounted device according to claim 1, further comprising: A beamformer controller that causes the first beamformer to be steered in a first direction and the second beamformer to be steered in a second direction.

4. The head-mounted device according to claim 3, wherein, When a user wears the head-mounted device, the first direction and the second direction are pointing towards the user's mouth.

5. The head-mounted device according to claim 3, wherein, The beamformer controller dynamically changes the first direction and the second direction.

6. The head-mounted device according to claim 1, wherein, The first beamformer signal is based on a first microphone array formed by a first front microphone and a first rear microphone, and the second beamformer signal is based on a second microphone array formed by a second front microphone and a second rear microphone.

7. The head-mounted device according to claim 6, wherein, The first microphone array and the second microphone array are wide-side arrays, end-fire arrays, or any combination thereof.

8. The head-mounted device according to claim 6, wherein, The first front microphone and the first rear microphone are located on a first plane, and the second front microphone and the second rear microphone are located on a second plane.

9. A method for improving the signal-to-noise ratio of a signal acquired using a head-mounted device, comprising: The noise content from the first beamformer signal and the noise content from the second beamformer signal are attenuated by a noise suppressor to generate a first noise-suppressed signal and a second noise-suppressed signal, respectively. The attenuation of the noise content from the first beamformer signal and the noise content from the second beamformer signal includes: The acoustic signals not included in the first beamformer signal are determined, wherein the noise content from the first beamformer signal includes acoustic signals not included in the second beamformer signal, and Determine the acoustic signals in the second beamformer signal that are not included in the first beamformer signal, wherein the noise content from the second beamformer signal includes acoustic signals not included in the first beamformer signal; and A clean signal is generated by the speech enhancer, comprising speech content from the first noise-suppressed signal and the second noise-suppressed signal, wherein generating the clean signal includes: The acoustic signal included in both the first beamformer signal and the second beamformer signal is determined, wherein the speech content is included in the acoustic signal included in both the first beamformer signal and the second beamformer signal.

10. The method according to claim 9, wherein, A first beamformer generates a first beamformer signal, and a second beamformer generates a second beamformer signal, wherein the first beamformer and the second beamformer are fixed beamformers.

11. The method of claim 9, further comprising: The beamformer controller causes the first beamformer to be steered in a first direction and the second beamformer to be steered in a second direction.

12. The method according to claim 11, wherein, The first direction and the second direction are directions pointing towards the user's mouth.

13. The method according to claim 11, wherein, The beamformer controller dynamically changes the first direction and the second direction.

14. The method according to claim 9, wherein, The first beamformer signal is based on a first microphone array formed by a first front microphone and a first rear microphone, and the second beamformer signal is based on a second microphone array formed by a second front microphone and a second rear microphone.

15. The method according to claim 14, wherein, The first microphone array and the second microphone array are wide-side arrays, end-fire arrays, or any combination thereof.

16. The method of claim 14, wherein, The first front microphone and the first rear microphone are located on a first plane, and the second front microphone and the second rear microphone are located on a second plane.

17. A non-transitory computer-readable medium having instructions stored thereon, the instructions, when executed by a processor, causing the processor to perform operations including: The noise content from the first beamformer signal and the noise content from the second beamformer signal are attenuated to generate a first noise suppression signal and a second noise suppression signal, respectively. in, Attenuation of noise content from the first beamformer signal and noise content from the second beamformer includes: The acoustic signals not included in the first beamformer signal are determined, wherein the noise content from the first beamformer signal includes acoustic signals not included in the second beamformer signal, and Determine the acoustic signals in the second beamformer signal that are not included in the first beamformer signal, wherein the noise content from the second beamformer signal includes acoustic signals not included in the first beamformer signal; and Generating a clean signal comprising speech content from the first noise-suppressed signal and the second noise-suppressed signal, wherein generating the clean signal includes: The acoustic signal included in both the first beamformer signal and the second beamformer signal is determined, wherein the speech content is included in the acoustic signal included in both the first beamformer signal and the second beamformer signal.

18. The non-transitory computer-readable medium according to claim 17, wherein, The processor performing the operation further includes: The first beamformer is turned in a first direction, and the second beamformer is turned in a second direction, the first direction and the second direction being directions pointing towards the user's mouth.

19. The non-transitory computer-readable medium according to claim 17, wherein, The processor performing the operation further includes: The first beamformer is steered in a first direction, and the second beamformer is steered in a second direction, wherein the first and second directions are dynamically changed.

20. The non-transitory computer-readable medium according to claim 17, wherein, A first beamformer generates a first beamformer signal, and a second beamformer generates a second beamformer signal, wherein the first beamformer and the second beamformer are fixed beamformers.

Citation Information

Patent Citations

  • Noise adaptive beamforming for microphone arrays

    CN102708874A

  • Stereo separation and directional suppression with omni-directional microphones

    CN109155884A