System and method for carrying out multi-direction sound reception and noise reduction by intelligent glasses microphone array

Through the dynamic optimization module of the smart glasses microphone array, the microphone directional weight is adjusted in real time, solving the problem of degradation in target voice signal pickup quality in complex dynamic noise environments, and achieving efficient voice signal clarity and noise reduction effect in multi-directional and multi-band noise environments.

CN120264197APending Publication Date: 2025-07-04SHENZHEN KAISHUODA DIGITAL CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510301947.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art cannot adjust the microphone directional weight in real time in complex dynamic noise environments, resulting in a decrease in the quality of target voice signal picking, especially in multi-person meetings or outdoor scenes, which is difficult to take into account both the clarity of voice signal and the noise reduction effect.

Method used

The smart glasses microphone array is adopted, including at least five different types of microphones arranged at different positions in the glasses frame. Combined with signal processing, noise modeling, voice enhancement and dynamic optimization modules, the pickup and noise suppression of the target voice signal are optimized in real time by dynamically adjusting the directional weight of the microphone array.

Benefits of technology

It significantly improves the clarity and fidelity of the voice signal, maintains the intelligibility of the target voice in complex noise environments, adapts to multi-directional and multi-band noise environments, and quickly responds to the movement and changes of noise sources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120264197A_ABST
    Figure CN120264197A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of sound reception, and discloses a system for multi-directional sound reception and noise reduction of a smart glasses microphone array, and the system comprises at least five microphones which are disposed at different positions of a smart glasses frame, and the first microphone is a unidirectional microphone, is located below the glasses frame, and is used for picking up the voice of a wearer; the second microphone and the third microphone are bi-directional microphones, are located right in front of the glasses frame and are used for picking up voice signals right in front of the wearer. The directivity weight of the microphone array is adjusted in real time through the dynamic optimization module, and the pickup effect of the target voice signal is enhanced preferentially. The voice enhancement module is matched to extract a target voice signal by using a Bayesian inference method, so that noise components in a mixed signal can be effectively separated, voice distortion is reduced, the definition and fidelity of the voice signal are remarkably improved, and the voice has better intelligibility in various complex noise environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sound collection, specifically to a system and method for multi-directional sound collection and noise reduction using a microphone array of smart glasses. Background Art

[0002] In recent years, intelligent voice interaction technology has been increasingly widely used in consumer electronic devices, especially in the field of wearable devices. Smart glasses have gradually become an important research direction due to their portability and intelligence. As an important part of smart glasses, voice acquisition and processing technology plays a key role in realizing functions such as voice interaction, real-time voice recording, and remote communication. In order to ensure the quality of voice signals, smart glasses usually need to effectively pick up target voices in complex multi-source noise environments while suppressing background noise, so as to achieve clear and natural voice output.

[0003] Currently, traditional voice acquisition and noise reduction systems mainly use a single microphone or a fixed-direction microphone array to pick up and process voice signals. These methods can meet certain voice acquisition requirements in simple noise environments, but when the noise sources are distributed complexly or change dynamically, their noise reduction effects will decrease significantly. At the same time, most of the noise suppression methods in the existing technology rely on fixed noise modeling algorithms and cannot flexibly adapt to multi-directional and multi-band noise environments. In addition, in a dynamic environment, traditional technologies cannot adjust the directional weights of microphones in real time, resulting in poor picking effects of voice signals in dynamic scenarios.

[0004] The existing technology has deficiencies in robustness and adaptability in dynamic noise environments. For example, when the noise sources are unevenly distributed or move quickly, traditional technologies cannot adjust the sound pickup direction in real time, resulting in a decrease in the picking quality of target voice signals. Especially in public speeches, multi-person meetings, or outdoor scenarios, it is difficult to balance the clarity of voice signals and the noise reduction effect. Therefore, an intelligent system that can adapt to complex noise environments in multiple directions and frequency bands and dynamically optimize the directional weights of the microphone array is needed to solve the problem of insufficient clarity in picking target voice signals in complex environments. Summary of the Invention

[0005] Aiming at the deficiencies of the existing technology, the present invention provides a system and method for multi-directional sound collection and noise reduction using a microphone array of smart glasses, which solves the problem that the existing technology cannot adjust the directional weights of microphones in real time to effectively pick up target voice signals in complex dynamic noise environments.

[0006] To achieve the above objectives, the present invention is realized through the following technical solutions: A system for multi-directional sound collection and noise reduction using a microphone array of smart glasses, comprising:

[0007] At least five microphones, arranged at different positions on the smart glasses frame, wherein:

[0008] The first microphone is a unidirectional microphone located below the glasses frame for picking up the voice of the wearer;

[0009] The second and third microphones are bidirectional microphones located directly in front of the glasses frame for picking up the voice signals directly in front of the wearer;

[0010] The fourth and fifth microphones are omnidirectional microphones respectively located at the rear ends of the temple arms for collecting ambient noise;

[0011] A signal processing module, connected to the microphone array, for receiving and performing analog-to-digital conversion, amplification, and filtering processing on the signals collected by the microphones;

[0012] A noise modeling module, connected to the signal processing module, for establishing a dynamic distribution model of the noise field;

[0013] A voice enhancement module, connected to the signal processing module, for extracting target voice signals based on a preset algorithm;

[0014] A noise reduction module, connected to the noise modeling module and the voice enhancement module, for performing distribution matching processing on the noise signals;

[0015] A dynamic optimization module, connected to the microphone array, for adjusting the sound collection directivity weights of the microphones.

[0016] Preferably, the signal processing module includes an analog-to-digital converter, a preamplifier, and a filter for converting analog signals into digital signals and performing amplification and filtering processing on the converted signals.

[0017] Preferably, the noise modeling module is based on the Markov random field theory and establishes a distribution model of the direction and intensity of noise sources in the noise field by calculating the conditional probability distributions of the signals in the microphone array.

[0018] Preferably, the voice enhancement module enhances the target voice signals through the following steps:

[0019] Separate the target voice signals and the noise signals;

[0020] Optimize the target voice signals based on the signal probability model.

[0021] Preferably, the noise reduction module adjusts the noise signal distribution to the target voice signal distribution by calculating the distribution difference between the noise signals and the target voice signals and through a specific signal distribution mapping.

[0022] Preferably, the dynamic optimization module adjusts the signal reception direction of the microphones by calculating the adjustment parameters of the microphone array directivity weights.

[0023] Preferably, the first microphone of the microphone array is a unidirectional microphone, the second and third microphones are bidirectional microphones, and the fourth and fifth microphones are omnidirectional microphones. Their specific parameters are used to meet the pickup requirements of signals in different directions.

[0024] A method for multi-directional sound collection and noise reduction by a microphone array of smart glasses, the steps of which include:

[0025] Step 1: Collect audio signals through the microphone array. The first microphone picks up the voice of the wearer, the second and third microphones pick up the voice signals directly in front, and the fourth and fifth microphones collect environmental noise signals;

[0026] Step 2: Use a noise modeling module to establish a dynamic distribution model of the noise field, where the direction and intensity of the noise signals are calculated from the conditional probability distribution;

[0027] Step 3: Use a voice enhancement module to extract the target voice signals from the observed signals, and the target voice signals and the noise signals are separated by a probability model;

[0028] Step 4: Use a noise reduction module to perform distribution matching processing on the noise signals, and adjust the distribution of the noise signals to the distribution of the target voice signals;

[0029] Step 5: Use a dynamic optimization module to adjust the sound pickup directivity weights of the microphones to adapt to the change in the direction of the signal source;

[0030] Step 6: Output the processed target voice signals to a terminal device for subsequent use.

[0031] Preferably, the signal processing module is combined with the noise reduction module to simultaneously complete signal amplification, filtering, analog-to-digital conversion, and matching adjustment of the noise signals.

[0032] Preferably, the dynamic optimization module dynamically generates microphone direction weight adjustment parameters by calculating the distribution of the current environmental noise direction in real time.

[0033] The present invention provides a system and method for multi-directional sound collection and noise reduction by a microphone array of smart glasses.

[0034] It has the following beneficial effects:

[0035] 1. The present invention adjusts the directional weights of the microphone array in real time through a dynamic optimization module, giving priority to enhancing the pickup effect of the target voice signal. When used in conjunction with a voice enhancement module that uses the Bayesian inference method to extract the target voice signal, it can effectively separate the noise components in the mixed signal, reduce voice distortion, thereby significantly improving the clarity and fidelity of the voice signal, and making the voice more intelligible in various complex noise environments.

[0036] 2. The present invention uses a noise reduction module combined with the optimal transport theory to dynamically model and adjust the distribution of the noise signal. By suppressing environmental noise in different frequency bands, the module can quickly reduce noise interference in a multi-directional and multi-source noise environment, while maintaining the integrity and quality of the target voice, and is applicable to complex and changing real-world scenarios.

[0037] 3. By using a noise modeling module to obtain the direction and intensity distribution information of noise sources in the environment in real time, and combining with a dynamic optimization module to dynamically update the weight allocation of the microphone array, the present invention can adapt to the movement of noise sources, changes in the direction of the target voice, and dynamic changes in environmental noise, ensuring high flexibility and adaptability of the system, especially having obvious advantages in dynamic scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 is a flowchart of the method of the present invention;

[0039] Figure 2 is a schematic diagram of the first hardware layout of the present invention;

[0040] Figure 3 is a schematic diagram of the second hardware layout of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0042] Please refer to the attached Figure 1 - attached Figure 3 , the embodiments of the present invention provide a system and method for multi-directional sound pickup and noise reduction using an intelligent glasses microphone array, including:

[0043] Microphone type and position:

[0044] The first microphone is a unidirectional microphone, located at a position near the wearer's mouth under the glasses frame, mainly picking up the voice signal of the wearer. Its characteristic is that the receiving range is concentrated, and it can effectively suppress noise interference from the sides and rear.

[0045] The second and third microphones are bidirectional microphones, which are respectively arranged on both sides of the front part of the glasses frame. These two microphones are mainly used to collect the voice signals directly in front of the wearer, such as the voice of the conversation partner.

[0046] The fourth and fifth microphones are omnidirectional microphones, which are respectively located at the rear ends of the temple arms. They are responsible for collecting the wide-area noise in the environment for subsequent noise modeling and suppression.

[0047] The distribution of the microphones has a reasonable division of directionality and function. The first microphone is close to the wearer's mouth, which can effectively reduce the attenuation problem of voice signals. The bidirectional design of the second and third microphones gives them strong sound pickup ability in collecting voice directly in front. The position design of the fourth and fifth microphones can comprehensively cover the source directions of environmental noise, which is beneficial to improving the accuracy of noise modeling.

[0048] Signal processing module

[0049] In the present invention, the signal processing module is one of the core components for the intelligent glasses microphone array to perform multi-directional sound collection and noise reduction. This module is mainly used to process the original signals from the microphone array and provide high-quality signal inputs for the subsequent noise modeling module, voice enhancement module, and noise reduction module. Generally, the signal processing module needs to complete operations such as analog-to-digital conversion, amplification, and filtering of the signals. These operations can ensure that the collected audio signals have sufficient quality and stability so that the subsequent algorithms can accurately analyze and process them. In addition, the signal processing module works in cooperation with the dynamic optimization module and can adjust the input weights of the microphone signals according to the optimization parameters, thereby further improving the overall system performance.

[0050] As an option, after the signal processing module is connected to the microphone array, the received signals will directly pass through an analog-to-digital converter to convert the analog signals into digital signals. The converted signals will enter an amplifier for gain compensation, and at the same time, a filter is used to filter the high-frequency noise and low-frequency interference in the signals. Specifically, the analog-to-digital conversion of the signals is the basic step of the entire signal processing, and its accuracy directly affects the subsequent signal quality.

[0051] In this embodiment, the analog-to-digital conversion unit of the signal processing module adopts a high-precision 16-bit analog-to-digital converter (ADC). Generally, the resolution of this analog-to-digital converter can meet the processing requirements of most audio signals, and the sampling rate of the conversion is usually set to 44.1 kHz or 48 kHz, and these rates can cover most of the voice signal frequency ranges. In some embodiments, the sampling rate can be adjusted according to the complexity of the environmental noise. For example, a higher sampling rate can be selected in a high-noise environment to capture more signal details.

[0052] Specifically, the output signal of the analog-to-digital conversion is subjected to gain compensation through a preamplifier. The gain value of the preamplifier can be adjusted between 5 dB and 20 dB to adapt to the acquisition sensitivities of different microphones and the variations in signal strength. As an option, an automatic gain control (AGC) circuit can be employed to automatically adjust the gain according to the amplitude of the input signal, thereby ensuring that the output signal has a high dynamic range.

[0053] In one possible implementation, the amplified signal enters a filter module for preprocessing. The filter module typically includes a high-pass filter and a low-pass filter for removing interference components in the signal. Specifically, the cut-off frequency of the high-pass filter can be set to 100 Hz to remove low-frequency noise in the environment, such as air conditioner noise or mechanical vibration noise. The cut-off frequency of the low-pass filter can be set between 10 kHz and 12 kHz to attenuate high-frequency interference, such as high-frequency noise generated by electronic devices.

[0054] In some embodiments, the filter module can also implement band-pass filtering. Generally, the passband range of the band-pass filter is from 300 Hz to 3.4 kHz, and this frequency range can cover the frequency characteristics of most speech signals. As an option, the passband range of the band-pass filter can be adjusted according to the actual application scenario. For example, in an indoor environment, the high-frequency boundary of the passband can be appropriately reduced to better suppress background noise.

[0055] In this embodiment, the output signal of the signal processing module is split into two paths: one path is directly input into the noise modeling module for subsequent noise direction estimation and intensity distribution modeling; the other path is input into the speech enhancement module for separating and extracting the target speech signal. In this signal splitting mode, the module can simultaneously provide high-quality input data for multiple subsequent algorithms, thereby ensuring the overall coordination of the system.

[0056] To further improve the flexibility of signal processing, the signal processing module also supports a dynamic weight adjustment function. As a possible implementation, the dynamic weight adjustment is supported by parameters provided by a dynamic optimization module. By adjusting the gain weights of each microphone signal, the signal in the target direction is preferentially amplified. Specifically, the core of the dynamic weight adjustment is to calculate an optimization coefficient matrix based on the amplitude differences of the input signals of each microphone, and this matrix represents the weight distribution ratio of each microphone. The calculation formula for the optimization coefficient is as follows:

[0057]

[0058] Where:

[0059] W i : The weight of the i-th microphone;

[0060] A i : The amplitude of the input signal of the i-th microphone;

[0061] N: The total number of microphones.

[0062] Through the above optimization calculation, the input signal intensity of each microphone can be adjusted in real time to ensure that the proportion of the target direction signal in the overall system input is larger.

[0063] In practical applications, the signal processing module of the present invention can also be optimized in cooperation with different external hardware or algorithms. For example, an adaptive noise reduction algorithm can be integrated into the microphone array to further optimize the input quality of the signal processing module. In some environments, such as scenes with high wind noise or sudden noise, the signal processing module can dynamically weaken the influence of these short-term noises through the gain suppression function, so as to provide a more stable signal input for the subsequent modules.

[0064] Noise modeling module

[0065] In the present invention, the noise modeling module is directly connected to the signal processing module, and is mainly used to analyze the environmental noise signals from the microphone array and establish a dynamic distribution model of the noise field. Through the effective estimation of the noise direction and intensity by the modeling module, it provides necessary input support for the subsequent speech enhancement module and noise reduction module. Generally, this module needs to be able to process multi-source noise signals in a dynamic environment in real time and accurately describe the characteristics of the noise field through a mathematical model. As a core function, the noise modeling module adopts an algorithm based on Markov random field (MRF), and combines multi-channel signals collected by microphones for modeling, which is suitable for complex non-uniform noise scenarios.

[0066] Specifically, the noise modeling module can extract spatial feature information from multi-microphone input signals, including the direction distribution and intensity distribution of the noise. Generally, the correlation between noise sources can be characterized by the correlation of signals from adjacent microphones. In a possible implementation, the noise modeling module establishes a dynamic model of the noise field by calculating the conditional probability distribution of the noise signal, and subsequent modules can separate and suppress the noise according to this model.

[0067] In this embodiment, the core of the noise modeling module is based on the Markov random field theory. Markov random field is a mathematical model that can describe the mutual relationship between local and global characteristics, and is suitable for modeling the directionality and intensity of noise signals. Generally, the signal input of the microphone array can be regarded as a set of nodes, and each node represents the noise signal collected by a microphone. Assume that this node set is N = {N1, N2,..., NM}, where M is the number of microphones.

[0068] In a Markov random field, the noise signal satisfies the locality principle, that is, the state of any node is only related to its neighboring nodes. Specifically, its conditional probability distribution can be expressed as:

[0069]

[0070] where N i represents the noise signal of the i-th microphone, and N -i represents the noise signals of all other microphones, and neighborhoodi is the neighborhood of the i-th node.

[0071] In some embodiments, the difference in noise signals between neighboring nodes can be described by the following formula:

[0072]

[0073] where:

[0074] N i , N j are the noise signal intensities of adjacent microphones respectively; σ is the correlation parameter of the noise signal, which is usually set according to the experimental environment.

[0075] To further describe the global distribution of the noise field, the noise modeling module expands the above conditional probability into a joint probability distribution. Generally, this joint distribution can be expressed as:

[0076]

[0077] where:

[0078] Z is the normalization factor, which is used to ensure that the sum of the probability distribution is 1; E represents the set of adjacent edges between microphone nodes. Through the above joint distribution formula, the noise modeling module can accurately describe the dynamic distribution characteristics of the noise source, including its directionality and intensity changes.

[0079] Through the above joint distribution formula, the noise modeling module can accurately describe the dynamic distribution characteristics of the noise source, including its directionality and intensity changes.

[0080] In a possible implementation, the noise modeling module adjusts the dynamic model of the noise field in real time by continuously updating the conditional probability of neighboring nodes. For example, when the signal intensity collected by the microphone changes, the module will recalculate the probability distribution of each node and provide more accurate noise direction information according to the updated model. In some embodiments, the time interval of this dynamic adjustment process can be set between 10 milliseconds and 100 milliseconds according to the speed of environmental changes.

[0081] Specifically, the noise modeling module can also combine the spectral characteristics of the signal to model the distribution of noise signals in different frequency bands. For example, in the low-frequency band (100 Hz to 300 Hz), the module will give priority to focusing on the intensity distribution of background noise; while in the high-frequency band (2 kHz to 8 kHz), the module will focus on analyzing the directional characteristics of sharp noise.

[0082] In this embodiment, the output results of the noise modeling module include the direction distribution matrix and the intensity distribution vector of the noise field. The direction distribution matrix D represents the angular distribution of the noise source in space and can be calculated by the following formula:

[0083]

[0084] Where:

[0085] x i , y i and x j , y j respectively represent the spatial coordinates of microphones i and j. The intensity distribution vector S represents the noise intensity collected by each microphone and can be obtained by weighted averaging of the node signals.

[0086] The intensity distribution vector represents the noise intensity collected by each microphone and can be obtained by weighted averaging of the node signals.

[0087] Generally, the output results of the noise modeling module will be directly input into the noise reduction module for distribution adjustment and directional filtering of the noise signal. In addition, the direction distribution matrix can also be called by the dynamic optimization module to adjust the directional weights of the microphones. In a possible implementation, when the noise direction changes rapidly, the noise modeling module will update the dynamic optimization parameters through the direction distribution calculated in real time, so as to ensure that the system can quickly respond to environmental changes.

[0088] Voice enhancement module

[0089] In the present invention, the voice enhancement module is closely combined with the signal processing module and the noise modeling module and is an important part of the intelligent glasses microphone array for multi-directional sound collection and noise reduction. This module is mainly used to separate the target voice signal from the mixed signal and enhance it to improve the quality and clarity of the voice signal. Generally, the voice enhancement module analyzes the input signal through a probability model and an optimization algorithm to separate the target voice signal and the environmental noise signal. In the subsequent noise reduction module, the enhanced target voice signal will be further processed to ensure that the finally output voice signal has high accuracy.

[0090] As an option, the voice enhancement module can work in cooperation with the output data of the noise modeling module to optimize the extraction effect of the target voice by using the noise direction and intensity distribution information. Specifically, the module constructs a probability model of the target voice and the noise signal through the Bayesian inference method, and makes an optimal estimation of the target voice in combination with the observed signal. In one possible implementation, the module can also dynamically adjust the enhancement parameters according to the characteristics of the input signal to adapt to the changes in different environmental noises.

[0091] In this embodiment, the voice enhancement module decomposes and optimizes the input mixed signal based on the Bayesian inference theory. Assume that the mixed signal S is composed of the target voice signal S clean and the noise signal N superimposed, that is:

[0092] S = S clean + N

[0093] Where:

[0094] S: The observed signal collected by the microphone;

[0095] S clean : The target voice signal;

[0096] N: The environmental noise signal.

[0097] Generally, the goal of voice enhancement is to extract S clean from the mixed signal S. To achieve this goal, the module first constructs a probability model of the observed signal according to the noise distribution information provided by the noise modeling module. Specifically, the probability density function of the observed signal can be expressed as:

[0098]

[0099] Where:

[0100] σ 2 : The variance of the noise, representing the intensity of the noise signal.

[0101] In some embodiments, the prior probability distribution of the target voice signal is assumed to be a normal distribution, and its probability density function is:

[0102]

[0103] Where:

[0104] Z clean : The normalization factor, used to ensure that the sum of the probability densities is 1;

[0105] λ: The prior intensity parameter of the target voice signal, usually set according to the environmental conditions.

[0106] To achieve voice enhancement, the module calculates the posterior probability of the target voice signal through Bayes' formula:

[0107] PS clean |S ∝ P(S|S clean PS clean

[0108] Where:

[0109] P(S|S clean : Likelihood function of the observed signal;

[0110] PS clean : Prior distribution of the target voice signal.

[0111] As an option, the voice enhancement module uses the maximum a posteriori (MAP) estimation method to extract the target voice signal, that is:

[0112]

[0113] In a possible implementation, the module realizes the extraction of the target voice signal by optimizing the following objective function:

[0114]

[0115] The goal is to minimize the objective function to obtain the optimal target voice signal Generally, this optimization problem can be solved by the gradient descent method or other numerical optimization algorithms.

[0116] Specifically, the input signal of the voice enhancement module comes from the output of the signal processing module and contains multi-channel audio data collected by the microphone array. In some embodiments, the module will preferentially process signals with stronger directivity, such as the wearer's voice signal picked up by the first microphone. To ensure the robustness of the optimization result, the module will also combine the noise direction distribution data output by the noise modeling module and assign different weights to each microphone input signal. As an option, the following weight assignment strategy can be adopted:

[0117]

[0118] Where:

[0119] W i : Weight of the input signal of the i-th microphone;

[0120] D i : Distance between the i-th microphone and the noise source;

[0121] N: Total number of microphones;

[0122] α: Adjustment parameter used to control the distribution range of weights.

[0123] Through the above weight assignment, the voice enhancement module can preferentially extract signals close to the target voice direction and simultaneously suppress noise signals far from the target direction.

[0124] In some embodiments, the voice enhancement module can also adjust and optimize parameters according to the dynamic characteristics of the input signal. For example, when the noise intensity is high, the module can increase σ 2 to improve the enhancement ability for the target voice; while when the noise direction is relatively concentrated, the module can reduce α to enhance the weight of the directional signal. Under this dynamic adjustment strategy, the module can adapt to different environmental conditions, thus ensuring the stability of voice enhancement.

[0125] In summary, the voice enhancement module separates and enhances the target voice signal from the mixed signal through Bayesian inference and weight assignment strategies, and at the same time dynamically adjusts and optimizes parameters by combining the environmental information provided by the noise modeling module, providing a high-quality input signal for the subsequent noise reduction module. The design of the module fully considers the complexity in practical applications, and its algorithms and implementation methods can effectively cope with changing environmental conditions. The detailed description of this embodiment has covered the key technical details and implementation methods of the voice enhancement module, and enables those skilled in the art to reproduce the technology based on the disclosed content.

[0126] Noise reduction module

[0127] In the present invention, the noise reduction module closely cooperates with the noise modeling module and the voice enhancement module, and is an important part of the system to achieve environmental noise suppression. The main function of this module is to separate the noise signal from the target voice signal by modeling and optimizing the noise signal, thereby improving the purity of the voice signal. Generally, the noise reduction module uses a method based on the optimal transport theory to suppress the noise component in the mixed signal by adjusting and matching the noise distribution, ensuring the integrity of the target voice signal.

[0128] As an option, the noise reduction module can combine the direction and intensity distribution information provided by the noise modeling module, as well as the target voice signal generated by the voice enhancement module, to dynamically optimize the noise processing effect. Specifically, the noise reduction module finds the optimal mapping relationship between the noise signal distribution and the target voice signal distribution through a distribution difference measurement method. In a possible implementation, the module further combines frequency band characteristics to classify and match noises in different frequency ranges to adapt to the multi-source noise characteristics in complex environments.

[0129] In this embodiment, the noise reduction module realizes the suppression of the noise signal based on the optimal transport theory. Assuming that the target voice signal and the noise signal have their respective distribution characteristics, the distribution of the noise signal can be expressed as PN , the distribution of the target voice signal can be expressed as P S . Generally, the difference between the noise signal and the target voice signal is adjusted by means of distribution mapping. The module constructs an optimal mapping to adjust the distribution of the noise signal to the same form as the target voice signal, thereby reducing noise interference in the mixed signal.

[0130] As a possible implementation, this distribution mapping aims at minimizing the transmission cost and finding the best noise signal distribution adjustment scheme. In a simplified case, the goal of the mapping is to make the distribution of the noise signal as close as possible to that of the target voice signal, thereby suppressing the energy of the noise signal.

[0131] In some embodiments, the implementation of the noise reduction module incorporates a dynamic adjustment strategy. For example, the module can update the distribution characteristics of the noise signal according to the changes in the real-time input signal. When the direction of the noise signal changes, the module can quickly adapt to the environmental changes by recalculating the mapping parameters. In addition, the module can also distinguish and process noise signals in different frequency ranges in combination with the spectral characteristics of the signal. Specifically, the module will give priority to processing high-frequency noise components, such as electromagnetic interference signals, while adopting different distribution adjustment strategies for low-frequency background noise.

[0132] Specifically, the output signal of the noise reduction module is directly used for the subsequent dynamic optimization module, and the optimized noise suppression signal can further improve the directivity and clarity of the target voice signal. In a possible implementation, the module decomposes the noise signal by means of frequency band analysis and adjusts each frequency band of the noise signal separately. For example, for low-frequency noise in the range of 100 Hz to 500 Hz, the module reduces the interference of background noise by weakening its amplitude; for high-frequency noise in the range of 2 kHz to 10 kHz, the module gives priority to increasing the proportion of the target voice signal in this frequency band to improve the overall clarity of the voice signal.

[0133] Generally, the implementation of the noise reduction module does not depend on fixed parameter configurations, but can be optimized and adjusted according to the dynamic characteristics of the input signal. In some embodiments, the module combines the direction distribution information generated by the noise modeling module and further suppresses the noise signal by means of direction weight allocation. For example, when the noise direction is concentrated in a specific area, the module assigns a lower weight to the noise signal in that direction, thereby improving the pickup effect of the target direction signal.

[0134] In summary, the noise reduction module of the present invention achieves precise suppression of environmental noise signals through distribution matching, spectrum analysis, and weight allocation strategies. The module combines the output data of the noise modeling module and the speech enhancement module, and can dynamically optimize the processing method of noise signals in complex and changing environments. The disclosure of the above technical details is complete enough to provide a reference basis for those skilled in the art to implement, and to ensure the practicality and technological innovation of the present invention.

[0135] Dynamic Optimization Module

[0136] In the present invention, the dynamic optimization module directly cooperates with the noise reduction module and the microphone array, and is one of the core modules for realizing multi-directional sound collection and environmental adaptive noise reduction of smart glasses. The module optimizes the pickup efficiency of the target speech by adjusting the directional weights of the microphone array in real time, while suppressing noise interference from non-target directions. Generally, the dynamic optimization module dynamically allocates the sound collection weights of each microphone based on the noise direction and intensity distribution information provided by the noise modeling module, as well as the signal characteristics optimized by the noise reduction module, so as to achieve directional enhancement of the target speech.

[0137] As an option, the dynamic optimization module ensures that the system can quickly respond to changes in environmental noise in a complex dynamic environment by updating the microphone direction weights in real time. Specifically, the module can assign different weight values to the channel signals of the microphone array according to the moving direction of the noise source and the frequency characteristics of the target speech signal. In one possible implementation, the module can also dynamically optimize the weight allocation strategy by combining the overall input signal intensity distribution of the system.

[0138] In this embodiment, the core of the dynamic optimization module is to adjust the weight allocation of the microphone array according to the dynamic changes of the noise direction and the target speech direction. Assume that the directional weight of the i-th microphone in the microphone array is w i , the module calculates the difference between the noise source direction and the target speech direction and assigns the optimal weight value. Generally, the weight allocation can be expressed as:

[0139]

[0140] Where:

[0141] w i : The weight activity of the i-th microphone;

[0142] D i : The distance between the i-th microphone provided by the noise modeling module and the noise source;

[0143] α: Adjustment parameter, controlling the sensitivity of the weight distribution;

[0144] N: The total number of the microphone array.

[0145] The above weight assignment method can ensure that microphones in the target direction obtain higher weights while suppressing noise interference from non-target directions.

[0146] As an option, the dynamic optimization module can also adjust the weights of different frequency bands in combination with the spectral characteristics of the input signal. Specifically, when the target voice signal is concentrated in a specific frequency range, the module will preferentially assign weights to the signal channels in that frequency range. For example, in an environment with more high-frequency noise, the module will enhance the weights of the mid-low frequency signals of the target voice, thereby improving the intelligibility of the voice. In some embodiments, the module dynamically updates the weight assignment parameters by real-time monitoring of the signal energy distribution to adapt to changes in the spectral characteristics of the environment.

[0147] In a possible implementation, the dynamic optimization module also analyzes the input signal intensity of each microphone and adjusts the weight assignment according to the amplitude value of the signal. For example, when the target voice signal picked up by a certain microphone is strong, its weight will increase accordingly, while the weight of the microphone that picks up more noise components will decrease. This dynamic adjustment strategy can be described by the following formula:

[0148]

[0149] Where:

[0150] A i : The amplitude of the input signal of the i-th microphone;

[0151] N: The total number of the microphone array.

[0152] Generally, the adjustment process of the dynamic optimization module is real-time, and its response speed can meet the requirements of a rapidly changing noise environment. For example, in the scenario of a moving noise source, the module can ensure the dynamics of the weights by updating the frequency. In some embodiments, the time interval for weight update can be set between 10 milliseconds and 50 milliseconds, and the specific time interval is adjusted according to the real-time requirements of the target application.

[0153] To further improve the adaptability of the dynamic optimization module, the module can also correct the weight assignment in combination with the processing results of the noise reduction module. For example, when the signal output by the noise reduction module still contains strong noise residues, the dynamic optimization module can optimize the weight assignment strategy by recalculating the noise direction distribution weights, thereby further reducing the interference of noise components.

[0154] Specifically, the output result of the dynamic optimization module directly acts on the directivity weight control unit of the microphone array to adjust the signal input gain of each microphone. Through the above optimization method, the system can achieve a dynamic balance between multi-directional speech acquisition and environmental noise suppression, thus ensuring high-quality reception of the target speech signal.

[0155] Embodiment 1:

[0156] To verify the actual application effect of the dynamic optimization module of the present invention in the multi-directional sound collection and noise reduction system of smart glasses, the present invention will be specifically described below in combination with an office meeting scenario. In this embodiment, the microphone array of the smart glasses is used to pick up the voice of the wearer, and at the same time suppress the surrounding interference sounds in a complex noise environment to improve the directivity of sound pickup and the clarity of the voice signal.

[0157] Scene description:

[0158] A user wearing smart glasses is participating in a group meeting in an open office environment. There are multiple noise sources in this environment, including the sound of nearby keyboard tapping (high-frequency noise, frequency range 2 kHz to 4 kHz), the noise of the air conditioner running (low-frequency noise, frequency range 100 Hz to 300 Hz), and the voice conversations of other colleagues (mid-frequency noise, frequency range 500 Hz to 2 kHz). The user needs to conduct a remote voice conference through the smart glasses to interact with other online participants in voice, and at the same time, the voice of the user needs to be clearly transmitted during the meeting.

[0159] 1. Signal acquisition and preprocessing

[0160] The microphone array of the smart glasses collects the voice of the wearer and the environmental noise signal. Specifically, it includes:

[0161] The first microphone (single-directional) mainly picks up the voice signal of the wearer;

[0162] The second and third microphones (bi-directional) pick up the voice signal directly in front of the wearer to supplement the voice direction information;

[0163] The fourth and fifth microphones (omnidirectional) collect the wide-area environmental noise signal.

[0164] The collected signals are preprocessed by the signal processing module for analog-to-digital conversion, amplification, and filtering, filtering out the background noise signals below 100 Hz and above 10 kHz. At this time, the data input to the dynamic optimization module is the mixed signal preliminarily processed by the noise reduction module.

[0165] 2. Dynamic weight allocation

[0166] In this embodiment, the dynamic optimization module calculates the weight assignment matrix of the microphone array based on the noise direction distribution and intensity information provided by the noise modeling module, in combination with the target direction of the wearer's voice. The specific steps are as follows:

[0167] The keyboard tapping sound is modeled as high-frequency noise coming from the rear right of the user (about 135° direction), and the weight is reduced;

[0168] The air conditioner noise is modeled as low-frequency noise coming from the upper left of the user (about 45° direction), and the weight is moderately suppressed;

[0169] The conversations of other colleagues are modeled as medium-frequency noise coming from directly in front of the user (about 90° direction), which partially overlaps with the target voice direction, and the weight is partially adjusted.

[0170] The dynamic optimization module assigns the directional weight w of each microphone through the following formula i as follows:

[0171]

[0172] where D i represents the directional difference between the noise source and the i-th microphone, and α is the directional adjustment parameter. In this embodiment, α is set to 2 to enhance the significance of the directional difference. The first microphone (close to the user's mouth) is given the highest weight, the second and third microphones are assigned medium weights to enhance the voice signal directly in front, while the weights of the fourth and fifth microphones are significantly reduced to suppress the influence of ambient noise.

[0173] 3. Dynamic spectrum optimization

[0174] The dynamic optimization module combines the spectrum analysis results of the noise reduction module and adjusts the weights for noise signals in different frequency bands. In this embodiment, the module detects that the energy of the keyboard tapping sound in the high-frequency band (2 kHz to 4 kHz) is strong, so the weight of this frequency band is suppressed. At the same time, in the mid-frequency band (500 Hz to 2 kHz), the module enhances the pickup weight of the target voice signal, and the specific calculation is as follows:

[0175]

[0176] where f target is the center frequency of the target voice signal (about 1 kHz), and β is the frequency band adjustment parameter. In this embodiment, β is set to 1.5 to achieve a smooth distribution of weights.

[0177] 4. Real-time weight update

[0178] During the meeting, due to the dynamic changes of the noise source, the dynamic optimization module updates the weight assignment in real time at intervals of 50 milliseconds.

[0179] For example, when the conversation sound of the colleague directly in front of the user increases, the module automatically adjusts the weights of the microphone array, further improving the pickup priority of the first, second, and third microphones, and reducing the weights of the fourth and fifth microphones to suppress the interference of ambient noise.

[0180] 5. Output the optimized signal

[0181] The optimized signal is transmitted in real time to the remote conferencing system through the voice transmission module built in the smart glasses. In this signal, the clarity of the target voice signal is significantly improved, while the noise components from the keyboard, air conditioner, and other conversation sounds are effectively suppressed.

[0182] Result analysis

[0183] In this embodiment, the introduction of the dynamic optimization module significantly improves the directivity enhancement effect of the target voice signal. Through the real-time dynamic adjustment of the weights of the microphone array, the voice of the wearer is clearly restored in the remote conference. The specific results are as follows:

[0184] The average energy of ambient noise (keyboard tapping sound, air conditioner noise) is reduced by about 85%;

[0185] The signal-to-noise ratio of the target voice signal is increased by about 12 dB;

[0186] The dynamic response time of the system is less than 50 milliseconds, and it can adapt to a rapidly changing noise environment.

[0187] In summary, this embodiment verifies the practical application effect of the dynamic optimization module of the present invention in a complex noise environment. The module combines directional weight allocation, spectrum optimization, and real-time update strategies to achieve effective suppression of multi-directional noise and accurate extraction of the target voice signal, and has broad practical application value.

[0188] Embodiment 2:

[0189] To verify the application effect of the dynamic optimization module of the present invention in an outdoor complex noise environment, the present invention will be specifically described below in combination with a public speaking scenario. In this embodiment, the smart glasses microphone array is used to pick up the voice of the speaker, and at the same time dynamically suppress the ambient noise from multiple directions, including wind noise, traffic noise, and the conversation sounds of the audience, to ensure the clarity and transmission quality of the voice signal.

[0190] Scene description

[0191] A certain user wears smart glasses and gives a public speech outdoors. There are the following noise interferences in the environment:

[0192] Gust noise from the left (about -45° direction), which is low-frequency noise (frequency range 50 Hz to 200 Hz);

[0193] Traffic noise from the right side (at about a 60° direction), which belongs to broadband noise (frequency range from 100 Hz to 2 kHz);

[0194] The conversation sound from the audience directly in front (at about 30° to 45° direction), which belongs to medium-frequency noise (frequency range from 300 Hz to 1.5 kHz).

[0195] The user needs to perform real-time voice recording through smart glasses and synchronously transmit the optimized voice signal to a remote live broadcast system.

[0196] Implementation steps

[0197] 1. Signal acquisition and preprocessing

[0198] The microphone array of the smart glasses simultaneously collects the wearer's voice signal and environmental noise signal. The distribution and functions of the microphone array are as follows:

[0199] The first microphone (single-directional) is located below the wearer's mouth and is dedicated to picking up the wearer's voice;

[0200] The second and third microphones (bi-directional) are located directly in front of the glasses frame and are used to pick up the conversation sound of the audience directly in front and the target voice signal;

[0201] The fourth and fifth microphones (omnidirectional) are located at the rear end of the temple and are used to collect background noise signals (including wind noise and traffic noise).

[0202] The collected mixed signal is preprocessed by the signal processing module to complete analog-to-digital conversion, gain amplification, and filtering operations, filtering out the ultra-low-frequency noise below 50 Hz and the high-frequency interference noise above 10 kHz. The data after signal processing is directly input into the dynamic optimization module.

[0203] 2. Dynamic weight allocation

[0204] The dynamic optimization module combines the noise direction and intensity distribution data provided by the noise modeling module and adjusts the weight allocation according to the real-time environment. The specific weight allocation strategy is as follows:

[0205] The wind noise signal on the left side (-45° direction) is identified as the main low-frequency noise source, and its corresponding weight is significantly reduced;

[0206] The traffic noise on the right side (60° direction) is modeled as broadband noise, and the weight of the corresponding microphone is moderately reduced;

[0207] The conversation sound directly in front (30° to 45°) has a certain interference to the target voice, and the weight is partially suppressed;

[0208] The first microphone has a strong directivity and is located at the wearer's mouth, and its weight is dynamically optimized to the highest value.

[0209] The dynamic optimization module completes the weight calculation through the following weight assignment formula:

[0210]

[0211] Where:

[0212] w i : The directivity weight of the i-th microphone;

[0213] D i : The distance or directivity difference from the i-th microphone to the noise source;

[0214] α: The weight adjustment parameter (set to 1.8 in this embodiment to enhance the directivity difference). The weight of the first microphone is significantly higher than that of other microphones. The second and third microphones are assigned medium weights to assist in picking up the wearer's voice signal. The weights of the fourth and fifth microphones are significantly reduced to maximize the suppression of background noise.

[0215] 3. Dynamic spectrum optimization

[0216] Combined with the spectrum analysis results of the noise reduction module, the dynamic optimization module further adjusts the weight assignment and adopts a differential optimization strategy for noise signals in different frequency bands.

[0217] For low-frequency wind noise (50Hz to 200Hz): The weight is significantly reduced, and filtering processing is used to suppress its interference;

[0218] For broadband traffic noise (100Hz to 2kHz): In this frequency band, the weights of the fourth and fifth microphones are dynamically lowered, while retaining the directivity signal of the first microphone;

[0219] For mid-frequency audience conversation sound (300Hz to 1.5kHz): By increasing the weights of the first, second, and third microphones in this frequency band, the pickup effect of the wearer's voice signal is further enhanced.

[0220] In this embodiment, the time interval of dynamic spectrum optimization is set to 50 milliseconds, ensuring the real-time nature of weight update.

[0221] 4. Real-time weight adjustment and environmental adaptation

[0222] During the speech, the dynamic optimization module adjusts the weights in real time according to the movement and intensity change of the noise source. For example, when the audience conversation sound spreads from the front (30° to 45° direction) to the front right (60° direction), the module updates the directivity distribution weight through real-time monitoring, and further reduces the weight of the microphone in the front right to maintain the high-fidelity pickup of the target voice signal.

[0223] In addition, when the wind speed suddenly increases, resulting in an increase in wind noise on the left side, the dynamic optimization module detects a sudden surge in low-frequency noise energy through spectral analysis and promptly reduces the weight of the microphone in that direction, avoiding the coverage of the target speech by low-frequency noise.

[0224] 5. Output the optimized signal

[0225] The optimized target speech signal is transmitted to the remote live broadcast system in real time, and the background noise is significantly suppressed. The speaker's voice is clear and natural, without obvious low-frequency wind noise or high-frequency traffic noise interference, ensuring that the remote audience can fully understand the speech content.

[0226] Result analysis

[0227] This embodiment verifies the application effect of the dynamic optimization module in a complex outdoor noise environment. The experimental results show that:

[0228] The average energy of the wind noise is suppressed by more than 90%;

[0229] The traffic noise energy is reduced by approximately 75%;

[0230] The signal-to-noise ratio of the target speech signal is increased by 14 dB;

[0231] The response time of the system under dynamic noise changes is less than 50 milliseconds, enabling it to adapt to a rapidly changing environment.

[0232] In summary, this embodiment demonstrates the performance advantages of the dynamic optimization module in a dynamic multi-source noise environment. The module combines a weight allocation strategy, spectral optimization, and real-time adjustment functions to achieve high-quality extraction of the target speech signal and effective suppression of multi-directional noise, making it suitable for practical application scenarios such as public speeches and outdoor live broadcasts, and having high practical value and promotional significance.

[0233] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made therein without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A system for multi-directional sound collection and noise reduction by a smart glasses microphone array, characterized in that, Comprising: At least five microphones, arranged at different positions on the smart glasses frame, where: The first microphone is a unidirectional microphone, located below the glasses frame, for picking up the wearer's voice; The second microphone and the third microphone are bidirectional microphones, located directly in front of the glasses frame, for picking up the voice signals directly in front of the wearer; The fourth microphone and the fifth microphone are omnidirectional microphones, respectively located at the rear ends of the temple arms, for collecting ambient noise; A signal processing module, connected to the microphone array, for receiving and performing analog-to-digital conversion, amplification, and filtering processing on the signals collected by the microphones; A noise modeling module, connected to the signal processing module, for establishing a dynamic distribution model of the noise field; A voice enhancement module, connected to the signal processing module, for extracting target voice signals based on a preset algorithm; A noise reduction module, connected to the noise modeling module and the voice enhancement module, for performing distribution matching processing on the noise signals; A dynamic optimization module, connected to the microphone array, for adjusting the sound collection directivity weights of the microphones.

2. The system for multi-directional sound collection and noise reduction by the intelligent glasses microphone array according to claim 1, characterized in that The signal processing module includes an analog-to-digital converter, a preamplifier, and a filter, for converting analog signals into digital signals and performing amplification and filtering processing on the converted signals.

3. The system for multi-directional sound collection and noise reduction using the intelligent glasses microphone array according to claim 1, wherein The noise modeling module, based on the Markov random field theory, establishes a direction and intensity distribution model of the noise source in the noise field by calculating the conditional probability distribution of each signal in the microphone array.

4. The system for multi-directional sound collection and noise reduction by the intelligent glasses microphone array according to claim 1, characterized in that, The voice enhancement module enhances the target voice signal through the following steps: Separating the target voice signal and the noise signal; Performing optimization processing on the target voice signal based on the signal probability model.

5. The system for multi-directional sound collection and noise reduction by the intelligent glasses microphone array according to claim 1, characterized in that, The noise reduction module adjusts the noise signal distribution to the target voice signal distribution by calculating the distribution difference between the noise signal and the target voice signal and through a specific signal distribution mapping.

6. The system for multi-directional sound collection and noise reduction by the intelligent glasses microphone array according to claim 1, characterized in that, The dynamic optimization module adjusts the signal reception direction of the microphones by calculating the adjustment parameters of the microphone array directivity weights.

7. The system for multi-directional sound collection and noise reduction of the intelligent glasses microphone array according to claim 1, characterized in that The first microphone of the microphone array is a unidirectional microphone, the second microphone and the third microphone are bidirectional microphones, and the fourth microphone and the fifth microphone are omnidirectional microphones, and their specific parameters are used to meet the pickup requirements of signals in different directions.

8. The method for multi-directional sound collection and noise reduction by the intelligent glasses microphone array, according to any one of claims 1-7, characterized in that, Its steps include: Step 1: Collecting audio signals through the microphone array, the first microphone picking up the wearer's voice, the second microphone and the third microphone picking up the voice signals directly in front, and the fourth microphone and the fifth microphone collecting ambient noise signals; Step 2: Using the noise modeling module to establish a dynamic distribution model of the noise field, where the direction and intensity of the noise signal are calculated from the conditional probability distribution; Step 3: Using the voice enhancement module to extract the target voice signal in the observed signal, and the target voice signal and the noise signal are separated by the probability model; Step 4: Using the noise reduction module to perform distribution matching processing on the noise signals and adjusting the distribution of the noise signals to the distribution of the target voice signals; Step 5: Using the dynamic optimization module to adjust the sound collection directivity weights of the microphones to adapt to the direction change of the signal source; Step 6: Outputting the processed target voice signal to the terminal device for subsequent use.

9. The system and method for multi-directional sound collection and noise reduction of the intelligent glasses microphone array according to claim 1, characterized in that, The signal processing module is combined with the noise reduction module to simultaneously complete signal amplification, filtering, analog-to-digital conversion, and matching adjustment of the noise signal.

10. The system and method for multi-directional sound collection and noise reduction of the intelligent glasses microphone array according to claim 1, characterized in that The dynamic optimization module dynamically generates the direction weight adjustment parameters of the microphone by calculating the distribution of the current ambient noise direction in real time.

Citation Information

Cited By

  • Intelligent glasses voice recognition method and system based on multi-channel voice enhancement

    CN120526758A

  • Intelligent glasses

    CN120871437A