A method, apparatus and storage medium for processing an audio signal

CN122845987APending Publication Date: 2026-09-29SHENZHEN XINGUODU JISUAN TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610661647.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-13
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0003]当前常规波束成形收音方案多采用线性、环形等规则化的麦克风阵列布局,在工作过程中,预先设定固定的收音方位作为波束指向角度,仅依据规则阵列的理论间距与平面位置关系求解各路麦克风的声波时延与相位偏移,进而计算得出固定的导向矢量;随后运用通用自适应波束成形算法迭代求解波束成形权重,将所生成的固定权重直接对麦克风阵列同步采集的原始音频信号进行加权叠加处理,仅能针对预先设定的固定方位实现有限的音频信号增强与侧向环境噪声抑制,整个过程不会跟随使用者姿态变化进行波束方向的自适应调整

Benefits of technology

以佩戴者面部朝向确定收音聚焦的目标方向,可随佩戴者头部姿态变化动态调整波束指向,使波束始终对准佩戴者发声区域,克服了固定波束方向无法自适应跟随人脸朝向的弊端,依托麦克风阵列实际三维坐标位置与朝向角度预先构建不规则阵列几何模型,适配可穿戴设备麦克风不规则的实际布设结构,降低导向矢量的计算误差,提升波束指向精准度。在此基础上结合目标方向与导向矢量自适应求解波束成形权重,对麦克风原始采集信号进行加权增强处理,提升目标人声拾取增益,强化非目标方向环境噪声的抑制能力,在佩戴姿态多变、环境噪声复杂的实际场景中,仍可稳定输出高清晰度的增强音频信号,提升可穿戴设备定向收音质量。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845987A_ABST
    Figure CN122845987A_ABST
Patent Text Reader

Abstract

The application discloses a processing method and device of an audio signal and a storage medium, and is used for realizing high-quality audio pickup and noise suppression under an irregular array geometry condition. The method comprises the following steps: determining a face direction of a wearer as a target direction of sound collection focusing; calculating a steering vector of a microphone array of a wearable device based on the target direction through an irregular array geometry model, the irregular array geometry model being a model established in advance based on three-dimensional coordinate positions and orientation angles of the microphone array; calculating beamforming weights by using an adaptive beamforming algorithm based on the steering vector and the target direction; and generating an enhanced beamforming audio signal according to the beamforming weights and original signals collected by the microphone array.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of signal processing technology, and in particular to a method, apparatus and storage medium for processing audio signals. Background Technology

[0002] Currently, most smart wearable devices are equipped with microphone arrays for audio pickup, and beamforming is a key technology for achieving directional sound pickup, environmental noise reduction, and human voice enhancement. In real-world scenarios such as daily calls, on-the-go recording, and near-field audio interaction, the device's wearing posture varies, and the surrounding environment is noisy, which places stringent demands on the accuracy of the microphone array's beam pointing, its ability to track human voices, and its robustness in pickup under irregular deployments.

[0003] Current conventional beamforming sound reception solutions mostly adopt regular microphone array layouts such as linear and circular arrays. During operation, a fixed reception direction is preset as the beam pointing angle. The acoustic delay and phase shift of each microphone are calculated based solely on the theoretical spacing and planar position relationship of the regular array, thereby deriving a fixed steering vector. Subsequently, a general adaptive beamforming algorithm is used to iteratively solve the beamforming weights. The generated fixed weights are directly applied to the weighted superposition of the original audio signals synchronously acquired by the microphone array. This can only achieve limited audio signal enhancement and lateral environmental noise suppression for the preset fixed direction. The entire process does not adaptively adjust the beam direction according to changes in the user's posture.

[0004] However, using only a fixed preset beam direction cannot dynamically adjust the sound focusing direction according to changes in the wearer's facial orientation and head posture. The beam is difficult to align with the sound-emitting area in real time. Furthermore, the lack of an irregular array model that incorporates the microphone's actual three-dimensional coordinates and orientation angle results in large errors in the guide vector calculation and low beam pointing accuracy. Simultaneously, the beam weight adaptability is poor, failing to fit the actual structure of irregularly shaped microphones in wearable devices. When posture changes, the voice pickup gain attenuates significantly, and the environmental noise suppression effect deteriorates, making it difficult to meet the actual needs of high-definition directional sound pickup in dynamic wearing scenarios. Summary of the Invention

[0005] This application discloses an audio signal processing method, apparatus, and storage medium for achieving high-quality audio pickup and noise suppression under irregular array geometry conditions.

[0006] The first aspect of this application discloses a method for processing audio signals, including:

[0007] The direction in which the wearer's face is turned is determined as the target direction for sound focusing; Based on the target direction, the guiding vector of the microphone array of the wearable device is calculated through an irregular array geometric model, which is a model pre-established based on the three-dimensional coordinate position and orientation angle of the microphone array; Based on the guide vector and the target direction, an adaptive beamforming algorithm is used to calculate the beamforming weights; An enhanced beamforming audio signal is generated based on the beamforming weights and the original signal acquired by the microphone array.

[0008] Optionally, calculating the steering vector of the microphone array of the wearable device based on the target direction using an irregular array geometric model includes: Based on the target direction and the three-dimensional coordinates of each microphone in the microphone array of the wearable device, the sound wave propagation delay corresponding to each microphone is calculated. The sound wave propagation delay is the time it takes for the audio signal to propagate from the target direction to each microphone in the microphone array. The propagation delay of the sound wave is converted into a phase shift at the corresponding frequency; A guide vector is constructed based on the phase offset, and the guide vector is a complex vector containing the phase offset components of each microphone.

[0009] Optionally, the step of calculating beamforming weights using an adaptive beamforming algorithm based on the guide vector and the target direction includes: The gain compensation factor corresponding to the target direction is retrieved from the pre-stored directional gain calibration table, and the gain compensation is performed on the steering vector based on the gain compensation factor to obtain the compensated steering vector; Based on the compensation steering vector, a fixed beamformer and a blocking matrix are constructed respectively; Static constraint weights are generated based on the fixed beamformer, and a noise reference signal is obtained based on the blocking matrix. The noise reference signal does not include the audio signal acquired in the target direction. The variable weight coefficients of the adaptive filter are updated based on the noise reference signal; The static constraint weights are fused with the updated variable weight coefficients to obtain the beamforming weights.

[0010] Optionally, the method further includes: The signal output from the bone conduction piezoelectric microphone at the preset position is used as the audio reference signal; Based on the audio reference signal, the current audio is determined to be the wearer's audio; The wearer's audio is separated from the ambient noise based on the audio reference signal.

[0011] Optionally, the method further includes: If the energy of the audio reference signal exceeds a preset threshold, it is determined that the wearer is in an audio activity state, and the update of the variable weight coefficient is stopped.

[0012] Optionally, the method further includes: The beam pattern is calculated based on the beamforming weight vector and the steering vector, and the beam pattern represents the audio signal gain in different directions; The average noise power is obtained by integrating the noise power across the entire angular direction of the beam pattern. The directivity factor is obtained based on the power gain in the target direction of the beam pattern and the average noise power, and the directivity factor is converted into a directivity index. The performance of the enhanced beamforming audio signal is evaluated based on the directivity index, and the evaluation results are output.

[0013] Optionally, the method further includes: A noise feature library is established based on the noise type and noise intensity of the environment in which the microphone array is located. The noise types include steady-state noise, non-steady-state burst noise, and reverberant noise. Identify the dominant noise type in the current environment based on the noise feature library; The parameters of the adaptive beamforming algorithm are adaptively adjusted based on the identified dominant noise type.

[0014] A second aspect of this application provides an audio signal processing apparatus, comprising: A targeting unit is used to determine the wearer's facial orientation as the target direction for sound focusing; The first calculation unit is used to calculate the guiding vector of the microphone array of the wearable device based on the target direction using an irregular array geometric model, wherein the irregular array geometric model is a model pre-established based on the three-dimensional coordinate position and orientation angle of the microphone array. The second calculation unit is used to calculate the beamforming weights based on the guide vector and the target direction using an adaptive beamforming algorithm; The output unit is used to generate an enhanced beamforming audio signal based on the beamforming weights and the original signal acquired by the microphone array.

[0015] A third aspect of this application provides an audio signal processing apparatus, comprising: Processor, memory, input / output units, and bus; The processor is connected to memory, input / output units, and a bus; The memory holds a program, which the processor calls to execute, as in the first aspect and any optional method of the first aspect.

[0016] The fourth aspect of this application provides a computer-readable storage medium on which a program is stored, which, when executed on a computer, performs the methods of the first aspect and any optional methods of the first aspect.

[0017] As can be seen from the above technical solutions, this application has the following advantages: The target direction for sound focusing is determined by the wearer's facial orientation. The beam direction can be dynamically adjusted according to changes in the wearer's head posture, ensuring the beam is always aligned with the wearer's vocal area. This overcomes the drawback of fixed beam directions, which cannot adaptively follow facial orientation. An irregular array geometry model is pre-constructed based on the actual three-dimensional coordinates and orientation angle of the microphone array, adapting to the irregular layout of wearable device microphones, reducing the calculation error of the guide vector, and improving beam pointing accuracy. Furthermore, by adaptively solving the beamforming weights using the target direction and guide vector, the original microphone signal is weighted and enhanced, improving the gain of target voice pickup and strengthening the suppression of ambient noise from non-target directions. Even in real-world scenarios with varying wearing postures and complex environmental noise, it can still stably output high-definition enhanced audio signals, improving the directional sound pickup quality of wearable devices. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 A schematic flowchart of an embodiment of an audio signal processing method provided in this application; Figure 2 A schematic flowchart illustrating an embodiment of calculating the guide vector of a microphone array provided in this application; Figure 3 A schematic diagram of an embodiment for calculating beamforming weights provided in this application; Figure 4 A schematic flowchart illustrating an embodiment of distinguishing between ambient noise and wearer audio provided in this application; Figure 5 A schematic flowchart of an embodiment for evaluating the performance of enhanced beamforming audio signals provided in this application; Figure 6 A schematic flowchart of an embodiment of the adaptive beamforming algorithm parameters provided in this application; Figure 7A structural diagram of an embodiment of an audio signal processing apparatus provided in this application; Figure 8 This is a structural diagram of another embodiment of an audio signal processing apparatus provided in this application. Detailed Implementation

[0020] This application provides an audio signal processing method, apparatus, and storage medium for achieving high-quality audio pickup and noise suppression under irregular array geometry conditions.

[0021] It should be noted that the processing provided in this application can be performed by a wearable device, such as smart glasses.

[0022] Please see Figure 1 This application provides an embodiment of an audio signal processing method, comprising: 101. Determine the direction in which the wearer's face is facing as the target direction for sound focusing; The built-in IMU sensor collects the wearer's head posture data and obtains the wearer's head orientation angle φ_head in real time. This head orientation angle value directly reflects the wearer's face orientation in the horizontal direction.

[0023] The direction of the main lobe of the beam is linked to the wearer's head orientation φ_head. Using this head orientation angle as a reference, the target direction of beamforming is automatically set. Specifically, the horizontal azimuth angle of the wearer's face is mapped to the center direction of the main lobe of the beam, so that the main direction of sound focusing is consistent with the horizontal orientation of the wearer's face, ensuring that the sound reception sensitivity of the microphone array is concentrated in the direction of the user's voice.

[0024] 102. Based on the target direction, calculate the guiding vector of the microphone array of the wearable device through an irregular array geometric model. The irregular array geometric model is a model pre-established based on the three-dimensional coordinate position and orientation angle of the microphone array. An irregular array geometric model is pre-established based on the three-dimensional coordinate position and orientation angle of the microphone array in the wearable device. Specifically, during the device production or calibration stage, the three-dimensional position parameters and signal reception orientation parameters of each microphone in the device coordinate system are collected and stored to construct an irregular array geometric model. This irregular array geometric model can reflect the spatial distribution and signal reception characteristics of the microphone array under actual wearing conditions.

[0025] Once the direction of the target sound source is determined, the irregular array geometric model is invoked, and the steering vector of the microphone array is calculated based on the target sound source direction. Specifically, according to the relative position of the target direction and each microphone, the difference in the propagation path of the target sound source signal to each microphone is obtained, and a steering vector adapted to the current target direction is generated accordingly. This steering vector is generated entirely based on the actual array structure of the device and the target direction, and can truly reflect the propagation differences of the target signal in the array.

[0026] For example, smart glasses have four microphones on the left and right temples and the front of the frame. The three-dimensional coordinates and orientation parameters of the four microphones in the glasses cavity coordinate system are pre-stored. When the target direction is the direction of the user's mouth, the irregular array geometric model is called to calculate the guide vector that is adapted to the target direction.

[0027] Calculate the array steering vector for the target direction θ For an irregular array, the m-th component of the steering vector is:

[0028] in, θ is the gain compensation factor for the m-th microphone in the θ direction, used to compensate for pattern distortion caused by microphone orientation inconsistency and scattering effects of the eyeglass structure; The propagation delay of the sound wave from the θ direction to the m-th microphone is calculated from the three-dimensional coordinate position of the microphone.

[0029] 103. Based on the guide vector and target direction, an adaptive beamforming algorithm is used to calculate the beamforming weights; Based on the steering vector obtained in step 102, and combined with the target direction, an adaptive beamforming algorithm is used to calculate the beamforming weights. This adaptive beamforming algorithm is a generalized sidelobe canceller (GSC) structure based on irregular array geometric constraints. The blocking matrix is ​​reconstructed according to the actual array geometry of the wearable device's microphone array to compensate for the radiation pattern distortion caused by the irregular arrangement. Specifically, taking smart glasses as an example, the original signal collected by the smart glasses' microphone array is first preprocessed, and the covariance matrix is ​​calculated based on the preprocessed signal. Combining the steering vector and target direction constraints, the reconstructed blocking matrix is ​​used to suppress interference signals from non-target directions, and the beamforming weights are obtained. This beamforming weight calculation process uses the lossless reception of the target direction signal as a constraint, and utilizes the spatial distribution information of at least four microphones and the reconstructed blocking matrix to effectively compensate for the radiation pattern distortion caused by the irregular array arrangement, ensuring that the weights can stably achieve the effects of target signal enhancement and interference suppression.

[0030] 104. Generate an enhanced beamforming audio signal based on the beamforming weights and the original signal acquired by the microphone array.

[0031] The beamforming weights calculated in step 103 are weighted and fused with the original signal collected by the microphone array of the wearable device to generate an enhanced beamforming audio signal.

[0032] Each microphone signal is linearly combined with its corresponding beamforming weight to obtain a single-channel enhanced audio signal. Simultaneously, the combined signal undergoes amplitude adjustment and phase correction to ensure distortion-free output. In the generated enhanced beamforming audio signal, the speech signal in the target direction is effectively enhanced by the GSC structure based on irregular array geometric constraints and the reconstructed blocking matrix, while interference signals in non-target directions are suppressed. This also avoids signal distortion caused by irregular array arrangement, resulting in a significantly improved signal-to-noise ratio.

[0033] For example, when a user makes a call in an outdoor environment while wearing smart glasses with no fewer than four irregularly distributed microphones, the generated enhanced beamforming audio signal will show an increase in the user's voice signal strength of more than 15dB, a decrease in background noise intensity of more than 20dB, and no signal distortion caused by pattern distortion, resulting in a significant improvement in call clarity.

[0034] In this embodiment, the target direction for sound focusing is determined by the wearer's facial orientation. The beam direction can be dynamically adjusted according to changes in the wearer's head posture, ensuring the beam is always aligned with the wearer's vocal area. This overcomes the drawback of fixed beam directions failing to adaptively follow facial orientation. An irregular array geometric model is pre-constructed based on the actual three-dimensional coordinates and orientation angle of the microphone array, adapting to the irregular layout of wearable device microphones, reducing the calculation error of the guide vector, and improving beam pointing accuracy. Furthermore, by adaptively solving the beamforming weights using the target direction and guide vector, the original microphone signal is weighted and enhanced, improving the target voice pickup gain and strengthening the suppression of ambient noise from non-target directions. Even in real-world scenarios with varying wearing postures and complex environmental noise, a stable high-definition enhanced audio signal can still be output, improving the directional sound reception quality of the wearable device.

[0035] In step 102 above, the steering vector of the microphone array of the wearable device is calculated based on the target direction using an irregular array geometric model. (See [link to relevant documentation]). Figure 2 , Figure 2 One embodiment of the computational microphone array steering vector provided in this application includes: 201. Based on the target direction and the three-dimensional coordinates of each microphone in the microphone array of the wearable device, calculate the sound wave propagation delay corresponding to each microphone. The sound wave propagation delay is the time it takes for the audio signal to propagate from the target direction to each microphone in the microphone array. Taking smart glasses as an example, based on the device coordinate system of the smart glasses, the three-dimensional coordinate parameters of no less than four microphones in the microphone array are pre-stored. These three-dimensional coordinate parameters are the actual installation position coordinates of each microphone on the device body, which can truly reflect the spatial distribution of the microphone array when the user is wearing it. When the direction of the target sound source is determined, the target direction parameter and the three-dimensional coordinate parameters of each microphone are called to calculate the sound wave propagation delay required for the audio signal to travel from the direction of the target sound source to each microphone.

[0036] First, the propagation direction vector of the incident sound wave is determined based on the direction of the target sound source. Then, the projected distance of each microphone relative to the origin of the device coordinate system in this propagation direction is calculated. This projected distance is applied to the speed of sound in air to obtain the sound wave propagation delay of the audio signal from the target sound source direction to the corresponding microphone. This sound wave propagation delay can accurately reflect the difference in the order in which the target sound source signal arrives at each microphone.

[0037] The formula for calculating the sound wave propagation delay for each microphone is as follows:

[0038] Where c = 343 m / s is the speed of sound (coordinate unit is mm, speed of sound is converted to 343000 mm / s). This represents the time lead of the sound wave to the m-th microphone relative to the origin. A positive value indicates that the sound wave arrives at the microphone first, and a negative value indicates that it arrives later. The incident angle of the target sound source; The unit vector in the direction of the target; = Let be the position vector of the m-th microphone in the device coordinate system.

[0039] Numerical examples ( =0, sound waves coming from directly in front): =(-65×0+35×1) / 343000=+102μs; =(-72×0+(-35)×1) / 343000=-102μs; =(65×0+35×1) / 343000=+102μs; =(72×0+(-35)×1) / 343000=-102μs.

[0040] Taking the four-microphone array of smart glasses as an example, for signals directly in front, the top row of microphones receives the signal 1 / 3 first (+102μs), and the bottom row of microphones receives the signal 2 / 4 later (-102μs), which is symmetrical from left to right - this perfectly matches the physical position of the four microphones.

[0041] When the target sound source direction is the direction of the user's mouth, the sound wave propagation delay corresponding to each microphone is calculated based on the three-dimensional coordinates of the target direction and the microphone. The microphone closer to the user's mouth has a smaller sound wave propagation delay, while the microphone farther away from the user's mouth has a larger sound wave propagation delay.

[0042] 202. Convert the sound wave propagation delay into a phase shift at the corresponding frequency; 203. Construct a steering vector based on phase offset. The steering vector is a complex vector containing the phase offset components of each microphone.

[0043] In step 202, after determining the target center frequency based on the operating frequency band of the audio signal, the phase offset corresponding to each microphone is calculated based on the sound wave propagation delay and the target center frequency. The phase offset calculation uses the signal phase at the origin of the device coordinate system as a reference. The difference in sound wave propagation delay will be directly converted into different phase offsets. This phase offset can reflect the phase difference when the target sound source signal arrives at each microphone.

[0044] When the target center frequency of the audio signal is 1kHz, the time delay difference is converted into the corresponding phase shift based on the sound wave propagation delay corresponding to each microphone. The smaller the sound wave propagation delay, the closer the corresponding phase shift is to the reference phase. The larger the sound wave propagation delay, the greater the difference between the corresponding phase shift and the reference phase.

[0045] In step 203, based on the phase offsets of each microphone obtained in step 202, a steering vector for the microphone array is constructed. Specifically, the phase offset of each microphone is converted into a complex phase component, which is represented in terms of unit amplitude and corresponding phase offset. The complex phase components of all microphones are arranged in array order to construct a complex vector containing the phase offset components of each microphone, i.e., the steering vector.

[0046] Each component This indicates that at frequency f, the m-th microphone experiences a propagation delay. The resulting phase shift.

[0047] This steering vector can accurately reflect the phase distribution characteristics of the target sound source signal in the microphone array. Taking a four-microphone array of smart glasses as an example, four complex phase components are generated based on the phase offsets corresponding to the four microphones. These four complex phase components are arranged in a preset order to construct a four-dimensional complex vector steering vector. This steering vector can adapt to the direction of the target sound source and accurately characterize the phase difference of the target signal propagation in the array.

[0048] In this embodiment, the sound wave propagation delay is calculated based on the actual three-dimensional coordinates of the microphone array and the target direction. This perfectly matches the actual spatial distribution of the irregular array of smart glasses and the user's wearing status, truly restoring the difference in the propagation path of the target sound source signal in the array, and avoiding the adaptation deviation of the traditional regular array model for asymmetric layout microphones.

[0049] By converting the time-domain propagation delay into the frequency-domain phase offset and constructing a steering vector based on the phase offset, a precise mapping from signal propagation differences to array phase characteristics is achieved. This steering vector can accurately characterize the phase distribution characteristics of the target sound source signal in an irregular array, solving the problem of radiation pattern distortion caused by irregular microphone array arrangement, and ensuring the stability of the beamforming algorithm and the target signal enhancement effect.

[0050] In step 103 above, the beamforming weights are calculated using an adaptive beamforming algorithm based on the guide vector and the target direction. Please refer to [link to relevant documentation]. Figure 3 , Figure 3 One embodiment of calculating beamforming weights provided in this application includes: 301. Query the gain compensation factor corresponding to the target direction from the pre-stored directional gain calibration table, and perform gain compensation on the steering vector based on the gain compensation factor to obtain the compensated steering vector; Taking smart glasses as an example, the smart glasses frame and temples are equipped with no fewer than four microphones arranged in an irregular pattern. During the factory calibration stage, the device completes acoustic gain calibration for sound sources with different incident angles, establishes and stores a directional gain calibration table, which records the gain compensation factors corresponding to different spatial orientations, in order to correct the sound source gain deviation caused by the irregular array structure and the obstruction of the glasses shell.

[0051] After determining the target direction corresponding to the current voice, the pre-stored directional gain calibration table is retrieved, and the gain compensation factor corresponding to the target direction is obtained by matching and querying. The amplitude gain of the steering vector obtained in the previous step is corrected in combination with the gain compensation factor to offset the signal gain inconsistency caused by irregular microphone arrangement, installation position differences, and acoustic obstruction of the shell. Thus, the corrected compensated steering vector is obtained, which can better fit the spatial acoustic response characteristics of smart glasses under actual wearing conditions and ensure the array response consistency of target sound sources under different incident directions.

[0052] Furthermore, considering that differences in user wearing posture and head shape may lead to deviations in the factory-calibrated gain compensation factor, this embodiment introduces an online calibration mechanism based on bone conduction VAD. During daily use, the bone conduction VAD is used to confirm the wearer's speech period. The wearer's voice is used as a reference signal with a known direction (directly in front), and the least mean square (LMS) algorithm is used to update the gain compensation factor of each microphone online. The update rate of the calibration process is limited by the duration of speech activity, and convergence is usually completed after accumulating 10 minutes of effective use. After calibration, the pointing accuracy of the main lobe of the beam can be improved from ±5° of the factory calibration to ±2°, correcting the acoustic response deviation caused by individual wearing differences. This ensures that the gain compensation factor always matches the user's actual wearing conditions, avoiding adaptation errors between different users with fixed factory parameters.

[0053] 302. Based on the compensation steering vector, construct a fixed beamformer and a blocking matrix respectively; A generalized sidelobe canceller (GSC) architecture based on irregular array geometric constraints is adopted. After obtaining the compensation steering vector, the fixed beamformer and the blocking matrix are constructed based on the compensation steering vector accurately corrected in step 301. The fixed beamformer is constructed with the constraint of distortion-free reception in the target direction to ensure that the human voice signal in the target direction is completely preserved without phase distortion. At the same time, the blocking matrix is ​​re-customized and constructed in combination with the actual irregular geometric arrangement characteristics of the smart glasses microphone array. The blocking matrix is ​​constrained and optimized according to the spatial position relationship of the array, and is specifically used to compensate for the inherent radiation pattern distortion defects of the irregular array, weaken the beam distortion and sidelobe rise problems caused by the irregular array arrangement, so that both the fixed beamformer and the blocking matrix are adapted to the unique irregular microphone array structure of this device.

[0054] The output expression of a fixed beamformer (FBF) is:

[0055] in, These are the static constraint weights based on the compensation guide vector; This is the original input signal vector of the microphone array.

[0056] 303. Static constraint weights are generated based on a fixed beamformer, and a noise reference signal is obtained based on a blocking matrix. The noise reference signal does not include the audio signal acquired in the target direction. A constant static constraint weight is generated by the completed fixed beamformer. This static constraint weight is used to lock the main lobe of the beam in the target direction, ensuring that the target human voice signal is output stably and will not be distorted by changes in the external noise environment. At the same time, the original audio signals from multiple microphones are blocked by a customized reconstructed blocking matrix to filter out and remove the effective human voice signal incident in the target direction, while retaining non-target direction interference signals such as environmental noise and lateral interference, and finally obtaining a pure noise reference signal that does not contain audio components in the target direction.

[0057] 304. Updating the variable weight coefficients of the adaptive filter based on the noisy reference signal; After obtaining the noise reference signal, it is input into an adaptive filter, and an adaptive iterative algorithm (such as normalized least mean square) is used to update the variable weight coefficients inside the filter in real time. During the iterative update process, the minimum output power of the noise reference signal is continuously used as the convergence criterion, and the variable weight coefficients are continuously optimized to track the real-time changing environmental noise. This allows for dynamic adjustment of the suppression intensity for different types of interference noise, such as outdoor wind noise, ambient human voices, and background traffic noise, continuously canceling residual interference components. This further adapts the smart glasses to the complex and ever-changing daily use environment and improves the algorithm's ability to adaptively follow dynamic noise.

[0058] 305. The static constraint weights are fused with the updated variable weight coefficients to obtain the beamforming weights.

[0059] The static constraint weights obtained in the preceding steps are weighted and fused with the variable weight coefficients optimized in real time. By combining the fixed beam constraint characteristics and adaptive noise suppression characteristics, the final output is a beamforming weight suitable for irregular microphone arrays. The static constraint weights ensure distortion-free and stable enhancement of the target human voice, while the variable weight coefficients are responsible for dynamically suppressing various environmental interference noises. The fusion of the two retains the stable beam directivity of the generalized sidelobe canceller structure and compensates for the defects of directional distortion of irregular arrays and abrupt changes in environmental noise. This allows the final synthesized beamforming weight to simultaneously ensure the fidelity of the target speech and the ability to suppress external interference.

[0060] In this embodiment, the guide vector is gain compensated using a pre-stored directional gain calibration table. This corrects the sound source gain deviation caused by irregular microphone array arrangement, acoustic occlusion of the shell, and individual wearing posture differences in wearable devices, enabling the compensated guide vector to accurately reflect the spatial acoustic response characteristics of the device under actual operating conditions. The fixed beamformer and blocking matrix constructed based on the compensated guide vector not only lock the main lobe of the target direction beam through static constraint weights, ensuring distortion-free and stable enhancement of the target speech signal, but also effectively filter out the audio components in the target direction through the blocking matrix adapted to the geometric constraints of the irregular array, obtaining a clean noise reference signal. This solves the problems of pattern distortion and sidelobe rise in traditional regular array blocking matrices under irregular arrangement. By integrating static constraint weights with dynamically updated variable weight coefficients, the fidelity of the target speech signal and the adaptive suppression capability of environmental noise are taken into account. It can dynamically adjust the suppression intensity for different interference noises such as wind noise, ambient human voices, and background traffic noise in the complex and ever-changing daily use scenarios of wearable devices, improving the beamforming algorithm's tracking capability for dynamic noise and overall anti-interference performance.

[0061] In practical use, wearable devices are prone to aliasing between environmental noise and the wearer's voice signal. Conventional beamforming algorithms update weight coefficients without reference signal constraints, making them susceptible to beam deviation from the target direction due to noise interference, thus affecting the voice enhancement effect. Furthermore, continuously adjusting weights while the wearer is speaking can distort the target voice waveform. Therefore, this embodiment introduces an audio reference signal from a bone conduction piezoelectric microphone to distinguish between the wearer's voice and environmental noise, and controls the timing of weight updates based on the user's audio activity status. Please refer to [link to relevant documentation]. Figure 4 , Figure 4 One embodiment of the present application for distinguishing between ambient noise and wearer audio includes: 401. Use the signal output from the bone conduction piezoelectric microphone at the preset position as the audio reference signal; Taking smart glasses as an example, a bone conduction piezoelectric microphone is fixedly installed at a preset position in the smart glasses. This microphone uses the piezoelectric sensing principle and can directly collect the vibration signal transmitted by the bones when the user speaks. Its signal transmission path is completely enclosed inside the human body tissue and is not affected by external airborne noise, shell wind noise, environmental human voices, etc. The output bone conduction signal only contains the wearer's own voice component, which can truly and purely reflect the wearer's voice state. Moreover, it has a very strong correlation with the user's voice component in the mixed signal collected by the air conduction microphone array. Therefore, it is used as a dedicated audio reference signal to distinguish the wearer's audio from environmental noise.

[0062] For example, a piezoelectric bone conduction microphone (resonant frequency 2kHz, sensitivity -30dB) is placed at the left temple to pick up the bone conduction vibration signal when the wearer speaks through skin contact. This signal has the following characteristics: it is not affected by ambient noise propagating through the air; the spectrum is concentrated in the 300Hz~3400Hz speech frequency band; and the SNR is >30dB when the wearer speaks.

[0063] 402. Determine that the current audio is the wearer's audio based on the audio reference signal; The system acquires audio reference signals from the bone conduction piezoelectric microphone in real time. By performing simple yet accurate analysis of the time and frequency domain characteristics of this signal, it can quickly determine whether the current audio is emitted by the wearer. Specifically, the audio reference signal undergoes short-time frame processing, calculating the short-time energy and zero-crossing rate of each frame. Combined with a preset wearer voice characteristic threshold, a preliminary determination is made. Since bone conduction signals only exhibit significant energy peaks when the user speaks, and the spectrum is concentrated in the 300Hz-3kHz human voice frequency band, while environmental noise cannot be transmitted to the bone conduction microphone through the bones, the system can accurately determine whether the current audio is the wearer's audio when the short-time energy of the audio reference signal exceeds the preset characteristic threshold and the spectrum matches the human voice distribution; otherwise, it is determined to be audio without a wearer.

[0064] 403. Separating wearer audio from ambient noise based on audio reference signals; The audio reference signal undergoes preprocessing, including DC removal, filtering, and gain normalization, to eliminate weak interference from the bone conduction signal itself and ensure its purity. The preprocessed bone conduction reference signal is then synchronized with multiple original mixed signals acquired by an air conduction microphone array. The cross-correlation coefficient between each mixed signal and the bone conduction reference signal is calculated. Based on the cross-correlation coefficient, the wearer's audio component and environmental noise component are initially distinguished. Components with a cross-correlation coefficient greater than a preset threshold (e.g., 0.7) are identified as audio components strongly correlated with the wearer's speech, while those with a cross-correlation coefficient less than the preset threshold are identified as environmental noise components. Building upon this, an adaptive separation model based on the reference signal is constructed. Using the bone conduction reference signal as a constraint, an iterative optimization algorithm further extracts weak wearer speech components masked by strong noise from the mixed signal. Simultaneously, the separated environmental noise components are initially suppressed to ensure that the separated wearer's audio signal is free from noise pollution and waveform distortion, and that the separated environmental noise components accurately reflect the interference of the current environment.

[0065] For example, when a user wears smart glasses and makes a call in a noisy outdoor environment (such as traffic noise or crowd noise), the wearer's voice is masked by a large amount of environmental noise in the mixed signal collected by the air conduction microphone array. Through correlation analysis between the bone conduction reference signal and the mixed signal, the wearer's voice component is accurately extracted, and environmental noise such as traffic noise and crowd noise is separated. Even in strong noise environments with a signal-to-noise ratio of less than 10dB, effective separation of the two can be achieved, avoiding beamforming weight calculation deviations caused by noise and voice aliasing. At the same time, this separation process balances real-time performance and accuracy, adapts to the lightweight hardware characteristics of smart glasses, and does not increase the device's computational burden excessively.

[0066] 404. If the energy of the audio reference signal exceeds the preset threshold, it is determined that the wearer is in an audio activity state, and the variable weight coefficients are stopped from being updated.

[0067] It monitors short-term energy changes of bone conduction audio reference signals in real time, and pre-sets a reasonable energy threshold based on the static noise level of smart glasses and the typical energy range when the user speaks normally. This threshold can be adaptively fine-tuned according to the user's usage habits to ensure the accuracy of the judgment.

[0068] The preset threshold is typically set to 2-3 times the static noise energy of the bone conduction signal. This avoids misjudgments caused by slight vibrations (such as chewing or turning the head) while accurately capturing the user's vocal movements. When the short-term energy of the audio reference signal continuously exceeds the preset threshold, it is determined that the wearer is currently in an audio activity state, i.e., speaking, making a call, or issuing a voice command. At this time, if the variable weight coefficients of the adaptive filter continue to be updated, it will cause dynamic fluctuations in the beamforming weights, resulting in distortion and phase shift of the target speech signal waveform, affecting the clarity of calls and voice interactions. Therefore, a weight update stop command is triggered, controlling the adaptive filter to stop updating the variable weight coefficients, maintaining the current beamforming weights stable, and ensuring that the wearer's speech signal can be output through beamforming processing without distortion.

[0069] When the energy of the audio reference signal continuously falls below the preset threshold and remains below the preset duration (e.g., 100ms), it is determined that the wearer has stopped speaking and is in a non-audio activity state. At this time, the update of the variable weight coefficient is resumed, so that the adaptive filter can continuously track the changes in the current environmental noise, dynamically optimize the weight coefficient, continuously suppress environmental noise, and ensure that the beamforming algorithm can always maintain good anti-interference performance during non-speech periods.

[0070] For example, when a user makes a call while wearing smart glasses, if the bone conduction signal energy exceeds a preset threshold at the moment of speaking, the system immediately stops weight updates to avoid voice distortion. When the user pauses to breathe, the bone conduction signal energy drops back below the threshold, weight updates resume, and background noise is continuously suppressed. This ensures the integrity and clarity of the voice during the call while also achieving continuous noise suppression during non-voice periods, thus improving the audio experience of smart glasses in complex environments.

[0071] In this embodiment, by placing bone conduction piezoelectric microphones at preset locations on the wearable device, their output signals are used as dedicated audio reference signals. The inherent characteristic of bone conduction signals—transmitted solely by the vibration of human bones and unaffected by external air noise—allows for the acquisition of a reference signal containing only the wearer's own vocal components. This avoids the discrimination errors caused by the mixing of environmental noise and the wearer's voice at the source. Relying on this audio reference signal, the currently valid audio can be accurately identified as the wearer's own audio, further achieving effective separation of the wearer's audio from environmental noise, clearly distinguishing between valid speech components and external interference noise components.

[0072] Simultaneously, by detecting the audio reference signal energy in real time and comparing it with a preset threshold, the system accurately determines the wearer's audio activity state. During the wearer's vocalization, the update process of the variable weight coefficients is promptly stopped, avoiding the continuous iteration of weights in conventional beamforming algorithms when speech is present. This prevents beam pointing deviation and distortion of the target speech waveform caused by noise interference. During the intervals when the wearer is not speaking, the weight coefficients can be updated adaptively, continuously tracking and suppressing dynamic environmental noise. This ensures both the complete and faithful output of the wearer's speech signal and the adaptive suppression capability for environmental noise in complex scenarios.

[0073] To quantify the enhancement effect and spatial anti-interference capability of the processing method on the audio signal, a beam pattern and directivity index evaluation mechanism is further introduced after generating the enhanced beamforming audio signal to systematically verify the beamforming performance. Please refer to [link / reference]. Figure 5 , Figure 5 An embodiment for evaluating the performance of enhanced beamforming audio signals provided in this application includes: 501. Calculate the beam pattern based on the beamforming weight vector and the steering vector. The beam pattern represents the audio signal gain in different directions. The formula for calculating the beam pattern is:

[0074] in, It is the beamformer weight vector (calculated by the SpeexDSPMDF or MVDR algorithm); It is a guide vector; This indicates the conjugate transpose.

[0075] Beam pattern Modulus square This refers to the power gain at frequency f and direction θ, which intuitively characterizes the gain characteristics of the beamformer for audio signals in different spatial directions.

[0076] The beam pattern generated by the combined action of the beamforming weight vector and the steering vector, in the target direction Satisfying the distortion-free constraint This ensures that the target signal passes through uninterrupted, while exhibiting low gain characteristics in non-target directions, clearly demonstrating the spatial directivity and interference suppression effect of the beamformer.

[0077] For example, in the context of smart glasses calls, the beam pattern forms a high-gain main lobe in the direction of the user's mouth and a low-gain side lobe in the direction of ambient noise incidence, such as the sides and rear, which intuitively reflects the signal enhancement and noise suppression capabilities of the algorithm.

[0078] 502. Integrate the noise power across the entire beam pattern to obtain the average noise power; Based on the beam pattern, the noise power across all angular directions is integrated to evaluate the overall noise suppression capability of the beamformer. Specifically, the average noise power remaining after isotropic diffused noise passes through the beamformer is calculated using the following formula:

[0079] This integral operation represents the average power remaining after a uniform noise field from all directions passes through the beamformer. The smaller the value, the narrower the beam, the better the spatial selectivity, and the stronger the suppression effect on noise from non-target directions.

[0080] For example, in the scenario of outdoor calls using smart glasses, the system performs full-angle noise power integration on the beam pattern and obtains a low average noise power value, indicating that the algorithm has a strong suppression effect on background noise that is uniformly distributed in the environment.

[0081] 503. Based on the power gain and average noise power in the target direction of the beam pattern, the directivity factor is obtained, and the directivity factor is converted into the directivity index; First, extract the target direction from the beam pattern. Power gain on Combined with the average noise power calculated in step 502 The directional factor is obtained through ratio calculation, and the formula is as follows:

[0082] Since the power gain in the target direction satisfies the distortion-free constraint Therefore, the directional factor can be simplified to The ratio of the signal power in the target direction to the average noise power in all directions directly reflects the overall performance of the beamformer in terms of signal enhancement and noise suppression.

[0083] The directional factor is converted into a directional index in decibels using the following formula:

[0084] It facilitates a direct comparison of performance differences under different algorithms or operating conditions.

[0085] For example, if the average noise power is 0.1, the directivity factor is 10, and the corresponding directivity index is 10dB, indicating that the target signal power is 10 times the average noise power, and the beamformer has good spatial anti-interference capability.

[0086] 504. Evaluate the performance of enhanced beamforming audio signals based on the directivity index and output the evaluation results.

[0087] Based on the calculated directivity index, the performance of enhanced beamforming audio signals is systematically evaluated, and the evaluation results are output. During the evaluation process, different directivity index thresholds corresponding to different performance levels can be preset, for example: A directional index ≥10dB is considered excellent. A value of 5dB ≤ Directional Index < 10dB is considered good. Directional index <5dB indicates a need for optimization.

[0088] The system automatically matches the corresponding performance level based on the actual calculation results and generates an evaluation report that includes the directional index value, performance level, and optimization suggestions. This evaluation result can be used for algorithm iteration and optimization, performance comparison under different operating conditions, or as feedback on user experience, providing quantitative support for the continuous optimization of beamforming algorithms for smart glasses.

[0089] For example, when a user is making a call while wearing smart glasses in a noisy environment, the calculated directivity index is 12dB, which is considered excellent, indicating that the enhanced beamforming audio signal has a high signal-to-noise ratio and clear call quality. If the directivity index is 4dB, it is determined that optimization is needed, and the algorithm parameter adjustment process can be automatically triggered to improve the spatial selectivity and noise suppression capability of the beamformer.

[0090] In this embodiment, a beam pattern is constructed using the beamforming weight vector and the steering vector. This allows for a direct representation of the audio signal gain characteristics of the beamformer in different spatial directions, accurately reflecting the distortion-free response of the main lobe in the target direction and the noise suppression effect of the side lobes in non-target directions. Integrating the noise power across the entire beam pattern yields the average noise power, which objectively quantifies the beamformer's overall suppression capability against diffused noise fields. A smaller value indicates stronger spatial selectivity and better anti-interference performance, providing reliable benchmark data for calculating the directivity factor. Based on this, a directivity factor is constructed using the target direction power gain and average noise power, and converted into a directivity index. This index is then presented in decibels as a direct representation of the ratio of the target signal power to the average noise power across all directions. This achieves standardized, comparable, and quantitative evaluation of beamforming performance, facilitating performance difference analysis and iterative optimization under different algorithms and operating conditions.

[0091] Meanwhile, an online calibration mechanism for directional gain factors is introduced. To address the factory calibration deviation caused by differences in user wearing posture and head shape, the bone conduction VAD is used to confirm the wearer's speech period. Using the user's voice as a known directional reference signal, the directional gain factors of each microphone are updated online. After calibration, the main lobe pointing accuracy is significantly improved, correcting the acoustic response deviation caused by individual wearing differences and avoiding the adaptation error of fixed factory parameters.

[0092] Because different users experience variations in their wearing environments, noise types, and intensities, fixed-parameter adaptive beamforming algorithms cannot adapt to complex acoustic scenarios. Furthermore, the environmental noise characteristics of the same user can change under different usage scenarios. Using fixed parameters for beamforming can easily lead to poor noise suppression and unstable target speech enhancement, thus affecting the clarity and signal-to-noise ratio of the speech signal. Therefore, by constructing a noise feature library and identifying the dominant noise type, beamforming algorithm parameters can be adaptively adjusted to adapt to different noise scenarios and varying user conditions. Please refer to [link to relevant documentation]. Figure 6 , Figure 6 An embodiment of the adaptive beamforming algorithm parameters provided in this application includes: 601. Establish a noise feature library based on the noise type and noise intensity of the environment in which the microphone array is located. The noise types include steady-state noise, non-steady-state burst noise, and reverberant noise. The pre-established noise feature library categorizes different environmental noises based on their statistical characteristics and spectral features, primarily including three types: steady-state noise, unsteady-state sudden noise, and reverberant noise. Steady-state noise includes continuous noise with a stable spectrum, such as air conditioner noise and background fan noise; unsteady-state sudden noise includes noise with short-term energy abrupt changes, such as keyboard typing and vehicle horn sounds; and reverberant noise is multipath interference noise formed by reflections within the indoor space. Before leaving the factory, the equipment undergoes multi-scenario acoustic calibration to collect various noise samples, extracting their spectral characteristics, short-term energy distribution, zero-crossing rate, and other parameters, which are then stored in the noise feature library. Simultaneously, during daily use, new environmental noise data is continuously collected to supplement and update the feature library, ensuring that it covers various typical usage scenarios.

[0093] 602. Identify the dominant noise type in the current environment based on the noise feature database; While the microphone array acquires audio signals, the characteristic parameters of the current environmental noise are extracted in real time and compared with samples in the noise feature library to identify the dominant noise type corresponding to the current environment. In specific implementation, the acquired audio signal is first processed by frame segmentation, and features such as spectrum, short-time energy, and autocorrelation coefficient of each frame are extracted. Similarity matching is performed with the feature templates of various types of noise in the noise feature library, and the category with the highest matching degree is selected as the dominant noise type of the current environment.

[0094] For example, when a user is making a call on an outdoor street while wearing smart glasses, the background noise spectrum shows a continuous and stable low-frequency distribution, which highly matches the steady-state traffic sound characteristics in the noise feature library. Therefore, the current dominant noise type is identified as steady-state noise. When the user is in an indoor multi-person conversation scenario, the background noise contains human voice segments with short-term energy changes, and the dominant noise type is identified as non-steady-state burst noise.

[0095] 603. Adaptively adjust the parameters of the adaptive beamforming algorithm based on the identified dominant noise type.

[0096] Based on the dominant noise type identified in step 602, the key parameters of the adaptive beamforming algorithm are dynamically adjusted to achieve targeted noise suppression.

[0097] When the dominant noise is steady-state noise, the algorithm adjusts the step size parameter of the filter to reduce the weight update rate and avoid over-suppression caused by the noise being stationary, while enhancing the tracking ability of non-stationary speech signals. When the dominant noise is non-stationary burst noise, the algorithm increases the adaptive step size to accelerate the update speed of the filter weights and quickly suppress the energy peak of the burst noise. When the dominant noise is reverberant noise, the spatial filtering parameters of beamforming are adjusted to enhance the suppression ability of multipath reflection signals and increase the gain of the direct sound component.

[0098] For example, when the current environment is identified as steady-state traffic noise, the step size of the adaptive filter is set to a small value, so that the algorithm focuses on suppressing continuous low-frequency background noise while preserving the high-frequency details of the user's speech; when sudden keyboard noise is identified, the step size parameter is automatically increased, and the filter responds quickly to the noise change, strengthening the suppression effect within the short window of noise occurrence, and avoiding noise interference with the user's speech signal.

[0099] In this embodiment, by constructing a noise feature library that includes steady-state noise, non-steady-state sudden noise, and reverberant noise, it can comprehensively cover various typical environmental noise conditions encountered in the daily use of wearable devices. Feature deposition is completed based on noise type and noise intensity dimensions, providing a complete benchmark for accurate identification of environmental noise. Calling the noise feature library to match and judge the current environment can quickly identify the dominant noise type in the scene, effectively distinguishing the spectral characteristics, energy change patterns, and spatial propagation characteristics of different noises, avoiding the limitations of a single noise reduction strategy in complex and mixed noise scenarios. Based on this, the core parameters of the adaptive beamforming algorithm are adaptively adjusted according to the dominant noise type. Differentiated filtering convergence step sizes, spatial constraints, and weight update mechanisms are adopted for different noise characteristics, dynamically adapting the algorithm's working mode according to the environmental noise type. This improves noise suppression capabilities and speech enhancement effects in complex and variable scenarios while ensuring the integrity and distortion-free speech of the target wearer.

[0100] The adaptive beamforming algorithm further combines beam tracking mode with multi-person dialogue scenario enhancement strategies to achieve voice enhancement optimization in all scenarios.

[0101] The beam tracking mode is jointly controlled by the IMU motion state and the bone conduction VAD to achieve accurate tracking and stable output of the main lobe beam. Natural Tracking Mode: When the IMU detects slow head rotation (angular velocity < 30° / s), the main lobe of the beam smoothly follows the head yaw angle change. The tracking process uses a first-order low-pass filter with a time constant of 200ms. This smooth transition avoids audio artifacts caused by sudden changes in beam direction, ensuring that the voice signal remains at high gain reception when the user slowly turns their head, thus improving the naturalness of the wearing experience.

[0102] Locked Mode: When the bone conduction VAD detects that the wearer is speaking or in a call, the beam direction is locked in the current direction and does not follow head rotation. This design is based on a user behavior model. During a conversation, the user is usually facing the other party, and small head movements will not cause beam shift, thus avoiding distortion of the voice signal due to beam movement and ensuring stable voice enhancement during the call.

[0103] Scanning mode: When the IMU detects rapid head rotation (angular velocity > 60° / s) and the bone conduction VAD does not detect speech, it determines that the user is turning their head to find a sound source. At this time, the beam follows the head direction at a rate of 120° / s and simultaneously starts the sound source localization based on signal power (SRP-PHAT algorithm). After the head rotation ends, the beam is automatically locked to the direction of the strongest sound source, realizing rapid sound source localization and beam alignment.

[0104] In multi-person dialogue scenarios, the spatial diversity advantage of irregular arrays is utilized for enhancement processing. The specific steps are as follows: (1) Based on the 4-channel DMIC signal, the SRP-PHAT algorithm is used to perform sound source localization scanning in 5° steps within the range of -90° to +90° on the horizontal plane, identify the direction angle of all active sound sources, and obtain the spatial location information of each speaker.

[0105] (2) For each detected sound source direction, calculate the corresponding beam weight vector to form multiple beam outputs. Each beam points to a different sound source direction to achieve synchronous enhancement of multi-directional speech signals.

[0106] (3) Distinguish the wearer’s speech from other speakers’ speech by using bone conduction reference signals: The wearer’s speech is extracted by the combined enhancement of the main beam (following the head direction) and bone conduction reference signals, with a signal-to-noise ratio improvement of up to 12dB; the speech of other speakers is extracted by auxiliary beams in their respective directions, thus achieving the separation of different speakers’ speech.

[0107] (4) The separated multiple voices are sent to the speech recognition engine to realize simultaneous transcription of multiple speakers, support speech recognition and content recording in multi-person dialogue scenarios, and improve the practical functions of smart glasses in scenarios such as meetings and multi-person conversations.

[0108] Please see Figure 7 This application provides an embodiment of an audio signal processing apparatus, comprising: The determining unit 701 is used to determine the wearer's facial orientation as the target direction for sound focusing; The first calculation unit 702 is used to calculate the guiding vector of the microphone array of the wearable device based on the target direction using an irregular array geometric model. The irregular array geometric model is a model pre-established based on the three-dimensional coordinate position and orientation angle of the microphone array. Optionally, the first computing unit 702 is also used for: Based on the target direction and the three-dimensional coordinates of each microphone in the microphone array of the wearable device, calculate the sound wave propagation delay corresponding to each microphone. The sound wave propagation delay is the time it takes for the audio signal to propagate from the target direction to each microphone in the microphone array. Convert the propagation delay of sound waves into a phase shift at the corresponding frequency; A guide vector is constructed based on the phase offset, and the guide vector is a complex vector containing the phase offset components of each microphone.

[0109] The second calculation unit 703 is used to calculate beamforming weights based on the guide vector and the target direction using an adaptive beamforming algorithm; Optionally, the second computing unit 703 is also used for: The gain compensation factor corresponding to the target direction is queried from the pre-stored directional gain calibration table, and the gain compensation is performed on the steering vector based on the gain compensation factor to obtain the compensated steering vector; Based on the compensation steering vector, a fixed beamformer and a blocking matrix are constructed respectively; Static constraint weights are generated based on a fixed beamformer, and a noise reference signal is obtained based on a blocking matrix. The noise reference signal does not include the audio signal acquired in the target direction. The variable weight coefficients of the adaptive filter are updated based on the noisy reference signal; The beamforming weights are obtained by fusing the static constraint weights with the updated variable weight coefficients.

[0110] Output unit 704 is used to generate an enhanced beamforming audio signal based on beamforming weights and the raw signal acquired by the microphone array.

[0111] Optionally, a signal calibration unit 705 is also included, for: The signal output from the bone conduction piezoelectric microphone at the preset position is used as the audio reference signal.

[0112] Optionally, a signal determination unit 706 is also included, for: The current audio is determined to be the wearer's audio based on the audio reference signal.

[0113] Optionally, a signal separation unit 707 is also included, for: The wearer's audio is separated from ambient noise based on an audio reference signal.

[0114] Optionally, a stop unit 708 is also included, for: If the energy of the audio reference signal exceeds a preset threshold, it is determined that the wearer is in an audio activity state, and the update of the variable weight coefficients is stopped.

[0115] Optionally, a third computing unit 709 is also included, for: The beam pattern is calculated based on the beamforming weight vector and the steering vector. The beam pattern represents the audio signal gain in different directions.

[0116] Optionally, an integration unit 710 is also included, for: The average noise power is obtained by integrating the noise power across the entire beam pattern.

[0117] Optionally, a conversion unit 711 is also included, for: The directivity factor is obtained based on the power gain and average noise power in the target direction of the beam pattern, and then the directivity factor is converted into the directivity index.

[0118] Optionally, an evaluation unit 712 is also included, for: The performance of enhanced beamforming audio signals is evaluated based on the directivity index, and the evaluation results are output.

[0119] Optionally, a building unit 713 is also included, for: A noise feature library is established based on the noise type and intensity of the environment in which the microphone array is located. The noise types include steady-state noise, non-steady-state burst noise, and reverberant noise.

[0120] Optionally, it also includes an identification unit 714, used for: Identify the dominant noise type in the current environment based on the noise feature library.

[0121] Optionally, an adjustment unit 715 is also included for: The parameters of the adaptive beamforming algorithm are adaptively adjusted based on the identified dominant noise type.

[0122] For detailed implementation methods, please refer to... Figures 1 to 6 Examples are not detailed here.

[0123] Please see Figure 8 This application provides a power consumption control device for palmprint recognition, comprising: The processor 801, memory 802, input / output unit 803, and bus 804.

[0124] The processor 801 is connected to the memory 802, the input / output unit 803, and the bus 804.

[0125] The memory 802 stores a program, and the processor 801 calls the program to execute it, such as... Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 as well as Figure 6 The method in the middle.

[0126] This application provides a computer-readable storage medium on which a program is stored, and when the program is executed on a computer, it performs the following... Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 as well as Figure 6 The method in the middle.

[0127] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0128] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0129] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0130] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0131] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. An audio signal processing method, applied to wearable devices, characterized in that, The processing method includes: The direction in which the wearer's face is turned is determined as the target direction for sound focusing; Based on the target direction, the guiding vector of the microphone array of the wearable device is calculated through an irregular array geometric model, which is a model pre-established based on the three-dimensional coordinate position and orientation angle of the microphone array; Based on the guide vector and the target direction, an adaptive beamforming algorithm is used to calculate the beamforming weights; An enhanced beamforming audio signal is generated based on the beamforming weights and the original signal acquired by the microphone array.

2. The processing method according to claim 1, characterized in that, The step of calculating the guiding vector of the microphone array of the wearable device based on the target direction using an irregular array geometric model includes: Based on the target direction and the three-dimensional coordinates of each microphone in the microphone array of the wearable device, the sound wave propagation delay corresponding to each microphone is calculated. The sound wave propagation delay is the time it takes for the audio signal to propagate from the target direction to each microphone in the microphone array. The propagation delay of the sound wave is converted into a phase shift at the corresponding frequency; A guide vector is constructed based on the phase offset, and the guide vector is a complex vector containing the phase offset components of each microphone.

3. The processing method according to claim 1, characterized in that, The step of calculating beamforming weights using an adaptive beamforming algorithm based on the guide vector and the target direction includes: The gain compensation factor corresponding to the target direction is retrieved from the pre-stored directional gain calibration table, and the gain compensation is performed on the steering vector based on the gain compensation factor to obtain the compensated steering vector; Based on the compensation steering vector, a fixed beamformer and a blocking matrix are constructed respectively; Static constraint weights are generated based on the fixed beamformer, and a noise reference signal is obtained based on the blocking matrix. The noise reference signal does not include the audio signal acquired in the target direction. The variable weight coefficients of the adaptive filter are updated based on the noise reference signal; The static constraint weights are fused with the updated variable weight coefficients to obtain the beamforming weights.

4. The processing method according to claim 1, characterized in that, The method further includes: The signal output from the bone conduction piezoelectric microphone at the preset position is used as the audio reference signal; Based on the audio reference signal, the current audio is determined to be the wearer's audio; The wearer's audio is separated from the ambient noise based on the audio reference signal.

5. The processing method according to claim 3 or 4, characterized in that, The method further includes: If the energy of the audio reference signal exceeds a preset threshold, it is determined that the wearer is in an audio activity state, and the update of the variable weight coefficient is stopped.

6. The processing method according to claim 1, characterized in that, The method further includes: The beam pattern is calculated based on the beamforming weight vector and the steering vector, and the beam pattern represents the audio signal gain in different directions; The average noise power is obtained by integrating the noise power across the entire angular direction of the beam pattern. The directivity factor is obtained based on the power gain in the target direction of the beam pattern and the average noise power, and the directivity factor is converted into a directivity index. The performance of the enhanced beamforming audio signal is evaluated based on the directivity index, and the evaluation results are output.

7. The processing method according to any one of claims 1 to 4, characterized in that, The method further includes: A noise feature library is established based on the noise type and noise intensity of the environment in which the microphone array is located. The noise types include steady-state noise, non-steady-state burst noise, and reverberant noise. Identify the dominant noise type in the current environment based on the noise feature library; The parameters of the adaptive beamforming algorithm are adaptively adjusted based on the identified dominant noise type.

8. An audio signal processing apparatus, characterized in that, The device includes: A targeting unit is used to determine the wearer's facial orientation as the target direction for sound focusing; The first calculation unit is used to calculate the guiding vector of the microphone array of the wearable device based on the target direction using an irregular array geometric model, wherein the irregular array geometric model is a model pre-established based on the three-dimensional coordinate position and orientation angle of the microphone array. The second calculation unit is used to calculate the beamforming weights based on the guide vector and the target direction using an adaptive beamforming algorithm; The output unit is used to generate an enhanced beamforming audio signal based on the beamforming weights and the original signal acquired by the microphone array.

9. An audio signal processing apparatus, characterized in that, The device includes: Processor, memory, input / output units, and bus; The processor is connected to the memory, the input / output unit, and the bus; The memory stores a program, which the processor invokes to perform the audio signal processing method as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a program stored thereon, the program performing, when executed on a computer, a method for processing an audio signal as described in any one of claims 1 to 7.