A multi-level microphone array speech noise reduction processing method and system

CN122511276APending Publication Date: 2026-08-04SHENZHEN NADRAY INNOVATIONS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN NADRAY INNOVATIONS TECH CO LTD
Filing Date
2026-05-29
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

[0004]本发明提供了一种多层次式麦克风阵列语音降噪处理方法及系统,本发明实现了识别语义置信度对信号层噪声抑制的直接反哺控制,使识别置信度高的帧段得以保留、置信度低的帧段得到抑制,解决了信号处理层与语音识别层割裂运行、无法协同优化的核心技术问题

Benefits of technology

[0015]本发明提供的技术方案中,从信号层与识别层两个维度协同解决了复杂噪声环境下的语音降噪识别问题。在信号采集层面,将麦克风阵列按间距划分为近场层、中场层和远场层,各层次独立执行声源方向时延补偿与短时能量归一化加权求和,使不同空间尺度的噪声抑制能力形成梯度互补覆盖,克服了单层阵列无法兼顾近端与远端噪声抑制的缺陷。在语音识别层面,针对各层次音频特性独立配置自动语音识别引擎并行解码,结合声学模型得分与语言模型得分线性加权生成单词置信度,再经跨层次时间对齐、时间槽划分与字错率倒数归一化投票权重加权的后验概率投票机制,从多层次识别假设中筛选出置信度最高的最优输出词,有效消解了单一引擎在噪声条件下产生的识别歧义。在协同降噪层面,以各时间槽最高与次高后验概率之差映射为频谱增益系数,对各层次波束成形增强信号逐采样点执行增益加权后,再按线性功率比归一化融合,实现了识别语义置信度对信号层噪声抑制的直接反哺控制,使识别置信度高的帧段得以保留、置信度低的帧段得到抑制,解决了信号处理层与语音识别层割裂运行、无法协同优化的核心技术问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122511276A_ABST
    Figure CN122511276A_ABST
Patent Text Reader

Abstract

This invention relates to the field of speech denoising technology and discloses a multi-layered microphone array speech denoising processing method and system. The method includes: acquiring multi-layered beamforming enhancement signals through a microphone array; performing speech recognition on the beamforming enhancement signals at each layer to obtain word hypothesis sequences and word confidence levels; time-aligning the word hypothesis sequences and dividing them into time slots on a unified time axis, performing posterior probability calculation to obtain the optimal output word and posterior probability pair for each time slot; calculating the spectral gain coefficient of each time slot and performing signal fusion on the multi-layered beamforming enhancement signals to output a denoised speech signal. This method achieves direct feedback control of semantic confidence on signal layer noise suppression, enabling the retention of frames with high confidence and the suppression of frames with low confidence, thus solving the core technical problem of the disconnected operation and inability to coordinate optimization between the signal processing layer and the speech recognition layer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech noise reduction technology, and in particular to a multi-layered microphone array speech noise reduction processing method and system. Background Technology

[0002] In multi-microphone voice acquisition scenarios, existing technologies typically employ a single-layer microphone array to beamform the acquired signal before outputting a single enhanced audio stream to the speech recognition system. However, a single-layer array can only perform noise suppression at a fixed spatial scale, failing to simultaneously suppress near-endpoint interference and far-end diffuse noise, resulting in high residual noise components in the beamformed output signal.

[0003] In speech recognition, current technologies generally employ a single recognition engine to decode a single input signal and output a unique recognition hypothesis sequence. When the signal-to-noise ratio of the input signal is low, the recognition result of a single engine is affected by residual noise, resulting in a large amount of ambiguity. This ambiguous information cannot be further utilized and ultimately accumulates directly into recognition error, leading to a significant drop in recognition accuracy in complex noisy environments. In existing technologies, the signal processing layer and the speech recognition layer operate independently. The output of the recognition layer no longer feeds back into the signal processing layer, and the noise suppression decisions of the signal layer are not guided by the semantic confidence of the recognition layer. This fragmented operation prevents the formation of a collaborative optimization loop, making it impossible to simultaneously guarantee both speech quality and recognition accuracy in complex noisy environments. Summary of the Invention

[0004] This invention provides a multi-layered microphone array speech noise reduction processing method and system. This invention realizes direct feedback control of the recognition semantic confidence on the noise suppression of the signal layer, so that frames with high recognition confidence are retained and frames with low confidence are suppressed, thus solving the core technical problem that the signal processing layer and the speech recognition layer operate separately and cannot be optimized in a coordinated manner.

[0005] In a first aspect, the present invention provides a multi-layered microphone array speech noise reduction processing method, the multi-layered microphone array speech noise reduction processing method comprising: Multi-layered beamforming enhancement signals are acquired using a microphone array; Speech recognition was performed on beamforming enhancement signals at each level to obtain word hypothesis sequences and word confidence scores; The word hypothesis sequence is time-aligned and divided into time slots on a unified time axis. Based on the word confidence and the time slot, the posterior probability is calculated to obtain the optimal output word and posterior probability pair for each time slot. The spectral gain coefficient of each time slot is calculated based on the posterior probability, and the multi-level beamforming enhancement signal is fused according to the spectral gain coefficient to output a noise-reduced speech signal.

[0006] In conjunction with the first aspect, in a first implementation of the first aspect of the present invention, the step of acquiring multi-layered beamforming enhancement signals through a microphone array includes: The microphone array is divided into a near-field layer, a mid-field layer, and a far-field layer, and the raw audio signals of multiple layers are acquired synchronously through the microphones in the near-field layer, the mid-field layer, and the far-field layer. The original audio signals of the multi-level layers are subjected to source direction delay compensation and weighted summation respectively to obtain multi-level beamforming enhancement signals.

[0007] In conjunction with the first aspect, in a second implementation of the first aspect of the present invention, the step of performing source direction delay compensation and weighted summation on the multi-level original audio signals to obtain multi-level beamforming enhancement signals includes: Cross-correlation is performed on the original audio signals of adjacent microphone pairs within each level to obtain the source direction delay difference of each microphone relative to the reference microphone. Integer sample alignment compensation is then performed on the original audio signals of each microphone within each level according to the source direction delay difference to obtain the delay compensation signal of each level. Calculate the short-time energy of the delay compensation signal at each level, and calculate the weighting coefficients based on the short-time energy; Based on the weighting coefficients, the time delay compensation signals at each level are weighted and summed to obtain multi-level beamforming enhancement signals.

[0008] In conjunction with the first aspect, in a third implementation of the first aspect of the present invention, the step of performing speech recognition on the beamforming enhancement signals at each level to obtain word hypothesis sequences and word confidence scores includes: The beamforming enhancement signals of each level are input into an automatic speech recognition engine that is independently configured for the audio characteristics of each level. The beamforming enhancement signals are decoded in parallel to obtain the word hypothesis sequence corresponding to each level and the acoustic model score and language model score of each word. The acoustic model score and the language model score are linearly weighted to obtain the word confidence of each word in each word hypothesis sequence.

[0009] In conjunction with the first aspect, in the fourth implementation of the first aspect of the present invention, the step of inputting the beamforming enhancement signals of each level into an automatic speech recognition engine independently configured for the audio characteristics of each level, and performing parallel decoding on the beamforming enhancement signals to obtain the word hypothesis sequence corresponding to each level and the acoustic model score and language model score of each word includes: The beamforming enhancement signals at each level are framed and the Mel frequency cepstral coefficients and their first and second differences are extracted to obtain the frame-level acoustic feature sequences at each level. The frame-level acoustic feature sequences of each level are input into an automatic speech recognition engine configured independently for the audio characteristics of each level. The posterior probability of phoneme units is calculated through the acoustic model in the automatic speech recognition engine, and Viterbi beam search decoding is performed through the language model in the automatic speech recognition engine to obtain the candidate word sequences of each level and the time boundaries of each candidate word. The acoustic model score is calculated based on the temporal boundaries of each candidate word and the posterior probability of the phoneme unit, and the language model score is calculated based on the conditional probabilities of adjacent words in the candidate word sequence.

[0010] In conjunction with the first aspect, in the fifth implementation of the first aspect of the present invention, the step of time-aligning the hypothetical word sequence and dividing it into time slots on a unified time axis, and performing posterior probability calculation based on the word confidence and the time slots to obtain the optimal output word and posterior probability pair for each time slot includes: Based on the total duration of the shortest word hypothesis sequence, word hypothesis sequences at each level are mapped to a unified time axis, and time slots are divided on the unified time axis. Each candidate word is assigned to the corresponding time slot according to the interval where the midpoint of the time boundary is located, thus obtaining the candidate word set for each time slot. The voting weights are calculated based on the word error rates of the automatic speech recognition engines at each level on the validation set. Then, based on the voting weights and the word confidence, a posterior probability weighted voting is performed on the candidate word set for each time slot to obtain the slot posterior probability of each candidate word in each time slot. The highest and second-highest slot posterior probabilities among the slot posterior probabilities of each candidate word in each time slot are taken as the posterior probability pair, and the candidate word corresponding to the highest slot posterior probability is taken as the optimal output word of the corresponding time slot.

[0011] In conjunction with the first aspect, in the sixth implementation of the first aspect of the present invention, the step of calculating voting weights based on the word error rates of the automatic speech recognition engines at each level on the verification set, and performing posterior probability weighted voting on the candidate word sets of each time slot based on the voting weights and the word confidence, to obtain the slot posterior probability of each candidate word in each time slot, includes: Voting weights are calculated based on the word error rates of the automatic speech recognition engines at each level on the validation set; Based on the voting weights and word confidence of each candidate word, a cross-level weighted summation is performed on each candidate word in the candidate word set within each time slot to obtain the weighted summation result; The slot posterior probability of each candidate word in each time slot is calculated based on the sum of the weighted summation result and the sum of the weighted summation results of all candidate words in the corresponding time slot.

[0012] In conjunction with the first aspect, in the seventh implementation of the first aspect of the present invention, the step of calculating the spectral gain coefficient of each time slot based on the posterior probability pair, and performing signal fusion on the multi-level beamforming enhancement signal according to the spectral gain coefficient to output a denoised speech signal includes: The spectral gain coefficient for each time slot is calculated based on the posterior probability pair. Based on the spectral gain coefficient, gain weighting is performed on the beamforming enhancement signals of each level within the frame segment of the corresponding time slot to obtain the gain weighted signals of each level. Calculate the linear power ratio of the short-time power of the gain-weighted signal at each level to the background noise power in the current frame, and calculate the fusion weight of each level based on the sum of the linear power ratio and the linear power ratios of all levels. Based on the fusion weights, the gain-weighted signals of each level are weighted and summed to output a denoised speech signal.

[0013] In conjunction with the first aspect, in the eighth implementation of the first aspect of the present invention, the step of performing gain weighting on the beamforming enhancement signals of each level within the frame segment of the corresponding time slot according to the spectral gain coefficient to obtain the gain-weighted signals of each level includes: Based on the product of the start and end times of each time slot and the sampling rate, the start and end times of each time slot are converted into the corresponding start and end sampling point positions in the beamforming enhancement signal of each level, so as to obtain the frame sampling range of each time slot in the beamforming enhancement signal of each level. The spectral gain coefficient of each time slot is multiplied point by point by all sampling points within the sampling range of the frame segment in the beamforming enhancement signal of each level. Point-by-point gain weighting is then performed on the corresponding frame segments of all time slots in the beamforming enhancement signal of each level to obtain the gain weighted signal of each level.

[0014] Secondly, the present invention provides a multi-layered microphone array speech noise reduction processing system, the multi-layered microphone array speech noise reduction processing system comprising: The acquisition module is used to acquire multi-layer beamforming enhancement signals through a microphone array; The speech recognition module is used to perform speech recognition on beamforming enhancement signals at each level to obtain word hypothesis sequences and word confidence scores. The calculation module is used to time-align the word hypothesis sequence and divide it into time slots on a unified time axis, and perform posterior probability calculation based on the word confidence and the time slot to obtain the optimal output word and posterior probability pair for each time slot. The signal fusion module is used to calculate the spectral gain coefficient of each time slot based on the posterior probability pair, and to perform signal fusion on the multi-level beamforming enhancement signal according to the spectral gain coefficient to output a noise-reduced speech signal.

[0015] The technical solution provided by this invention addresses the problem of speech denoising and recognition in complex noisy environments through a collaborative approach from both the signal and recognition layers. At the signal acquisition level, the microphone array is divided into near-field, mid-field, and far-field layers based on spacing. Each layer independently performs source direction delay compensation and short-time energy normalization weighted summation, enabling complementary gradient coverage of noise suppression capabilities at different spatial scales. This overcomes the limitation of single-layer arrays in simultaneously addressing near- and far-field noise suppression. At the speech recognition level, an automatic speech recognition engine is independently configured for parallel decoding based on the audio characteristics of each layer. Word confidence is generated by linearly weighting the acoustic model score and language model score. Then, a posterior probability voting mechanism, using cross-layer time alignment, time slot division, and inverse normalization of word error rate voting weights, selects the optimal output word with the highest confidence from multi-layer recognition hypotheses, effectively resolving recognition ambiguities caused by a single engine under noisy conditions. At the collaborative noise reduction level, the difference between the highest and second-highest posterior probabilities of each time slot is mapped as the spectral gain coefficient. After gain weighting is performed on each sampling point of the beamforming enhancement signal at each level, it is then fused by normalization according to the linear power ratio. This realizes the direct feedback control of the recognition semantic confidence on the noise suppression of the signal layer, so that frames with high recognition confidence are retained and frames with low confidence are suppressed. This solves the core technical problem that the signal processing layer and the speech recognition layer operate separately and cannot be optimized collaboratively.

[0016] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.

[0017] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of an embodiment of the multi-layered microphone array speech noise reduction processing method of the present invention; Figure 2 This is a schematic diagram illustrating the acquisition of beamforming enhancement signals in an embodiment of the present invention; Figure 3 This is a schematic diagram of speech recognition in an embodiment of the present invention; Figure 4 This is a schematic diagram illustrating the calculation of the optimal output word and posterior probability pair in an embodiment of the present invention; Figure 5 This is a schematic diagram of signal fusion in an embodiment of the present invention; Figure 6 This is a schematic diagram of an embodiment of the multi-layered microphone array voice noise reduction processing system of the present invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] The terms "comprising" and "having," and any variations thereof, used in the embodiments of this invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the steps or units listed, but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.

[0021] To facilitate understanding of this embodiment, a multi-layered microphone array speech noise reduction method disclosed in this embodiment of the invention will first be described in detail. For example... Figure 1 As shown, this method includes the following steps: 101. Acquire multi-layered beamforming enhancement signals through a microphone array; 102. Perform speech recognition on the beamforming enhancement signals at each level to obtain the word hypothesis sequence and word confidence. 103. Align the word hypothesis sequence with time and divide it into time slots on a unified time axis. Calculate the posterior probability based on the word confidence and the time slot to obtain the optimal output word and posterior probability pair for each time slot. 104. Calculate the spectral gain coefficient of each time slot based on the posterior probability pair, and perform signal fusion on the multi-level beamforming enhancement signal according to the spectral gain coefficient to output a noise-reduced speech signal.

[0022] In one specific embodiment, such as Figure 2 The process of executing step 101 may specifically include the following steps: 1011. Divide multiple microphones in the microphone array into a near-field layer, a mid-field layer, and a far-field layer, and synchronously acquire multi-layer raw audio signals through the microphones in the near-field layer, mid-field layer, and far-field layer; 1012. Perform source direction delay compensation and weighted summation on the original audio signals of multiple levels respectively to obtain multi-level beamforming enhancement signals.

[0023] Specifically, an array hardware structure and synchronous acquisition link are established around spatial layer acquisition. Multiple microphones in the same array are divided into near-field, mid-field, and far-field layers according to their physical positions and spacing differences. The three types of microphones synchronously output multi-layer raw audio signals under the same sampling clock. Different layers correspond to different spatial scale noise suppression capabilities. The near-field layer is more suitable for suppressing near-endpoint interference, the mid-field layer plays a transitional coverage role, and the far-field layer is more suitable for forming a wider spatial integral suppression effect on diffuse noise. Therefore, the three-layer parallel acquisition establishes a gradient complementary spatial perception foundation.

[0024] Within each layer, a reference microphone is selected. Then, the time delay information related to the direction of arrival of the sound source is estimated around the adjacent microphones within that layer. Based on the estimation results, compensation is performed on the original audio signals of each channel to achieve time alignment of the target speech across multiple channels within the same layer as much as possible. After alignment, weighting coefficients are calculated for the effective energy distribution of each channel within the current analysis frame. Based on the weights of each channel, the layer is summed and output to form near-field beamforming enhancement signals, mid-field beamforming enhancement signals, and far-field beamforming enhancement signals, respectively.

[0025] In one specific embodiment, the process of performing step 1012 may specifically include the following steps: Cross-correlation is performed on the original audio signals of adjacent microphone pairs within each level to obtain the source direction delay difference of each microphone relative to the reference microphone. Integer sample alignment compensation is then performed on the original audio signals of each microphone within each level according to the source direction delay difference to obtain the delay compensation signal of each level. Calculate the short-time energy of the delay compensation signal at each level, and calculate the weighting coefficients based on the short-time energy; Based on the weighting coefficients, the time delay compensation signals at each level are weighted and summed to obtain multi-level beamforming enhancement signals.

[0026] Specifically, the raw audio signals within each layer enter the sound source direction time delay estimation link. Within the near-field, mid-field, and far-field layers, cross-correlation is performed on the acquired signals of adjacent microphone pairs. The time difference of the sound wave arriving at different microphone positions is extracted using the sample delay corresponding to the cross-correlation peak. For the m-th group of adjacent microphone pairs in the l-th layer, the sample delay is denoted as... Then according to the sampling rate Convert the sample delay into a time delay estimate: ;

[0027] in, This represents the estimated time delay of the m-th adjacent microphone pair in the l-th layer, in seconds. This represents the number of delayed sampling points corresponding to the peak cross-correlation value, in units of sampling points. This represents the system sampling rate. The average delay within the layer is obtained by taking the arithmetic mean of the delay estimates for all adjacent microphone pairs in the same layer. Combined with the average spacing between adjacent microphones in this layer Calculate the incident angle of the sound source from the speed of sound c: ;

[0028] in, This represents the incident angle of the sound source in the l-th layer; The expression represents the average delay estimates of all adjacent microphone pairs in the l-th layer; c represents the speed of sound. This represents the average distance between adjacent microphones in the l-th layer. After processing, each layer obtains spatial incidence parameters consistent with the direction of the target sound source.

[0029] After obtaining the incident direction angles for each layer, the theoretical time delay difference of each microphone channel relative to the reference microphone is calculated, and integer sample alignment compensation is performed on the original audio signal accordingly. For the m-th microphone in the l-th layer, the theoretical time delay difference relative to the reference microphone is: ;

[0030] in, This represents the theoretical time delay difference between the m-th microphone in layer l and the reference microphone, in seconds. This represents the projected distance of the m-th microphone in the l-th layer relative to the reference microphone along the array axis, in meters. represents the incident angle of the sound source at the corresponding layer; c represents the speed of sound. The reference microphone can be uniformly selected as the first microphone at each layer to ensure consistent time reference across all channels. Convert to the corresponding number of delayed sampling points, and use Integer sample shift compensation is performed to ensure that speech components from the target direction arrive as simultaneously as possible on multiple channels of this layer, thereby reducing phase mismatch caused by propagation path differences. After compensation, the aligned signal of the m-th microphone in the l-th layer is denoted as... Each layer thus forms a set of time delay compensation signals.

[0031] After time alignment of signals from each layer, weighting coefficients are calculated based on the energy state of each channel in the current frame, ensuring that channels with better signal quality have a higher weighting in the in-layer synthesis result. This is done for the time delay compensation signal of the m-th microphone in layer l. Calculate the short-time energy within the current analysis frame: ;

[0032] in, This represents the short-time energy of the m-th microphone in the l-th layer in the current frame; This represents the amplitude of the delay compensation signal at time t; This represents the sampling set corresponding to the current short-time analysis window. After the short-time energy is calculated, the weighting coefficients are obtained using an intra-layer normalization method. ;

[0033] in, This represents the normalized weight of the m-th microphone channel in layer l. This weight is a dimensionless ratio, and the sum of the weights of all channels is 1. After calculating the weights, a weighted sum is performed on all delay compensation signals in layer l to obtain the beamforming enhancement signal for that layer. ;

[0034] in, This represents the beamforming enhancement output of the l-th layer at time t. Speech components that are already time-aligned in the target direction will mutually enhance each other during the summation process, while misaligned noise or interference components are difficult to superimpose in phase. Simultaneously, high-energy channels with higher signal-to-noise ratios receive greater weight, while the contributions of low-energy channels, those that are heavily obstructed, or those severely submerged by noise are actively suppressed. Ultimately, the near-field layer enhancement signal, mid-field layer enhancement signal, and far-field layer enhancement signal are output respectively.

[0035] In one specific embodiment, such as Figure 3 The process of executing step 102 may specifically include the following steps: 1021. Input the beamforming enhancement signals of each level into an automatic speech recognition engine that is independently configured for the audio characteristics of each level, and perform parallel decoding on the beamforming enhancement signals to obtain the word hypothesis sequence corresponding to each level and the acoustic model score and language model score of each word. 1022. The acoustic model score and the language model score are linearly weighted to obtain the word confidence of each word in each word hypothesis sequence.

[0036] Specifically, the near-field, mid-field, and far-field beamforming enhancement signals are input into automatic speech recognition engines that match the audio characteristics of each layer, and decoding is performed simultaneously in a parallel structure. The beamforming enhancement signals at different layers differ in spatial coverage, residual noise distribution, and target speech fidelity. By configuring recognition engines separately, each engine becomes complementary in terms of acoustic modeling structure or training data, thereby reducing the systematic bias that a single recognition engine might exhibit when facing complex noise. At the start of parallel decoding, the beamforming enhancement signals of each layer are framed. After framing, Mel-frequency cepstral features are extracted from each frame. The feature sequences of each layer are input into the corresponding automatic speech recognition engine, where the internal acoustic model performs frame-level acoustic scoring. This is then combined with language model constraints for Viterbi beam search decoding. After the search, each engine outputs N-best=10 candidate word hypotheses for each layer's input.

[0037] The scores of the original acoustic model and the original language model are uniformly represented as normalized logarithmic posterior probabilities to eliminate comparison distortion caused by inconsistencies in units. Among these, words... Acoustic model score Defined as the word at the corresponding time boundary The frame-to-frame average of the log-posterior probability of the acoustic model for each frame, for each word. Language model score Defined as log-conditional probability under a binary language model After both are uniformly entered into the logarithmic probability domain, the word confidence score is calculated according to a linear weighted relationship. ,in, This represents the word confidence score given by the nth recognition engine for the kth word when processing the lth layer input; Represents the interpolation coefficients of the language model; Representing words Acoustic model score; Representing words The language model score.

[0038] In one specific embodiment, the process of performing step 1021 may specifically include the following steps: The beamforming enhancement signals at each level are framed and the Mel frequency cepstral coefficients and their first and second differences are extracted to obtain the frame-level acoustic feature sequences at each level. The frame-level acoustic feature sequences of each level are input into an automatic speech recognition engine that is independently configured for the audio characteristics of each level. The posterior probability of phoneme units is calculated through the acoustic model in the automatic speech recognition engine, and the Viterbi beam search decoding is performed through the language model in the automatic speech recognition engine to obtain the candidate word sequences of each level and the time boundary of each candidate word. The acoustic model score is calculated based on the temporal boundaries and posterior probabilities of each candidate word, and the language model score is calculated based on the conditional probabilities of adjacent words in the candidate word sequence.

[0039] Specifically, beamforming enhancement signals at each level are processed into frames according to unified short-time analysis parameters, ensuring that the speech approximately meets short-time stationarity conditions within a single frame, while maintaining appropriate overlap between adjacent frames. This allows for continuous characterization of speech transition segments, consonant transient segments, and vowel steady-state segments. After framing, Mel-frequency cepstral coefficients and their first and second differences are extracted from each frame of enhanced speech, forming a 39-dimensional frame-level acoustic feature sequence. Thirteen dimensions describe the spectral envelope of the current frame, thirteen first-order differences characterize the dynamic change trend between adjacent frames, and thirteen second-order differences further characterize the acceleration features of the dynamic changes.

[0040] After the frame-level acoustic feature sequences at each layer are fed into the corresponding automatic speech recognition engine, the engine internally uses an acoustic model to estimate the posterior probability of each frame's features. This involves estimating the posterior probability distribution of the current frame's acoustic features based on the matching relationship between the current frame's acoustic features and the statistical patterns of each phoneme unit, outputting the posterior probability distribution of the current frame belonging to different phoneme units. This posterior probability distribution is then fed into the decoding search chain, participating in candidate word path evaluation along with the language model. Subsequently, the automatic speech recognition engine calls the language model to perform Viterbi beam search decoding. During the search process, it evaluates the degree of local acoustic matching based on the posterior probability of the phoneme units output by the acoustic model, while simultaneously limiting the search path based on the linguistic connections between candidate words. This weakens unreasonable word sequences while retaining paths that better conform to the linguistic context. After the Viterbi beam search, each layer of the recognition engine outputs a corresponding candidate word sequence, and simultaneously records the start and end times of each candidate word on a unified timeline, forming temporal boundary information.

[0041] In one specific embodiment, such as Figure 4 The process of executing step 103 may specifically include the following steps: 1031. Based on the total duration of the shortest word hypothesis sequence, map the word hypothesis sequences at each level to a unified time axis, divide the time axis into time slots, and assign each candidate word to the corresponding time slot according to the interval where the midpoint of the time boundary is located, so as to obtain the candidate word set of each time slot. 1032. Calculate the voting weight based on the word error rate of the automatic speech recognition engine at each level on the validation set, and perform posterior probability weighted voting on the candidate word set of each time slot based on the voting weight and word confidence to obtain the slot posterior probability of each candidate word in each time slot. 1033. Take the highest and second-highest slot posterior probabilities among the slot posterior probabilities of each candidate word in each time slot as the posterior probability pair, and take the candidate word corresponding to the highest slot posterior probability as the optimal output word of the corresponding time slot.

[0042] Specifically, select the sequence with the shortest total duration from all hypothetical word sequences, and denote this shortest total duration as... ,in, The baseline total duration represents the time taken along a unified timeline; for any nth recognition engine, lth layer input, and ith candidate sequence, the original total duration is denoted as . Then calculate the scaling factor according to the proportional mapping relationship: ;

[0043] in, This represents the time scaling factor used by the nth recognition engine when mapping the i-th candidate sequence to the unified time axis under the l-th layer input. This represents the original total duration of the candidate sequence. After scaling, the start and end time boundaries of all words in each candidate sequence will be linearly mapped to the interval. Internally, time slots are divided with a fixed step size. For each candidate word, based on the midpoint of the mapped time boundary... The candidate word is assigned to the corresponding time slot based on the time interval it falls into; among which, This indicates the starting time of the k-th candidate word. This represents the end time of the k-th candidate word. If no word midpoint falls within a certain time slot of a candidate sequence, then an empty word is filled in that time slot. Each candidate sequence has a uniform placeholder format in all time slots, and each time slot will form a set of candidate words contributed by different levels and different recognition paths.

[0044] Voting weights are calculated based on the word error rates of each level of automatic speech recognition engines on the validation set. This ensures that combinations of engine levels with superior recognition performance receive a higher weight in slot voting, while combinations with weaker recognition stability contribute less. Combinations with lower word error rates have larger reciprocals, resulting in higher normalized voting weights. Within each time slot, posterior probability-weighted voting is performed by combining the candidate word set with the existing word confidence scores of each word. This means that within each time slot, the support strength of different engine level combinations for candidate words is not directly and equally weighted, but rather the global reliability of the combination itself and the local confidence score of the current candidate word within that slot are considered simultaneously. After normalization and weighting, a slot-specific posterior probability distribution is formed for each candidate word.

[0045] Extract the optimal output word and the posterior probability pair required for collaborative denoising from the probability distribution of each time slot. For any time slot... First, select the candidate word with the highest posterior probability from all candidate words for that slot: ;

[0046] in, Indicates time slot The optimal output word, Represents the set of candidate words for time slots Any candidate word in the list, This represents the set of candidate words for the s-th time slot. Indicate candidate words In the time slot The posterior probability of a slot within a given slot. Simultaneously, retain the highest posterior probability within that slot. Posterior probability of the second highest trough These two quantities together constitute the posterior probability pair. Meanwhile, when the posterior probability of the slot corresponding to the selected optimal output word is lower than the confidence threshold... When this time slot is reached, mark it as a low-confidence slot and output an empty word. .in, This represents the slot-level confidence threshold, with a value of 0.4. Each time slot will yield an optimal output word and a set of posterior probability pairs. Then, the optimal output words from all non-empty word time slots are concatenated in chronological order to form the final recognized text.

[0047] In one specific embodiment, the process of performing step 1032 may specifically include the following steps: Voting weights are calculated based on the word error rates of the automatic speech recognition engines at each level on the validation set; Based on the voting weights and the word confidence of each candidate word, a cross-level weighted summation is performed on each candidate word in the candidate word set within each time slot to obtain the weighted summation result; The posterior probability of each candidate word in each time slot is calculated based on the sum of the weighted sum and the sum of the weighted sum of all candidate words in the corresponding time slot.

[0048] Specifically, let the word error rate be % when the nth automatic speech recognition engine processes the lth layer input. The corresponding voting weight is calculated as follows: ;

[0049] in, This represents the normalized voting weight of the nth automatic speech recognition engine under the input of the lth layer. This represents the word error rate of the corresponding combination on the validation set. The lower the word error rate, the better the recognition effect. After the reciprocal transformation, the combination with better performance corresponds to a larger value. After the entire set is normalized, the sum of the voting weights of all combinations satisfies 1.

[0050] Collect the current time slot The candidate set consists of all candidate words. Then for any candidate word Iterate through all engine and layer combinations. If a combination provides the candidate word in the current time slot, its voting weight and the candidate word's confidence score are combined and included in the candidate word's support strength. If a combination does not provide the candidate word in the current time slot, its contribution to the candidate word is recorded as zero. The slot posterior probability is calculated as follows: ;

[0051] in, Indicate candidate words In the time slot The posterior probability of the slot within, This represents the voting weight of the nth automatic speech recognition engine under the input of the lth layer. Let j represent the indicator function, and j be the set of candidate words C. s The traversal index of candidate words, w j C represents s The j-th candidate word, s n,l,j This indicates that when the nth automatic speech recognition engine processes the lth layer input, it evaluates the candidate word w. j The corresponding word confidence score is 1 if the candidate word belongs to the output of the corresponding engine hierarchical combination in that time slot, and 0 otherwise. This indicates that when the nth automatic speech recognition engine processes the lth layer input, it is in the time slot. The corresponding candidate word results, This represents the word confidence score corresponding to the candidate word. This represents the result after restoring the word confidence score from the logarithmic probability domain to the linear probability domain. According to this relationship, the numerator is actually the current candidate word. The cross-level weighted summation result within this time slot is the total support strength formed by the combination of each engine level based on long-term recognition performance and current term confidence.

[0052] The weighted sum of all candidate words in the current time slot is further normalized to obtain the slot posterior probability for each candidate word. The numerator is used to form the support level for a single candidate word, and then the sum of the support levels of all candidate words in the current time slot is accumulated in the denominator to complete the conversion from "local support strength" to "relative posterior probability". If there is an empty word in the current time slot... Then, empty words participate in the normalization process according to the unified settings, and their confidence scores are taken as... ,in This setting is designed to ensure that empty words can participate in the competition in low-confidence scenarios, but will not dominate the voting results in the slot due to excessively large values.

[0053] In one specific embodiment, such as Figure 5 The process of executing step 104 can specifically include the following steps: 1041. Calculate the spectral gain coefficient of each time slot based on the posterior probability pair; 1042. Based on the spectral gain coefficient, gain weighting is performed on the beamforming enhancement signals of each level within the frame segment of the corresponding time slot to obtain the gain weighted signals of each level. 1043. Calculate the linear power ratio of the short-time power of the gain-weighted signal at each level to the background noise power in the current frame, and calculate the fusion weight of each level based on the sum of the linear power ratio and the linear power ratio of all levels. 1044. Based on the fusion weights, perform weighted summation on the gain weighted signals of each level to output the denoised speech signal.

[0054] Specifically, a confidence level is established to feed back the control quantity based on the posterior probability of each time slot, so that the slot discrimination result given by the identification layer can be directly fed into the gain adjustment link of the signal layer. For any time slot... First take the highest slot and then test the probability. Posterior probability of the second highest trough And calculate the difference between the two. ,in, This represents the posterior probability difference for the s-th time slot, with a value ranging from 0 to 1. This represents the highest posterior probability within that time slot; This represents the second-highest posterior probability within that time slot. A larger difference indicates stronger voting consistency in the recognition results within the current time slot, suggesting the target speech component is more likely to dominate; a smaller difference indicates stronger recognition ambiguity within the current time slot, making noise interference more likely to perturb the recognition results. Based on this, the posterior probability difference is mapped to a spectral gain coefficient: ;

[0055] in, This represents the spectral gain coefficient corresponding to the s-th time slot; This represents the lower limit of gain, with a value of 0.1. This value is used to avoid speech continuity being impaired when low-confidence frames are completely suppressed. This represents the gain curve steepness control coefficient, with a value of 0.1. This value controls the rate of increase when the posterior probability difference propagates to gain changes, ensuring that even small posterior differences can result in a identifiable gain adjustment amplitude. After mapping, each time slot receives a gain control value corresponding to the recognition consistency.

[0056] The spectral gain coefficient is applied to the beamforming enhancement signals at each layer, ensuring that the near-field, mid-field, and far-field layers receive consistent confidence adjustment simultaneously within the same time slot. The corresponding frame position of each time slot in the beamforming enhancement signal is determined based on the time slot boundaries, and then the spectral gain coefficient of the corresponding time slot is applied to that frame segment, forming the gain-weighted signal for each layer. For the l-th layer beamforming enhancement signal... In terms of time slots Within the corresponding time domain interval, the signal after gain adjustment is represented as follows: ,in, This represents the output signal of the l-th layer after gain adjustment. Let represent the original beamforming enhancement signal of the l-th layer, and t represent the time variable. Time slots with higher identification consistency are retained at a relatively large amplitude, while time slots with lower identification consistency are moderately suppressed. Therefore, the signal layer noise suppression process begins to exhibit a cooperative characteristic driven by the identification results.

[0057] Based on the ratio of effective power to background noise power in the current frame for each layer, the contribution ratios of the near-field, mid-field, and far-field layers in the final output are dynamically determined. The linear power ratio is calculated for the l-th layer: ;

[0058] in, This represents the linear power ratio of the current frame in layer l; This represents the short-time power of the current frame in layer l; This represents the estimated background noise power at layer l. (Short-time power) The calculation relationship is as follows: ;

[0059] in, This represents the sample set corresponding to the current 25ms analysis frame. Background noise power. A sliding minimum statistical method is used for estimation. This involves statistically analyzing the short-time power of all 25ms subframes within a 5-second sliding window ending at the current time, and taking the 5th percentile as the estimated background noise power value. Simultaneously, to prevent excessively low power ratios in a particular layer from causing extreme imbalances in inter-layer weights, [further details are needed]. Apply a lower bound truncation, making The value is no less than 0.01, which ensures that all layers retain basic participation capabilities during fusion. The fusion weight of each layer is determined by the ratio of its own linear power ratio to the sum of the linear power ratios of all layers. Therefore, the layer with the higher power ratio receives a larger fusion percentage in the current frame.

[0060] Based on the fusion weights, the gain-weighted signals at each level are summed using a weighted average, and the denoised speech signal is output. The output relationship is as follows: ;

[0061] in, This represents the final denoised speech signal. This represents the normalized fusion weights of the l-th layer. This represents the gain-weighted signal of the l-th layer. This represents all levels of numbers participating in the normalization summation.

[0062] In one specific embodiment, the process of performing step 1042 may specifically include the following steps: Based on the product of the start and end times of each time slot and the sampling rate, the start and end times of each time slot are converted into the corresponding start and end sampling point positions in the beamforming enhancement signal of each level, so as to obtain the frame sampling range of each time slot in the beamforming enhancement signal of each level. The spectral gain coefficient of each time slot is multiplied point by point by all sampling points within the frame sampling range of the beamforming enhancement signal at each level. Point-by-point gain weighting is then performed on the corresponding frame segments of all time slots in the beamforming enhancement signal at each level to obtain the gain weighted signal at each level.

[0063] Specifically, for any time slot Read the time slot The start and end times, then the time slot The start and end times are respectively related to the sampling rate. Multiply to obtain the time slot The start and end sampling points in each layer of beamforming enhancement signal are determined, thereby identifying the time slot in the l-th layer of beamforming enhancement signal. The corresponding frame sampling range. This represents the s-th time slot. The system sampling rate is represented by 'l', and the hierarchy number is represented by 'l'. This represents the beamforming enhancement signal of the l-th layer. Since all three layers of signals are generated under a unified sampling clock, the same time slot has a consistent time domain boundary when mapped to the three layers of signals, only corresponding to different data segment positions in the signal buffers of different layers.

[0064] The spectral gain coefficients corresponding to each time slot are written into the point-by-point weighted link, enabling each layer of beamforming enhancement signal to complete sample-level amplitude modulation within the corresponding time slot range. For the l-th layer beamforming enhancement signal... In the time slot The following gain relationship is applied within the corresponding time domain interval: ;

[0065] in, This represents the output signal after the time slot gain adjustment of the l-th layer is completed. Indicates time slot The corresponding spectral gain coefficient, Let represent the original beamforming enhancement signal of the l-th layer, and t represent the time variable. According to this relationship, in the time slot... The same method is used for all corresponding sampling points. The gain is multiplied point-by-point with the corresponding sampled value, thus maintaining a consistent gain adjustment direction for each sampled point within a time slot. However, different time slots employ different degrees of amplitude suppression or retention based on the posterior differences in recognition. If the recognition consistency of a certain time slot is high, then the corresponding... A larger amplitude indicates a higher retention rate of the speech sampling amplitude within that time slot; conversely, a lower recognition consistency in a particular time slot indicates a lower retention rate. If the amplitude is small, the sampling amplitude within that time slot will be suppressed overall, thus achieving the reverse constraint of the recognition layer on the signal layer. Point-by-point gain weighting is continuously performed on all frame segments according to the time slot sequence. Starting from the first time slot, the corresponding start and end sampling point positions and spectral gain coefficients are read sequentially. Then, all samples of the l-th layer beamforming enhancement signal within that sampling interval are multiplied point-by-point. Processing continues to the next time slot until all time slots have been processed. The near-field layer, mid-field layer, and far-field layer each obtain their respective gain-weighted signals. , and .

[0066] The above describes the multi-layered microphone array speech noise reduction processing method in the embodiments of the present invention. The following describes the multi-layered microphone array speech noise reduction processing system in the embodiments of the present invention. Please refer to [link / reference]. Figure 6One embodiment of the multi-layered microphone array speech noise reduction processing system of the present invention includes: Acquisition module 601 is used to acquire multi-layer beamforming enhancement signals through a microphone array; The speech recognition module 602 is used to perform speech recognition on beamforming enhancement signals at each level to obtain word hypothesis sequences and word confidence scores. The calculation module 603 is used to align the word hypothesis sequence in time and divide it into time slots on a unified time axis. Based on the word confidence and time slot, it performs posterior probability calculation to obtain the optimal output word and posterior probability pair for each time slot. The signal fusion module 604 is used to calculate the spectral gain coefficient of each time slot based on the posterior probability pair, and to perform signal fusion on the multi-level beamforming enhancement signal according to the spectral gain coefficient, and output the noise-reduced speech signal.

[0067] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0068] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0069] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multi-layered microphone array speech noise reduction processing method, characterized in that, include: Multi-layered beamforming enhancement signals are acquired using a microphone array; Speech recognition was performed on beamforming enhancement signals at each level to obtain word hypothesis sequences and word confidence scores; The word hypothesis sequence is time-aligned and divided into time slots on a unified time axis. Based on the word confidence and the time slot, the posterior probability is calculated to obtain the optimal output word and posterior probability pair for each time slot. The spectral gain coefficient of each time slot is calculated based on the posterior probability, and the multi-level beamforming enhancement signal is fused according to the spectral gain coefficient to output a noise-reduced speech signal.

2. The multi-layered microphone array speech noise reduction processing method according to claim 1, characterized in that, The acquisition of multi-layered beamforming enhancement signals via a microphone array includes: The microphone array is divided into a near-field layer, a mid-field layer, and a far-field layer, and the raw audio signals of multiple layers are acquired synchronously through the microphones in the near-field layer, the mid-field layer, and the far-field layer. The original audio signals of the multi-level layers are subjected to source direction delay compensation and weighted summation respectively to obtain multi-level beamforming enhancement signals.

3. The multi-layered microphone array speech noise reduction processing method according to claim 2, characterized in that, The process of performing source direction delay compensation and weighted summation on the multi-level original audio signals to obtain multi-level beamforming enhancement signals includes: Cross-correlation is performed on the original audio signals of adjacent microphone pairs within each level to obtain the source direction delay difference of each microphone relative to the reference microphone. Integer sample alignment compensation is then performed on the original audio signals of each microphone within each level according to the source direction delay difference to obtain the delay compensation signal of each level. Calculate the short-time energy of the delay compensation signal at each level, and calculate the weighting coefficients based on the short-time energy; Based on the weighting coefficients, the time delay compensation signals at each level are weighted and summed to obtain multi-level beamforming enhancement signals.

4. The multi-layered microphone array speech noise reduction processing method according to claim 1, characterized in that, The process of performing speech recognition on beamforming enhancement signals at each level to obtain word hypothesis sequences and word confidence scores includes: The beamforming enhancement signals of each level are input into an automatic speech recognition engine that is independently configured for the audio characteristics of each level. The beamforming enhancement signals are decoded in parallel to obtain the word hypothesis sequence corresponding to each level and the acoustic model score and language model score of each word. The acoustic model score and the language model score are linearly weighted to obtain the word confidence of each word in each word hypothesis sequence.

5. The multi-layered microphone array speech noise reduction processing method according to claim 4, characterized in that, The process involves inputting beamforming enhancement signals at each level into an automatic speech recognition engine independently configured for the audio characteristics of each level, and performing parallel decoding on the beamforming enhancement signals to obtain the word hypothesis sequence corresponding to each level, as well as the acoustic model score and language model score of each word, including: The beamforming enhancement signals at each level are framed and the Mel frequency cepstral coefficients and their first and second differences are extracted to obtain the frame-level acoustic feature sequences at each level. The frame-level acoustic feature sequences of each level are input into an automatic speech recognition engine configured independently for the audio characteristics of each level. The posterior probability of phoneme units is calculated through the acoustic model in the automatic speech recognition engine, and Viterbi beam search decoding is performed through the language model in the automatic speech recognition engine to obtain the candidate word sequences of each level and the time boundaries of each candidate word. The acoustic model score is calculated based on the temporal boundaries of each candidate word and the posterior probability of the phoneme unit, and the language model score is calculated based on the conditional probabilities of adjacent words in the candidate word sequence.

6. The multi-layered microphone array speech noise reduction processing method according to claim 5, characterized in that, The process of aligning the hypothetical word sequence with time and dividing it into time slots on a unified time axis, and calculating the posterior probability based on the word confidence and the time slot to obtain the optimal output word and posterior probability pair for each time slot includes: Based on the total duration of the shortest word hypothesis sequence, word hypothesis sequences at each level are mapped to a unified time axis, and time slots are divided on the unified time axis. Each candidate word is assigned to the corresponding time slot according to the interval where the midpoint of the time boundary is located, thus obtaining the candidate word set for each time slot. The voting weights are calculated based on the word error rates of the automatic speech recognition engines at each level on the validation set. Then, based on the voting weights and the word confidence, a posterior probability weighted voting is performed on the candidate word set for each time slot to obtain the slot posterior probability of each candidate word in each time slot. The highest and second-highest slot posterior probabilities among the slot posterior probabilities of each candidate word in each time slot are taken as the posterior probability pair, and the candidate word corresponding to the highest slot posterior probability is taken as the optimal output word of the corresponding time slot.

7. The multi-layered microphone array speech noise reduction processing method according to claim 6, characterized in that, The step of calculating voting weights based on the word error rates of the automatic speech recognition engines at each level on the validation set, and performing posterior probability-weighted voting on the candidate word sets for each time slot based on the voting weights and the word confidence, to obtain the slot posterior probability of each candidate word in each time slot, includes: Voting weights are calculated based on the word error rates of the automatic speech recognition engines at each level on the validation set; Based on the voting weights and word confidence of each candidate word, a cross-level weighted summation is performed on each candidate word in the candidate word set within each time slot to obtain the weighted summation result; The slot posterior probability of each candidate word in each time slot is calculated based on the sum of the weighted summation result and the sum of the weighted summation results of all candidate words in the corresponding time slot.

8. The multi-layered microphone array speech noise reduction processing method according to claim 1, characterized in that, The process of calculating the spectral gain coefficient for each time slot based on the posterior probability pair, and then performing signal fusion on the multi-level beamforming enhancement signal according to the spectral gain coefficient to output a denoised speech signal, includes: The spectral gain coefficient for each time slot is calculated based on the posterior probability pair. Based on the spectral gain coefficient, gain weighting is performed on the beamforming enhancement signals of each level within the frame segment of the corresponding time slot to obtain the gain weighted signals of each level. Calculate the linear power ratio of the short-time power of the gain-weighted signal at each level to the background noise power in the current frame, and calculate the fusion weight of each level based on the sum of the linear power ratio and the linear power ratios of all levels. Based on the fusion weights, the gain-weighted signals of each level are weighted and summed to output a denoised speech signal.

9. The multi-layered microphone array speech noise reduction processing method according to claim 8, characterized in that, The step of performing gain weighting on the beamforming enhancement signals of each level within the corresponding time slot frame segment according to the spectral gain coefficient, to obtain the gain-weighted signals of each level, includes: Based on the product of the start and end times of each time slot and the sampling rate, the start and end times of each time slot are converted into the corresponding start and end sampling point positions in the beamforming enhancement signal of each level, so as to obtain the frame sampling range of each time slot in the beamforming enhancement signal of each level. The spectral gain coefficient of each time slot is multiplied point by point by all sampling points within the sampling range of the frame segment in the beamforming enhancement signal of each level. Point-by-point gain weighting is then performed on the corresponding frame segments of all time slots in the beamforming enhancement signal of each level to obtain the gain weighted signal of each level.

10. A multi-layered microphone array speech noise reduction processing system, characterized in that, A method for performing multi-layered microphone array speech noise reduction processing as described in any one of claims 1-9, comprising: The acquisition module is used to acquire multi-layer beamforming enhancement signals through a microphone array; The speech recognition module is used to perform speech recognition on beamforming enhancement signals at each level to obtain word hypothesis sequences and word confidence scores. The calculation module is used to time-align the word hypothesis sequence and divide it into time slots on a unified time axis, and perform posterior probability calculation based on the word confidence and the time slot to obtain the optimal output word and posterior probability pair for each time slot. The signal fusion module is used to calculate the spectral gain coefficient of each time slot based on the posterior probability pair, and to perform signal fusion on the multi-level beamforming enhancement signal according to the spectral gain coefficient to output a noise-reduced speech signal.