Intelligent sound box voice noise reduction method based on edge computing
Patent Information
- Application Number
- CN202611268583.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-20
- Publication Date
- 2026-09-29
AI Technical Summary
[0005]本发明的一个目的在于提出基于边缘计算的智能音箱语音降噪方法,针对复杂声场中持续噪声与瞬态干扰动态特征不同、语音重叠时噪声记忆易受污染、固定深度推理增加边缘运算且路径切换易引起增益抖动的问题,提出由多通道复数频谱形成目标语音波束和阻塞参考波束,基于语音存在概率、归一化谱通量、分频带能量比及阵列相干性,以具有不同更新系数的背景噪声存储单元和瞬态噪声存储单元建立双噪声记忆,在语音与瞬态噪声重叠时冻结背景记忆并启用语音频带保护图,再以跨帧统一控制向量选择不同处理层数的因果增强路径并约束复数抑制增益的技术方案,由此形成面向持续噪声、瞬态干扰和语音保护的协同控制,并在本地重构供唤醒或指令识别使用的增强语音
[0048]1、背景噪声存储单元与瞬态噪声存储单元采用有序更新系数,并由语音存在、谱变化、频带能量及阵列相干性共同确定状态,使持续噪声描述能够稳定更新,同时使瞬态噪声描述能够响应突发变化。
Smart Images

Figure CN122842601A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio signal processing, and more particularly to a method for voice noise reduction in smart speakers based on edge computing. Background Technology
[0002] Smart speakers typically collect far-field speech through microphone arrays and perform echo cancellation, beamforming, noise suppression, and speech recognition locally or in the cloud. Continuous noise from kitchen appliances, air conditioners, and other equipment exhibits slow-changing characteristics, while collision sounds, tableware noises, and sudden bursts of noise from multiple people during living room gatherings show short-term energy abrupt changes and spatial coherence variations. Furthermore, the content played through the speaker is coupled to the microphone via the room's acoustic path. Existing processing workflows often employ single noise estimation or a fixed update rate, which can easily allow transient interference to contaminate background noise estimation, or attenuate target speech harmonics and formants when speech overlaps with transient interference.
[0003] Edge devices are limited in computing power and power consumption. If all audio frames always call the enhancement network of the same depth, it will increase unnecessary computation under stable sound field. If the processing path is switched only according to instantaneous energy, it may cause path jitter and gain discontinuity. Existing solutions also have difficulty coordinating noise memory update, speech band protection, gain attack release and smooth path switching with the same state information, thus affecting the stability of wake-up and command recognition input.
[0004] Therefore, there is a need for a method to reduce voice noise in smart speakers that can address the shortcomings of existing technologies. Summary of the Invention
[0005] One objective of this invention is to propose a speech noise reduction method for smart speakers based on edge computing. Addressing the issues of differing dynamic characteristics between continuous noise and transient interference in complex sound fields, the susceptibility of noise memory to contamination during speech overlap, and the increased edge computation and gain jitter caused by path switching in fixed-depth inference, this invention proposes a method that uses multi-channel complex spectra to form a target speech beam and a blocking reference beam. Based on speech presence probability, normalized spectral flux, frequency band energy ratio, and array coherence, a dual-noise memory is established using background noise storage units and transient noise storage units with different update coefficients. When speech overlaps with transient noise, the background memory is frozen and a speech band protection map is activated. Furthermore, a cross-frame unified control vector is used to select causal enhancement paths with different processing layers and constrain complex suppression gain. This forms a collaborative control approach for continuous noise, transient interference, and speech protection, and locally reconstructs enhanced speech for wake-up or command recognition.
[0006] This invention provides a method for voice noise reduction in smart speakers based on edge computing, comprising: S1, simultaneously acquiring multi-channel far-field audio from at least two microphones, and obtaining multi-channel complex spectra through frame-by-frame time-frequency transformation; S2, forming a target speech beam complex spectrum and a blocking reference beam complex spectrum from the multi-channel complex spectra, and generating an acoustic feature set; S3, updating the background noise description with a first update coefficient and updating the transient noise description with a second update coefficient greater than the first update coefficient according to the blocking reference beam complex spectrum, the acoustic feature set, and the unified control vector of the previous frame, freezing the update of the background noise description and generating a speech band protection map when the speech overlaps with the transient noise, obtaining noise state data and updating the unified control vector of the previous frame accordingly. S4. Input the target speech beam complex spectrum, the acoustic feature set, and the noise state data into the causal enhancement network of the local edge audio processor. The unified control vector of the current frame is used to select one of the following inference paths: a first inference path, a second inference path, or a transient refinement inference path. The second inference path adds a causal temporal processing layer compared to the first inference path, outputting a complex suppression gain and its confidence level. The complex suppression gain is constrained according to the unified control vector of the current frame and the speech band protection map, and the unified control vector of the next frame is updated according to the confidence level. S5. The constrained complex suppression gain is applied to the target speech beam complex spectrum and an inverse time-frequency transform is performed to obtain locally enhanced speech.
[0007] Optionally, S1 includes: further including: obtaining a speaker playback reference, and aligning the audio sampling sequence of each microphone with the speaker playback reference according to a unified sampling clock;
[0008] The echo estimate of each microphone channel is obtained by adaptive echo path filtering based on the speaker playback reference, and the echo estimate is subtracted from the corresponding audio sampling sequence;
[0009] The audio sampling sequences of each channel after subtraction are framed and subjected to short-time Fourier transform using the same frame length, frame shift, and analysis window. The frame length, frame shift, and analysis window are recorded as local configuration parameters in the local edge audio processor, and the multi-channel complex spectrum with time index and frequency index aligned is output.
[0010] Optionally, S2 includes: weighting the multi-channel complex spectrum according to the steering vector of the microphone array to synthesize the target speech beam complex spectrum, and generating the blocking reference beam complex spectrum using a blocking matrix that has zero response to the steering vector;
[0011] The probability of speech presence is generated based on the energy relationship between the complex spectrum of the target speech beam and the complex spectrum of the blocking reference beam. The normalized spectral flux is generated based on the ratio of the amplitude spectrum difference between adjacent audio frames to the sum of the amplitude spectra of the previous audio frame. The frequency band energy ratio is generated based on the ratio of the energy of each frequency band to the energy of the entire frequency band under the preset frequency band division. The array coherence is generated based on the cross power spectrum and the power spectrum of each microphone channel pair.
[0012] In the first audio frame or the first frame after the state is reset, the amplitude spectrum of the previous audio frame is initialized to the amplitude spectrum of the current audio frame.
[0013] The input source, time index, frequency index, and normalization parameter of each feature are associated with the acoustic feature set as feature records.
[0014] Optionally, S3 includes: updating the background noise description includes: setting the background noise storage unit and the background state;
[0015] The speech probability threshold and the spectral flux threshold are determined by the speech presence probability in the historical audio frames without wake-words and the quantile of the normalized spectral flux, respectively. The array coherence is weighted and aggregated at the corresponding frequency points in the neighborhood of the target direction to obtain the coherence statistics in the target direction. The coherence threshold is determined by the quantile of the coherence statistics in the target direction.
[0016] The current blocking reference power spectrum is obtained by squared the modulus of the complex spectrum of the blocking reference beam;
[0017] When the unified control vector of the previous audio frame allows background updates, and the probability of speech presence is less than the speech probability threshold, the normalized spectral flux is less than the spectral flux threshold, and the target direction coherence statistics are less than the coherence threshold, the first update coefficient is multiplied by the current blocking reference power spectrum and added to a product of the first update coefficient and the previous background noise description to obtain the current background noise description, and the current background noise description is written into the background noise storage unit.
[0018] Write the update permission state of the background noise storage unit and the data source of the threshold into the background state;
[0019] Furthermore, updating the transient noise description includes: setting a transient noise storage unit and a transient state;
[0020] The energy change threshold and the coherence change threshold are determined by the quantiles of the energy difference between adjacent frames in the historical audio frames without wake word and the difference in the coherence statistics of the target direction, respectively.
[0021] Three judgment values are generated based on the over-limit result of the normalized spectral flux relative to the spectral flux threshold, the over-limit result of the energy change between the current audio frame and the previous audio frame relative to the energy change threshold, and the over-limit result of the change between the current target direction coherence statistics and the previous target direction coherence statistics relative to the coherence change threshold.
[0022] When at least two judgment values indicate that the limit has been exceeded, the transient state is set to the entering state, and the current blocking reference power spectrum and the previous transient noise description are weighted using the second update coefficient to obtain the current transient noise description, and the current transient noise description is written into the transient noise storage unit.
[0023] In the first audio frame or the first frame after the state reset, the energy of the previous audio frame and the coherence statistics of the previous target direction are initialized to their respective current observation values, the transient state is initialized to the local initial state, and the transient noise description is initialized to the current blocking reference power spectrum.
[0024] When any input feature is missing or exceeds the effective range of the local configuration, an abnormal fallback is performed to keep the transient state in the state of the most recent valid audio frame and to limit the complex suppression gain after constraint to be no less than the lower limit of the safe gain.
[0025] Furthermore, generating the audio band protection map includes: setting a freeze flag;
[0026] When the transient state is in the entering state and the probability of the speech is not less than the speech probability threshold, the freeze flag and the speech band protection map are set to valid and the background noise storage unit update is paused. Harmonic protection regions are generated based on the fundamental frequency candidates and integer multiples of the target speech beam complex spectrum, and formant protection regions are generated based on the frequency band where the spectral envelope peak is located. The harmonic protection regions and the formant protection regions are merged into the speech band protection map.
[0027] If the conditions of the transient state being in the entry state and the probability of the speech being present being not less than the speech probability threshold are not met, the frequency protection flag of the speech band protection map will be set to invalid.
[0028] The number of frames exiting the decision and the noise description deviation are used as local verification configuration parameters.
[0029] If at most one of the three judgment values in an audio frame that continuously reaches the exit judgment frame number indicates an over-limit, the transient state is set to the exit state, and the transient noise description is attenuated frame by frame according to the locally configured noise description release coefficient until the transient noise description is not greater than the sum of the deviation between the background noise description and the noise description, and then the freeze flag is released.
[0030] Record the entry state, the exit state, and the corresponding determination reason.
[0031] Optionally, S4 includes: the causal reinforcement network includes a shared encoder, a reusable hidden state, a first inference path, a second inference path, the transient refined inference path, a complex gain output layer, and the same gain smoother;
[0032] The first inference path invokes the shared encoder and the complex gain output layer; the second inference path invokes at least one causal timing processing layer based on the first inference path; and the transient refinement inference path invokes a transient frequency band attention layer based on the second inference path.
[0033] Each inference path reads the reusable hidden state of the previous audio frame and provides the updated hidden state for the next audio frame to read;
[0034] In the first audio frame or the first frame after state reset, the reusable hidden state is initialized to a zero state, the previous background noise description is initialized to the current blocking reference power spectrum, the background state is initialized to an update-allowed state, the transient state is initialized to a non-entry state, the freeze flag is initialized to invalid, and the gain confidence of the previous audio frame is initialized to the locally configured initial confidence. The unified control vector of the previous audio frame is composed of the initialized noise state data and the initial confidence.
[0035] When the transient state is an entering state, the transient refined inference path is selected. When the background state indicates that updates are allowed and the gain confidence of the previous audio frame is not less than the confidence threshold, the first inference path is selected. In other states, the second inference path is selected, and the path identifier is recorded in the unified control vector.
[0036] Furthermore, constraining the complex suppression gain includes setting the lower limit of the gain of the effective frequency points covered by the speech band protection map as a first lower limit of gain, and setting the lower limit of the gain of the remaining frequency points as a second lower limit of gain that is less than the first lower limit of gain.
[0037] The confidence threshold is determined by the gain error quantile of the verified audio set. When the gain confidence of the previous audio frame contained in the unified control vector of the current audio frame is less than the confidence threshold, a first upper limit of allowable gain change is adopted. When the gain confidence of the previous audio frame is not less than the confidence threshold, a second upper limit of allowable gain change that is greater than the first upper limit of allowable gain change is adopted.
[0038] When the transient state switches from a non-entry state to an entry state, the complex suppression gain is updated using a gain attack coefficient; when the transient state switches from an entry state to an exit state, the complex suppression gain is updated using a gain release coefficient that is smaller than the gain attack coefficient.
[0039] When the path identifier is switched, the gain before and after the switch is passed through the same gain smoother, and the constraint verification result is whether the constrained complex suppressed gain meets the first lower gain limit, the second lower gain limit and the corresponding upper limit of allowable gain change.
[0040] Furthermore, the selection of inference path also includes: generating a computational budget state based on the local edge audio processor's processor utilization rate, single-frame processing time, and remaining power within a preset statistical period;
[0041] When the computational budget state indicates that the single-frame processing time has reached the frame shift corresponding time or the remaining battery power is less than the battery power threshold, the transient refinement inference path is maintained when the transient state is in the entry state and the probability of speech presence is not less than the speech probability threshold. When the transient state is in the entry state and the probability of speech presence is less than the speech probability threshold, the transient refinement inference path is reversed to the second inference path. In other states, the second inference path is reversed to the first inference path.
[0042] When the computational budget state is restored to a state where the single-frame processing time is less than the frame shift corresponding time and the remaining power is not less than the power threshold, and multiple consecutive audio frames satisfy the selection conditions of the first inference path, the second inference path, or the transient refined inference path, the target path is restored to be determined according to the corresponding selection conditions.
[0043] The path identifier is updated only when the target path is consistent across multiple consecutive audio frames, and the reason for rollback, the reason for recovery, the path identifier before update, and the path identifier after update are written into the path switching record.
[0044] Optionally, S5 includes: multiplying the constrained complex suppression gain by the frequency-point complex spectrum of the target speech beam to obtain the enhanced complex spectrum;
[0045] The enhanced complex spectrum is subjected to inverse short-time Fourier transform and overlapped summation using a synthesis window that matches the analysis window used in the frame-segmented time-frequency transformation of step S1 to obtain the local enhanced speech;
[0046] When the freeze flag is valid, the local enhanced speech, the speech presence probability, and the transient state are sent to the wake-up recognizer or the instruction recognizer, and the time index, path identifier, and constraint verification result of the recognition input frame are recorded.
[0047] The beneficial effects of this invention are:
[0048] 1. The background noise storage unit and the transient noise storage unit adopt ordered update coefficients, and the state is jointly determined by the presence of speech, spectral changes, frequency band energy and array coherence, so that the continuous noise description can be updated stably, while the transient noise description can respond to sudden changes.
[0049] 2. When speech overlaps with transient noise, the background noise storage unit is frozen and the corresponding speech band protection map is enabled for harmonic and formant protection. Combined with confidence, gain lower limit and attack release constraints, the risk of target speech being falsely suppressed and background noise memory being contaminated by speech can be reduced.
[0050] 3. Unified control vector coordinates noise memory, inference path and gain constraints across frames. The causal enhancement network selects paths with different processing layers based on shared encoding and hidden states, and outputs through the same gain smoother, which helps reduce discontinuity in path switching and adjusts the average edge computation under a stable sound field. Attached Figure Description
[0051] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0052] Figure 1 This is a flowchart of the edge computing-based voice noise reduction method for smart speakers according to the present invention.
[0053] Figure 2 This is a flowchart of the S3 noise state management and unified control vector update of the present invention. Detailed Implementation
[0054] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0055] refer to Figures 1-2A speech noise reduction method for smart speakers based on edge computing includes: S1, simultaneously acquiring multi-channel far-field audio from at least two microphones, and obtaining multi-channel complex spectra through frame-by-frame time-frequency transformation; S2, forming a target speech beam complex spectrum and a blocking reference beam complex spectrum from the multi-channel complex spectrum, and generating an acoustic feature set; S3, updating the background noise description with a first update coefficient and updating the transient noise description with a second update coefficient greater than the first update coefficient according to the blocking reference beam complex spectrum, the acoustic feature set, and the unified control vector of the previous frame, freezing the update of the background noise description and generating a speech band protection map when the speech overlaps with the transient noise, obtaining noise state data and updating the unified control vector of the previous frame accordingly. S4. Input the target speech beam complex spectrum, the acoustic feature set, and the noise state data into the causal enhancement network of the local edge audio processor. The unified control vector of this frame selects one of the following inference paths: a first inference path, a second inference path, or a transient refinement inference path. The second inference path adds a causal temporal processing layer compared to the first inference path, outputting the complex suppression gain and its confidence level. The complex suppression gain is constrained according to the unified control vector of this frame and the speech band protection map, and the unified control vector of the next frame is updated according to the confidence level. S5. Apply the constrained complex suppression gain to the target speech beam complex spectrum and perform an inverse time-frequency transform to obtain the locally enhanced speech.
[0056] In this specific embodiment, S1 includes:
[0057] In this specific embodiment, far-field audio acquisition is performed by a circular array with four digital microphones, a speaker playback reference acquisition interface, and a local edge audio processor. The four microphones use the array center as a common reference point and are provided with a unified sampling clock of 16kHz and 24bit by the same audio codec. The acquisition record includes device identifier, channel identifier, monotonically increasing sampling sequence number, and local timestamp. The speaker playback reference record uses the same sampling sequence number system. The acquisition interface submits a set of 160 channel data every 10ms. When a change in sampling sequence number is detected, the set of data is discarded and the most recent complete set of data for each channel is marked as to be resynchronized input.
[0058] The local edge audio processor uses the speaker playback reference as a fixed time base and calculates an 80ms sliding cross-correlation with the sampling sequences of the four microphone channels. After deducting the geometric propagation delay from the speaker to each microphone recorded in the array calibration area, four acquisition link offsets are obtained. The median of the four offsets is used as the common acquisition offset. When the correlation peak is consistent with the zero-delay common offset for three consecutive windows and the residual of each channel relative to the median does not exceed one sampling point, the common offset is written into the time alignment record. Only the four microphone channels are shifted in the same direction by integer sampling points while keeping the playback reference unchanged. The remaining part of the common offset that is less than one sampling point is processed using the same set of third-order Lagrange fractions. Delay filter compensation ensures that the physical relative time delay between array channels remains unchanged. When the correlation peak is less than the 0.35 threshold obtained by playing a sweep signal in a quiet room before deployment, echo path updates are prohibited. If a previous valid common offset exists, that offset is maintained. If no previous valid value exists after the device's first startup, sampling clock reset, or buffer clearing, the deployment common offset saved in the array calibration area with the configuration version is read as the deterministic initial value. If the deployment calibration record is also unavailable, a fixed 0-sampling-point common offset is used, the alignment source is recorded as a safe initial value, and the current frame is marked as pending resynchronization. Echo path updates are still frozen until the correlation peak meets the consistency condition for three consecutive windows. After this, the newly measured common offset overwrites the initial value, and the pending resynchronization flag is cleared.
[0059] Integer sampling offset according to Confirmed, among which This represents the integer sample offset of the m-th channel relative to the playback reference, in sample points. A positive value indicates that the microphone channel lags behind the playback reference and requires all four channels to be moved forward by a common offset, while a negative value indicates that all four channels need to be moved backward by a common offset. This represents the microphone channel index, ranging from 1 to 4. Indicates the sampling index within the alignment window and takes to , This indicates a candidate offset limited to sampling points between -32 and +32. This indicates the starting sampling index of the 80ms alignment window. This indicates that the alignment window length is 1280 points. This represents the normalized magnitude of the m-th channel before alignment. This represents the playback reference normalized amplitude at the corresponding candidate offset. This indicates the candidate offset when the cross-correlation sum is maximized. If the offset range of three consecutive windows exceeds 1 point, it is determined that the clock has not yet stabilized and continues to be cached without being submitted downstream.
[0060] The aligned playback reference drives the normalized minimum mean square adaptive echo path filter for each microphone channel, and the echo estimation is compared with the residual signal according to... and Calculation, where This represents the echo estimate of the m-th channel at sampling point q, expressed in normalized amplitude. This represents the microphone channel index, ranging from 1 to 4. Indicates the sampling point index under a unified sampling clock. This represents the number of echo path filter taps, and in this specific embodiment, it is set to 2048. This represents the dimensionless coefficient of the ℓth tap in the m-th channel at sampling point q. This indicates the tap index, ranging from 0 to 2047. This represents the playback reference normalized amplitude with a delay of ℓ sampling points. This represents the normalized amplitude of the residual signal in the m-th channel after subtracting the echo estimate. This represents the normalized amplitude of the m-th input channel after alignment;
[0061] Echo path updates use S2 to write the frame-level speech presence probability to the cross-frame buffer after the previous complete audio frame ends. The first frame or the first frame after state reset initializes this buffer value to 1 to freeze the update. The frame to which the current sample belongs only reads the buffer value with the time index of the previous frame, thus not relying on the S2 result that has not yet been generated for the current frame. When the current read value is lower than 0.2 and the playback reference power is higher than the deployment calibration of -45dBFS, the update proceeds according to... Update the echo path, where This represents the dimensionless echo path coefficient used at the next sampling point. This represents the dimensionless update step size, which is set to 0.12 in this specific implementation. This represents a smoothing amount to prevent the reference power normalization factor from reaching zero, and is set to... , This indicates the index for tap summation, ranging from 0 to 2047. The normalized amplitude of the playback reference is delayed by u sampling points. The smoothing amount and the reference sum of squares together form a dimensionless reference energy. The coefficients are frozen when the playback reference power is insufficient, the buffer time index is discontinuous, or near-end speech is detected to avoid echo path drift caused by double talk.
[0062] The four-channel residual signal after echo subtraction is subjected to frame-by-frame time-frequency transformation using a 320-point square root Hanning analysis window, a 160-point frame shift, and a 512-point Fast Fourier Transform. The complex spectrum is then processed according to... Generate, where This represents the complex spectrum of the m-th channel, frame t, and frequency point k, with the amplitude consistent with the input normalized amplitude. Indicates the audio frame time index. This represents a frequency index, ranging from 0 to 256. Indicates the index of the intra-frame sampling point. This indicates the analysis window length is 320 points. This indicates the number of points in the Fast Fourier Transform, which is 512 points. This indicates a frame shift of 160 points. This represents the echo subtraction residual signal corresponding to the sampling point of the m-th channel. This represents the dimensionless window coefficient of the square root Hanning analysis window. Represents the imaginary unit. Represents pi (π). Indicates complex exponentiation;
[0063] Before frame division, the clipping ratio is calculated for each channel. ,in This represents the dimensionless clipping ratio of the m-th channel in the t-th frame. This indicates the length of the 320-point analysis window. Indicates the index of the intra-frame sampling point. This is an indicator function that takes the value 1 if the condition is true and 0 otherwise. This represents the normalized amplitude of the corresponding residual signal. This represents the dimensionless clipping threshold relative to the normalized full scale of 1. In this specific implementation, the clipping calibration is fixed at 0.98 based on the deployment of the 24-bit acquisition link and written into the configuration version. When the clipping ratio of any valid channel exceeds 0.02, the abnormal flag of the frame is set to clipped and the echo path update is frozen, but the spectrum is still reserved for subsequent modules to decide whether to continue using it based on the validity mask.
[0064] The local configuration area stores the above parameters with fields such as configuration version, sampling rate, frame length, frame shift, number of fast Fourier transform points, analysis window check value, and time alignment record version. In this specific implementation, version number AUD-FRAME-01 is used. When the device starts up, the sampling clock is reset, or the configuration check fails, the cross-frame buffer is cleared and the time index of the first frame is set to zero. Only frames with the same sampling sequence number range and frequency index range of the four channels are written into the multi-channel complex spectrum ring buffer.
[0065] Each frame record in the circular buffer includes a time index, sampling start and end sequence numbers, 257 one-sided complex frequency points for the four channels, a valid playback reference record identifier, an actual playback activity identifier, an echo path version, and an anomaly identifier. The valid playback reference record identifier only indicates that the record is consistent with the time range of the current frame, has a continuous sampling sequence number, and has a finite sample value. It is valid regardless of whether the speaker is actually playing. Valid reference records during silent periods are saved as zero or near-zero value samples collected during the corresponding time period. The actual playback activity identifier is determined based on whether the frame average power of the same valid reference record is higher than the -45dBFS threshold set by the deployment calibration. If it is higher than the threshold, it is set to play; if it is lower than the threshold, it is set to no play. The two cannot be substituted for each other. If a single microphone channel is missing two consecutive frames, the channel is marked as invalid and the submission of new frames to the subsequent beamforming module is suspended. It will resume from the first frame after the state is reset after the four channels pass the clock consistency check again.
[0066] Frame submission criteria according to Calculation, where This indicates that the dimensionless overall valid identifier of the t-th frame is either 0 or 1. This indicates a valid channel identifier where the sampling sequence number of the m-th microphone channel is consecutive and aligned, and is set to 1. This indicates a valid reference flag that takes the value 1 when the playback reference record matches the time range of the current frame. This indicates a spectral quality flag that is set to 1 when the Fast Fourier Transform output has no non-numerical values and the frequency index is complete. When the product is 1, the frame is written to the readable queue; when the product is 0, the failure field is written to the exception record and the system waits for the next complete frame.
[0067] The above processing yields a multi-channel complex spectrum with strictly aligned time and frequency indices, and its buffer address, frame time index, effective channel mask, and analysis window configuration version are passed as the output descriptor of S1 to the beam and acoustic feature generation module of S2.
[0068] In this specific embodiment, S2 includes:
[0069] S2 reads the multi-channel complex spectrum corresponding to the output descriptor of S1, and obtains the microphone's three-dimensional coordinates, target azimuth, sound speed, and array steering vector version from the array calibration area. In this specific implementation, the front of the smart speaker is taken as the target direction, and the sound speed is set to 346m / s at 25 degrees Celsius. After each successful wake-up positioning, the target azimuth is updated with the arrival time difference of the last five frames. When the positioning confidence is less than 0.6, the previous valid target azimuth is used.
[0070] The beamforming module organizes the four channel complex values at the same time and frequency position into a column vector, and then... Forming the complex spectrum of the target speech beam, where This represents the complex spectrum of the target speech beam at the k-th frequency point in the t-th frame. This represents a complex column vector consisting of the spectra of the four channels. This represents the complex steering vector at the k-th frequency, representing the target direction. The denominator represents the conjugate transpose of the guiding vector. The dimensionless energy normalization factor represents the steering vector, which is calculated from the array coordinates, target azimuth angle, and sound velocity and stored in the array calibration area.
[0071] The blocking module generates a blocking reference using a three-column blocking matrix obtained through Gram-Schmidt orthogonalization, satisfying... and ,in This represents the complex blocking matrix that exhibits zero response at the k-th frequency to the target direction. This indicates its conjugate transpose. This represents the complex vector of the three-channel blocking reference beam at frequency k in frame t. Indicates and A three-dimensional complex zero vector of the same dimension, where each element is dimensionless zero. Its origin lies in the orthogonal constraint of the blocking matrix on the target guidance vector. The matrix is recalculated each time the target azimuth angle changes by more than 5 degrees, with the orthogonal residual being less than... As a valid condition, if the residual does not meet the condition, the state update of the current frame is stopped and the previous valid matrix is used.
[0072] The feature generation module first calculates the target beam power and the blocking reference average power, then... and Speech generation is probabilistic, where This represents the dimensionless proportion of the target beam energy in the total energy of the target beam and the blocking reference. This represents the squared mean of the three-channel blocking reference mode, expressed in normalized power. Represents the smoothing amount of unit power and is set as follows: , This indicates the probability of speech occurring at a given frequency point, ranging from 0 to 1. and The dimensionless proportional calibration points representing the no-speech and high-confidence speech, respectively, are determined by the 10% and 90% quantiles of the no-wake word set and the proximal wake word set. Calibration records are only available on [date / time]. If the strict sequence is not met, the previous valid calibration pair will be used. If the device is started for the first time and there is no previous valid record, it will fall back to the preset dimensionless calibration pairs 0.15 and 0.65 and write the fallback source. This means restricting the input to 0 to 1;
[0073] Record the target speech beam amplitude spectrum into the previous frame buffer, and then... Calculate the normalized spectral flux, where Let represent the dimensionless normalized spectral flux of the t-th frame. This represents the amplitude of the target speech beam at the k-th frequency point in the t-th frame. This indicates the amplitude of the corresponding frequency point in the previous frame. This indicates a one-sided frequency index, and the numerator and denominator of this formula are summed over the same complete set of frequency points from 0 to 256. Represents the smoothing amount in the same unit as the amplitude and set to . , This means that only positive amplitude changes are retained. In the first frame or after the state is reset, the current amplitude spectrum is written to both the current record and the previous frame buffer, so that the spectral flux is initialized to zero.
[0074] The preset frequency band division uses four non-overlapping frequency bands that completely cover the frequency range of 0 to 256: 0 to 300Hz, 300 to 1000Hz, 1000 to 3000Hz, and 3000 to 8000Hz. First, according to... Calculate the total power of the target speech beam, then press Generate the frequency band energy ratio, where This represents the total power of the target speech beam at frequencies 0 to 256 in frame t, expressed in normalized power. This represents the dimensionless energy ratio of the b-th frequency band in the t-th frame. This indicates a frequency band index, ranging from 1 to 4. This represents the set of frequency indices corresponding to the b-th frequency band, where the sum of the four energy ratios is strictly 1; when Below the positive smoothing threshold in units of frame power The ratio calculation is not performed and all four energy ratios are marked as invalid. This represents a strictly positive effective threshold for total power and is set to... The frequency band division and index boundaries are stored in the feature configuration version AUD-FEAT-01;
[0075] Array coherence for six different microphone channels to press Calculation, where This represents the dimensionless coherence between channel a and channel b at frequency k in frame t, and is restricted to between 0 and 1. and Indicates the index of different microphone channels. This represents the exponentially smoothed cross-power spectrum of the channel pair. and This represents the exponentially smoothed autopower spectrum of the corresponding channel; all three values are in the unit of normalized power. Represents a smoothed quantity on the order of the square of power and is set as follows: Mutual power and self power are updated with a weight of 0.2 for this frame.
[0076] The acoustic feature set stores the frequency point speech existence probability, frame-level speech existence probability, normalized spectral flux, four frequency band energy ratios, six sets of array coherence, target azimuth angle, and validity mask using time index as the primary key. The frame-level speech existence probability is taken as the energy-weighted average of the frequency point probabilities within 300 to 3000 Hz. If the effective power of any denominator is less than ten times the smoothing amount, the corresponding feature is set to invalid instead of being substituted with a zero value to represent the observation.
[0077] Each feature record is simultaneously written to the input buffer address, time index, frequency index, normalization parameter version, array calibration version, and previous frame buffer version. After completion, the target speech beam complex spectrum, the blocking reference beam complex spectrum, and the read-only handle of the acoustic feature set are passed to S3 as output of S2, and the read-only handle of the target speech beam complex spectrum is provided to S4 in parallel.
[0078] In this specific embodiment, S3 includes:
[0079] S3 reads the complex spectrum and acoustic feature set of the blocking reference beam from S2, and reads the background freeze update qualification, previous transient state, freeze flag, path flag and previous gain confidence from the unified control vector of the previous frame. The background freeze update qualification is a cross-frame field that is only changed by initialization, freeze setting and freeze release. The threshold judgment result of this frame is a frame-by-frame field that only records whether the current frame is actually updated and the reason for not updating. The two are independent of each other. The background noise storage unit and the transient noise storage unit both store the power spectrum with time index and frequency index as keys. The background state also stores the threshold version, the most recent update time, the background freeze update qualification, the threshold judgment result of this frame and the reason for not updating. The transient state also stores the entry or exit reason and the continuous exit count.
[0080] During the deployment and calibration phase, from a cumulative total of no less than 30 minutes of wake-word-free historical audio, only frames with valid playback reference record identifiers and no actual playback activity identifiers are selected. Manually annotated speech frames are further excluded. Empirical distributions are formed for frame-level speech existence probability, normalized spectral flux, energy change, and target direction coherence statistics change. Both the valid playback reference record identifier and the actual playback activity identifier directly follow the definition of S1 and are recorded using the same time index. The former is used to prove the completeness and time alignment of the reference record, while the latter is used to exclude frames actually played by the speaker. Therefore, frames with complete reference records during periods of silence can be used as... The calibration samples are as follows: the speech probability threshold is taken as the 80th quantile of the speech presence probability, the spectral flux threshold is taken as the 85th quantile of the normalized spectral flux, and the energy change threshold and coherence change threshold are taken as the 90th quantile of their respective absolute changes. In both the calibration and operation phases, the sum of the blocking reference power spectrum at frequencies from 0 to 256 is used as the total frame energy, and the same logarithmic smoothing, validity, and previous frame continuity rules are adopted. The threshold record includes the room configuration identifier, quantile, number of sample frames, effective range, and version AUD-THR-01. A new candidate version is generated only when the cumulative number of newly added effective audio without wake-word reaches 10 minutes and the recognizer wakes up without error.
[0081] Target direction coherence statistics according to Calculation, where This represents the dimensionless target direction coherence statistic in frame t, ranging from 0 to 1. This represents the set of frequency indices corresponding to the main lobe in the target direction within the range of 300 to 3000 Hz. This indicates that the dimensionless weight at the k-th frequency point obtained by normalizing the steering vector response has a weight sum of 1. This represents the mean coherence of the six microphone channels. The coherence threshold is the 85th percentile of this statistic in the history frames without wake-words and is recorded in the same threshold record.
[0082] After weighted aggregation of the array coherence at corresponding frequency points in the target direction neighborhood to obtain the target direction coherence statistics, the mean squared modulus of the complex vector of the three-channel blocking reference beam is used as the current blocking reference power spectrum, and then... Calculation, where This represents the current blocking reference power spectrum at frequency k in frame t, expressed in normalized power. This represents the complex vector of the three-channel blocking reference beam output by S2. The L2 norm of a complex vector is... The number of blocking reference channels is a dimensionless positive integer and is fixed at 3 in this specific implementation. Its source is the three-dimensional orthogonal complement space formed by the four-channel array relative to a single target guide vector.
[0083] If the background freeze update qualification in the current frame's unified control vector is allowed, and the frame-level speech presence probability is less than the speech probability threshold, the normalized spectral flux is less than the spectral flux threshold, and the target direction coherence statistic is less than the coherence threshold, then according to... Update the background noise description, where This describes the current background noise level, expressed in normalized power. This indicates a description of the background noise in the previous frame. The first update coefficient is a dimensionless number, which is calibrated to 0.02 by the stable noise verification set in this specific embodiment. When the threshold condition is not met, only the previous background noise description is maintained and the condition combination that has not been updated in this frame is recorded. The background freeze update qualification for the next frame is not cleared.
[0084] Transient Detector First Press Calculate the total blocking reference energy for the current frame, then press Calculate the absolute logarithmic difference of energy relative to the previous consecutive frames, where and These represent the sum of the blocking reference power spectra at frequencies 0 to 256 in the current frame and the previous consecutive frame, respectively, in normalized power. This represents the dimensionless log-absolute difference of energy in frame t. Represents a strictly positive smooth quantity in the same unit as the total energy and is set as follows: , Represent the natural logarithm; then calculate the absolute difference of the coherence statistic in the target direction, and normalize the spectral flux to exceed the spectral flux threshold. Three Boolean decision values are generated when the energy change threshold and the absolute difference of the coherence statistic exceeds the coherence change threshold, respectively. When at least two of the three decision values are true, the transient state is set to the entering state, and then... Update the transient noise description, where This describes the current transient noise, expressed in normalized power. This indicates a description of transient noise in the previous frame. This represents the second update coefficient, which is dimensionless. In this specific embodiment, it is calibrated to 0.35 and satisfies... If the entry condition is not met, the current transient description is retained for use in the exit determination.
[0085] In the first frame after the state reset, the current blocking reference total energy, target direction coherence statistics, and current blocking reference power spectrum, calculated using the same caliber, are written into the previous frame energy, previous frame coherence statistics, and background and transient noise storage units, respectively. Background freeze update qualification is set to allowed, and the current frame threshold determination result is initialized to not executed. The transient state is initialized to a local initial state, which is a non-entry state. Simultaneously, the continuous exit count is cleared. When the time index is discontinuous, the absolute logarithmic difference of energy is not calculated; the previous frame energy buffer is reconstructed according to the first frame rules. The validity of acoustic features is checked separately for each feature, including the probability of speech presence, coherence statistics, and absolute logarithmic difference. The effective range of the difference is 0 to 1. The effective range of the normalized spectral flux and the absolute logarithmic difference of energy is a non-negative finite value and does not exceed the upper bound of the 99.9% quantile of the same type of no-wake-word calibrated distribution in the threshold record. Abnormal backoff is triggered only when the input is missing, a non-numerical or infinite value appears, or the effective range of each is exceeded. When triggered, the most recent effective transient state and the two types of noise description are maintained, the consecutive exit count is not increased, and the abnormal backoff flag and a safety gain lower limit of 0.25 are written into the unified control vector of this frame for S4 to execute. Subsequently, the abnormal backoff flag is removed, the normal judgment is restored, and the restored time index is written into the background state only when the acoustic feature is effective for three consecutive frames and the time index is continuous.
[0086] When the transient state is entered and the probability of the presence of frame-level speech is not less than the speech probability threshold, the freeze flag is set to valid, the background freeze update qualification is set to prohibited, and the background noise storage unit update is paused. The fundamental frequency detection performs consistency screening on the normalized autocorrelation peak and cepstral peak of the target speech beam in the range of 80 to 350 Hz. When the peak difference does not exceed 5 Hz and the confidence level is not less than 0.65, the fundamental frequency and its integer multiples of 8000 Hz are expanded into a harmonic protection region of one frequency point on each side. At the same time, in the logarithmic spectrum envelope of 300 to 3400 Hz, the frequency band where the local peak value is higher than the adjacent valley value by 6 dB is retained as the formant protection region.
[0087] The speech band protection map stores the union of the harmonic protection region and the formant protection region using a 257-bit frequency point mask, and records the fundamental frequency, spectral envelope version, generation time index, and valid identifier. When the fundamental frequency candidate does not meet the consistency condition, only the formant protection region is used. When neither has a valid result, the freeze identifier is maintained, the protection map is marked as having no valid frequency points, and the protection map empty backoff identifier is set in the unified control vector, so that S4 adopts a safety gain lower limit of 0.25 for all frequency points without forging speech bands. When a subsequent frame generates at least one valid protection frequency point or the freeze identifier is released, the protection map empty backoff identifier is cleared. When the transient state is not entered or the probability of frame-level speech existence is lower than the threshold, all frequency point protection identifiers are set to invalid.
[0088] The number of exit judgment frames and the noise description deviation are used as local verification configuration parameters. In this specific implementation, the number of exit judgment frames is set to 8 frames, and the noise description deviation is set to 20% of the corresponding background noise description. When at most one judgment value exceeds the limit in each of the 8 consecutive frames, the transient state is set to the exit state, and then... Release frame by frame, of which This represents the dimensionless noise description release factor, set to 0.85. This indicates that the maximum power value among the two items is obtained point-by-point; the comparison is released according to... Execution per frequency point, where This represents the dimensionless release factor, which describes the background noise plus a 20% bias. This represents the strictly positive release comparison floor value in the same unit as power and is fixed at 1. Normalized power, which is derived from the deployment resolution lower limit of the noise state storage format; when all 0 to 256 frequency points meet the release comparison, the transient noise description of each frequency point is first written back to the corresponding background noise description, then the freeze flag is cleared, the background freeze update qualification is restored to allowed, and the continuous exit count is cleared, so that the release is determined within a limited frame when the background power is zero.
[0089] The noise status data records the current background noise description, current transient noise description, background status, transient status, freeze flag, voice band protection map, three judgment values, continuous exit count, threshold version, abnormal backoff flag and protection map empty backoff flag, and merges it with the path flag and gain confidence in the unified control vector of the previous frame into the unified control vector of this frame. The unified control vector is written to the cross-frame status cache with a fixed field order and cyclic redundancy check value.
[0090] After completing the state writing, S3 passes the noise state data, speech band protection map and read-only handle of the unified control vector of this frame to S4, synchronously registers the frame-level speech existence probability, transient state and freeze flag in the output index for S5 to call, and updates the total energy of this frame and the target direction coherence statistics to the comparison benchmark of the next frame.
[0091] In this specific embodiment, S4 includes:
[0092] S4 reads the target speech beam complex spectrum and acoustic feature set from S2, the noise state data from S3, and the unified control vector for the current frame from the local edge audio processor. It also reads the amplitude limiting security request with the target time index equal to the current frame from the independent cross-frame security request cache, and the actual application gain record with the time index equal to the previous consecutive frame from the S5 actual application gain write-back cache, verifying the consistency of each object according to the time index. The network input concatenates the target beam real part, imaginary part, amplitude logarithmic value, frequency point speech presence probability, frequency band energy ratio, array coherence mean, background noise logarithmic power, transient noise logarithmic power, guard map frequency point identifier, and various feature validity masks into features for each frequency point. The background noise logarithmic power and transient noise logarithmic power... The power is first calculated by taking the maximum value of the logarithmic smoothing base value between the corresponding non-negative power and the strictly positive and same power units, and then calculating the natural logarithm. The two smoothing base values are fixed by the feature configuration version AUD-FEAT-01 according to the deployment power quantization resolution and remain consistent during training and operation. When the power is missing, non-finite, or not higher than the base value, the natural logarithm of the base value is used as the determined padding value, the validity mask for this type is set to 0, and the base value version is recorded. When three consecutive frames show finite power that is higher than the base value, the measured logarithmic value is used again and the mask is set to 1, thus ensuring that the network input is still finite when there is silence, zero power, or numerical underflow. The four frame-level frequency band energy ratios are mapped according to the unique frequency band to which the current frequency point belongs, that is, only the k-th frequency point is read. ,in This represents the frequency band index corresponding to frequency point k in feature configuration version AUD-FEAT-01; when any sub-band energy ratio or array coherence mean is invalid, the mean of the same frequency point in the training set is used. As a fixed fill value, the corresponding validity mask is set to 0. When valid, the measured value is used and the mask is set to 1. The mask is broadcast to 257 frequency points in a fixed channel order. The same fill value, mapping and mask layout version is used for training and operation, so that invalid features are normalized to form a fixed zero center value and do not spoof valid observations.
[0093] Each input feature is... Normalization, in which This represents the dimensionless network input value of the f-th feature at the k-th frequency point in the t-th frame. Indicates the feature category index. Represents the eigenvalues before normalization. and Let represent the mean and standard deviation of features of the same type and frequency in the training set, respectively, with the units consistent with the original features. The smoothing amount represents the same unit as the original feature and is strictly greater than 0, preferably taking the median of the corresponding standard deviation. When the median is 0, the strict positive unit matching lower bound preset in the feature configuration is taken and the lower bound version is recorded. This means limiting the standardized results to between -5 and +5;
[0094] The causal reinforcement network employs a shared one-dimensional frequency convolutional encoder, a lightweight state adapter, a causal temporal processing layer consisting of two layers of gated recurrent units, a transient frequency band attention layer, a complex gain output layer, and a gain confidence output head. The shared encoder outputs a 257x64 feature tensor per frame, and the reusable hidden state consists of two layers of 128-dimensional vectors. All three inference paths read the two layers of reusable hidden states with the time index of the previous consecutive frame. The first inference path fuses the shared encoding result with the two layers of previous frame hidden states using the lightweight state adapter and sends the fused result to the complex gain output layer. The second inference path calls the first causal temporal processing layer based on this. The transient refinement inference path then calls the second causal temporal processing layer and the transient frequency band attention layer with the speech band protection map as the query mask.
[0095] The invoked causal timing processing layer is in accordance with Update hidden state, where This represents the 128-dimensional dimensionless hidden state vector of layer l in frame t. This represents the index of the causal time-series processing layer, and is either 1 or 2. This represents the mapping of the gated recurrent unit after training of the l-th layer. This represents the dimensionless input tensor of the current frame of the l-th layer, obtained from the shared encoder or the output of the previous layer. This indicates that the layer has a reusable hidden state with the time index of the previous consecutive frame. Layers not called by the selected path do not perform gated cyclic unit mapping. Instead, the lightweight state adapter reads the shared encoding result and the hidden state of the previous consecutive frame of the layer at the same time. It outputs the fusion feature for use by the complex gain layer and generates the corresponding 128-dimensional update state for this frame. It writes the state to the time index of this frame. If the state time index is not continuous or the state age exceeds 1 frame, the state of the layer is first reset to the zero vector and then the adaptation of this frame is performed. This ensures that any path actually consumes the state of the previous frame and the state read by the subsequent path always corresponds to the previous consecutive frame.
[0096] Network training used a smart speaker array to collect clean speech, continuous appliance noise, collision-type transient noise, and speaker playback mixed data in 12 room configurations. The training samples covered input signal-to-noise ratios from -5dB to 20dB. The target complex gain was obtained by limiting the amplitude of the complex ratio of the clean target speech spectrum to the mixed target beam spectrum to 0 to 1.05. The loss was composed of the complex spectrum mean square error, time-domain scale-invariant signal-to-noise ratio loss, and guard band oversuppression penalty weighted at 0.5, 0.4, and 0.1. The validation set was isolated from the room configurations and used to determine the gain error quantile threshold.
[0097] Complex gain output head according to The original complex suppression gain is formed, where This represents the dimensionless primitive complex suppression gain at the k-th frequency point in the t-th frame. and These represent the dimensionless inactive real part and the dimensionless inactive imaginary part generated by the complex gain output head, respectively. Indicates hyperbolic tangent activation. Represents the imaginary unit, and simultaneously... Forming network credibility, among which This represents the dimensionless network confidence level of the t-th frame. This represents the dimensionless frame-level inactive value generated by the credibility output header. This represents the temperature scaling parameter for the validation set and is a positive dimensionless number. Represents a logical function;
[0098] The complex gain output layer provides the original complex suppression gain and gain confidence, with the confidence level determined according to... Calibration, among which This represents the dimensionless gain confidence level of the t-th frame, and a larger value indicates a smaller prediction error. This represents the dimensionless original complex gain mean absolute error, weighted by the frequency energy, used in verification. This indicates the dimensionless 95th percentile of the error in the verification audio set, which is strictly greater than 0. When the statistical quantile is 0, the dimensionless positive lower limit is used. As And write the rollback source in the calibration record. This means that the result is restricted to 0 to 1. During runtime, the confidence output header is fitted to this calibration value by temperature scaling and the current value is written into the unified control vector.
[0099] The first frame or the first frame after state reset initializes the two reusable hidden states to zero vectors, initializes the previous background noise description to the current blocking reference power spectrum, initializes the background state to an update-allowed state, initializes the transient state to a non-entry state, initializes the freeze flag to invalid, and sets the previous frame gain confidence to the median confidence in the verification set of 0.72; at the same frequency point, it initializes the previous smooth gain amplitude to a dimensionless number 1, the previous phase field to 0 radians and sets the phase state valid flag to 0, initializes the complex gain verified in the most recent frame to a unit complex number 1, initializes the amplitude limiting security request to invalid and writes it into the initialization version, exempts the upper limit of ordinary inter-frame variation in the first frame but still executes the lower limit and upper limit projection of the current frame gain, and uses the unit complex number re-trimmed by the lower limit of the current frame gain as the deterministic security output when the first verification fails; these fields form the unified control vector of the previous frame, and after each inference path completes the calculation of the current frame, it writes the hidden state of the current frame and its time index obtained by the actual calling layer or lightweight state projection.
[0100] The path selector first executes the state rules defined in the claims and maintains candidate paths, actual execution paths, and formal recording paths respectively: When the budget is unconstrained, as long as the transient state is "entering," both the candidate path and the actual execution path are immediately set to transient refinement inference paths during the duration of this state, no longer subject to the frame-level speech presence probability as an additional condition. This state takes precedence over five consecutive frames of hysteresis in the normal path, and the formal recording path synchronously records the transient refinement inference path actually called in this frame; When the budget is constrained and the transient state is "entering," the transient refinement inference path is still executed if the frame-level speech presence probability is not lower than the speech probability threshold, and if it is lower than the threshold, it uniquely backslides to the second inference path; the transient state exits. If the background freeze update qualification is allowed and the gain confidence of the previous frame is not less than the confidence threshold of 0.65, then the first inference path is selected. In other states, the second inference path is selected. The actual execution path and the formal recording path are updated only if the ordinary candidate path is consistent for five consecutive frames. Before confirmation, the current formal recording path continues to be executed. If the short transient ends before the two frames exit confirmation, the already determined transient refined inference path or the budget-constrained second path completes the frames that have been entered and continues to execute the audio band protection and gain constraint. The three path identifiers and the frame-by-frame update order are all written into the unified control vector of this frame.
[0101] The lower limit of the gain amplitude for the frequency points covered by the voice band protection map is set to the first lower limit of 0.55, and the second lower limit of the gain for the remaining frequency points is set to 0.12. The second lower limit of the gain is only raised to the safe lower limit of 0.25 when the abnormal backoff flag or the protection map empty backoff flag is valid. The clipping safety request no longer raises the lower limit of the gain, but instead reduces the upper limit of the candidate gain amplitude through the clipping attenuation stage. The first lower limit of the gain of 0.55 in the protection band has higher priority than the second lower limit of the gain. The upper limit of the clipping gain for each frequency point must not be lower than the corresponding lower limit of the gain. The upper limit of the allowable change and the constraint verification still have the final constraint priority. The final gain of the previous consecutive frame in the S5 actual application gain write-back record is used as the reference for read-only inter-frame changes. The final gain of the previous frame that has been verified and actually submitted in S4 is used only when the write-back record does not exist. It must not be overwritten before the constraint verification of the current frame is completed. When the clipping safety request switches from invalid to valid or the request version is increased, only the independent clipping attenuation stage and the request consumption flag are updated. They are not reset or rewritten. The candidate gain is still projected onto the allowable variation range centered on the output of the previous real frame, considering the complex phase state, two reusable hidden states, or the most recently passed gain. Only after the output and constraint verification of the current frame are passed is the final gain and time index atomically updated to the candidate state of the next frame. The upper limit of allowable gain variation is determined by the confidence level of the previous frame. When it is lower than 0.65, the first allowable variation upper limit is 0.45, and when it is higher than or equal to 0.65, the second allowable variation upper limit is 0.60. The variation upper limit represents the dimensionless maximum absolute difference of the complex gain amplitude of the real adjacent frames. Since the maximum transition amount of the legal lower limit from 0.12 to 0.55 is 0.43 and is less than the first allowable variation upper limit, any legal lower limit switch can be satisfied simultaneously with the corresponding allowable variation upper limit.
[0102] Gain lower limit Calculation, where This represents the dimensionless lower limit of the gain at frequency k in frame t. This indicates a valid audio band protection image mask, and the value is either 0 or 1. This indicates the lower limit of the dimensionless first gain amplitude, which is 0.55. This indicates the lower limit of the dimensionless second gain amplitude; it is 0.12 when both the abnormal backoff indicator and the protection chart empty backoff indicator are invalid, and 0.25 when either is valid; the upper limit of the clipping gain is... Calculation, where This represents the upper limit of the dimensionless candidate gain amplitude at the k-th frequency point in the t-th frame. This indicates the integer clipping attenuation level carried by the clipping security request with target time index t, ranging from 0 to 16. The constant 1.05 represents the upper limit of the dimensionless candidate gain amplitude reference without clipping attenuation, consistent with the upper limit of the amplitude cutoff of the training target gain. The constant 0.05 represents the step size of the dimensionless gain amplitude reduction for each additional clipping attenuation level, and is calculated according to... Both are fixed, and the gain configuration is fixed by version AUD-GAIN-01 based on the limiting rate and voice fidelity results of the verification audio. This indicates taking the minimum of the two values; the upper limit of the allowed variation is... Calculation, where This indicates the upper limit of the dimensionless allowable gain variation in this frame. Indicates the gain confidence level of the previous frame. This is an indicator function that takes the value 1 if the condition is true and 0 otherwise.
[0103] First, suppress the gain of the original complex number. Create upper and lower limit clipping candidates, then... and Construct the feasible interval for this frame, and in Press at time Perform inter-frame projection, where This represents the dimensionless candidate gain amplitude after initial upper and lower limit clipping. This represents the original complex suppression gain at the k-th frequency point of the t-th frame output by the network. and These represent the dimensionless lower limit of gain and the upper limit of clipping gain for the corresponding frequency point in this frame, respectively. and These represent the dimensionless lower and upper feasible bounds that simultaneously satisfy the upper and lower limits of the current frame and the upper limit of the change in adjacent frames, respectively. and These represent taking the maximum and minimum values of the two options, respectively. This represents the dimensionless gain magnitude after projection onto the feasible region. The value must always represent the final gain amplitude that was actually applied in the previous consecutive frame and verified by the current frame's aperture, and must not be reset by a clipping safety request. This represents the dimensionless upper limit of the allowable gain variation selected based on the confidence level of the previous frame. This indicates that the input is restricted to a lower bound of l and an upper bound of u. The above feasibility condition means that the feasible lower bound is not greater than the feasible upper bound. The control state is the gain constraint state for each frequency point of the current frame. The controlled object is the candidate gain amplitude of 257 frequency points. The control action is to sequentially perform initial upper and lower bound clipping, feasible interval construction, and inter-frame projection in each audio frame. The control objective is to output an executable gain while satisfying the lower bound of speech fidelity, the upper bound of clipping safety, and the change rate limit of the current frame. When any input is invalid, the feasible lower bound of any frequency point is greater than the feasible upper bound, or the projection result fails the aperture verification of the current frame, the invalid silence block is entered and the S5 reconstruction branch is triggered. The candidate value is not written to the historical state. In other cases, The data is fed into the same smoother, and the frame number, frequency point, action result, and actual application gain are recorded for reading in the next consecutive frame.
[0104] Same gain smoother Output amplitude, where This represents the dimensionless final gain amplitude after smoothing and reprojection onto the feasible region of the current frame. This means limiting the smoothing result to a dimensionless feasible lower bound within the current frame. and feasible upper bound between, This represents the dimensionless smoothing coefficient for this frame. When the transient state switches from non-entry to entry, a gain attack coefficient of 0.75 is used; when switching from entry to exit, a gain release coefficient of 0.18 (value less than the attack coefficient) is used; for other states, 0.35 is used. The phase constraint chain first determines the original complex gain amplitude. Press at time Extract the original phase; when this condition is not met and the phase state is valid, let... When this condition is not met and it is the first frame, the first frame after state reset, or the phase state is invalid, then... Then press Obtain the principal phase difference and then... Obtain the constrained phase, and finally press Output constrained complex gain; where This represents a dimensionless phase threshold that is strictly greater than 0, taken as twice the gain output quantization resolution and fixed by the gain configuration version AUD-GAIN-01. This represents the phase of the original complex gain at frequency k in frame t within the left-closed, right-open principal value interval from negative pi to positive pi, expressed in radians. This indicates that only when the magnitude of the complex number is greater than... When taking the complex principal value phase, This indicates the phase difference between adjacent frames after wrapping. This represents adding or subtracting integer multiples of pi and uniquely maps to the left-closed, right-open interval from negative to positive pi. This means that when the positive difference exceeds one-quarter of the positive pi, it is intercepted in the positive direction as one-quarter of the positive pi; when the negative difference is less than one-quarter of the negative pi, it is intercepted in the negative direction as one-quarter of the negative pi. Indicates the constrained phase and the unit is radians; when the first frame, the first frame after state reset, or the phase state valid flag is 0, directly set... And set the valid phase state flag to 1. This indicates the final dimensionless constrained complex suppressed gain. Represents the imaginary unit. This indicates complex exponential operation, where the gain before and after the path switch both enter the same amplitude and phase smoothing chain;
[0105] The constraint verification module first checks on a frequency-by-frequency basis whether the lower feasible bound is greater than the upper feasible bound, and then checks the final gain amplitude. Not less than Not higher than And simultaneously located by and Within a defined closed interval, the amplitude difference between adjacent frames does not exceed The system checks if the attack coefficient is greater than the release coefficient. If all conditions are met, the original verification state is set to pass. If any check fails, the original verification state is set to fail. First, the complex gain amplitude of the most recently verified frame is reprojected onto the same closed interval as a backoff candidate. If no historical pass value exists, the first frame's safe output projected onto the same feasible interval is used. Then, the backoff candidate is repeatedly checked, including values not exceeding [a certain threshold]. All the above constraint checks are performed; when the secondary check passes, the rollback execution status is set to executed, the rollback secondary verification status is set to passed, and the complex gain is output by enumerating the status "rollback passed"; when the feasible interval is empty or the secondary check still fails, the rollback secondary verification status is set to failed, and no identifiable gain is output to S5, but an invalid placeholder instruction is output; the failure frequency, upper and lower limit sources, feasible interval, original verification status, rollback execution status, rollback secondary verification status, and path identifier are all written into the same constraint verification record;
[0106] The computational budget state records the average processor utilization, the 95th percentile of single-frame processing time, and the remaining battery power within a 100-frame statistical period. In this specific implementation, the average processor utilization of 85%, the 95th percentile of the processing time of the most recent 100 frames (10ms), and 15% of the remaining battery power are used as unified budget boundaries. When any statistical boundary is crossed, the budget is set to be limited. In the budget-limited state, when a transient state is entered and the probability of speech presence is not lower than the speech probability threshold, the transient refined inference path is maintained. When a transient state is entered but the probability of speech presence is lower than the threshold, the process falls back to the second inference path. In other states, the process falls back from the second inference path to the first inference path.
[0107] Budget-constrained label Generate, where This indicates the budget constraint flag for the statistical period in which the t-th frame belongs, and it takes the value 0 or 1. This represents the 95th percentile of the processing time for a single frame in the most recent 100 frames, in milliseconds (ms). This indicates that the S1 frame shift corresponds to a duration of 10ms. This represents the dimensionless mean of processor utilization over the last 100 frames. This indicates a dimensionless processor utilization threshold of 0.85. Indicates the remaining battery percentage. This indicates a power threshold of 15%. This indicates that the value is 1 if the condition is true and 0 otherwise. This represents a logical OR condition with three boundary conditions.
[0108] Budget constraints are lifted when two consecutive 100-frame statistical cycles meet the following conditions: the 95th percentile of processing time is less than 9ms, the average processor utilization is less than 0.80, and the remaining battery power is not less than 15%. The same path conditions given by the state rules are met for 20 consecutive frames. The formal path identifier is only updated if the candidate path is consistent for 5 consecutive frames. If the first statistical cycle after the switch crosses the budget boundary again, the path before the switch is restored and maintained for at least 100 frames. After that, restoration is only allowed again when there are no constraints for two consecutive statistical cycles. The reason for the rollback, the reason for the restoration, the path identifier before the update, the path identifier after the update, the statistical cycle number, the 95th percentile of processing time, the average processor utilization, and the remaining battery power are written to the path switching record at the same time to avoid the number of processing layers fluctuating between adjacent frames.
[0109] S4 ultimately writes the constrained complex suppression gain or invalid placeholder instruction, current gain confidence, reusable hidden state, formal path identifier, original verification state, rollback execution state, rollback secondary verification state, consumed clipping security request version, and phase state validity identifier into the unified control vector of this frame. This vector is stored in the cross-frame state buffer as the source of the unified control vector for the next frame. Only when the state is "passed" or "rollback passed" is the corresponding constrained complex suppression gain and its time-frequency index passed to S5; otherwise, an invalid placeholder instruction is passed.
[0110] In this specific embodiment, S5 includes:
[0111] S5 reads the constrained complex suppression gain of S4 and the complex spectrum of the target speech beam of S2 according to the time index and frequency index. When the configuration versions, frame time indexes and 257 single-sided frequency points of the two objects are consistent, it performs frequency-point-by-frequency complex multiplication. If any index is inconsistent, the frame is marked as invalid before the reconstruction queue consumption stage, and the whole block silence rule is used as the highest priority abnormal output: the 160-point output sampling interval corresponding to the frame does not perform effective contribution frame set reconstruction, the overlapping items of the current frame and adjacent frames falling into the interval are all set to zero, and linear fade-out and fade-in are performed on the nearest effective block and the next effective block at 16 points before and after the interval. Then, the 160-point silence placeholder block with invalidation is submitted, the sampling sequence number is advanced according to the original time index and the reason for frame loss is recorded, so that the next effective frame can still resume the supply with continuous sampling sequence number.
[0112] Enhanced complex spectrum according to Calculation, where This represents the enhanced complex spectrum at frequency k in the t-th frame, which has the same dimensions as the complex spectrum of the target speech beam, and its amplitude is equal to... , This represents the dimensionless final constrained complex suppression gain of the S4 output. This represents the complex spectrum of the target speech beam output by S2. Complex multiplication acts on both amplitude and phase while maintaining the conjugate symmetric frequency relationship.
[0113] The time-domain reconstruction module expands the 257 single-sided frequency points into a 512-point conjugate symmetric spectrum, performs a 512-point inverse fast Fourier transform, and truncates 320 time-domain sampling points with the same length as the S1 analysis window. It then uses a square-root Hanning synthesis window paired with the square-root Hanning analysis window for overlapping and addition. For a given output sampling point q that does not belong to the silent occupancy interval, a set of valid contribution frames is defined. , and according to Reconstruction, in which This represents the set of frame indices that simultaneously satisfy both the window support range and the frame validity condition. This represents the time index of the audio frames in the set. This indicates that the time-domain frame length is fixed at 320 points. This indicates a valid reconstruction identifier where the configuration version, time index, and frequency index of frame t are all consistent, and the S4 constraint status is either passed or rollback passed, at which point the value is 1. This represents the local augmented speech normalized amplitude at sampling point q. This represents the normalized amplitude obtained by the inverse transform of frame t, and its unit is . Consistent The dimensionless coefficients representing the position of the composite window within the corresponding frame. This indicates that the frame shift is set to 160 points in S1. This represents the sum of the products of the analysis window and the synthesis window on the same set of valid contribution frames, and is a dimensionless normalization factor. Indicates the output sampling point index under a unified sampling clock; invalid or missing frames are not included. However, its corresponding 160-point sampling interval directly bypasses this reconstruction formula and forces the output to zero, and cannot be covered by contributions from adjacent valid frames;
[0114] Normalization factor according to The calculation, the addition of overlapping numerators, and the strict use of the same formula are all performed. ,in This indicates the analysis window coefficient used in S1. Represents the paired synthesis window coefficients. This indicates that for a given q, a window of 0 to 319 points supports and Frame index, This represents a dimensionless smoothing quantity used to prevent the normalization factor at boundary positions from becoming zero, and is set to... When the window configuration versions are inconsistent, the reconstruction is stopped and a state reset is requested. In this specific embodiment, the window combination satisfies the constant overlap addition condition in the steady-state overlap region.
[0115] The reconstruction quality check calculates the first-order difference root mean square value of 16 sampling points at the connection of adjacent output blocks and compares it with the 99th percentile of 0.08 in the audio without transient verification. If the threshold is exceeded, the block is not directly submitted. Instead, the complex gain that passed the constraint verification in the most recent frame is used as an alternative candidate. First, it is reprojected according to the lower limit of the gain corresponding to the speech protection map of this frame, the upper limit of the clipping gain, and the allowable variation range based on the actual application gain of the previous consecutive frame. Then, the same amplitude, phase, attack release, and full constraint secondary verification as S4 are performed. If the secondary verification fails, the process immediately enters the aforementioned band. The 160-point block of silent safety path with invalid identifiers is recalculated with alternative gain only when the verification is successful. The actual alternative gain, the reason for the alternative, the configuration version, and the time index of this frame are atomically written into the S5 actual application gain write-back buffer, overwriting the S4 candidate record and providing the next frame's S4 with priority as the real smoothing benchmark of the previous frame. If the recalculation still exceeds the connection threshold, a reconstruction anomaly identifier is written and the effective blocks before and after are connected with a linear fade-in and fade-out with a length of 16 points. The threshold, the number of recalculations, and the fade-in and fade-out lengths are all stored in the reconstruction configuration version AUD-OLA-01.
[0116] The reconstructed local enhanced speech is written into a 16kHz audio circular queue in 160-point output blocks. The output record includes time index, start and end sampling sequence number, formal path identifier, gain confidence, freeze flag, original verification status, backoff execution status, and backoff secondary verification status. When any frame is missing or an invalid S4 placeholder instruction is received, a 160-point silence placeholder block is immediately submitted to the recognizer according to the aforementioned highest priority whole block silence rule and marked as invalid input. Normalization and reconstruction of the sampling interval according to the remaining valid frames is prohibited. At the same time, the sampling sequence number is advanced so that the recognizer does not include the zero-padding interval in the wake-up confidence and the next valid block maintains the sequence number continuity.
[0117] When the freeze flag is valid, the local enhanced speech output block, along with the frame-level speech presence probability and transient state, is sent to the wake-up recognizer. If the wake-up recognizer is already in an active session, the signal is sent to the instruction recognizer. At the same time, the time index of the recognition input frame, the formal path identifier, the constraint verification result, the number of valid frequency points of the protection graph, and the recognition session identifier are recorded. When the freeze flag is invalid, the flow is still continuously supplied to the wake-up recognizer but no transient overlap protection tag is attached.
[0118] The recognition interface only accepts valid audio blocks with a constraint status of "pass" or "backoff pass" and continuous output sampling sequence numbers. When the original verification status fails but the secondary verification status passes after backoff, the audio block is reconstructed using the backoff gain output by S4, enqueued in the "backoff pass" status, and the substitution flag is written into the recognition input record. When the secondary verification still fails after backoff, only 160-point silence placeholder blocks with invalid flags are submitted to maintain continuous sampling sequence numbers and are not sent to the recognition calculation. The wake-up score and instruction text returned by the recognizer are only used as terminal consumption results and do not provide feedback to modify the enhanced speech of this frame.
[0119] The input record is identified by using a combined key consisting of the session identifier and the time index. The fields also include the probability of voice presence, transient state, formal path identifier, original verification state, rollback execution state, secondary verification state after rollback, reconstruction anomaly identifier, sampling sequence number range, and queue write time. The writing and audio block enqueuing are completed in the same transaction. If the record is successfully written but the audio block enqueuing fails, the record is rolled back. If the audio block enqueuing is successful but the record writing fails, the block is marked as unrecognizable and skipped by the queue consumer.
[0120] Before writing to the audio queue, the normalized amplitude is limited to between -1 and +1 and quantized into a 24-bit signed pulse code modulation sample value. Simultaneously, the number of amplitude-limited sample points in every 160-point output block is accumulated. If the amplitude-limiting ratio exceeds 0.01, an amplitude anomaly is written to the recognition input record without changing the time index. If three consecutive output blocks exceed this ratio, S5 atomically writes the amplitude-limiting security request field, target frame time index, request version, trigger block time index, and clipping attenuation level to the independent cross-frame security request buffer. The target frame is fixed as the next consecutive audio frame after the third over-limited block. The amplitude-limiting security request is denoted as... ,in This field represents a dimensionless clipping security request identifier with a target time index of t+1, and can be either 0 or 1. t represents the audio frame time index of the third consecutive out-of-limit output block, and t+1 represents the next consecutive audio frame after that block. A status value of 1 indicates that S5 has atomically written a valid request to the independent cross-frame security request cache, while a status value of 0 indicates that the request is invalid, has been consumed and cleared, or has been cleared by a device reset. The source of the write is the S5's determination result that the clipping ratio of three consecutive output blocks exceeds 0.01. The initial clipping attenuation level is 1. For each additional output block when the limit is exceeded, the clipping attenuation level of subsequent target frames is increased by 1, up to a maximum of 16. S4 only reads this field when the current frame time index is the same as or later than the target request index and the request version has not been consumed. According to S4's specifications, the upper limit of the candidate gain amplitude is reduced, the lower limit of the second gain is not increased, and the frame is not reset. Or any inter-frame variation reference; after the clipping ratio of five consecutive output blocks is restored to within 0.01, S5 reduces the clipping attenuation level by 1 for each safe output block from the next consecutive frame until it reaches 0 and writes it to the new version, so that the gain limit is restored to a maximum of 0.05 per frame. When the clipping attenuation level is returned to zero, Set to 0; the same applies when resetting the device. If the clipping attenuation level reaches 16 and the limiting ratio still exceeds 0.01, the block is marked as constrained and enters the 160-point block silent safety path with invalid flags. When the abnormal backoff flag or the protection diagram empty backoff flag is still valid, S4 independently maintains the second gain lower limit of 0.25, and the protection frequency band lower limit of 0.55, the clipping gain upper limit, the change upper limit, and the constraint verification rules are always executed with priority.
[0121] This results in enhanced speech that is reconstructed within the local edge audio processor. The audio queue address, sampling sequence number range, recognition input record address, and unified control vector version are output by the S5 terminal and read sequentially by the smart speaker's wake-up recognition or command recognition process.
[0122] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
[0123] This invention uses the noise information provided by the blocking reference beam, the target speech beam, spatial acoustic features, and dual noise storage state together for causal enhancement. It uses the unified control vector of the previous audio frame to control the current noise memory, and then uses the current noise state and the gain confidence of the previous frame to control the inference path and gain constraints of the current frame, so that noise estimation, network inference, and speech protection maintain a causal closed loop.
[0124] This invention alters the edge audio processing structure by storing background and transient noise with different update coefficients, freezing speech overlap, using effective or invalid speech band protection maps, and employing multiple inference paths with different processing layers. Furthermore, it coordinates state switching using a unified control vector and the same gain smoother, thereby enabling continuous noise tracking, transient suppression, speech protection, and edge operation adjustment to work together.
Claims
1. A method for voice noise reduction in smart speakers based on edge computing, characterized in that, include: S1. Multi-channel far-field audio is simultaneously acquired by at least two microphones, and multi-channel complex spectra are obtained through frame-by-frame time-frequency transformation; S2. The target speech beam complex spectrum and the blocking reference beam complex spectrum are formed from the multi-channel complex spectra, and an acoustic feature set is generated; S3. Based on the blocking reference beam complex spectrum, the acoustic feature set, and the unified control vector of the previous frame, the background noise description is updated with a first update coefficient, and the transient noise description is updated with a second update coefficient greater than the first update coefficient. When the speech and transient noise overlap, the update of the background noise description is frozen and a speech band protection map is generated to obtain noise state data and update the unified control vector of the previous frame accordingly to obtain the current frame. S4. Input the target speech beam complex spectrum, acoustic feature set, and noise state data into the causal enhancement network of the local edge audio processor. Select one of the following inference paths to execute: the first inference path, the second inference path, or the transient refinement inference path, based on the unified control vector of this frame. The second inference path adds a causal temporal processing layer compared to the first inference path, outputting the complex suppression gain and its confidence level. Constrain the complex suppression gain according to the unified control vector of this frame and the speech band protection map, and update the unified control vector of the next frame according to the confidence level. S5. Apply the constrained complex suppression gain to the target speech beam complex spectrum and perform inverse time-frequency transform to obtain the locally enhanced speech.
2. The edge computing-based voice noise reduction method for smart speakers according to claim 1, characterized in that, Step S1 further includes: obtaining a speaker playback reference; aligning the audio sampling sequences of each microphone with the speaker playback reference according to a unified sampling clock; obtaining echo estimates for each microphone channel based on the speaker playback reference through adaptive echo path filtering, and subtracting the echo estimates from the corresponding audio sampling sequences; performing frame division and short-time Fourier transform on the subtracted audio sampling sequences of each channel using the same frame length, frame shift, and analysis window, wherein the frame length, frame shift, and analysis window are recorded as local configuration parameters in the local edge audio processor, and outputting the multi-channel complex spectrum aligned with the time index and frequency index.
3. The edge computing-based voice noise reduction method for smart speakers according to claim 1, characterized in that, S2 includes: The target speech beam complex spectrum is synthesized by weighting the multi-channel complex spectrum based on the steering vector of the microphone array, and the blocking reference beam complex spectrum is generated using a blocking matrix that has zero response to the steering vector. The speech presence probability is generated based on the energy relationship between the target speech beam complex spectrum and the blocking reference beam complex spectrum. The normalized spectral flux is generated based on the ratio of the amplitude spectrum difference between adjacent audio frames to the sum of the amplitude spectra of the previous audio frame. The frequency band energy ratio is generated based on the ratio of the energy of each frequency band to the energy of the entire frequency band under a preset frequency band division. The array coherence is generated based on the cross power spectrum and the power spectrum of each microphone channel pair. In the first audio frame or the first frame after the state is reset, the amplitude spectrum of the previous audio frame is initialized to the amplitude spectrum of the current audio frame; the input source, time index, frequency index and normalization parameter of each feature are associated with the acoustic feature set as feature records.
4. The edge computing-based voice noise reduction method for smart speakers according to claim 1, characterized in that, Step S3, updating the background noise description, includes: setting a background noise storage unit and a background state; determining a speech probability threshold and a spectral flux threshold based on the speech presence probability and the quantile of the normalized spectral flux in the historical audio frames without wake-words, respectively; weighted aggregation of the array coherence in the target direction neighborhood corresponding frequency points to obtain a target direction coherence statistic, and determining a coherence threshold based on the quantile of the target direction coherence statistic; obtaining the current blocking reference power spectrum based on the square of the modulus of the blocking reference beam complex spectrum; when the unified control vector of the previous audio frame allows background updates, and the speech presence probability is less than the speech probability threshold, the normalized spectral flux is less than the spectral flux threshold, and the target direction coherence statistic is less than the coherence threshold, multiplying the first update coefficient by the current blocking reference power spectrum and adding it to the product of the first update coefficient and the previous background noise description to obtain the current background noise description, and writing the current background noise description into the background noise storage unit; and writing the update allowance state of the background noise storage unit and the data source of the threshold into the background state.
5. The edge computing-based voice noise reduction method for smart speakers according to claim 4, characterized in that, Step S3, updating the transient noise description, includes: setting a transient noise storage unit and a transient state; determining the energy change threshold and coherence change threshold based on the quantiles of the energy difference between adjacent frames in the wake-word-free historical audio frames and the difference in the coherence statistics of the target direction; generating three judgment values based on the exceeding results of the normalized spectral flux relative to the spectral flux threshold, the exceeding results of the energy change between the current audio frame and the previous audio frame relative to the energy change threshold, and the exceeding results of the change in the coherence statistics of the current target direction and the previous target direction relative to the coherence change threshold; when at least two judgment values indicate exceeding the threshold, setting the transient state to an entering state, and using the first... The two update coefficients are used to weight the current blocking reference power spectrum and the previous transient noise description to obtain the current transient noise description, and the current transient noise description is written into the transient noise storage unit. In the first audio frame or the first frame after the state reset, the energy of the previous audio frame and the coherence statistics of the previous target direction are initialized to their respective current observation values, the transient state is initialized to the local initial state, and the transient noise description is initialized to the current blocking reference power spectrum. When any input feature is missing or exceeds the effective range of the local configuration, an abnormal fallback is performed to keep the transient state in the state of the most recent effective audio frame and to limit the complex suppression gain after constraint to be no less than the lower limit of the safe gain.
6. The edge computing-based voice noise reduction method for smart speakers according to claim 5, characterized in that, Step S3, generating the speech band protection map, includes: setting a freeze flag; when the transient state is in the entry state and the probability of speech presence is not less than the speech probability threshold, setting the freeze flag and the speech band protection map to be valid and pausing the background noise storage unit update; generating a harmonic protection region based on the fundamental frequency candidate and integer multiples of the complex spectrum of the target speech beam, and generating a formant protection region based on the frequency band where the spectral envelope peak is located; merging the harmonic protection region and the formant protection region into the speech band protection map; when the transient state is not in the entry state and the probability of speech presence is not less than the speech probability threshold, the freeze flag and the speech band protection map are set to be valid and the background noise storage unit update is paused; when the transient state is ... paused; when the transient state is in the entry state and the probability of speech presence is not less than the speech probability threshold, the freeze flag and the speech band protection map are paused; when the transient state is in the entry state and the probability of speech presence is not less than the speech probability threshold, the freeze flag and the speech band protection map are paused; when the transient state is in the entry state and the probability of speech presence is not less than the speech probability threshold, the freeze flag and the speech band protection map are paused; when the transient state is in the entry state and the probability of speech presence is not less than the speech probability threshold, the freeze flag and the speech band When the probability is not less than the speech probability threshold, the frequency protection flag of the speech band protection map is set to invalid; the number of exit judgment frames and the noise description deviation are used as local verification configuration parameters; when at most one of the three judgment values in an audio frame that continuously reaches the number of exit judgment frames indicates an over-limit, the transient state is set to the exit state, and the transient noise description is attenuated frame by frame according to the locally configured noise description release coefficient until the transient noise description is not greater than the sum of the background noise description and the noise description deviation, and then the freeze flag is released; the entry state, the exit state and the corresponding judgment reason are recorded.
7. The edge computing-based voice noise reduction method for smart speakers according to claim 6, characterized in that, The causal enhancement network in step S4 includes a shared encoder, a reusable hidden state, a first inference path, a second inference path, a transient refinement inference path, a complex gain output layer, and a common gain smoother. The first inference path invokes the shared encoder and the complex gain output layer. The second inference path invokes at least one causal timing processing layer based on the first inference path. The transient refinement inference path invokes a transient frequency band attention layer based on the second inference path. Each inference path reads the reusable hidden state of the previous audio frame and provides the updated hidden state for the next audio frame to read. In the first audio frame or the first frame after state reset, the reusable hidden state is initialized to... In state zero, the previous background noise description is initialized to the current blocking reference power spectrum, the background state is initialized to an update-allowed state, the transient state is initialized to a non-entry state, the freeze flag is initialized to invalid, and the gain confidence of the previous audio frame is initialized to the locally configured initial confidence. The unified control vector of the previous audio frame is composed of the initialized noise state data and the initial confidence. When the transient state is an entry state, the transient refinement inference path is selected. When the background state indicates that updates are allowed and the gain confidence of the previous audio frame is not less than the confidence threshold, the first inference path is selected. In other states, the second inference path is selected, and the path identifier is recorded in the unified control vector.
8. The edge computing-based voice noise reduction method for smart speakers according to claim 7, characterized in that, The constraint on the complex suppression gain in step S4 includes: setting the gain lower limit of the effective frequency points covered by the voice band protection map as a first gain lower limit, and setting the gain lower limit of the remaining frequency points as a second gain lower limit that is less than the first gain lower limit; determining the confidence threshold by the gain error quantile in the verification audio set; using a first allowable gain change upper limit when the gain confidence of the previous audio frame contained in the unified control vector of the current audio frame is less than the confidence threshold, and using a second allowable gain change upper limit that is greater than the first allowable gain change upper limit when the gain confidence of the previous audio frame is not less than the confidence threshold; updating the complex suppression gain with a gain attack coefficient when the transient state switches from a non-entry state to an entry state, and updating the complex suppression gain with a gain release coefficient that is less than the gain attack coefficient when the transient state switches from an entry state to an exit state; when the path identifier is switched, ensuring that the gain before and after the switch passes through the same gain smoother, and using whether the constrained complex suppression gain satisfies the first gain lower limit, the second gain lower limit, and the corresponding allowable gain change upper limit as the constraint verification result.
9. The edge computing-based voice noise reduction method for smart speakers according to claim 8, characterized in that, S5 includes: The constrained complex suppression gain is multiplied by the frequency-point complex spectrum of the target speech beam to obtain the enhanced complex spectrum; The enhanced complex spectrum is subjected to inverse short-time Fourier transform and overlapped summation using a synthesis window that matches the analysis window used in the frame-segmented time-frequency transformation of step S1 to obtain the local enhanced speech; When the freeze flag is valid, the local enhanced speech, the speech presence probability, and the transient state are sent to the wake-up recognizer or the instruction recognizer, and the time index, path identifier, and constraint verification result of the recognition input frame are recorded.
10. The edge computing-based voice noise reduction method for smart speakers according to claim 7, characterized in that, Step S4, selecting the inference path, further includes: generating a computational budget state based on the processor utilization rate, single-frame processing duration, and remaining battery power of the local edge audio processor within a preset statistical period; when the computational budget state indicates that the single-frame processing duration reaches the duration corresponding to the frame shift or the remaining battery power is less than the battery power threshold, maintaining the transient refined inference path when the transient state is in the entry state and the probability of speech presence is not less than the speech probability threshold; falling back from the transient refined inference path to the second inference path when the transient state is in the entry state and the probability of speech presence is less than the speech probability threshold; falling back from the second inference path to the first inference path in other states; when the computational budget state recovers to the condition that the single-frame processing duration is less than the duration corresponding to the frame shift and the remaining battery power is not less than the battery power threshold, and multiple consecutive audio frames satisfy the selection conditions of the first inference path, the second inference path, or the transient refined inference path, resuming the determination of the target path according to the corresponding selection conditions; updating the path identifier only when the target paths of multiple consecutive audio frames are consistent, and writing the fallback reason, recovery reason, path identifier before update, and path identifier after update into the path switching record.