Array microphone noise reduction recording method based on cascade noise reduction and blind source separation

By employing collaborative signal processing, weighted kernel function blind source separation, and adaptive cascaded noise reduction models, combined with lossless coding based on speech feature perception, the problems of insufficient stability of multi-channel signals, mixed signal separation accuracy, and coding adaptability of array microphones were solved, enabling high-definition recording in far-field multi-noise environments.

CN121237113APending Publication Date: 2025-12-30SHANGHAI RONGDA DIGITAL TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511338173.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

Existing array microphone noise reduction recording technology has shortcomings in terms of multi-channel signal consistency, mixed signal separation accuracy, noise reduction comprehensiveness, and coding adaptability, making it difficult to meet the high-definition recording requirements in far-field, high-noise, and multi-scenario environments.

Method used

A collaborative signal processing algorithm is used to maintain the amplitude and phase consistency of the channel recording signal. Combined with a frame loss and audio distortion prevention mechanism, a weighted kernel function blind source separation algorithm, a target-oriented adaptive cascaded noise reduction model, and a speech feature perception lossless coding algorithm, high-quality acquisition, accurate separation, and high-fidelity storage of multi-channel signals are achieved.

Benefits of technology

Significantly improves multi-channel signal consistency and recording integrity, enhances mixed signal separation accuracy, achieves dual optimization of noise reduction comprehensiveness and far-field voice quality, balances recording quality and storage efficiency, and outputs high-definition recordings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121237113A_ABST
    Figure CN121237113A_ABST
Patent Text Reader

Abstract

The invention relates to an array microphone noise reduction recording method based on cascade noise reduction and blind source separation, which belongs to the technical field of voice signal processing and recording, and comprises the following steps: configuring a multi-channel array microphone, ensuring that the amplitude and phase of a channel signal are consistent, and collecting an original multi-channel voice signal; a weighted kernel function blind source separation algorithm is adopted, signal-to-noise ratio distribution characteristics of signals are extracted, kernel function weights are given, and target voice and interference signal components are obtained through decoupling of an independent component analysis model; executing target-oriented adaptive cascade noise reduction, locking the voice of a keynote speaker through directional pickup, reducing noise, filtering out reverberation, and enhancing the voice of a far-field target by combining a voice mask neural network with a far-field pickup algorithm in sequence; and processing the target voice through voice feature perception lossless coding and storing the target voice. According to the invention, stable acquisition of multi-channel signals, accurate separation of mixed signals and layered suppression of noise reverberation are realized, the signal-to-noise ratio and definition of far-field voice are significantly improved, and the method is suitable for single-person speaking or multi-person dialogue scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of speech signal processing and recording technology, specifically relating to an array microphone noise reduction recording method based on cascaded noise reduction and blind source separation. Background Technology

[0002] With the development of speech signal processing technology, array microphones, due to their multi-channel signal acquisition capabilities, are widely used in scenarios such as far-field recording, conference recording, and multi-person dialogue capture. Their core requirement is to achieve high-definition speech recording in complex environments, ensuring the signal-to-noise ratio and intelligibility of the speech signal. However, existing array microphone noise reduction recording technology still suffers from several technical bottlenecks, making it difficult to meet the high-quality requirements of practical applications. Specific shortcomings are as follows: The consistency of multi-channel signals and the integrity of recordings are difficult to guarantee. Most current multi-channel array microphones only achieve signal acquisition through simple hardware synchronization, lacking a targeted collaborative signal processing mechanism. As a result, due to differences in hardware sensitivity and transmission link loss, the recording signal amplitude deviation and phase delay are prone to occur in each channel, which seriously affects the accuracy of subsequent multi-channel signal fusion and processing. At the same time, some recording solutions are not equipped with effective anti-frame loss and audio distortion monitoring and compensation mechanisms. Under the influence of signal transmission load fluctuations or environmental interference, recording frame loss, sound disconnection or distortion often occur, resulting in incomplete recording content. The accuracy of mixed signal separation is insufficient. In existing technologies, blind source separation algorithms are a common means of separating speech from interference signals. However, traditional blind source separation often uses fixed kernel function parameters and does not dynamically adjust the weights based on the signal-to-noise ratio distribution characteristics of the original signal. In low signal-to-noise ratio regions (such as scenes with dense environmental noise), due to the lack of targeted weight enhancement, it is difficult to effectively distinguish between target speech and interference signals. This results in a large amount of noise remaining in the target speech after separation, or speech components being mixed into the interference signal, affecting the basic effect of subsequent noise reduction processing. Noise reduction processing has significant limitations, resulting in poor far-field speech quality. Existing noise reduction solutions mostly focus on suppressing a single type of noise, or can only handle steady-state noise (such as air conditioner background noise, fan noise) and non-steady-state noise (such as sudden door closing sound, clapping sound) separately. However, they lack a cascaded collaborative design of "noise suppression-reverberation filtering-far-field enhancement": on the one hand, it is difficult to simultaneously achieve efficient suppression of both steady-state and non-steady-state noise, easily leading to the problem of "suppressing one type of noise but leaving another type of noise"; on the other hand, for speech reverberation tails in open environments, traditional filtering algorithms mostly use fixed window parameters, which cannot adaptively match reverberation delay, easily leading to incomplete reverberation filtering or damage to speech details; in addition, in far-field scenarios, existing enhancement algorithms mostly rely on simple gain adjustment, without combining speech features to accurately distinguish residual noise from target speech, resulting in low far-field speech clarity and poor intelligibility, making it difficult to meet the high-definition recording requirements of single-person or multi-person dialogue scenarios. The encoding and storage are not well adapted to speech features. Existing recording solutions mostly use general lossless encoding strategies, which are not optimized for features such as the pitch period and harmonic distribution of speech signals. For areas where speech energy is concentrated (such as vowel pronunciation segments), the fidelity processing is not enhanced, which may lead to the loss of key speech details. The lack of redundant compression for silent or weak segments can easily lead to a waste of storage resources and make it difficult to balance recording quality and storage efficiency. In summary, existing array microphone noise reduction recording technology has shortcomings in terms of multi-channel signal stability, mixed signal separation accuracy, noise reduction comprehensiveness, and coding adaptability. It cannot meet the high-definition recording requirements in far-field, high-noise, and multi-scenario environments. A technical solution that can overcome the above bottlenecks is needed. Summary of the Invention

[0003] To address the aforementioned problems in the existing technology, this invention provides an array microphone noise reduction recording method based on cascaded noise reduction and blind source separation; The objective of this invention can be achieved through the following technical solutions: A noise reduction recording method for array microphones based on cascaded noise reduction and blind source separation, characterized by the following steps: S1: Configure a multi-channel array microphone. The multi-channel array microphone maintains the amplitude and phase coordination consistency of the channel recording signal through a cooperative signal processing algorithm, monitors the channel recording process through an anti-frame loss and audio distortion mechanism, and collects the original multi-channel voice signal in the environment. S2: The original multi-channel speech signal is preprocessed based on the weighted kernel function blind source separation algorithm. First, the signal-to-noise ratio distribution features are extracted and assigned weights to the kernel function. Then, the original multi-channel speech signal is decoupled through the independent component analysis model to obtain the target speech signal component and the interference signal component. S3: Perform a target-oriented adaptive cascaded noise reduction model on the target speech signal component: First, based on preset adjustable pickup range parameters, lock the speech signal emitted by the speaker through directional pickup technology. Then, use an adaptive multimodal noise reduction algorithm to suppress steady-state noise and non-steady-state noise in the speech signal. Subsequently, use an acoustic signal adaptive processing algorithm to filter out the speech reverberation tail in the noise-suppressed output speech signal. Finally, combine a speech masking neural network and a far-field pickup algorithm. The speech masking neural network distinguishes the residual noise and target speech in the reverberation-filtered output speech signal, calculates the energy expectation of the target speech, and enhances the target speech in the far-field environment based on the far-field pickup algorithm. S4: Apply a speech feature-aware lossless encoding algorithm to the target speech signal: extract the pitch period and harmonic distribution features of the target speech signal, apply a high-fidelity encoding strategy to the speech energy concentration area, store the encoded target speech signal, and generate a noise-reduced recording process file.

[0004] Specifically, the collaborative signal processing algorithm includes a dynamic amplitude compensation sub-algorithm and a phase calibration sub-algorithm: the dynamic amplitude compensation sub-algorithm collects the amplitude difference of the channel recording signal in real time, analyzes the cause of the difference, generates a dynamic scaling factor, and adjusts the signal gain of the amplitude channel through the dynamic scaling factor; the phase calibration sub-algorithm presets a reference channel, uses the phase of the reference channel as a reference, calculates the phase delay between other channels and the reference channel in real time, and generates a reverse compensation signal based on the phase delay, which is then superimposed on the corresponding channel.

[0005] Specifically, the anti-frame loss and audio distortion mechanism includes a signal buffer pool and a frame loss detection and retransmission sub-algorithm: the signal buffer pool has a preset buffer space, which temporarily stores real-time recording signals within a continuous time period; the frame loss detection and retransmission sub-algorithm monitors frame loss during signal transmission in real time, and determines the frame loss situation by comparing the signal integrity flags before and after transmission. When frame loss is detected and affects the continuity of recording, a retransmission request is immediately triggered, and the non-frame-loss signals stored before the frame loss period are retrieved from the signal buffer pool and added to the transmission link.

[0006] Specifically, the weighting method of the weighted kernel function blind source separation algorithm is as follows: First, the signal-to-noise ratio distribution characteristics of the original multi-channel speech signal are analyzed in multiple dimensions. In the time domain, the fluctuation frequency and amplitude of the energy of the original multi-channel speech signal are monitored. In the frequency domain, the energy ratio of the original multi-channel speech signal to the noise signal within the frequency band is calculated. By fusing the time domain and frequency domain features, low signal-to-noise ratio regions and high signal-to-noise ratio regions are divided. Then, a nonlinear weighting function is used to calculate the kernel function weight. The weight value of this function is dynamically adjusted in combination with the noise energy ratio in the corresponding region. The weight is differentiated based on the actual characteristics of the original multi-channel speech signal.

[0007] Specifically, the independent component analysis model integrates a gradient descent optimization sub-algorithm: the gradient descent optimization sub-algorithm takes reducing the mutual information of the separated signals as the objective function, and optimizes and reduces the correlation between the target speech signal component and the interference signal component; during iteration, the parameters of the independent component analysis model are adjusted according to the currently calculated mutual information value to gradually optimize the separation effect, and the iteration stops when the difference in mutual information calculated in two adjacent iterations is reduced to a preset threshold.

[0008] Specifically, the directional sound pickup technology is combined with a spatial spectrum estimation algorithm: the spatial spectrum estimation algorithm first collects the phase information of the channel speech signal, and calculates the spatial azimuth angle of the source of the channel speech signal by comparing the phase differences between channels; then it compares the spatial azimuth angle with the azimuth angle range corresponding to the preset adjustable sound pickup range, and when the spatial azimuth angle is within the azimuth angle range, it is determined to be the speech signal of the target speaker, and the target speaker speech signal is locked by enhancing the channel signal gain.

[0009] Specifically, the adaptive multimodal noise reduction algorithm includes a noise type classification sub-algorithm: the noise type classification sub-algorithm first extracts the time-domain and frequency-domain features of the speech signal, generates feature analysis results, and classifies the noise into steady-state noise and non-steady-state noise based on the feature analysis results; for the steady-state noise, a frequency-domain notch filtering strategy is adopted, setting a filtering notch in the frequency band where the noise is concentrated to weaken the steady-state noise in the frequency band; for the non-steady-state noise, a time-domain threshold suppression strategy is adopted, setting a signal amplitude threshold, and when the speech signal exceeds the threshold, it is determined to be non-steady-state noise, and the speech signal is suppressed by a threshold filter.

[0010] Specifically, the acoustic signal adaptive processing algorithm includes a reverberation delay estimation sub-algorithm and a dynamic filtering sub-algorithm: the reverberation delay estimation sub-algorithm first identifies the direct wave and the first reflected wave in the speech signal. The direct wave propagates directly from the sound source to the microphone, and the first reflected wave propagates to the microphone after one reflection, lagging behind the direct wave. The reverberation delay is obtained by calculating the time difference between the direct wave and the first reflected wave. The dynamic filtering sub-algorithm adjusts the length of the filtering window according to the reverberation delay. The window length needs to cover the duration of the reverberation tail to completely filter out the reverberation tail.

[0011] Specifically, the speech masking neural network adopts an architecture combining a bidirectional long short-term memory network layer and an attention mechanism: the bidirectional long short-term memory network layer can simultaneously extract the historical context features and future context features of the speech signal, extracting the temporal sequence pattern of the speech signal from the beginning to the current moment in the forward direction, and extracting the temporal sequence pattern of the speech signal from the current moment to the end in the backward direction, capturing the temporal characteristics of the speech signal through bidirectional feature fusion; the attention mechanism focuses on the frequency bands where human speech is mainly distributed, and by strengthening the feature weights of the frequency bands, the speech masking neural network prioritizes the frequency bands when distinguishing residual noise from human speech.

[0012] Specifically, the far-field sound pickup algorithm includes a speech energy compensation sub-algorithm and a frequency equalization sub-algorithm: the speech energy compensation sub-algorithm is based on the attenuation law of far-field speech propagation, that is, the law that the speech energy gradually weakens as the propagation distance increases. It first analyzes the distance of the sound source in the current recording environment, then calculates the degree of energy attenuation based on the sound source distance, and generates a corresponding energy compensation coefficient. The overall energy of the far-field speech is improved through the energy compensation coefficient. The frequency equalization sub-algorithm, in view of the characteristic that the high-frequency signal attenuates significantly in far-field propagation, identifies the high-frequency region in the far-field speech and adjusts the signal strength of the high-frequency region through gain adjustment.

[0013] Specifically, the speech feature-aware lossless coding algorithm includes a feature-driven coding redundancy optimization sub-algorithm: first, the energy distribution characteristics of the target speech signal are analyzed to identify speech energy concentration areas and silence areas. The speech energy concentration areas are the core carriers of the speech content, while the silence areas contain no effective speech and only include weak background noise. For the speech energy concentration areas, a high-fidelity coding strategy is adopted to preserve the detailed features of the signal, and for the silence areas, a coding redundancy compression strategy is adopted to remove redundant information in the signal.

[0014] Specifically, the speech quality detection and feedback step involves: using a speech quality perception evaluation algorithm to detect the quality of the target speech signal; analyzing the clarity, naturalness, and noise residue of the target speech to comprehensively determine the speech quality level; when the quality of the target speech signal does not meet the preset standard, the quality evaluation result is fed back to the noise reduction stage, and the noise reduction operation is repeated until the speech quality is detected to meet the preset standard, before proceeding to the encoding and storage stage.

[0015] The beneficial effects of this invention are as follows: Significantly improves the consistency and integrity of multi-channel signals. This invention addresses the amplitude deviation and phase delay issues caused by hardware differences and transmission losses in each channel through a collaborative signal processing algorithm, ensuring high coordination of amplitude and phase in multi-channel recording signals, providing a high-quality foundation for subsequent signal processing. Simultaneously, an anti-frame loss and audio distortion mechanism monitors the recording process in real time, avoiding frame loss, disconnection, and audio distortion caused by signal transmission fluctuations or environmental interference, ensuring the integrity of the recording content and solving the pain point of "original signal distortion affecting subsequent processing" in existing technologies.

[0016] This invention significantly improves the accuracy of mixed signal separation. It employs a weighted kernel function blind source separation algorithm, overcoming the limitations of traditional fixed kernel function parameters. By dynamically assigning weights to the kernel function based on the signal-to-noise ratio (SNR) distribution characteristics of the original signal—strengthening weights in low SNR regions with dense noise—it accurately distinguishes target speech from interference signals. Furthermore, combined with an independent component analysis (ICA) model, it efficiently decouples the mixed signals, effectively reducing noise residue in the separated target speech while preventing speech components from being mixed into the interference signal. This lays the foundation for a "high-purity target signal" in subsequent noise reduction processing, resulting in a significant improvement in separation accuracy compared to traditional blind source separation techniques.

[0017] This invention achieves dual optimization of comprehensive noise reduction and far-field speech quality. It constructs a collaborative processing system of "directional pickup - noise suppression - reverberation removal - far-field enhancement" through a target-oriented adaptive cascaded noise reduction model: directional pickup technology locks onto the speaker's voice within an adjustable range, reducing irrelevant interference from a spatial perspective; the adaptive multimodal noise reduction algorithm can simultaneously and efficiently suppress both steady-state and non-steady-state noise, avoiding the problem of "suppressing one type of noise while leaving another"; the acoustic signal adaptive processing algorithm can dynamically adjust parameters to match reverberation delay, thoroughly filtering out reverberation tails without damaging speech details; and the speech masking neural network combined with the far-field pickup algorithm can accurately distinguish residual noise from the target speech. Through energy expectation calculation and gain optimization, it significantly improves the clarity and intelligibility of far-field speech, perfectly adapting to scenarios such as single-person presentations and multi-person dialogues. The far-field speech signal-to-noise ratio and intelligibility are significantly improved compared to existing technologies.

[0018] Balancing recording quality and storage efficiency, this invention's speech feature-aware lossless coding algorithm overcomes the limitations of traditional general-purpose coding. It optimizes coding strategies based on features such as the pitch period and harmonic distribution of speech signals: high-fidelity coding is used for areas with concentrated speech energy to preserve key speech details such as vowels and consonants to the maximum extent, avoiding sound quality loss; redundant compression is applied to silent or weak segments to reduce storage resource consumption. This achieves an optimal balance between sound quality and efficiency while ensuring high-definition recording quality.

[0019] In summary, this invention overcomes the bottlenecks of existing technologies in terms of multi-channel signal stability, mixed signal separation accuracy, noise reduction comprehensiveness, and coding adaptability. It can stably output high-definition recordings in far-field, high-noise, and multi-scenario environments, and has strong practicality and promotional value. Attached Figure Description

[0020] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings.

[0021] Figure 1 This is a flowchart illustrating an array microphone noise reduction recording method based on cascaded noise reduction and blind source separation according to the present invention. Figure 2 This is a diagram illustrating the speech enhancement effect of an array microphone noise reduction recording method based on cascaded noise reduction and blind source separation according to the present invention. Detailed Implementation

[0022] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided.

[0023] Please see Figure 1-2 A noise reduction recording method for array microphones based on cascaded noise reduction and blind source separation, comprising the following steps: A noise reduction recording method for array microphones based on cascaded noise reduction and blind source separation, characterized by the following steps: S1: Configure a multi-channel array microphone. The multi-channel array microphone maintains the amplitude and phase coordination consistency of the channel recording signal through a cooperative signal processing algorithm, monitors the channel recording process through an anti-frame loss and audio distortion mechanism, and collects the original multi-channel voice signal in the environment. S2: The original multi-channel speech signal is preprocessed based on the weighted kernel function blind source separation algorithm. First, the signal-to-noise ratio distribution features are extracted and assigned weights to the kernel function. Then, the original multi-channel speech signal is decoupled through the independent component analysis model to obtain the target speech signal component and the interference signal component. S3: Perform a target-oriented adaptive cascaded noise reduction model on the target speech signal component: First, based on preset adjustable pickup range parameters, lock the speech signal emitted by the speaker through directional pickup technology. Then, use an adaptive multimodal noise reduction algorithm to suppress steady-state noise and non-steady-state noise in the speech signal. Subsequently, use an acoustic signal adaptive processing algorithm to filter out the speech reverberation tail in the noise-suppressed output speech signal. Finally, combine a speech masking neural network and a far-field pickup algorithm. The speech masking neural network distinguishes the residual noise and target speech in the reverberation-filtered output speech signal, calculates the energy expectation of the target speech, and enhances the target speech in the far-field environment based on the far-field pickup algorithm. S4: Apply a speech feature-aware lossless encoding algorithm to the target speech signal: extract the pitch period and harmonic distribution features of the target speech signal, apply a high-fidelity encoding strategy to the speech energy concentration area, store the encoded target speech signal, and generate a noise-reduced recording process file.

[0024] Specifically, the collaborative signal processing algorithm includes a dynamic amplitude compensation sub-algorithm and a phase calibration sub-algorithm: the dynamic amplitude compensation sub-algorithm collects the amplitude difference of the channel recording signal in real time, analyzes the cause of the difference, generates a dynamic scaling factor, and adjusts the signal gain of the amplitude channel through the dynamic scaling factor; the phase calibration sub-algorithm presets a reference channel, uses the phase of the reference channel as a reference, calculates the phase delay between other channels and the reference channel in real time, and generates a reverse compensation signal based on the phase delay, which is then superimposed on the corresponding channel.

[0025] After activating the multi-channel array microphone, the recording signals of each channel are acquired in real time for 5-10 consecutive sampling cycles. The signals of each channel are then imported into amplitude analysis. By comparing the average amplitude of each channel signal, if the amplitude of a certain channel is consistently lower than that of other channels and the fluctuation is stable, it is determined to be a hardware sensitivity difference. If the amplitude fluctuation changes with the transmission time, it is determined to be a transmission link loss. Dynamic proportional coefficients are generated according to the cause of the difference: for hardware sensitivity differences, a fixed proportional coefficient is calculated according to "target amplitude / current amplitude" (ensuring that the amplitude deviation of each channel is reduced to a negligible range after adjustment); for transmission link loss, a dynamic coefficient is generated according to "real-time loss rate reverse derivation" (updated every 2 sampling cycles). The compensation coefficient is superimposed on the signal amplification circuit of the corresponding channel, and the signal gain of the low amplitude channel is gradually adjusted until the amplitude of each channel remains stable and consistent within 3 consecutive sampling cycles.

[0026] Specifically, the anti-frame loss and audio distortion mechanism includes a signal buffer pool and a frame loss detection and retransmission sub-algorithm: the signal buffer pool has a preset buffer space, which temporarily stores real-time recording signals within a continuous time period; the frame loss detection and retransmission sub-algorithm monitors frame loss during signal transmission in real time, and determines the frame loss situation by comparing the signal integrity flags before and after transmission. When frame loss is detected and affects the continuity of recording, a retransmission request is immediately triggered, and the non-frame-loss signals stored before the frame loss period are retrieved from the signal buffer pool and added to the transmission link.

[0027] Specifically, the weighting method of the weighted kernel function blind source separation algorithm is as follows: First, the signal-to-noise ratio distribution characteristics of the original multi-channel speech signal are analyzed in multiple dimensions. In the time domain, the fluctuation frequency and amplitude of the energy of the original multi-channel speech signal are monitored. In the frequency domain, the energy ratio of the original multi-channel speech signal to the noise signal within the frequency band is calculated. By fusing the time domain and frequency domain features, low signal-to-noise ratio regions and high signal-to-noise ratio regions are divided. Then, a nonlinear weighting function is used to calculate the kernel function weight. The weight value of this function is dynamically adjusted in combination with the noise energy ratio in the corresponding region. The weight is differentiated based on the actual characteristics of the original multi-channel speech signal.

[0028] Multi-dimensional analysis of signal-to-noise ratio (SNR) characteristics: The original multi-channel speech signal is divided into 50ms frames. The energy fluctuation frequency (the number of times the energy exceeds twice the average value per second) and fluctuation amplitude (the difference between the maximum and minimum energy) of each frame are calculated. Frames with high fluctuation frequency and large amplitude are identified as candidates for low SNR regions. Fourier transform is performed on each frame to divide it into multiple frequency bands (covering human speech and common noise frequency bands). The ratio of "speech energy (based on speech segment feature recognition)" to "noise energy (non-speech segment energy)" in each frequency band is calculated. Frequency bands with a ratio less than 1 are identified as low SNR frequency bands. The time-domain and frequency-domain analysis results are fused, and the region where "time-domain candidates + low frequency-domain ratios" overlap is identified as a low SNR region, while the remaining regions are identified as high SNR regions. Nonlinear weighting function calculation and weight assignment stage: An S-shaped nonlinear weighting function is selected, with the function input being "noise energy percentage within the region" (noise energy / (speech energy + noise energy)) and the output being the weight value; Dynamic weight adjustment: For low signal-to-noise ratio regions, if the noise energy percentage exceeds 60%, a high weight value is output (ensuring that the signal in this region receives more computational resources during separation); For high signal-to-noise ratio regions, if the speech energy percentage exceeds 70%, a low weight value is output (to avoid over-processing); Weight application: The calculated weight values ​​are bound to the kernel function parameters of the corresponding regions. In the matrix operations of the blind source separation algorithm, the calculation coefficients of the signals in each region are adjusted according to the weight values ​​to enhance the signal separation gradient in low signal-to-noise ratio regions.

[0029] Specifically, the independent component analysis model integrates a gradient descent optimization sub-algorithm: the gradient descent optimization sub-algorithm takes reducing the mutual information of the separated signals as the objective function, and optimizes and reduces the correlation between the target speech signal component and the interference signal component; during iteration, the parameters of the independent component analysis model are adjusted according to the currently calculated mutual information value to gradually optimize the separation effect, and the iteration stops when the difference in mutual information calculated in two adjacent iterations is reduced to a preset threshold.

[0030] Specifically, the directional sound pickup technology is combined with a spatial spectrum estimation algorithm: the spatial spectrum estimation algorithm first collects the phase information of the channel speech signal, and calculates the spatial azimuth angle of the source of the channel speech signal by comparing the phase differences between channels; then it compares the spatial azimuth angle with the azimuth angle range corresponding to the preset adjustable sound pickup range, and when the spatial azimuth angle is within the azimuth angle range, it is determined to be the speech signal of the target speaker, and the target speaker speech signal is locked by enhancing the channel signal gain.

[0031] Specifically, the adaptive multimodal noise reduction algorithm includes a noise type classification sub-algorithm: the noise type classification sub-algorithm first extracts the time-domain and frequency-domain features of the speech signal, generates feature analysis results, and classifies the noise into steady-state noise and non-steady-state noise based on the feature analysis results; for the steady-state noise, a frequency-domain notch filtering strategy is adopted, setting a filtering notch in the frequency band where the noise is concentrated to weaken the steady-state noise in the frequency band; for the non-steady-state noise, a time-domain threshold suppression strategy is adopted, setting a signal amplitude threshold, and when the speech signal exceeds the threshold, it is determined to be non-steady-state noise, and the speech signal is suppressed by a threshold filter.

[0032] Temporal feature extraction and analysis: The directional pickup speech signal is divided into 10ms frames, and the peak amplitude and peak duration of each frame are calculated. If the peak amplitude of a certain frame suddenly exceeds three times the average amplitude of the previous five frames, and the peak duration is less than 200ms, it is identified as a non-steady-state noise feature (such as a sudden door slamming sound). Frequency domain feature extraction and analysis: Fourier transform is performed on each frame of the signal, and the energy values ​​in each frequency band are statistically analyzed. The energy fluctuation amplitude of 10 consecutive frames in the same frequency band is calculated. If the energy fluctuation amplitude of 10 consecutive frames in a certain frequency band is less than 10%, and the frequency band corresponds to a common steady-state noise frequency (such as the low-frequency band of air conditioner background noise), it is identified as a steady-state noise feature. Noise classification and targeted suppression: For signals identified as steady-state noise, the frequency domain notch filter module is activated—notch points are set in the frequency bands where noise is concentrated, and the noise energy in that frequency band is weakened by the filtering circuit, while retaining the speech signal in other frequency bands; For signals identified as non-steady-state noise, time domain threshold suppression is activated—a dynamic threshold is set (determined based on the average amplitude of the previous 5 frames of speech signal). When the signal amplitude exceeds the threshold, the gain of the signal in that frame is temporarily reduced by the threshold filter to suppress sudden noise. The threshold parameter is updated every 500ms to avoid affecting normal speech.

[0033] Specifically, the acoustic signal adaptive processing algorithm includes a reverberation delay estimation sub-algorithm and a dynamic filtering sub-algorithm: the reverberation delay estimation sub-algorithm first identifies the direct wave and the first reflected wave in the speech signal. The direct wave propagates directly from the sound source to the microphone, and the first reflected wave propagates to the microphone after one reflection, lagging behind the direct wave. The reverberation delay is obtained by calculating the time difference between the direct wave and the first reflected wave. The dynamic filtering sub-algorithm adjusts the length of the filtering window according to the reverberation delay. The window length needs to cover the duration of the reverberation tail to completely filter out the reverberation tail.

[0034] Specifically, the speech masking neural network adopts an architecture combining a bidirectional long short-term memory network layer and an attention mechanism: the bidirectional long short-term memory network layer can simultaneously extract the historical context features and future context features of the speech signal, extracting the temporal sequence pattern of the speech signal from the beginning to the current moment in the forward direction, and extracting the temporal sequence pattern of the speech signal from the current moment to the end in the backward direction, capturing the temporal characteristics of the speech signal through bidirectional feature fusion; the attention mechanism focuses on the frequency bands where human speech is mainly distributed, and by strengthening the feature weights of the frequency bands, the speech masking neural network prioritizes the frequency bands when distinguishing residual noise from human speech.

[0035] Network architecture setup: Input layer: Receives the reverberation-filtered speech signal and converts the signal into a Mel-frequency cepstral coefficient (MFCC) feature vector (covering key speech features); Bidirectional Long Short-Term Memory (Bi-LSTM) layer: Two Bi-LSTM units are set up. The forward LSTM unit extracts historical context features (such as the association between the previous frame of speech and the current frame) from the beginning to the end of the signal, and the backward LSTM unit extracts future context features (such as the association between the next frame of speech and the current frame) from the end to the beginning of the signal. The outputs of the two units are spliced ​​and fused to obtain complete temporal features. Attention mechanism layer: Focuses on the frequency bands where human speech is mainly distributed (which are crucial to speech intelligibility), and assigns higher weights to the feature vectors of these frequency bands (with reduced weights for other frequency bands) through a weight matrix, thereby strengthening the expression of key features; Output layer: Outputs a “speech mask” (a binary mask that distinguishes the target speech from residual noise, where 1 represents the target speech and 0 represents residual noise) through a fully connected layer. Network training and inference process: Training phase: A mixed dataset containing "clean speech + various residual noises" is used. The mixed signal is input into the network, and the network parameters are iteratively optimized with the goal of "minimizing the cross-entropy between the network output mask and the real mask (manually labeled)" until the model converges. Inference phase: The reverberation-filtered speech signal is input into the trained network, which outputs a speech mask. The target speech components are preserved and residual noise components are suppressed by multiplying the mask with the original signal point by point.

[0036] Specifically, the far-field sound pickup algorithm includes a speech energy compensation sub-algorithm and a frequency equalization sub-algorithm: the speech energy compensation sub-algorithm is based on the attenuation law of far-field speech propagation, that is, the law that the speech energy gradually weakens as the propagation distance increases. It first analyzes the distance of the sound source in the current recording environment, then calculates the degree of energy attenuation based on the sound source distance, and generates a corresponding energy compensation coefficient. The overall energy of the far-field speech is improved through the energy compensation coefficient. The frequency equalization sub-algorithm, in view of the characteristic that the high-frequency signal attenuates significantly in far-field propagation, identifies the high-frequency region in the far-field speech and adjusts the signal strength of the high-frequency region through gain adjustment.

[0037] Specifically, the speech feature-aware lossless coding algorithm includes a feature-driven coding redundancy optimization sub-algorithm: first, the energy distribution characteristics of the target speech signal are analyzed to identify speech energy concentration areas and silence areas. The speech energy concentration areas are the core carriers of the speech content, while the silence areas contain no effective speech and only include weak background noise. For the speech energy concentration areas, a high-fidelity coding strategy is adopted to preserve the detailed features of the signal, and for the silence areas, a coding redundancy compression strategy is adopted to remove redundant information in the signal.

[0038] Specifically, the speech quality detection and feedback step involves: using a speech quality perception evaluation algorithm to detect the quality of the target speech signal; analyzing the clarity, naturalness, and noise residue of the target speech to comprehensively determine the speech quality level; when the quality of the target speech signal does not meet the preset standard, the quality evaluation result is fed back to the noise reduction stage, and the noise reduction operation is repeated until the speech quality is detected to meet the preset standard, before proceeding to the encoding and storage stage.

[0039] In this embodiment, the scenario is applied to a medium-sized conference room measuring 15m × 10m × 3m, and the characteristics of the scenario are as follows: Noise environment: Includes steady-state noise (ceiling air conditioner operating noise, sound pressure level 45dB) and non-steady-state noise (occasional door closing sound, chair dragging sound, peak sound pressure level 65dB). Reverberation characteristics: The conference room walls are painted with ordinary latex paint and the floor is carpeted. The reverberation time (RT60) is approximately 0.5 seconds. Recording requirements: Three participants are positioned 3m directly in front of the microphone array (speaker A), 4m to the left front (participant B), and 4.5m to the right front (participant C). The recording must ensure clear audio from all three participants, free from noise and reverberation. Hardware configuration: Multi-channel array microphone: 4 channels (channels 1-4), sampling rate 48kHz, bit depth 24bit, channel spacing 15cm; Signal processing terminal: Equipped with a processor, supports floating-point operations, and has 8GB of memory; Storage device: SSD solid-state drive, 1TB capacity.

[0040] Specific implementation steps Multi-channel array microphone configuration and raw signal acquisition Cooperative signal processing algorithm execution: After the microphone is activated, the dynamic amplitude compensation sub-algorithm acquires signals from each channel for 8 consecutive sampling cycles (20ms per cycle): the amplitude of channels 1-3 is stable at 0.8V, while the amplitude of channel 4 is only 0.5V due to low hardware sensitivity, which is determined to be a hardware sensitivity difference; a fixed compensation coefficient of 1.6 is calculated according to "target amplitude 0.8V / current amplitude 0.5V", and the gain of channel 4 is adjusted to 1.6 times, and finally the amplitude deviation of each channel is reduced to below 0.02V; The phase calibration sub-algorithm selects channel 2 as the reference channel (phase fluctuation frequency 0.2Hz, fluctuation amplitude 0.1rad, the most stable among all channels), calculates and integrates the phase difference between channels 1-4 and channel 2: the cumulative delay of channel 1 is 0.3rad, the cumulative delay of channel 3 is 0.25rad, and the cumulative delay of channel 4 is 0.4rad. Inverse compensation signals (-0.3rad, -0.25rad, -0.4rad) are generated and injected into the corresponding channels. After compensation, the phase difference of each channel is stabilized within 0.05rad. Anti-frame loss and audio distortion mechanism configuration: The buffer pool capacity is preset to 2 seconds (maximum single transmission delay of 0.8 seconds, 2 times the delay duration), and "first-in, first-out" storage is adopted. Every 10ms, a segment of signal is divided and an integrity identifier containing "signal length 10ms + CRC32 check code" is added. During the recording process, when the transmission load fluctuation was detected, a single-segment signal frame was lost (identifier missing) on ​​channel 3. A retransmission request was immediately triggered, and the corresponding 10ms signal was retrieved from the buffer pool to supplement the signal. There was no disconnection or audio distortion. Raw signal acquisition: 10 minutes of raw multi-channel audio signal was continuously acquired and stored in WAV format, with a single channel data volume of approximately 5.76 GB (48 kHz × 24 bit × 600 s).

[0041] Weighted kernel function blind source separation preprocessing Signal-to-noise ratio characteristic analysis: Temporal analysis: The original signal is divided into 50ms frames, and the energy fluctuation of each frame is calculated: the number of times the energy exceeds twice the average value per second in the air conditioner noise segment (without speech) reaches 15 times, with a fluctuation amplitude of 0.6V, and is judged as a low signal-to-noise ratio candidate frame; the number of fluctuations per second in the speech segment is 3 times, with a fluctuation amplitude of 0.2V, and is a high signal-to-noise ratio candidate frame; Frequency domain analysis: Perform Fourier transform on each frame to divide it into three segments: 20-200Hz (low frequency), 200-3400Hz (voice band), and 3400-20000Hz (high frequency). The air conditioner noise is concentrated in the 20-200Hz range. The speech energy / noise energy ratio in this frequency band is 0.6 (<1), which is determined to be a low signal-to-noise ratio frequency band. Region division: By superimposing the time domain and frequency domain results, the "20-200Hz frequency band + frames with frequent time domain fluctuations" is identified as the low signal-to-noise ratio region, and the rest is the high signal-to-noise ratio region; Weighting and Separation Execution: An S-shaped nonlinear weighting function is selected. In the low signal-to-noise ratio region, noise energy accounts for 72%, and the output weight is 1.8; in the high signal-to-noise ratio region, speech energy accounts for 85%, and the output weight is 0.9. Initial separation matrix of the independent component analysis model. Learning step size η = 0.2 (initial mutual information) After 6 iterations, the mutual information difference (≤0.01), stop iteration, and separate 3 target speech signal components (corresponding to A, B, C) and 1 interference signal component (including air conditioner noise, door closing sound, and reverberation).

[0042] Target-oriented adaptive cascaded noise reduction execution Directional pickup lock: Preset adjustable pickup range corresponding to the azimuth angle interval: Speaker A (0°±10°), Participant B (-30°±10°), Participant C (35°±10°); Spatial spectrum estimation algorithm acquires the phase difference Δϕ=0.5rad between channel 1 and channel 2, channel spacing d=0.15m, and voice frequency... (Wavelength λ = 0.34 m), substituting into the formula θ = arcsin(Δϕ) λ / (2πd)) is used to calculate the azimuth angle of A to be 1.2° (within 0°±10°), the azimuth angle of B to be -28°, and the azimuth angle of C to be 32°. All of these fall within the corresponding intervals, increasing the gain of the three channels to 1.2 times and suppressing interference outside the interval. Adaptive multimodal noise reduction: Time domain analysis: The peak amplitude of the door closing sound frame is 1.2V (the average of the first 5 frames is 0.3V, which is more than 3 times). The peak lasts for 150ms (<200ms), which is determined to be non-steady-state noise. The dynamic threshold is set to 0.6V. When the limit is exceeded, the gain drops to 0.5 times, and the noise attenuation is 32dB. Frequency domain analysis: The energy fluctuation of air conditioner noise in the 20-200Hz frequency band is 8% (<10%) for 10 consecutive frames, which is determined to be steady-state noise. The notch filter frequency is set to 20-200Hz, and the attenuation is 28dB. Adaptive processing of acoustic signals: Reverberation delay estimation: Identifying the start timestamp of the direct wave =0.2s, first reflected wave =0.6s, calculate the reverberation delay τ=0.4s; Dynamic filtering: Window length L = 2 × 0.4 = 0.8 s, cutoff frequency After filtering, the reverberation tail was shortened from 0.5 seconds to 0.1 seconds, with a reverberation suppression of 22dB; Speech masking neural networks and far-field sound pickup: Neural network: It is trained using a dataset of "clean speech (100,000 conference voices) + residual noise (50,000 air conditioner / reverberation noises)". The Bi-LSTM layer has 64 units, and the attention mechanism focuses on the 200-3400Hz frequency band. The output mask accuracy during inference is 92%, which can distinguish residual noise from target speech. Far-field pickup: TDOA calculation shows that with a distance of 3m between A and the microphone, the energy attenuation coefficient α = (1 / 3)2 ≈ 0.11, and the compensation coefficient G = 9, the energy is increased from 0.2V to 1.8V. The high-frequency band (2000-3400Hz) experiences severe attenuation. With a gain coefficient of 1.8, the energy proportion of the high-frequency band increases from 15% to 30%, and the speech clarity is improved by 35%.

[0043] S4: Lossless Encoding and Storage of Speech Features Feature extraction: Fundamental period: A has a fundamental period of 120ms (vowel segment), B has 110ms, and C has 130ms, marked as the energy concentration region; Silent segment: When there is no speech, the frame energy is 0.02V (1 / 15 of the average energy of 0.3V), which is determined to be a silent segment; Differential coding: Energy concentration area: 24-bit encoding sampling precision, preserving pitch and harmonic characteristics, with a PESQ score of 3.8; Silent segment: Sampling precision reduced to 16-bit, compression ratio 1:2.5, storage usage reduced by 60%; Storage: The encoded 3-channel audio is stored as MP3 format (lossless compression) according to "timestamp + speaker ID". The total storage size of 10 minutes of recording is 120MB, generating a noise-reduced recording file. Implementation effect verification Signal quality: The processed speech signal-to-noise ratio improved from 12dB to 48dB, and the intelligibility improved from 65% to 98%. Completeness: No dropped frames or audio distortion throughout the recording; 100% recording integrity. Storage efficiency: Compared to general lossless encoding, storage usage is reduced by 45%, balancing sound quality and efficiency.

[0044] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. An array microphone noise reduction recording method based on cascaded noise reduction and blind source separation, characterized in that, The method comprises the following steps: S1: configuring a multi-channel array microphone, which maintains the amplitude and phase consistency of the channel recording signals through a cooperative signal processing algorithm, monitors the channel recording process through an anti-frame loss and tone breaking mechanism, and collects original multi-channel voice signals in the environment; S2: preprocessing the original multi-channel voice signals based on a weighted kernel function blind source separation algorithm, first extracting the signal-to-noise ratio distribution characteristics, giving the kernel function a weight, and then decoupling the original multi-channel voice signals through an independent component analysis model to obtain target voice signal components and interference signal components; S3: performing a target-oriented adaptive cascading noise reduction model on the target voice signal components: first, locking the voice signal emitted by the main speaker based on a preset adjustable pickup range parameter through directional pickup technology, then using an adaptive multi-modal noise reduction algorithm to suppress the stable noise and unstable noise in the voice signal; subsequently, using an acoustic signal adaptive processing algorithm to filter out the voice reverberation tail in the noise-reduced voice signal, and finally combining a voice mask neural network and a far-field pickup algorithm to distinguish the residual noise and target voice in the voice signal output after reverberation filtering through the voice mask neural network, calculate the energy expectation of the target voice, and enhance the target voice in the far-field environment based on the far-field pickup algorithm; S4: using a voice feature perception lossless encoding algorithm for the target voice signal: extracting the pitch period and harmonic distribution characteristics of the target voice signal, using a high-fidelity encoding strategy for the voice energy concentration area, storing the encoded target voice signal, and generating a noise reduction recording process file.

2. The method of claim 1, wherein, In S1, the cooperative signal processing algorithm includes a dynamic amplitude compensation sub-algorithm and a phase calibration sub-algorithm: the dynamic amplitude compensation sub-algorithm real-time collects the amplitude difference of the channel recording signal, first analyzes the reason for the difference, then generates a dynamic scaling factor, and adjusts the signal gain of the amplitude channel through the dynamic scaling factor; the phase calibration sub-algorithm presets a reference channel, takes the phase of the reference channel as the reference, real-time calculates the phase delay of other channels and the reference channel, and then generates a reverse compensation signal according to the phase delay and superimposes it to the corresponding channel.

3. The method of claim 1, wherein, In S1, the anti-frame loss and tone breaking mechanism includes a signal cache pool and a frame loss detection and retransmission sub-algorithm: the signal cache pool presets a cache space that temporarily stores real-time recording signals within a continuous time; the frame loss detection and retransmission sub-algorithm real-time monitors the frame loss in the signal transmission process, compares the signal integrity before and after transmission to determine the frame loss, and when the frame loss affects the recording continuity, triggers a retransmission request immediately, retrieves the non-frame loss signal stored before the frame loss period from the signal cache pool, and supplements it to the transmission link.

4. The method of claim 1, wherein, In S2, the weight assignment mode of the weighted kernel function blind source separation algorithm is: first, the signal-to-noise ratio distribution characteristics of the original multi-channel speech signal are analyzed in multiple dimensions, the fluctuation frequency and amplitude of the energy of the original multi-channel speech signal in the time domain are monitored, the energy ratio of the original multi-channel speech signal and the noise signal in the frequency band is calculated in the frequency domain, and the low signal-to-noise ratio region and the high signal-to-noise ratio region are divided through the fusion judgment of the time domain and frequency domain characteristics; then, the kernel function weight is calculated by using a nonlinear weighting function, the weight value of the function is dynamically adjusted in combination with the noise energy proportion in the corresponding region, and the weight is assigned based on the actual characteristic difference of the original multi-channel speech signal.

5. The method of claim 1, wherein, In S2, the independent component analysis model integrates a gradient descent optimization sub-algorithm: the gradient descent optimization sub-algorithm takes the mutual information of the separated signal as an objective function, and reduces the correlation between the target speech signal component and the interference signal component through optimization; in iteration, the independent component analysis model parameters are adjusted according to the current calculated mutual information value, and the separation effect is gradually optimized; when the difference between the mutual information calculated by adjacent two iterations is reduced to a preset threshold, the iteration is stopped.

6. The method of claim 1, wherein, In S3, the directional sound pickup technology combines a spatial spectrum estimation algorithm: the spatial spectrum estimation algorithm first collects the phase information of the channel speech signal, calculates the spatial azimuth of the channel speech signal source by comparing the phase difference between channels; then, the spatial azimuth is compared with the azimuth angle range corresponding to the preset adjustable pickup range, when the spatial azimuth is in the azimuth angle range, it is determined as the target speaker speech signal, and the target speaker speech signal is locked by enhancing the channel signal gain.

7. The method of claim 1, wherein, In S3, the adaptive multi-modal noise reduction algorithm includes a noise type classification sub-algorithm: the noise type classification sub-algorithm first extracts the time domain and frequency domain characteristics of the speech signal to generate a feature analysis result, and classifies the noise into stationary noise and non-stationary noise according to the feature analysis result; for the stationary noise, a frequency domain notch filtering strategy is adopted to set a filter notch in the frequency band where the noise is concentrated, to weaken the stationary noise in the frequency band; for the non-stationary noise, a time domain threshold suppression strategy is adopted, a signal amplitude threshold is set, when the speech signal exceeds the threshold, it is determined as the non-stationary noise, and the speech signal is suppressed by a threshold filter.

8. The method of claim 1, wherein, In S3, the acoustic signal adaptive processing algorithm includes a reverberation time delay estimation sub-algorithm and a dynamic filtering sub-algorithm: the reverberation time delay estimation sub-algorithm first identifies the direct wave and the first reflected wave in the speech signal, the direct wave propagates directly from the sound source to the microphone, and the first reflected wave propagates to the microphone after one reflection, lagging behind the direct wave, and the reverberation time delay is obtained by calculating the time difference between the direct wave and the first reflected wave; the dynamic filtering sub-algorithm adjusts the length of the filtering window according to the reverberation time delay, and the window length needs to cover the duration of the reverberation tail to completely filter out the reverberation tail.

9. The method of claim 1, wherein, In S3, the speech mask neural network adopts an architecture combining a bidirectional long short-term memory network layer and an attention mechanism: the bidirectional long short-term memory network layer can simultaneously extract historical and future context features of the speech signal, extract the timing regularity of the speech signal from the start to the current time in the forward direction, extract the timing regularity of the speech signal from the current time to the end in the backward direction, and capture the timing characteristics of the speech signal through bidirectional feature fusion; the attention mechanism focuses on the frequency band where human speech is mainly distributed, and by strengthening the feature weight of the frequency band, the speech mask neural network gives priority to the frequency band when distinguishing between residual noise and human speech.

10. The method of claim 1, wherein, In S3, the far-field sound pickup algorithm includes a speech energy compensation sub-algorithm and a frequency equalization sub-algorithm: the speech energy compensation sub-algorithm is based on the attenuation law of far-field speech propagation, i.e., the law that the speech energy gradually weakens with increasing propagation distance, first analyzes the sound source distance in the current recording environment, then calculates the energy attenuation degree according to the sound source distance, generates a corresponding energy compensation coefficient, and enhances the overall energy of the far-field speech through the energy compensation coefficient; The frequency equalization sub-algorithm identifies the high-frequency region in the far-field speech in view of the characteristic that high-frequency signals are significantly attenuated in far-field propagation, and adjusts the signal strength of the high-frequency region through gain.

11. The method of claim 1, wherein, In S4, the speech feature perception lossless encoding algorithm includes a feature-driven encoding redundancy optimization sub-algorithm: first analyze the energy distribution characteristics of the target speech signal, identify the speech energy concentration area and the silent area, the speech energy concentration area is the core carrying part of the speech content, and the silent area has no effective speech and only includes weak background noise; For the speech energy concentration area, a high-fidelity encoding strategy is adopted to retain the detailed features of the signal, and for the silent area, an encoding redundancy compression strategy is adopted to remove redundant information in the signal.

12. The method of claim 1, wherein, Voice quality detection feedback step: a speech quality perception evaluation algorithm is used to detect the quality of the target speech signal, the intelligibility, naturalness, and noise residual amount of the target speech are analyzed, and the speech quality level is comprehensively judged; when the quality of the target speech signal does not meet the preset standard, the quality evaluation result is fed back to the noise reduction link, the noise reduction operation is repeated, until the quality of the speech reaches the preset standard, and then the encoding and storage step is entered.

Citation Information

Cited By

  • Audio processing method, device and system based on multi-modal noise reduction and sound source separation

    CN121725807A

  • Audio enhancement system and method of underground broadcasting system

    CN122090815A