Intelligent speech noise reduction system and method based on feature recognition
By separating and synthesizing speech signals and related signals using feature recognition technology, the problem of speech signal distortion in existing technologies is solved, and efficient noise reduction is achieved in complex noise environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-14
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies directly remove non-speech signals associated with the speech signal during the noise reduction process, resulting in distortion of the speech information content and affecting the application scope and signal fidelity.
By using feature recognition technology, speech signals and non-speech signals are separated, signal similarity coefficients are calculated, and candidate signals with correlation are synthesized with speech signals to obtain the target signal, thus avoiding speech signal distortion.
In complex noisy environments, the correlation information of the speech signal is preserved to the greatest extent, signal distortion is avoided, and signal processing efficiency is improved.
Smart Images

Figure CN116682445B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of speech processing and relates to intelligent speech noise reduction technology based on feature recognition, specifically an intelligent speech noise reduction system and method based on feature recognition. Background Technology
[0002] Wireless voice noise reduction technology refers to the technique of extracting and enhancing useful voice signals from the noisy background and reducing noise interference when wireless voice signals are interfered with by various noises during transmission. Voice noise reduction at the signal receiver mainly involves signal analysis in the time domain, frequency domain, and other transform domains to identify the differences between voice and noise for noise reduction.
[0003] Currently used methods mainly include designing bandpass filters and noise compensation algorithms, which perform well in stable noise environments. Existing technologies have improved upon these methods by separating speech information into speech and non-speech signals. They automatically identify speech information by extracting audio features and designing classifier models, and then remove non-speech signals based on the identification results to achieve noise reduction. However, current technologies only retain the speech signal from the speech information, directly removing the associated non-speech signals. This leads to distortion of the expressed content of the speech information and limits its application scope.
[0004] This invention provides an intelligent speech noise reduction system and method based on feature recognition to solve the above-mentioned problems. Summary of the Invention
[0005] The present invention aims to at least solve one of the technical problems existing in the prior art; to this end, the present invention proposes an intelligent speech noise reduction system and method based on feature recognition, which is used to solve the technical problem that the prior art directly removes non-speech signals associated with speech signals, resulting in distortion of the content expressed by speech information.
[0006] To achieve the above objectives, a first aspect of the present invention provides an intelligent speech noise reduction system based on feature recognition, including a central control module, and a speech acquisition device and a database connected thereto; the central control module acquires speech information through the speech acquisition device, and identifies speech signals and non-speech signals in the speech information through signal preprocessing and feature extraction; the central control module separates the non-speech signals according to speech separation technology to obtain several candidate signals; calculates the similarity coefficient between the speech signal and the several candidate signals based on speech features; and determines whether the similarity coefficient is greater than a similarity threshold; if yes, the corresponding candidate signal is marked as the basic signal; otherwise, it is not marked; and synthesizes the basic signal and the speech signal based on speech synthesis technology to obtain the target signal.
[0007] Existing technologies for speech denoising either have a narrow range of applicability, mostly effective only in environments with stable noise levels, or their denoising methods are too simplistic and crude, extracting only the speech signal containing the speech content and discarding other speech signals. Clearly, existing technologies cannot simultaneously balance applicability and signal fidelity, thus affecting the effectiveness of speech denoising.
[0008] This invention, after acquiring speech information, divides it into speech signals and non-speech signals based on whether it contains speech content. Speech separation technology is used to separate the non-speech information, obtaining several candidate signals. The similarity between the candidate signals and the speech signal is compared; a high similarity indicates a correlation between the two. Finally, speech synthesis technology is used to synthesize the speech signal and the candidate signals with high similarity to obtain the denoised target signal. Clearly, this invention comprehensively considers the correlation between signals during the speech denoising process, thus minimizing distortion of the denoised speech signal.
[0009] The central control module of this invention communicates and / or is electrically connected to the voice acquisition device and the database, respectively. The voice acquisition device is used to acquire voice information that needs to be denoised. The database is used to store the processed signal separation model and speech synthesis model. The central control module is responsible for all data processing, and the database stores the data required for each data process of this invention. The signal separation model and speech synthesis model in this invention are constructed by training an artificial intelligence model based on reliable theoretical methods and data from the database. The artificial intelligence model includes a BP neural network model or an RBF neural network model.
[0010] Preferably, the step of identifying speech signals and non-speech signals in speech information through signal preprocessing and feature extraction includes: performing Hamming window processing on the speech information after framing to obtain the original signal; extracting features from the original signal based on the Mel cepstral coefficients and their first-order difference and sub-band energy distribution; and designing a classifier model to divide the original signal into speech signals and non-speech signals.
[0011] The following features are mainly used to extract speech and non-speech signals from speech information:
[0012] 1) Me1 cepstral coefficients (MFCC) and their first-order differences
[0013] The human auditory system is a unique nonlinear system, exhibiting varying sensitivities to signals of different frequencies. The differential cepstral parameter (MCF) divides the frequency axis non-uniformly, combining the auditory perception characteristics of the human ear with the speech generation mechanism. Standard FCC parameters only reflect the static characteristics of speech parameters, while the human ear is more sensitive to the dynamic features of speech, which are typically described using differential cepstral parameters.
[0014] 2) Subband energy distribution
[0015] Within a frame of an audio signal, the ratio of the power spectral energy of each sub-band to the total power spectral energy of the entire frame is different, thus forming a distribution called the sub-band energy distribution.
[0016] Preferably, the central control module separates non-speech signals according to speech separation technology, including: constructing a speech separation model based on speech separation technology; and separating non-speech signals through the speech separation model to obtain several candidate signals.
[0017] Speech separation techniques include spectral subtraction, Wiener filtering, nonnegative matrix factorization, and computational auditory scene analysis, etc. The appropriate speech separation technique can be selected to construct the speech separation model based on the specific application scenario of the invention. The speech signal in speech information can be understood as the core content after speech denoising, while non-speech signals may still contain signals related to the speech signal.
[0018] Based on the established speech separation model, speech separation processing is performed on non-speech signals to separate different types of information from the non-speech signals. At this point, it is unknown which signals are related to the speech signals. Therefore, these signals separated from the non-speech signals are marked as candidate signals.
[0019] Preferably, the step of calculating the similarity coefficient between the speech signal and several candidate signals based on speech features includes: setting several feature items, verifying the weight coefficient of each feature item in the similarity calculation through speech synthesis data; extracting and integrating feature items with weight coefficients greater than weight thresholds to obtain several speech features; analyzing the changing trends of the speech signal and several candidate signals based on the several speech features, and obtaining the similarity coefficient according to the consistency of the changing trends.
[0020] To determine the similarity between two sets of signals, certain features are required. This invention sets several feature items based on features commonly used in the field of speech recognition. For example, fundamental frequency features mainly include the fundamental frequency and its mean, range of variation, rate of variation, and standard deviation; energy features mainly include short-time average energy, short-time energy change rate, short-time average amplitude, average amplitude change rate, and short-time maximum amplitude; duration features mainly include speech rate and short-time average zero-crossing rate. The similarity calculation is verified using the set speech synthesis data, and some feature items are selected as speech features based on their effectiveness.
[0021] The speech synthesis data in this invention is pre-selected speech data from various fields, such as music, construction, and remote work. After signal separation in this speech synthesis data, it is known which signals have a certain correlation. For example, in the wireless transmission of music, the signal corresponding to the singer's performance is correlated with the signal of the instrumental accompaniment, but it is not correlated with the noise of traffic (assuming music is played through wireless headphones on a bus). The role of each feature in determining correlation is verified using speech synthesis data with known signal correlations, and then several speech features are selected and acquired.
[0022] Preferably, the step of analyzing the changing trends of the speech signal and several candidate signals based on several speech features includes: numbering the several speech features one by one to obtain speech feature i; analyzing the changing trends of the speech signal and the candidate signals under speech feature i, and determining the similarity QDi based on the degree of consistency of the changing trends; and calculating the similarity coefficient XDX between the speech signal and the candidate signal according to the formula XDX=∑(QZi×QDi).
[0023] In the process of similarity analysis between speech signals and candidate signals, the changing trend of the same speech feature between the speech signal and the candidate signal can be analyzed, and the similarity can be determined according to the degree of consistency of the changing trend; if the changing trend of a certain speech feature is completely consistent between the two, the similarity is 100%.
[0024] Based on several speech features, similarity scores between a speech signal and a candidate signal can be obtained. These scores, combined with corresponding weighting coefficients, allow for the calculation of a similarity coefficient between the two signals. A higher similarity coefficient indicates a stronger correlation between the two. It's important to note that the weighting coefficients of each speech feature in the similarity coefficient calculation formula are related to the weighting coefficients of several speech features selected during speech synthesis data verification. Specifically, the weighting coefficients of the speech features obtained during speech synthesis data verification are adjusted so that the sum of all speech feature weighting coefficients equals 1, which is then substituted into the formula for calculation.
[0025] Preferably, the synthesis of the basic signal and the speech signal based on the speech synthesis technology includes: constructing a speech synthesis model based on the speech synthesis technology; and synthesizing the speech signal and the labeled basic signal through the speech synthesis model to obtain the target signal.
[0026] Speech synthesis technology includes waveform concatenation synthesis, statistical parametric speech synthesis, or end-to-end neural network speech synthesis. Similar to speech separation models, an appropriate speech synthesis technology is selected based on the application scenario of this invention, and then a speech synthesis model is constructed to ensure the effectiveness of speech synthesis.
[0027] By comparing the similarity coefficient with a similarity threshold, a suitable base signal can be selected from the candidate signals. Based on the speech synthesis model, the speech signal is synthesized with the acquired base signal to obtain the desired target signal. The target signal is the denoised speech information, which includes not only the speech content but also associated auxiliary signals.
[0028] A second aspect of the present invention provides an intelligent speech denoising method based on feature recognition, comprising: acquiring speech information; identifying speech signals and non-speech signals in the speech information through signal preprocessing and feature extraction; wherein, signal preprocessing includes framing and windowing; separating non-speech signals according to speech separation technology to obtain several candidate signals; calculating a similarity coefficient between the speech signal and several candidate signals based on speech features; determining whether the similarity coefficient is greater than a similarity threshold; if yes, marking the corresponding candidate signal as a base signal; otherwise, not marking it; and synthesizing the base signal and the speech signal based on speech synthesis technology to obtain a target signal.
[0029] Compared to traditional speech denoising methods, this method can be applied in complex noisy environments. Compared to denoising methods that directly identify the speech signal and remove all other signals, this method can identify the signals associated with the speech signal together, thus minimizing the distortion of speech information.
[0030] Compared with the prior art, the beneficial effects of the present invention are:
[0031] 1. This invention calculates the similarity coefficient between a speech signal and several candidate signals based on speech features; extracts several basic signals with strong correlation to the speech signal from non-speech signals according to the similarity coefficient, and synthesizes the basic signals and the speech signal to obtain the target signal; this invention avoids distortion after speech signal noise reduction through similarity analysis.
[0032] 2. This invention verifies the weight coefficients of each feature item in similarity calculation using speech synthesis data, determines several speech features based on the weight coefficients, and judges the correlation between candidate signals and speech signals based on the several speech features. This invention can quickly select the basic signal based on several speech features, thereby improving signal processing efficiency. Attached Figure Description
[0033] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0034] Figure 1This is a schematic diagram of the working steps of the present invention;
[0035] Figure 2 This is a schematic diagram of the system principle of the present invention. Detailed Implementation
[0036] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0037] Please see Figures 1-2 The first aspect of this invention provides an intelligent speech noise reduction system based on feature recognition, including a central control module and a speech acquisition device connected thereto. The central control module acquires speech information through the speech acquisition device, and identifies speech signals and non-speech signals in the speech information through signal preprocessing and feature extraction. The central control module separates the non-speech signals according to speech separation technology to obtain several candidate signals. It calculates the similarity coefficient between the speech signal and the several candidate signals based on speech features, and determines whether the similarity coefficient is greater than a similarity threshold. If yes, the corresponding candidate signal is marked as the basic signal; otherwise, it is not marked. The basic signal and the speech signal are synthesized based on speech synthesis technology to obtain the target signal.
[0038] This example uses the scenario of a user listening to music with wireless headphones on the subway:
[0039] Step 1: The central control module collects voice information through the voice acquisition device, and identifies the voice signals and non-voice signals in the voice information through signal preprocessing and feature extraction.
[0040] When users listen to music with wireless headphones on the subway, the signal is affected by the surrounding environment, specifically by electromagnetic waves emitted by surrounding electromagnetic devices and the subway itself. These effects introduce noise into the music signal, resulting in poor music listening quality.
[0041] Speech signals include speech content. Refer to Chinese invention patent CN101404160B, which discloses a speech noise reduction method based on audio recognition. This method introduces pattern recognition principles into communication speech noise reduction, dividing the audio signal into speech and non-speech components. By extracting speech features and designing a classifier model, the input signal is automatically identified to determine the audio type. If it is noise, it is removed; if it is speech, it is retained and further processed. The process of extracting speech and non-speech signals from the speech signal can be found in the signal preprocessing, feature extraction, and audio signal classification described in this patent.
[0042] Returning to this embodiment, after extracting the speech signal and non-speech signal from the music signal, the music signal only includes the singer's performance, while accompaniment, harmony, etc. will be classified as non-speech signal.
[0043] The second step is to calculate the similarity coefficient between the speech signal and several candidate signals based on the speech features; and to determine whether the similarity coefficient is greater than the similarity threshold. If yes, the corresponding candidate signal is marked as the base signal; otherwise, it is not marked.
[0044] Based on the scenario in this embodiment, a suitable speech separation model is selected. The speech separation model is used to separate and integrate non-speech signals to obtain several candidate signals. For example, harmony will be integrated into one candidate signal, a musical accompaniment will be divided into one candidate signal, and electromagnetic noise during subway operation will be integrated into one candidate signal.
[0045] Several feature terms are pre-defined, such as fundamental frequency variation, short-time energy variation, short-time maximum amplitude, and short-time average zero-crossing rate. Multiple speech synthesis data points from the same scenario are selected, but the correlation between the signals in the speech synthesis data is known beforehand. Feature terms that exhibit good correlation are extracted as speech features. Here, correlation mainly refers to the degree to which the feature terms fit the changing trends of the two signals.
[0046] This method classifies non-speech signals using several speech features, with the classification criterion being the speech features of the speech signal. This can be accomplished using a classification model established by an artificial intelligence model. In this embodiment, classification is achieved by calculating a similarity coefficient. Specifically, several speech features are numbered one by one to obtain speech feature i; the changing trends of the speech signal and the candidate signal under speech feature i are analyzed, and the similarity QDi is determined based on the degree of consistency of the changing trends; the similarity coefficient XDX between the speech signal and the candidate signal is calculated using the formula XDX = ∑(QZi × QDi).
[0047] The resulting similarity coefficient is essentially a score representing the correlation; a higher correlation coefficient indicates a stronger correlation between the candidate signal and the speech signal. Combined with a set correlation threshold, this allows for the filtering of candidate signals, yielding several basic signals. Of course, it's also possible that all candidate signals have low correlation with the speech signal, preventing the selection of any basic signals; in this case, the speech signal can be used as the final target signal.
[0048] Step 3: Synthesize the basic signal and the speech signal based on speech synthesis technology to obtain the target signal.
[0049] The best-performing speech synthesis technology for this application scenario was selected, and a speech synthesis model was trained using a large amount of data. The speech synthesis model then synthesizes the speech signal with several basic signals to obtain the target signal. In this application scenario, the target signal is the music the user is listening to, including the singer's vocals and the accompanying music; other signals are directly discarded.
[0050] Some of the data in the above formula are calculated by removing dimensions and taking their numerical values. The formula is the closest to the real situation obtained by software simulation of a large amount of collected data. The preset parameters and preset thresholds in the formula are set by those skilled in the art according to the actual situation or obtained through simulation of a large amount of data.
[0051] The working principle of this invention is as follows: Speech information is acquired, and speech and non-speech signals are identified through signal preprocessing and feature extraction. Non-speech signals are separated using speech separation technology to obtain several candidate signals. A similarity coefficient between the speech signal and the candidate signals is calculated based on speech features. It is then determined whether the similarity coefficient is greater than a similarity threshold; if yes, the corresponding candidate signal is marked as the base signal; otherwise, no marking is performed. The base signal and the speech signal are synthesized using speech synthesis technology to obtain the target signal.
[0052] The above embodiments are only used to illustrate the technical methods of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of the present invention without departing from the spirit and scope of the technical methods of the present invention.
Claims
1. A feature-based intelligent speech noise reduction system, comprising a central control module, and a speech acquisition device and a database connected thereto; characterized in that: The central control module collects voice information through a voice acquisition device and identifies voice signals and non-voice signals in the voice information through signal preprocessing and feature extraction; among which, signal preprocessing includes framing and windowing; The central control module separates non-speech signals using speech separation technology to obtain several candidate signals; it calculates the similarity coefficient between the speech signal and the candidate signals based on speech features; and... Determine if the similarity coefficient is greater than the similarity threshold; if yes, mark the corresponding candidate signal as the base signal; otherwise, do not mark it; synthesize the base signal and the speech signal based on speech synthesis technology to obtain the target signal.
2. The intelligent speech noise reduction system based on feature recognition according to claim 1, characterized in that, The process of identifying speech signals and non-speech signals in speech information through signal preprocessing and feature extraction includes: After framing the speech information, Hamming windowing is applied to obtain the original signal; Feature extraction of the original signal is performed based on the Mel cepstral coefficients, their first-order difference, and sub-band energy distribution. A classifier model is designed to divide the original signal into speech signals and non-speech signals.
3. The intelligent speech noise reduction system based on feature recognition according to claim 1, characterized in that, The central control module separates non-speech signals using speech separation technology, including: Speech separation models are constructed based on speech separation techniques, including spectral subtraction, Wiener filtering, nonnegative matrix factorization, or computational auditory scene analysis. Non-speech signals are separated using a speech separation model to obtain several candidate signals.
4. The intelligent speech noise reduction system based on feature recognition according to claim 1, characterized in that, The calculation of the similarity coefficient between the speech signal and several candidate signals based on speech features includes: Several feature terms are set up, and the weight coefficients of each feature term in similarity calculation are verified through speech synthesis data. Feature terms with weight coefficients greater than weight thresholds are extracted and integrated to obtain several speech features. Based on several speech features, the changing trends of the speech signal and several candidate signals are analyzed, and the similarity coefficient is obtained according to the degree of consistency of the changing trends.
5. The intelligent speech noise reduction system based on feature recognition according to claim 4, characterized in that, The analysis of the changing trends of the speech signal and several candidate signals based on several speech features includes: Several speech features are numbered one by one to obtain speech feature i; the changing trends of the speech signal and the candidate signal under speech feature i are analyzed, and the similarity QDi is determined according to the degree of consistency of the changing trends. The similarity coefficient XDX between the speech signal and the candidate signal is calculated according to the formula XDX=∑(QZi×QDi); where QZi is the weight coefficient corresponding to speech feature i, and i is a positive integer.
6. The intelligent speech noise reduction system based on feature recognition according to claim 1, characterized in that, The synthesis of basic signals and speech signals based on speech synthesis technology includes: Speech synthesis models are constructed based on speech synthesis technology; among which, speech synthesis technology includes waveform splicing synthesis technology, statistical parametric speech synthesis technology, or end-to-end neural network speech synthesis technology; The target signal is obtained by synthesizing speech signals and labeled base signals using a speech synthesis model.
7. The intelligent speech noise reduction system based on feature recognition according to claim 1, characterized in that, The central control module is communicatively and / or electrically connected to the voice acquisition device and the database, respectively; the voice acquisition device is used to acquire voice information that needs to be noise-reduced; The database is used to store processed signal separation models and speech synthesis models; wherein, the signal separation models and speech synthesis models are constructed based on artificial intelligence models.
8. A feature-recognition-based intelligent speech denoising method, operating based on the feature-recognition-based intelligent speech denoising system according to any one of claims 1 to 7, characterized in that, include: The system collects speech information and identifies speech and non-speech signals through signal preprocessing and feature extraction. Signal preprocessing includes framing and windowing. Non-speech signals are separated using speech separation technology to obtain several candidate signals; similarity coefficients between the speech signals and the candidate signals are calculated based on speech features. Determine if the similarity coefficient is greater than the similarity threshold; if yes, mark the corresponding candidate signal as the base signal; otherwise, do not mark it; synthesize the base signal and the speech signal based on speech synthesis technology to obtain the target signal.
Citation Information
Patent Citations
Voice denoising method based on audio recognition
CN101404160B
Voice denoising method based on audio recognition
CN101404160A
Call noise reduction method based on voiceprint recognition, call noise reduction device and earphone
CN114724565A