Voice signal output method and device, electronic equipment and storage medium
By employing a sparsity processing method combining dictionary learning and principal component analysis, an enhanced speech power spectrum is generated and the target weights are iteratively calculated. This solves the quality and clarity issues of speech signals under noise pollution and improves the signal-to-noise ratio.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-18
- Publication Date
- 2026-03-13
AI Technical Summary
During the acquisition of speech signals, the raw audio signals are easily contaminated by channel noise, environmental noise and random noise, which leads to a decrease in the quality and clarity of the speech signals.
A sparsity processing method based on dictionary learning and principal component analysis is adopted to generate the power spectrum of enhanced speech, and the target weight is determined by iterative calculation, and finally the final speech enhancement signal is output.
It improves the signal-to-noise ratio of the speech signal, effectively suppresses noise, and enhances the quality and clarity of the speech signal.
Smart Images

Figure CN121662069A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of voice signal output technology, and in particular to a voice signal output method, a voice signal output device, an electronic device, and a readable storage medium. Background Technology
[0002] Audio and video conferencing is widely used in modern work and life. However, during the acquisition of voice signals, the raw audio signals are easily contaminated by channel noise, environmental noise, random noise, and other interference. Therefore, how to effectively suppress these noises and improve the quality and clarity of voice signals has always been a research hotspot in the field of audio processing. Summary of the Invention
[0003] The present invention provides a method, apparatus, electronic device, and readable storage medium for outputting voice signals to overcome or at least partially solve the above-mentioned problems.
[0004] To solve the above-mentioned technical problems, this application is implemented as follows: In a first aspect, embodiments of this application provide a method for outputting voice signals, including: Acquire the raw speech signal; The original speech signal is subjected to dictionary-based sparsity processing to generate a first enhanced speech power spectrum; The original speech signal is subjected to sparsification processing based on principal component analysis to generate a second enhanced speech power spectrum composed of row components and column components; Determine the initial weights of the first enhanced speech power spectrum and the second enhanced speech power spectrum, and perform iterative calculations on the first enhanced speech power spectrum and the second enhanced speech power spectrum based on the initial weights to determine the target weights; The final speech enhancement signal is output based on the target weights.
[0005] Optionally, the step of performing dictionary-based sparsity processing on the original speech signal to generate a first enhanced speech power spectrum includes: Determine the time-domain form of the original speech signal; The original speech signal in time domain form is decomposed into the effective speech signal in time domain form and the irrelevant noise signal in time domain form; A short-time Fourier transform is performed using the time-domain form of the effective speech signal and the time-domain form of the irrelevant noise signal to generate short-time Fourier transform results for the effective speech signal and the irrelevant noise signal; Extract the amplitude information of the short-time Fourier transform result, and calculate the power spectrum of the effective speech signal and the irrelevant noise signal based on the amplitude information; Based on the power spectra of the effective speech signal and the irrelevant noise signal, iterative training is performed using a generative dictionary learning method to generate an effective speech signal dictionary matrix; Based on the effective speech signal dictionary matrix and the pre-acquired test signal, a first enhanced speech power spectrum is generated.
[0006] Optionally, the step of performing sparsification processing on the original speech signal based on principal component analysis to generate a second enhanced speech power spectrum composed of row and column components includes: Perform a Fourier transform on the original speech signal to generate signal amplitude information and phase information; Based on the signal amplitude information and the phase information, the original power spectrum is generated; Perform multi-dimensional spectral median filtering on the original power spectrum to decompose it into row and column components; Based on the row components and the column components, calculate the time dimension mask and the frequency dimension mask; The time-dimensional mask, the frequency-dimensional mask, and the original power spectrum are element-wise multiplied to generate a row component power spectrum matrix and a column component power spectrum matrix. The second enhanced speech power spectrum is calculated based on the row component power spectral moments, the column component power spectral matrices, and the scaling factor.
[0007] Optionally, the initial weights include a first initial weight for a first enhanced speech power spectrum and a second initial weight for a second enhanced speech power spectrum, wherein the first enhanced speech power spectrum is the output of a first model and the second enhanced speech power spectrum is the output of a second model. The step of performing iterative calculations on the first enhanced speech power spectrum and the second enhanced speech power spectrum based on the initial weights to determine the target weights includes: Determine the maximum number of iterations and the threshold for perceptual speech quality assessment score; Calculate the first speech quality perception evaluation score of the first enhanced speech power spectrum at a preset number of iterations; Calculate the second speech quality perception evaluation score of the second enhanced speech power spectrum at a preset number of iterations; Based on the first speech quality perception evaluation score, the second speech quality perception evaluation score, the first initial weight, and the second initial weight, the third speech quality perception evaluation score is calculated after the first model and the second model are fused and output. When the third speech quality perception evaluation score is greater than the speech quality perception evaluation score threshold or the current iteration number reaches the maximum iteration number, the iteration ends and the target weight is generated.
[0008] Optionally, it also includes: When the third speech quality perception evaluation score is not greater than the speech quality perception evaluation score threshold and the current iteration number has not reached the maximum iteration number, calculate the speech quality perception evaluation score difference between the first speech quality perception evaluation score and the second speech quality perception evaluation score. Based on the difference in the voice quality perception assessment score and the preset single adjustment scaling factor, the first initial weight and the second initial weight are adjusted to generate weight coefficients for the next iteration.
[0009] Optionally, it also includes: A preset function that is inversely proportional to the current iteration number is determined as the first dynamic adjustment scaling factor; When the third speech quality perception evaluation score is not greater than the speech quality perception evaluation score threshold and the current iteration number has not reached the maximum iteration number, calculate the speech quality perception evaluation score difference between the first speech quality perception evaluation score and the second speech quality perception evaluation score. Based on the difference in the perceived speech quality assessment score and the first dynamic adjustment scaling factor, the first initial weight and the second initial weight are adjusted to generate weight coefficients for the next iteration.
[0010] Optionally, it also includes: A preset function related to the difference between the third speech quality perception assessment score and the speech quality perception assessment score threshold is determined as the second dynamic adjustment scaling factor; When the third speech quality perception evaluation score is not greater than the speech quality perception evaluation score threshold and the current iteration number has not reached the maximum iteration number, calculate the speech quality perception evaluation score difference between the first speech quality perception evaluation score and the second speech quality perception evaluation score. Based on the difference in the perceived speech quality assessment score and the second dynamic adjustment scaling factor, the first initial weight and the second initial weight are adjusted to generate weight coefficients for the next iteration.
[0011] Secondly, embodiments of this application provide a voice signal output device, including: The raw speech signal acquisition module is used to acquire the raw speech signal; The first enhanced speech power spectrum generation module is used to perform dictionary-based sparsity processing on the original speech signal to generate the first enhanced speech power spectrum. The second enhanced speech power spectrum generation module is used to perform sparsification processing based on principal component analysis on the original speech signal to generate a second enhanced speech power spectrum composed of row components and column components. The target weight determination module is used to determine the initial weights of the first enhanced speech power spectrum and the second enhanced speech power spectrum, and to perform iterative calculations on the first enhanced speech power spectrum and the second enhanced speech power spectrum based on the initial weights to determine the target weights; The speech enhancement signal output module is used to output the final speech enhancement signal based on the target weight.
[0012] Thirdly, embodiments of this application provide an electronic device including a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the method described in the first aspect.
[0013] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0014] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.
[0015] The embodiments of the present invention have the following advantages: In this embodiment of the invention, the following steps are taken: First, an original speech signal is acquired. Then, a dictionary-based sparsification process is performed on the original speech signal to generate a first enhanced speech power spectrum. Next, principal component analysis (PCA)-based sparsification process is performed on the original speech signal to generate a second enhanced speech power spectrum composed of row and column components. Initial weights for the first and second enhanced speech power spectra are determined, and iterative calculations are performed on the first and second enhanced speech power spectra based on the initial weights to determine target weights. Finally, the final enhanced speech signal is output based on the target weights. This method integrates two high-performance sparsification speech enhancement models—dictionary learning and PCA—to adjust the fusion weights, thereby improving the signal-to-noise ratio of the output signal. Attached Figure Description
[0016] Figure 1 This is a flowchart of the steps of a voice signal output method provided in an embodiment of the present invention; Figure 2 This is a flowchart illustrating a voice signal output method provided in an embodiment of the present invention; Figure 3 This is a flowchart illustrating a method for generating a first enhanced speech power spectrum provided in an embodiment of the present invention. Figure 4This is a flowchart illustrating a method for generating a second enhanced speech power spectrum provided in an embodiment of the present invention; Figure 5 This is a flowchart illustrating another voice signal output method provided in an embodiment of the present invention; Figure 6 This is a structural block diagram of a voice signal output device provided in an embodiment of the present invention; Figure 7 This is a hardware structure block diagram of an electronic device provided in an embodiment of the present invention; Figure 8 This is a schematic diagram of a computer-readable medium provided in an embodiment of the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0018] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0019] To enable those skilled in the art to better understand the embodiments of the present invention, some technical terms involved in the embodiments of the present invention will be explained below.
[0020] Principal Component Analysis (PCA) is a multivariate statistical analysis method that uses linear transformations of multiple variables to select a smaller number of important variables.
[0021] Generative Dictionary Learning (GDL): A speech enhancement method proposed by Christian D. Sigg et al., see Appendix References [1].
[0022] Perceptual Evaluation Speech Quality (PESQ) is an objective, full-reference method for evaluating speech quality. Its designation in the International Telecommunication Union is ITU-T P.862.
[0023] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0024] Reference Figure 1 The diagram illustrates a flowchart of a speech signal output method provided in an embodiment of the present invention, which may specifically include the following steps: Step 101: Obtain the original speech signal; Step 102: Perform dictionary-based sparsity processing on the original speech signal to generate a first enhanced speech power spectrum; Step 103: Perform sparsification processing based on principal component analysis on the original speech signal to generate a second enhanced speech power spectrum composed of row components and column components; Step 104: Determine the initial weights of the first enhanced speech power spectrum and the second enhanced speech power spectrum, and perform iterative calculations on the first enhanced speech power spectrum and the second enhanced speech power spectrum based on the initial weights to determine the target weights; Step 105: Output the final speech enhancement signal based on the target weights.
[0025] In a specific implementation, the embodiments of the present invention can acquire raw input data containing the target valid speech signal and various irrelevant noises such as channel noise and environmental noise, as the basis for all subsequent sparsity processing and speech enhancement.
[0026] The embodiments of the present invention can perform sparse processing based on dictionary learning on the original speech signal to generate a first enhanced speech power spectrum, so as to take advantage of the advantage of dictionary learning to extract signal structural features more accurately and generate an enhanced speech power spectrum based on dictionary model processing.
[0027] Dictionary learning, also known as generative dictionary learning (GDL), is a speech enhancement method that constructs a dictionary through iterative training. After sparsely encoding the signal, the atoms in the dictionary are updated column by column, which can extract the structural features of the signal more accurately.
[0028] Sparsity processing: By significantly reducing the dimensionality of the signal through sparsity constraints, the important spectral components in the sound spectrum are enhanced, and background noise and irrelevant harmonic components are effectively reduced.
[0029] The embodiments of the present invention can perform sparsification processing based on principal component analysis on the original speech signal to generate a second enhanced speech power spectrum composed of row components and column components. This takes advantage of the fact that principal component analysis is suitable for processing high-dimensional data to generate another enhanced speech power spectrum decomposed into row components and column components.
[0030] Principal Component Analysis (PCA): A multivariate statistical analysis method that uses linear transformations of multiple variables to select a smaller number of important variables.
[0031] Row and column components: These are the components of the second enhanced speech power spectrum. The row component is horizontally smooth in the time dimension and has time stability; the column component is vertically smooth in the frequency dimension and has a wide vertical bandwidth.
[0032] The embodiments of the present invention can determine the initial weights of the first enhanced speech power spectrum and the second enhanced speech power spectrum, and perform iterative calculations on the first enhanced speech power spectrum and the second enhanced speech power spectrum based on the initial weights to determine the target weights. By using the Perceptual Speech Quality Assessment Score (PESQ) as an evaluation index, the weight coefficients of the two models are dynamically adjusted during the sparsity iteration process to determine the optimal fusion weights (target weights) that take into account the advantages of the two models.
[0033] Perceptual Speech Quality Assessment Score (PESQ): This is an objective, full-reference speech quality assessment method, and its standardized designation in the International Telecommunication Union is ITU-T P.862.
[0034] Initial weights are the starting point for the iterative optimization process, representing the initial contributions assigned to the dictionary learning model and the principal component analysis model at the beginning of optimization. They set the basis for calculating the fused speech quality perception assessment score in the first iteration. For example, the initial weights are typically set to be equal, implying that, in the absence of further information, the contributions of both models to the enhancement are presumed to be equally important.
[0035] The target weights are the final product of iterative calculations, representing the weight coefficients of the final fusion model determined after multiple dynamic adjustments and optimizations. They represent the contribution ratio of the two sparse models to the optimal fusion effect in the current specific speech enhancement task. The target weights are key parameters for the final output enhanced speech signal. They ensure that the power spectra of the two enhanced speech signals are weighted and fused in the optimal proportion, enabling the speech enhancement performance of the fusion model to meet the preset requirements.
[0036] Iterative calculation: refers to the process of repeatedly performing weight adjustment and fusion score calculation until the fusion score meets the preset conditions, so as to finally determine the final target weight.
[0037] The embodiments of the present invention can output the final speech enhancement signal based on the target weights, and use the target weights determined through iterative optimization to perform final weighted fusion of the power spectra of the two enhanced models to output the final result.
[0038] This invention involves acquiring an original speech signal; performing dictionary-based sparsification on the original speech signal to generate a first enhanced speech power spectrum; performing principal component analysis-based sparsification on the original speech signal to generate a second enhanced speech power spectrum composed of row and column components; determining initial weights for the first and second enhanced speech power spectra; performing iterative calculations on the first and second enhanced speech power spectra based on the initial weights to determine target weights; and outputting the final enhanced speech signal based on the target weights. This achieves improved speech enhancement performance and signal-to-noise ratio of the output signal by integrating two high-performance sparsification speech enhancement models—dictionary learning and principal component analysis—and dynamically adjusting the fusion weights using a speech quality perception evaluation score.
[0039] For example, refer to Figure 2 , Figure 2 This is a flowchart illustrating a voice signal output method provided in an embodiment of the present invention; S1. Establish a sparsity processing model based on dictionary learning, perform sparsity processing on the original speech signal using dictionary learning, and obtain the enhanced speech power spectrum after sparsification. ; S2. Establish a sparsification processing model based on principal component analysis (PCA) to perform spectral sparsification processing on the original speech signal, and output an enhanced speech power spectrum composed of row and column components. ; S3. The power spectrum matrix output in step S1 Set initial weight coefficients The sparsed power spectrum matrix output in step S2 Set weight coefficients Set the maximum number of iterations. and speech quality perception assessment score threshold For the first In the next iteration, calculate respectively and In the Speech quality perception assessment score corresponding to the second sparsening iteration and Calculate the perceptual evaluation score of speech quality after fusing the two models. Determine the fusion speech quality perception assessment score Is it greater than the threshold? If greater than The iteration ends if the result is less than or equal to 1. Then compare the speech quality perception evaluation scores of the two models. and The weight coefficients of the two models are dynamically adjusted as follows: and Repeat S3 for the next iteration until... Greater than the speech quality perception assessment score threshold Or reach the maximum number of iterations. When the iteration ends, the weight coefficients of the fusion model at the end are obtained. and Input voice signal After iteratively training the fusion model, an inverse Fourier transform is performed to output the final speech enhancement signal. .
[0040] Based on the above embodiments, modified embodiments of the above embodiments are proposed. It should be noted that, in order to keep the description brief, only the differences from the above embodiments are described in the modified embodiments.
[0041] In an optional embodiment of the present invention, the original speech signal time-domain form can be determined to determine the representation of the original input signal containing effective speech and noise in the time dimension, in preparation for subsequent signal decomposition and frequency domain analysis.
[0042] Time domain form: refers to the way a signal is represented as it changes over time. The time domain form of the original speech signal can be the way the original speech signal changes over time.
[0043] In this embodiment of the invention, the original speech signal in the time domain can be decomposed into the effective speech signal in the time domain and the irrelevant noise signal in the time domain, so as to decouple the complex original input signal in the time domain, separate the target speech signal and the noise signal that needs to be suppressed, and provide basic data for the subsequent dictionary learning model to train their respective characteristics.
[0044] Effective speech signal in time domain form: The target speech signal represented in the time domain.
[0045] Irrelevant noise signal in time domain form: background noise, environmental interference, and other signals that need to be suppressed, represented in the time domain.
[0046] In this embodiment of the invention, a short-time Fourier transform can be performed using the time-domain form of the effective speech signal and the time-domain form of the irrelevant noise signal to generate a short-time Fourier transform result for the effective speech signal and the irrelevant noise signal, thereby converting the time-domain signal to the frequency domain for spectrum analysis and power calculation. This is an essential stage for feature extraction in sparse representation.
[0047] Short-time Fourier Transform (STFT): A signal analysis technique used to analyze the frequency content of a signal within a local time period, that is, to decompose a time-domain signal into a series of spectra.
[0048] In this embodiment of the invention, the amplitude information of the short-time Fourier transform result can be extracted, and the power spectrum of the effective speech signal and the irrelevant noise signal can be calculated based on the amplitude information to extract the energy distribution information (i.e., power spectrum) of the signal in the frequency domain, which can be used as the training data input for the dictionary learning model.
[0049] Power spectrum: A function that describes the distribution of signal energy or power in the frequency domain.
[0050] Amplitude information: The magnitude of the signal amplitude in the frequency domain.
[0051] In this embodiment of the invention, based on the power spectrum of the effective speech signal and the irrelevant noise signal, iterative training can be performed using generative dictionary learning to generate an effective speech signal dictionary matrix. By utilizing the iterative training mechanism of generative dictionary learning (GDL), the most concise and representative features of the effective speech signal are extracted and learned from the power spectrum, and finally an efficient effective speech signal dictionary matrix is constructed.
[0052] Generative dictionary learning (GDL): A speech enhancement method that constructs a dictionary through iterative training to extract and learn features from a signal.
[0053] Effective speech signal dictionary matrix: A matrix obtained by dictionary learning, where each column (atom) represents the basic building block of the effective speech signal, used for subsequent sparse coding and reconstruction.
[0054] In this embodiment of the invention, a first enhanced speech power spectrum can be generated based on the effective speech signal dictionary matrix and the pre-acquired test signal. The trained effective speech signal dictionary matrix is then used to perform sparse coding and reconstruction on the actual test signal (original noisy signal), thereby suppressing noise, highlighting effective speech components, and finally generating an enhanced speech power spectrum based on a dictionary learning model.
[0055] Test signal: refers to the original noisy speech signal to be enhanced.
[0056] First enhanced speech power spectrum: the frequency domain energy representation of the enhanced speech signal obtained after processing by the dictionary learning model.
[0057] This model example leverages the iterative training advantages of generative dictionary learning to construct an accurate and effective speech dictionary matrix, achieving efficient extraction of effective speech components and noise suppression from the original signal. This lays a high-performance foundation for subsequent multi-model fusion.
[0058] refer to Figure 3 , Figure 3 This is a flowchart illustrating a method for generating a first enhanced speech power spectrum provided in an embodiment of the present invention. For step S1, the steps of establishing the sparse speech enhancement model based on the dictionary learning model and calculating the sparse power spectrum of the model are explained in detail.
[0059] S11. Establish a sparse speech enhancement model based on a dictionary learning model, where the input original speech signal consists of valid speech signal and invalid noise, as shown in the following expression:
[0060] In the above formula, This is the time-domain form of the original speech signal. For an effective time-domain form of speech signals, This is the time-domain form of an irrelevant noise signal.
[0061] S12, convert the original time-domain signal Performing a short-time Fourier transform for frame-by-frame analysis yields the following expression:
[0062] In the above formula, , , These are the short-time Fourier transform results of the original speech signal, the valid speech signal, and the irrelevant noise signal, respectively. and The expression is as follows:
[0063] In the above formula, It is a time-domain window function. These are discrete frequency points. It is the duration of each frame of signal during sampling. It is the symbol for an imaginary number. It is the subscript of the signal frame number.
[0064] S13. Extract amplitude information from the short-time Fourier transform result of the signal in step S12, calculate the power spectrum, ignoring phase information, as shown in the following expression:
[0065] In the above formula, , , These are the amplitude spectra of the short-time Fourier transforms of the original speech signal, the valid speech signal, and the irrelevant noise signal, respectively. The power spectrum is obtained by squaring them.
[0066] S14. Train the model by using existing speech and noise signals through short-time Fourier transform. and The objective function for training the generative dictionary learning method is as follows:
[0067] In the above formula, Right now The amplitude spectrum of the effective speech signal after short-time Fourier transform. Right now The amplitude spectrum of the short-time Fourier transform of the signal with no noise is shown. and These are the dictionary matrices corresponding to the valid speech signal and the irrelevant noise signal, respectively. and These are the coefficient matrices corresponding to the effective speech signal and the irrelevant noise signal, respectively. and These are the first and second coefficient matrices corresponding to the effective speech signal and the irrelevant noise signal, respectively. List It is an abbreviation for constraint conditions.
[0068] S15. After training is complete, for the input test signal vector The enhanced speech power spectrum output after processing The expression is:
[0069] In the above formula, It is the input test signal vector The enhanced speech power spectrum output after sparsification processing based on dictionary learning method. It is a dictionary matrix of valid speech signals. It is the coefficient matrix corresponding to the valid speech signal.
[0070] In an optional embodiment of the present invention, a Fourier transform can be performed on the original speech signal to generate signal amplitude information and phase information, so as to convert the original speech signal in the time domain to the frequency domain and obtain the signal intensity (amplitude information) and time relationship (phase information) at different frequencies.
[0071] Fourier Transform: A mathematical tool used to transform a signal from the time domain to the frequency domain, revealing the frequency components of the signal.
[0072] Amplitude information: The magnitude of the Fourier transform result, representing the strength or energy of the signal at a specific frequency.
[0073] Phase information: The argument (angle) of the Fourier transform result indicates the starting position or relative time relationship of the signal at a specific frequency.
[0074] In this embodiment of the invention, an original power spectrum can be generated based on the signal amplitude information and the phase information. Based on the amplitude information obtained by Fourier transform, the energy distribution of the original speech signal in the frequency domain can be calculated to form an original power spectrum matrix, which serves as input data for subsequent principal component analysis (PCA) processing.
[0075] Original power spectrum: A function describing the distribution of the original signal energy in the frequency domain, usually calculated by squaring the amplitude information.
[0076] In this embodiment of the invention, multi-dimensional spectral median filtering can be performed on the original power spectrum to decompose it into row components and column components. Through multi-dimensional filtering, the original power spectrum can be decomposed into two uncorrelated components with different characteristics (time stability and broadband effect) to achieve signal dimensionality reduction and feature extraction.
[0077] Multidimensional spectral median filtering: a signal processing technique used to filter signals on a spectrogram to separate different structural characteristics of the signal.
[0078] Row component: A component that is smooth in the frequency dimension and has time stability.
[0079] Column component: A component that is smooth in the time dimension and has a broadband effect.
[0080] In this embodiment of the invention, time-dimensional masks and frequency-dimensional masks can be calculated based on the row components and column components. Based on the decomposed row components and column components, time-dimensional masks and frequency-dimensional masks for selectively retaining or suppressing signals can be calculated respectively. These masks will guide subsequent power spectrum reconstruction.
[0081] Mask: In signal processing, a matrix used to selectively preserve (usually a value of 1) or suppress (usually a value of 0) specific regions in the spectrum to achieve noise reduction.
[0082] In this embodiment of the invention, the time-dimensional mask, the frequency-dimensional mask, and the original power spectrum can be element-wise multiplied to generate a row component power spectrum matrix and a column component power spectrum matrix. By applying the calculated mask and performing element-wise multiplication, an enhanced power spectrum matrix containing only row component characteristics and column component characteristics can be separated and extracted from the original power spectrum.
[0083] Element-wise product: Element-wise multiplication between matrices, also known as the Hadamard product, is a mathematical operation that applies masks.
[0084] Row component power spectrum matrix and column component power spectrum matrix: After masking, only the enhanced matrices of the row component and column component characteristics in the original power spectrum are retained.
[0085] In this embodiment of the invention, a second enhanced speech power spectrum can be calculated based on the row component power spectrum matrix, the column component power spectrum matrix, and the scaling factor. The row component power spectrum matrix and the column component power spectrum matrix obtained after independent processing are linearly combined according to a preset scaling factor to generate the final second enhanced speech power spectrum, which is the output of this sub-step.
[0086] Scaling factor: A coefficient used to balance the contributions of the row component power spectrum matrix and the column component power spectrum matrix in the final combination.
[0087] In this invention, the original speech power spectrum is efficiently decomposed into two uncorrelated and structurally stable components by using a principal component analysis model, and selective reconstruction is performed by combining masking techniques, thereby achieving effective noise reduction and feature extraction of high-dimensional noisy signals.
[0088] For example, refer to Figure 4 , Figure 4 This is a flowchart illustrating a method for generating a second enhanced speech power spectrum provided in an embodiment of the present invention; For step S2, the steps of establishing the sparse speech enhancement model based on the principal component analysis model and calculating the sparse power spectrum of the model are explained in detail.
[0089] S21. Preprocess the input raw speech, including performing Fourier transform on the raw time-domain signal to obtain signal amplitude and phase information, and calculating the power spectrum of the amplitude vector after Fourier transform. S22. Perform multi-dimensional spectral median filtering on the power spectrum to decompose the travel components. Sum of components , and The expression is as follows:
[0090] In the above formula, This indicates taking the median value of the sampled signal. and These are the time dimensions of the original power spectrum. The data and frequency dimensions of frames are in The frequency band at that location, ,in This is the sliding window length for median filtering.
[0091] S23, Computing the mask for the time dimension Masks with frequency dimension :
[0092] In the above formula, and These are the row and column components obtained by median filtering, respectively. It is the frequency band of the signal power spectrum. It is the frame number subscript of the time dimension of the signal power spectrum.
[0093] S24. Mask based on the time dimension Frequency dimension mask and the original signal power spectrum The power spectrum of the line components can be obtained by calculation. Power spectrum of components The formula is as follows:
[0094] In the above formula, It is the power spectrum vector of the input signal. It is the row component power spectrum matrix. It is a column component power spectrum matrix. It is the identifier for performing a Hadamard product on a matrix. It is a mask in the time dimension. It is a mask in the frequency dimension.
[0095] S25. Based on the row component power spectrum matrix Sum of column component power spectrum matrix The enhanced speech power spectrum can then be calculated. The formula is as follows:
[0096] In the above formula, It is the enhanced speech power spectrum output after sparsification processing based on principal component analysis. It is the row component power spectrum matrix. It is horizontally smooth in the time dimension and has time stability. It is a column component power spectrum matrix. It has a wide bandwidth effect in the vertical frequency band, and it is vertically smooth in the frequency dimension. It is a scaling factor, and its value range is... The value is generally around 0.7.
[0097] In an optional embodiment of the present invention, the initial weights include a first initial weight for a first enhanced speech power spectrum and a second initial weight for a second enhanced speech power spectrum, wherein the first enhanced speech power spectrum is the output of a first model and the second enhanced speech power spectrum is the output of a second model.
[0098] In practical applications, the first model in this embodiment of the invention can be a dictionary learning model for outputting a first enhanced speech power spectrum; the second learning model can be a principal component analysis model for outputting a second enhanced speech power spectrum.
[0099] In this embodiment of the invention, a maximum number of iterations and a speech quality perception evaluation score threshold can be determined to set the termination conditions and performance goals of the iteration process. The maximum number of iterations is used to limit the consumption of computing resources and prevent the model from running indefinitely; the PESQ threshold sets the objective quality standard that the speech enhancement effect needs to achieve.
[0100] Maximum number of iterations: A preset upper limit on the number of iterations. When the current number of iterations reaches this value, the iteration will terminate regardless of whether the performance target is met.
[0101] Speech quality perception assessment score threshold: A preset standard value for the acceptable speech enhancement effect. When the fusion score exceeds this threshold, the enhancement effect is considered to meet the requirements.
[0102] In this embodiment of the invention, the first speech quality perception evaluation score of the first enhanced speech power spectrum at a preset number of iterations can be calculated to evaluate the speech quality level achieved by the enhanced power spectrum independently output by the dictionary learning model at the current number of iterations.
[0103] First speech quality perceived assessment score: An objective speech quality score calculated from the power spectrum output by the dictionary learning model in the i-th iteration, according to the ITU-T P.862 standard procedure.
[0104] In this embodiment of the invention, the second speech quality perception evaluation score of the second enhanced speech power spectrum at a preset number of iterations can be calculated to evaluate the speech quality level achieved by the enhanced power spectrum independently output by the principal component analysis model at the current number of iterations.
[0105] The second speech quality perception assessment score is an objective speech quality score calculated from the power spectrum output by the principal component analysis model in the i-th iteration, according to the ITU-T P.862 standard procedure.
[0106] In this embodiment of the invention, a third speech quality perception evaluation score can be calculated based on the first speech quality perception evaluation score, the second speech quality perception evaluation score, the first initial weight, and the second initial weight. This score is then used to calculate the overall speech quality achievable by fusing the two enhanced power spectra at this point, based on the weight coefficients of the current iteration and the independent PESQ scores of the two models. This score serves as the basis for determining whether to continue iteration and adjust the weights.
[0107] The third speech quality perception assessment score represents the overall speech quality score after the dictionary learning model and the principal component analysis model are weighted and fused under the current weights.
[0108] The third speech quality perception evaluation score being greater than the speech quality perception evaluation score threshold or the current iteration number reaching the maximum iteration number is the termination condition for the iteration. In this embodiment of the invention, when the third speech quality perception evaluation score is greater than the speech quality perception evaluation score threshold or the current iteration number reaches the maximum iteration number, the iteration ends and a target weight is generated. When the speech enhancement effect (fusion PESQ) meets the standard or reaches the preset maximum computational load, the iteration process stops and the current weight coefficient is locked as the final target weight.
[0109] Target weights: The weight coefficients finally determined at the end of the iteration process, representing the optimal contribution ratio of the two models to the final enhancement effect.
[0110] In this embodiment of the invention, when the third speech quality perception evaluation score is not greater than the speech quality perception evaluation score threshold and the current iteration number has not reached the maximum iteration number, the speech quality perception evaluation score difference between the first speech quality perception evaluation score and the second speech quality perception evaluation score can be calculated. This is to quantify the performance difference between the two models in the current iteration when the performance does not meet the standard but the calculation is allowed to continue, and to provide data basis for the next step of dynamic weight adjustment.
[0111] Speech quality perception evaluation score difference: refers to the absolute difference between the first speech quality perception evaluation score and the second speech quality perception evaluation score, used to measure which model performs better and is more worthy of increasing weight in the current iteration.
[0112] In this embodiment of the invention, the first initial weight and the second initial weight can be adjusted based on the difference in the speech quality perception evaluation score and a preset single-adjustment scaling factor to generate weight coefficients for the next iteration, thereby implementing a dynamic weight adjustment mechanism. According to the performance difference between the two models and a preset adjustment step size, the weights are tilted towards the model with better performance, thus generating new weight coefficients for the next iteration calculation.
[0113] Single-step scaling factor: A preset fixed adjustment step size used to control the magnitude of weight adjustment in each iteration. It determines the rate of change of the weight coefficients.
[0114] In this embodiment of the invention, PESQ is used as an objective evaluation index, and a dynamic adjustment mechanism is adopted to achieve complementary advantages and performance optimization of two sparse speech enhancement models: dictionary learning and principal component analysis.
[0115] For example, refer to Figure 5 , Figure 5 This is a flowchart illustrating another voice signal output method provided in an embodiment of the present invention; Regarding step S3 in Example 1, the method of iteratively fusing the sparse model based on dictionary learning and the sparse model based on principal component analysis in step S3 will be explained in detail: S31. The power spectrum matrix output in step S1 Set initial weight coefficients The sparsed power spectrum matrix output in step S2 Set weight coefficients , and The initial values for all elements are 1; the maximum number of iterations is set. , The value needs to be determined based on the device's performance; based on experience, it is generally set to... Set a threshold for the perceived speech quality assessment score. Based on the empirical speech quality perception assessment score threshold Setting it to around 3.5 to 4 is advisable.
[0116] S32, in the In the next iteration, calculate respectively and In the Speech quality perception assessment score corresponding to the second sparsening iteration and The calculation method for the voice quality perception assessment score refers to the standard procedure provided by the International Telecommunication Union with the standardization code ITU-T P.862; S33. Calculate the speech quality perception assessment score after fusing the two models. The calculation formula is as follows:
[0117] In the above formula, It is the first The weighting coefficients of the power spectrum output by the dictionary learning model in the next iteration. It is the first In the next iteration, the dictionary learning model outputs a speech quality perception score based on the power spectrum. It is the first The weighting coefficients of the power spectrum output by the principal component analysis model in the next iteration. It is the first The speech quality perception assessment score of the power spectrum output by the principal component analysis model in the next iteration.
[0118] S34. Determine the perceptual assessment score of fused speech quality. Is it greater than the threshold? If greater than The iteration ends if the result is less than or equal to 1. Then compare the speech quality perception evaluation scores of the two models. and And dynamically adjust the weight coefficients of the two models to and The specific method is to calculate the first... The speech quality perception assessment score of the output power spectrum of the dictionary learning model and the principal component analysis model at the next iteration. and absolute value of the difference Then the first In the next iteration, the weight coefficients and The expression is:
[0119]
[0120] In the above formula, and They are the first The weight coefficients of the dictionary learning model and the principal component analysis model in the next iteration. and They are the first In the next iteration, the dictionary learning model and the principal component analysis model output the power spectrum of the speech quality perception assessment score. yes and The absolute value of the difference It adjusts the scaling factor only once, and sets a fixed value for each iteration. The range of values for is The specific value needs to be adjusted by professionals based on the actual effect; S35. Repeat steps S32, S33, and S34 for the next iteration until... Greater than the speech quality perception assessment score threshold Or reach the maximum number of iterations. When the iteration ends, the weight coefficients of the fusion model at the end are obtained. and ; S36, Input voice signal The fusion model, after iterative training, undergoes an inverse Fourier transform to output the final speech enhancement signal. The output voice enhancement signal The expression is:
[0121] In the above formula, Indicates the inverse Fourier transform. and These are the weight coefficients of the dictionary learning model and the principal component analysis model after the iteration. and These are the enhanced speech power spectra of the dictionary learning model and the principal component analysis model after the iteration.
[0122] Optionally, it also includes: A preset function that is inversely proportional to the current iteration number is determined as the first dynamic adjustment scaling factor; When the third speech quality perception evaluation score is not greater than the speech quality perception evaluation score threshold and the current iteration number has not reached the maximum iteration number, calculate the speech quality perception evaluation score difference between the first speech quality perception evaluation score and the second speech quality perception evaluation score. Based on the difference in the perceived speech quality assessment score and the first dynamic adjustment scaling factor, the first initial weight and the second initial weight are adjusted to generate weight coefficients for the next iteration.
[0123] For example, the first dynamic adjustment scaling factor η (i) It can be designed in the following way: η (i) It can be designed as a function inversely proportional to i, for example: η (i) =C / i or η (i) =C*e -ki Where C and k are constants, the adjustment step size gradually decreases as the number of times the slugging increases.
[0124] Optionally, it also includes: A preset function related to the difference between the third speech quality perception assessment score and the speech quality perception assessment score threshold is determined as the second dynamic adjustment scaling factor; When the third speech quality perception evaluation score is not greater than the speech quality perception evaluation score threshold and the current iteration number has not reached the maximum iteration number, calculate the speech quality perception evaluation score difference between the first speech quality perception evaluation score and the second speech quality perception evaluation score. Based on the difference in the perceived speech quality assessment score and the second dynamic adjustment scaling factor, the first initial weight and the second initial weight are adjusted to generate weight coefficients for the next iteration.
[0125] For example, the second dynamically adjusted scaling factor η (i) It can be designed in the following way: η (i) It can be designed as a function relating to the difference between the third speech quality perception assessment score and the speech quality perception assessment score threshold, for example: η (i) =f(Third Speech Quality Perception Assessment Score - Speech Quality Perception Assessment Score Threshold) When the distance is large, η (i) Large, and vice versa.
[0126] Beneficial effects: Accelerated convergence: In the early stages of iteration, if the adjustment coefficient is large, it can approach the optimal weight more quickly, reducing the total number of iterations required.
[0127] Improved accuracy and stability: In the later stages of iteration, the adjustment coefficient is reduced, which avoids oscillations around the speech quality perception assessment score threshold caused by a fixed large step size, making the final weights more accurate and stable.
[0128] More adaptable: Compared to fixed coefficients, adaptive learning rates can better adapt to the needs of different noise environments or speech data for weight adjustment.
[0129] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.
[0130] Reference Figure 6The diagram shows a structural block diagram of a voice signal output device provided in an embodiment of the present invention, which may specifically include the following modules: The raw speech signal acquisition module 601 is used to acquire the raw speech signal; The first enhanced speech power spectrum generation module 602 is used to perform dictionary-based sparsity processing on the original speech signal to generate the first enhanced speech power spectrum. The second enhanced speech power spectrum generation module 603 is used to perform sparsification processing based on principal component analysis on the original speech signal to generate a second enhanced speech power spectrum composed of row components and column components. The target weight determination module 604 is used to determine the initial weights of the first enhanced speech power spectrum and the second enhanced speech power spectrum, and to perform iterative calculations on the first enhanced speech power spectrum and the second enhanced speech power spectrum based on the initial weights to determine the target weights; The speech enhancement signal output module 605 is used to output the final speech enhancement signal based on the target weight.
[0131] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.
[0132] In addition, embodiments of the present invention also provide an electronic device, such as... Figure 7 As shown, it includes a processor 701, a communication interface 702, a memory 703, and a communication bus 704, wherein the processor 701, the communication interface 702, and the memory 703 communicate with each other through the communication bus 704. Memory 703 is used to store computer programs; When the processor 701 executes the program stored in the memory 703, it implements any of the voice signal output methods described in the above embodiments: The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0133] The communication interface is used for communication between the aforementioned terminal and other devices.
[0134] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0135] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0136] like Figure 8 As shown, in another embodiment of the present invention, a computer-readable storage medium 801 is also provided, which stores instructions that, when executed on a computer, cause the computer to perform the voice signal output method described in the above embodiment.
[0137] This application also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described voice signal output method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0138] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0139] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0140] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0141] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A method for outputting a voice signal, characterized in that, include: Acquire the raw speech signal; The original speech signal is subjected to dictionary-based sparsity processing to generate a first enhanced speech power spectrum; The original speech signal is subjected to sparsification processing based on principal component analysis to generate a second enhanced speech power spectrum composed of row components and column components; Determine the initial weights of the first enhanced speech power spectrum and the second enhanced speech power spectrum, and perform iterative calculations on the first enhanced speech power spectrum and the second enhanced speech power spectrum based on the initial weights to determine the target weights; The final speech enhancement signal is output based on the target weights.
2. The method according to claim 1, characterized in that, The step of performing dictionary-based sparsity processing on the original speech signal to generate the first enhanced speech power spectrum includes: Determine the time-domain form of the original speech signal; The original speech signal in time domain form is decomposed into the effective speech signal in time domain form and the irrelevant noise signal in time domain form; A short-time Fourier transform is performed using the time-domain form of the effective speech signal and the time-domain form of the irrelevant noise signal to generate short-time Fourier transform results for the effective speech signal and the irrelevant noise signal; Extract the amplitude information of the short-time Fourier transform result, and calculate the power spectrum of the effective speech signal and the irrelevant noise signal based on the amplitude information; Based on the power spectra of the effective speech signal and the irrelevant noise signal, iterative training is performed using a generative dictionary learning method to generate an effective speech signal dictionary matrix; Based on the effective speech signal dictionary matrix and the pre-acquired test signal, a first enhanced speech power spectrum is generated.
3. The method according to claim 1, characterized in that, The step of performing sparsification processing based on principal component analysis on the original speech signal to generate a second enhanced speech power spectrum composed of row and column components includes: Perform a Fourier transform on the original speech signal to generate signal amplitude information and phase information; Based on the signal amplitude information and the phase information, the original power spectrum is generated; Perform multi-dimensional spectral median filtering on the original power spectrum to decompose it into row and column components; Based on the row components and the column components, calculate the time dimension mask and the frequency dimension mask; The time-dimensional mask, the frequency-dimensional mask, and the original power spectrum are element-wise multiplied to generate a row component power spectrum matrix and a column component power spectrum matrix. The second enhanced speech power spectrum is calculated based on the row component power spectral moments, the column component power spectral matrices, and the scaling factor.
4. The method according to claim 1, characterized in that, The initial weights include a first initial weight for a first enhanced speech power spectrum and a second initial weight for a second enhanced speech power spectrum, wherein the first enhanced speech power spectrum is the output of a first model and the second enhanced speech power spectrum is the output of a second model. The step of performing iterative calculations on the first enhanced speech power spectrum and the second enhanced speech power spectrum based on the initial weights to determine the target weights includes: Determine the maximum number of iterations and the threshold for perceptual speech quality assessment score; Calculate the first speech quality perception evaluation score of the first enhanced speech power spectrum at a preset number of iterations; Calculate the second speech quality perception evaluation score of the second enhanced speech power spectrum at a preset number of iterations; Based on the first speech quality perception evaluation score, the second speech quality perception evaluation score, the first initial weight, and the second initial weight, the third speech quality perception evaluation score is calculated after the first model and the second model are fused and output. When the third speech quality perception evaluation score is greater than the speech quality perception evaluation score threshold or the current iteration number reaches the maximum iteration number, the iteration ends and the target weight is generated.
5. The method according to claim 4, characterized in that, Also includes: When the third speech quality perception evaluation score is not greater than the speech quality perception evaluation score threshold and the current iteration number has not reached the maximum iteration number, calculate the speech quality perception evaluation score difference between the first speech quality perception evaluation score and the second speech quality perception evaluation score. Based on the difference in the voice quality perception assessment score and the preset single adjustment scaling factor, the first initial weight and the second initial weight are adjusted to generate weight coefficients for the next iteration.
6. The method according to claim 4, characterized in that, Also includes: A preset function that is inversely proportional to the current iteration number is determined as the first dynamic adjustment scaling factor; When the third speech quality perception evaluation score is not greater than the speech quality perception evaluation score threshold and the current iteration number has not reached the maximum iteration number, calculate the speech quality perception evaluation score difference between the first speech quality perception evaluation score and the second speech quality perception evaluation score. Based on the difference in the perceived speech quality assessment score and the first dynamic adjustment scaling factor, the first initial weight and the second initial weight are adjusted to generate weight coefficients for the next iteration.
7. The method according to claim 4, characterized in that, Also includes: A preset function related to the difference between the third speech quality perception assessment score and the speech quality perception assessment score threshold is determined as the second dynamic adjustment scaling factor; When the third speech quality perception evaluation score is not greater than the speech quality perception evaluation score threshold and the current iteration number has not reached the maximum iteration number, calculate the speech quality perception evaluation score difference between the first speech quality perception evaluation score and the second speech quality perception evaluation score. Based on the difference in the perceived speech quality assessment score and the second dynamic adjustment scaling factor, the first initial weight and the second initial weight are adjusted to generate weight coefficients for the next iteration.
8. A voice signal output device, characterized in that, include: The raw speech signal acquisition module is used to acquire the raw speech signal; The first enhanced speech power spectrum generation module is used to perform dictionary-based sparsity processing on the original speech signal to generate the first enhanced speech power spectrum. The second enhanced speech power spectrum generation module is used to perform sparsification processing based on principal component analysis on the original speech signal to generate a second enhanced speech power spectrum composed of row components and column components. The target weight determination module is used to determine the initial weights of the first enhanced speech power spectrum and the second enhanced speech power spectrum, and to perform iterative calculations on the first enhanced speech power spectrum and the second enhanced speech power spectrum based on the initial weights to determine the target weights; The speech enhancement signal output module is used to output the final speech enhancement signal based on the target weight.
9. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the method as described in claims 1-7.
10. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the method as described in claims 1-7.