Karaoke audio processing method based on human voice separation and restoration

Through short-time Fourier transform and non-negative matrix decomposition, an adaptive filter and phase recovery algorithm are designed to solve the problem of poor noise and audio synchronization in K song applications, and high-quality audio processing effects are achieved.

CN120510863AActive Publication Date: 2025-08-19BOSHILIAN (SHENZHEN) INFORMATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510818630.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-08-19
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

There are problems in existing karaoke applications such as background noise, audio signal distortion, and poor synchronism of vocals and accompaniment. The existing technology is difficult to effectively solve the overlapping frequency band processing between noise and audio, especially when multiple people karaokes affect the user experience.

Method used

The time frequency matrix is obtained through short-time Fourier transform, the noise and effective audio basis matrix are identified by non-negative matrix decomposition, overlap coefficients are calculated, adaptive filters are designed for real-time filtering, combined with user feedback to optimize the denoising effect, and improve audio quality through phase recovery and delay compensation.

Benefits of technology

Accurately identify and remove noise, restore lost phase information, improve the naturalness and synchronization of audio signals, provide a personalized karaoke experience, and significantly improve audio quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510863A_ABST
    Figure CN120510863A_ABST
Patent Text Reader

Abstract

The invention discloses a karaoke audio processing method based on human voice separation and restoration, and the method comprises the steps: obtaining a time-frequency matrix of a karaoke audio signal, and judging whether a noise component exists in the karaoke audio signal or not; identifying a noise basis matrix and an effective audio basis matrix, calculating overlapping coefficients of noise and effective audio on different frequency bands, and determining a cross frequency band; identifying transient noise and steady-state noise according to the noise basis matrix and the corresponding activation matrix; designing a filter, and filtering the noise of each frame in real time; judging lost phase information, interpolating a missing region phase by using a linear difference value, and compensating time delay through time-frequency matrix translation; and judging a de-noising effect and an audio repairing effect, and formulating and executing a karaoke audio signal processing improvement scheme. The problems of noise interference, audio distortion, phase loss, delay compensation and the like in the current karaoke application can be effectively solved, the audio quality is remarkably improved, and better and more personalized karaoke experience is provided for users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of new generation information technology, and in particular to a karaoke audio processing method based on voice separation and restoration. Background Art

[0002] With the popularity of karaoke entertainment applications, users' demands for audio quality are increasing, especially in terms of sound quality, clarity, audio delay, and synchronization between vocals and accompaniment. However, most current karaoke applications still face problems such as background noise, audio signal distortion, and poor synchronization between vocals and accompaniment. These problems prevent users from enjoying ideal audio effects during karaoke. Traditional audio denoising technology is usually based on simple frequency domain analysis and filtering methods, such as filtering the audio signal by setting a fixed energy threshold. This method is prone to audio distortion in practical applications, especially when the frequency bands of vocals and accompaniment overlap strongly, making denoising effectiveness difficult to guarantee. In the audio denoising process, the overlapping frequency bands of noise and valid audio, especially transient noise and steady-state noise, remain a difficult problem. In addition, existing audio restoration technologies often focus on single issues, such as phase recovery or delay compensation, but few comprehensive solutions can simultaneously address noise suppression, audio restoration, and synchronization between vocals and accompaniment. Phase information loss and delay compensation issues in existing technologies often result in a lack of naturalness in the restored audio signal. This is especially true during group karaoke sessions, where the asynchrony between the accompaniment and vocals can severely impact the user experience. To address these issues, existing solutions for audio denoising, restoration, and enhancement remain incomplete and unable to meet the high-quality audio demands of karaoke users. Therefore, to improve audio signal quality, a multi-dimensional karaoke audio processing method is urgently needed. Summary of the Invention

[0003] The present invention addresses the problems existing in the above-mentioned prior art and provides a karaoke audio processing method based on voice separation and restoration, which mainly includes:

[0004] The karaoke audio signal is acquired through the audio acquisition module, and the short-time Fourier transform algorithm is used to convert the karaoke audio signal from the time domain to the frequency domain to obtain the time-frequency matrix of the karaoke audio signal, and determine whether there is noise component in the karaoke audio signal;

[0005] Based on the time-frequency matrix of the karaoke audio signal, non-negative matrix decomposition is used to obtain the basis matrix and activation matrix, identify the noise basis matrix and the effective audio basis matrix, calculate the overlap coefficient of noise and effective audio in different frequency bands, and determine the cross-band;

[0006] According to the noise basis matrix and the corresponding activation matrix, the energy value, zero-crossing rate, spectrum flatness and statistical characteristics of each frame of the noise signal are determined, and transient noise and steady-state noise are identified;

[0007] Based on the time domain envelope characteristics, spectral flatness, statistical characteristics, basis matrix, cross-band and overlap coefficient of the noise signal, an adaptive filtering algorithm is used to design a filter. The noise of each frame is filtered in real time to obtain the denoised karaoke audio signal.

[0008] By comparing the amplitude spectrum and phase spectrum of the denoised karaoke audio signal, the lost phase information is determined, the envelope is calculated based on the amplitude spectrum, the phase of the missing area is interpolated using linear interpolation, and the time delay is compensated by time-frequency matrix translation;

[0009] Based on user feedback data, determine the denoising and audio restoration effects of karaoke audio signals, and formulate and implement a karaoke audio signal processing improvement plan.

[0010] Furthermore, the karaoke audio signal is obtained through the audio acquisition module, the karaoke audio signal is converted from the time domain to the frequency domain using a short-time Fourier transform algorithm, a time-frequency matrix of the karaoke audio signal is obtained, and whether there is a noise component in the karaoke audio signal is determined, including:

[0011] The karaoke audio signal is obtained through the audio acquisition module, and the audio signal is divided into several overlapping short-time frame sequences based on the preset frame length and preset frame shift; a Hamming window is applied to each frame, and the short-time Fourier transform algorithm is used to convert the karaoke audio signal from the time domain to the frequency domain, and the spectrum of each frame is calculated to obtain the time-frequency matrix of the karaoke audio signal, including the amplitude spectrum and phase spectrum; based on the amplitude spectrum of the karaoke audio signal, the power spectrum and spectrum flatness of each frame are calculated to determine the energy distribution of each frequency band; based on the energy distribution and spectrum flatness of each frequency band, the support vector machine algorithm is used for model training to construct a noise component discrimination model to determine whether there is noise component in the karaoke audio signal.

[0012] Furthermore, the method uses non-negative matrix decomposition to obtain a basis matrix and an activation matrix based on the time-frequency matrix of the karaoke audio signal, identifies the noise basis matrix and the effective audio basis matrix, calculates the overlap coefficient of noise and effective audio in different frequency bands, and determines the cross-band, including:

[0013] According to the time-frequency matrix of the karaoke audio signal, the amplitude spectrum is decomposed by non-negative matrix decomposition, the number of basis functions and the number of iterations are set, and the KL divergence is used as the loss function to decompose the amplitude spectrum of the karaoke audio signal. The basis matrix and the activation matrix are alternately optimized through the multiplication update rule, wherein the basis matrix represents the basic frequency components of the audio signal, and the activation matrix represents the activation degree of the basic frequency components in each frame; according to the column vector of the basis matrix, the K-maens clustering algorithm is used for clustering analysis to identify the noise basis matrix and the effective audio basis matrix, and the effective audio includes human voice and accompaniment; according to the activation matrix corresponding to the noise basis matrix and the activation matrix corresponding to the effective audio basis matrix, the cosine similarity calculation method is used to calculate the similarity between the activation value of the noise part and the activation value of the effective audio part in each frequency band, and determine the overlapping coefficient of the noise and effective audio in different frequency bands; if the overlapping coefficient is greater than the preset coefficient threshold, the frequency band is judged to be a cross-band.

[0014] Furthermore, the method of determining the energy value, zero-crossing rate, spectrum flatness, and statistical characteristics of each frame of the noise signal based on the noise basis matrix and the corresponding activation matrix, and identifying transient noise and steady-state noise, includes:

[0015] According to the noise basis matrix and the corresponding activation matrix, according to the short-time energy calculation formula Obtain the energy value of each frame of the noise signal and determine the time domain envelope characteristics of the noise signal, where E(t) is the audio signal energy at time t, X(t,f) is the amplitude spectrum value at time t and frequency f, F is the frequency range, and the time domain envelope characteristics are the energy fluctuation amplitude of the noise signal; based on the noise basis matrix, calculate the zero-crossing rate and spectrum flatness of each frame of the noise signal, where the zero-crossing rate is the number of times the signal passes through the zero point per unit time; calculate the energy value difference of each frame of the noise signal, and if there is a short-time frame sequence with an energy value difference greater than a preset difference threshold or a zero-crossing rate greater than a preset zero-crossing rate threshold, mark the noise of the frame as transient noise, otherwise mark it as steady-state noise; by comparing the energy value of each frame with the energy values of the previous and next frames, judge the duration and intensity of the transient noise; use statistical methods to calculate the mean, variance, kurtosis, and skewness of the noise signal energy to determine the statistical characteristics of the noise signal.

[0016] Furthermore, the method designs a filter using an adaptive filtering algorithm based on the time domain envelope characteristics, spectrum flatness, statistical characteristics, base matrix, cross-band and overlap coefficient of the noise signal, performs real-time filtering on the noise of each frame, and obtains a denoised karaoke audio signal, including:

[0017] According to the time domain envelope characteristics, spectral flatness, statistical characteristics, basis matrix, cross-frequency band and overlapping coefficient of the noise signal, an adaptive filtering algorithm is used to design the filter, and the initial parameters of the filter are determined to perform real-time filtering on the noise of each frame. The initial parameters of the filter include but are not limited to the length, bandwidth and gain coefficient of the filter; if a short-time frame of transient noise is identified in the noise signal, an energy adjustment factor is calculated according to the energy of the noise signal and the effective audio of the frame, and the denoising strength of each frame of the audio signal is adjusted based on the energy adjustment factor; if the energy adjustment factor is greater than the preset factor threshold, the gain coefficient of the filter is increased or the bandwidth is narrowed based on the energy adjustment factor; the parameters of the filter are dynamically adjusted by the time domain envelope characteristics, spectral flatness, statistical characteristics, basis matrix, cross-frequency band and overlapping coefficient of the noise signal in the karaoke audio signal obtained in real time; the noise component in the original karaoke audio signal is removed by the filter to obtain the denoised karaoke audio signal, and the time-frequency matrix of the denoised karaoke audio signal is generated.

[0018] The method further includes: if a short-term frame of transient noise is identified, calculating an energy adjustment factor based on the energy of the noise signal and the effective audio of the frame, and adjusting the denoising strength of each frame of the audio signal based on the energy adjustment factor, specifically including:

[0019] If a short-term frame of transient noise is identified, the energy difference ratio between the noise source and the effective audio source is calculated based on the energy of the noise signal and the effective audio in the frame. Among them, E n (t) is the energy of the noise signal, E a (t) is the energy of the effective audio; according to the energy difference ratio between the noise source and the effective audio source, the energy adjustment factor formula is used Calculate the energy adjustment factor α(t ) , where β and γ are adjustment parameters used to control the sensitivity of noise and audio balance, which are obtained by fitting historical data; according to the energy adjustment factor, the denoising strength of each frame of audio signal is dynamically adjusted.

[0020] Furthermore, the method comprises: determining the lost phase information by comparing the amplitude spectrum and phase spectrum of the denoised karaoke audio signal; calculating the envelope in combination with the amplitude spectrum; interpolating the phase of the missing region using linear difference; and compensating for the time delay by time-frequency matrix shifting.

[0021] By comparing the amplitude spectrum and phase spectrum of the denoised karaoke audio signal, it is determined whether there is any lost phase information; if there is any lost phase information in the denoised karaoke audio signal, the frequency range where the lost phase is located is determined by checking the missing area in the phase spectrum; at each frequency point, the instantaneous phase information of the signal is obtained by combining the logarithm of the amplitude spectrum with the phase spectrum, and the envelope of the signal is calculated based on the amplitude spectrum; based on the instantaneous phase and envelope information of the known adjacent frequency points, the linear difference method is used to interpolate the missing phase area to restore the lost phase information; based on the restored lost Phase information is combined with the amplitude spectrum of the denoised karaoke audio signal to reconstruct the time-frequency matrix of the repaired karaoke audio signal; according to the time-frequency matrix of the repaired karaoke audio signal, the time axis of the audio signal is matched using a dynamic time warping algorithm to calculate the time delay between the input and output of the karaoke audio signal; based on the delay calculation result, the time column in the time-frequency matrix of the repaired karaoke audio signal is shifted to obtain the time-frequency matrix of the karaoke audio signal after delay compensation; the inverse short-time Fourier transform is used to convert the time-frequency matrix back to the time domain signal to obtain the output signal of the karaoke audio.

[0022] Furthermore, the denoising effect and audio restoration effect of the karaoke audio signal are determined based on user feedback data, and a karaoke audio signal processing improvement plan is formulated and implemented, including:

[0023] By setting up a feedback mechanism, we obtain user feedback data on sound quality, audio delay, vocal clarity, and accompaniment synchronization after their karaoke experience, and judge the denoising effect and audio repair effect of the karaoke audio signal; if the denoising effect or audio repair effect of the karaoke audio signal does not meet the preset effect standards, we will formulate and implement a karaoke audio signal processing improvement plan based on user feedback data until the denoising effect and audio repair effect of the karaoke audio signal meet the preset effect standards. The processing improvement plan includes but is not limited to adjusting filter parameters, optimizing phase interpolation methods to fill in lost phase information, and readjusting delay compensation through dynamic time warping.

[0024] The technical solution provided by the embodiment of the present invention may have the following beneficial effects:

[0025] The present invention provides a karaoke audio processing method based on voice separation and restoration. This method accurately distinguishes the frequency bands of noise and valid audio components, particularly in complex situations where the noise and audio signal frequencies overlap. This method is particularly capable of addressing the complex problem of overlapping noise and audio components. By calculating the overlap coefficient to determine the crossover frequency band, the method accurately removes noise while preserving the core components of the audio signal. The present invention distinguishes between transient and steady-state noise and provides the ability to dynamically adjust the denoising intensity, enabling timely optimization of the denoising effect based on noise changes. By comparing the amplitude spectrum and phase spectrum, the present invention restores lost phase information, ensuring the naturalness and synchronization of the audio signal. By integrating user feedback for real-time optimization and dynamically adjusting filter parameters and restoration strategies, the audio processing effect can be improved in real time based on user needs, providing a personalized audio restoration solution. The present invention provides a karaoke audio processing method based on voice separation and restoration. This method effectively addresses the problems of noise interference, audio distortion, phase loss, and delay compensation that exist in current karaoke applications, significantly improving audio quality and providing users with a higher-quality, personalized karaoke experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 This is a flow chart of a karaoke audio processing method based on voice separation and restoration of the present invention;

[0027] Figure 2 Schematic diagram of a karaoke audio processing method based on voice separation and restoration of the present invention;

[0028] Figure 3 This is another schematic diagram of a karaoke audio processing method based on voice separation and restoration of the present invention. DETAILED DESCRIPTION

[0029] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0030] like Figure 1-3 In this embodiment, a karaoke audio processing method based on voice separation and restoration may specifically include:

[0031] Step S101: obtain a karaoke audio signal through an audio acquisition module, convert the karaoke audio signal from the time domain to the frequency domain using a short-time Fourier transform algorithm, obtain a time-frequency matrix of the karaoke audio signal, and determine whether there is a noise component in the karaoke audio signal.

[0032] The karaoke audio signal is acquired through the audio acquisition module, and the audio signal is divided into several overlapping short-time frame sequences based on the preset frame length and preset frame shift. A Hamming window is applied to each frame, and the short-time Fourier transform algorithm is used to convert the karaoke audio signal from the time domain to the frequency domain. The spectrum of each frame is calculated to obtain the time-frequency matrix of the karaoke audio signal, including the amplitude spectrum and phase spectrum. Based on the amplitude spectrum of the karaoke audio signal, the power spectrum and spectrum flatness of each frame are calculated to determine the energy distribution of each frequency band. Based on the energy distribution and spectrum flatness of each frequency band, the support vector machine algorithm is used for model training to construct a noise component discrimination model to determine whether there is a noise component in the karaoke audio signal.

[0033] For example, a karaoke audio signal is being processed. This signal contains human voice and background noise. The audio signal is obtained through the audio acquisition module with a sampling rate of 16kHz and a duration of 5 seconds. This means that the collected audio signal has a total of 80,000 sampling points. According to the preset frame length and frame shift, the entire audio signal is divided into multiple overlapping short-time frame sequences. If the frame length is set to 512 sampling points and the frame shift is set to 256 sampling points. In this way, each frame contains 512 sampling points, and the frames overlap by 256 sampling points, so about 320 frames can be obtained. The signal of each frame will be processed by the Hamming window function to make it transition smoothly at the boundary, thereby reducing spectral leakage. The time domain signal of each frame is converted into a frequency domain signal using a short-time Fourier transform. The spectrum of a certain frame may show that the low-frequency part is mainly composed of frequency components of 100Hz, 200Hz and 300Hz, while the high-frequency part may be mainly composed of frequency components of 1kHz and 2kHz. After obtaining the spectrum of each frame, the power spectrum of each frame is calculated by squaring the amplitude of the spectrum. If for a certain frame signal, the calculated amplitude spectrum shows that the frequency component is 50Hz with an amplitude of 2.0, 200Hz with an amplitude of 1.5, 1kHz with an amplitude of 3.0, and 2kHz with an amplitude of 0.5, then each frequency component in the power spectrum is the square of these amplitude values, namely 4.0, 2.25, 9.0, and 0.25 respectively. These power spectrum values represent the energy distribution of each frequency component, that is, the energy contribution of the frame at different frequencies. If the spectrum flatness of 50Hz, 200Hz, 1kHz, and 2kHz is calculated to be 0.4, 0.2, 0.5, and 0.8 respectively, and the historical data of the energy distribution and spectrum flatness of each frequency band is obtained, the support vector machine algorithm is used to train the model based on the energy distribution and spectrum flatness of each frequency band to construct a noise component discrimination model. Based on the trained noise component discrimination model, when a new audio signal is input, the noise component discrimination model will determine in real time whether there is a noise component in the audio signal.

[0034] Step S102: Based on the time-frequency matrix of the karaoke audio signal, non-negative matrix decomposition is used to obtain the basis matrix and activation matrix, identify the noise basis matrix and the effective audio basis matrix, calculate the overlapping coefficients of noise and effective audio in different frequency bands, and determine the cross-band.

[0035] According to the time-frequency matrix of the karaoke audio signal, the amplitude spectrum is decomposed using a non-negative matrix. The number of basis functions and the number of iterations are set. The KL divergence is used as the loss function to decompose the amplitude spectrum of the karaoke audio signal. The basis matrix and the activation matrix are alternately optimized through the multiplication update rule, where the basis matrix represents the basic frequency components of the audio signal and the activation matrix represents the activation degree of the basic frequency components in each frame. According to the column vectors of the basis matrix, the K-maens clustering algorithm is used for cluster analysis to identify the noise basis matrix and the effective audio basis matrix. The effective audio includes human voice and accompaniment. According to the activation matrix corresponding to the noise basis matrix and the activation matrix corresponding to the effective audio basis matrix, the cosine similarity calculation method is used to calculate the similarity of the activation value of the noise part and the activation value of the effective audio part in each frequency band, and determine the overlap coefficient of the noise and effective audio in different frequency bands. If the overlap coefficient is greater than the preset coefficient threshold, the frequency band is judged to be a cross-band.

[0036] For example, there is a karaoke audio signal with a duration of 5 seconds and a sampling rate of 16kHz, that is, a total of 80,000 sampling points. The audio signal is divided into multiple overlapping short-time frame sequences, each frame length is 512 sampling points, and the frame shift is 256 sampling points, so that about 320 frames can be obtained. The audio signal is converted from the time domain to the frequency domain by short-time Fourier transform to obtain a time-frequency matrix, which contains the amplitude spectrum and phase spectrum of the audio signal. Using non-negative matrix decomposition, the number of basis functions is set to 10, that is, the frequency components of the audio signal will be decomposed into 10 basis functions, and the number of iterations is set to 1000 times. Through non-negative matrix decomposition, two matrices are obtained, where the basis matrix represents the basic frequency components of the audio signal, each column represents a fundamental frequency component, and the activation matrix represents the activation degree of each fundamental frequency component in each frame. For example, the first column in the basis matrix may represent a low-frequency component, the second column represents a medium-frequency component, and the third column represents a high-frequency component, while the activation matrix represents the intensity of each frequency component in each frame. After non-negative matrix decomposition, if the frequency components of the basis matrix are as follows: the first column has a frequency of 100 Hz and a fundamental frequency component with an amplitude of 0.8; the second column has a frequency of 500 Hz and a fundamental frequency component with an amplitude of 0.5; the third column has a frequency of 1000 Hz and a fundamental frequency component with an amplitude of 0.3; and the fourth column has a frequency of 1500 Hz and a fundamental frequency component with an amplitude of 0.2, then a K-means clustering algorithm is used to perform cluster analysis based on the column vectors of the basis matrix. The purpose is to divide these fundamental frequency components into a noise basis matrix and a valid audio basis matrix. Using K = 2 as the number of clusters, the K-means algorithm will divide these fundamental frequency components into two categories: one category is valid audio components with fundamental frequency components at frequencies of 100 Hz, 500 Hz, and 10,000 Hz; the other category is noise audio components with a fundamental frequency component at 1500 Hz. In different columns of the activation matrix corresponding to the noise basis matrix and the activation matrix corresponding to the effective audio basis matrix, the activation value of each frequency component will change over time. In order to measure the overlap of the noise and effective audio basis matrices in a specific frequency band, the activation value in each frequency band is selected, that is, the row of the corresponding activation matrix. If the activation value of the noise basis matrix in the 1000Hz frequency band in a certain time frame is [0.6, 0.7, 0.4, 0.3], and the activation value of the effective audio basis matrix is [0.7, 0.6, 0.3, 0.4], and their cosine similarity is calculated to be 0.8, then the overlap coefficient between the noise and effective audio in the 1000Hz frequency band is 0.8, which is greater than the preset coefficient threshold of 0.75. This means that there is a strong overlap between the noise and effective audio components in this frequency band, and therefore this frequency band is a crossover band.

[0037] Step S103 : determining the energy value, zero-crossing rate, spectrum flatness and statistical characteristics of each frame of the noise signal according to the noise basis matrix and the corresponding activation matrix, and identifying transient noise and steady-state noise.

[0038] According to the noise basis matrix and the corresponding activation matrix, according to the short-time energy calculation formula The energy value of each frame of the noise signal is obtained, and the time-domain envelope characteristics of the noise signal are determined. Here, E(t) is the audio signal energy at time t, X(t,f) is the amplitude spectrum at time t and frequency f, and F is the frequency range. The time-domain envelope characteristics represent the amplitude of the noise signal's energy fluctuations. Based on the noise basis matrix, the zero-crossing rate and spectral flatness of each frame of the noise signal are calculated. The zero-crossing rate is the number of times the signal passes through zero per unit time. The energy difference of each frame of the noise signal is calculated. If there is a short-term frame sequence with an energy difference greater than a preset difference threshold or a zero-crossing rate greater than a preset zero-crossing rate threshold, the noise in that frame is marked as transient; otherwise, it is marked as steady-state noise. The energy value of each frame is compared with the energy values of the previous and next frames to determine the duration and intensity of the transient noise. Statistical methods are used to calculate the mean, variance, kurtosis, and skewness of the noise signal energy to determine the statistical characteristics of the noise signal.

[0039] For example, there is a karaoke audio signal with a sampling rate of 16kHz and a duration of 3 seconds, which means that the audio signal has a total of 48,000 sampling points. The audio signal is divided into multiple overlapping frames through short-time Fourier transform. The frame length of each frame is 512 sampling points and the frame shift is 256 sampling points. Therefore, about 200 frames will be obtained in this audio. The amplitude spectrum |X(t,f)| of each frame is calculated through the time-frequency matrix, and the short-time energy calculation formula is used Calculate the energy of the frame, where E(t) is the audio signal energy at time t, X(t,f) is the amplitude spectrum at time t and frequency f, F is the frequency range, and the time envelope characteristic is the energy fluctuation amplitude of the noise signal. If the amplitude spectrum of the first frame, that is, at t = 1, is |X(1,50)|=1.5, |X(1,200)|=0.8, |X(1,1000)|=2.0, and |X(1,3000)|=0.5, then the total energy of the first frame is 7.14. The amplitude spectra for the second frame, at t = 2, are |X(2, 50)|=2.0, |X(2, 200)|=1.2, |X(2, 1000)|=1.8, and |X(2, 3000)|=0.6, resulting in a total energy of 8.04. Based on the noise basis matrix, the zero-crossing rate and spectral flatness of the noise signal are calculated. The zero-crossing rate is the number of times the signal passes through zero per unit time. The first frame has a zero-crossing rate of 10 and a spectral flatness of 0.872, while the second frame has a zero-crossing rate of 25 and a spectral flatness of 0.906. The change in the zero-crossing rate can be used to determine whether the signal is transient noise. Since the second frame has a higher zero-crossing rate, indicating a more dramatic change, it can be labeled as transient noise. By calculating the energy difference between two adjacent frames, we found that the energy of the first frame was 7.14, while the energy of the second frame was 8.04. The energy difference between the first and second frames was 0.90. If the preset energy difference threshold is 0.5, this difference is greater than the threshold, so the second frame can be considered transient noise. Finally, statistical methods were used to calculate the mean, variance, kurtosis, and skewness of the noise signal, resulting in a mean of 7.6, a variance of 0.92, a kurtosis of 4.1, and a skewness of 0.3.

[0040] Step S104, based on the time domain envelope characteristics, spectrum flatness, statistical characteristics, base matrix, cross-band and overlap coefficient of the noise signal, an adaptive filtering algorithm is used to design a filter, and the noise of each frame is filtered in real time to obtain a denoised karaoke audio signal.

[0041] According to the time domain envelope characteristics, spectral flatness, statistical characteristics, base matrix, cross-band and overlapping coefficient of the noise signal, an adaptive filtering algorithm is used to design the filter, and the initial parameters of the filter are determined to perform real-time filtering on the noise of each frame. The initial parameters of the filter include but are not limited to the length, bandwidth and gain coefficient of the filter. If a short-time frame of transient noise is identified in the noise signal, the energy adjustment factor is calculated according to the energy of the noise signal and the effective audio of the frame, and the denoising strength of each frame of the audio signal is adjusted based on the energy adjustment factor. If the energy adjustment factor is greater than the preset factor threshold, the gain coefficient of the filter is increased or the bandwidth is reduced based on the energy adjustment factor. The parameters of the filter are dynamically adjusted by the time domain envelope characteristics, spectral flatness, statistical characteristics, base matrix, cross-band and overlapping coefficient of the noise signal in the karaoke audio signal obtained in real time. The noise component in the original karaoke audio signal is removed by the filter to obtain the denoised karaoke audio signal, and the time-frequency matrix of the denoised karaoke audio signal is generated.

[0042] For example, during the processing, the amplitude spectrum and power spectrum of each frame have been calculated, and the time domain envelope characteristics of the noise signal are obtained to be 0.8, the spectrum flatness is 0.3, and the statistical characteristics include a mean of 1.5, a variance of 0.2, a kurtosis of 3.5, and a skewness of 0.1. Certain columns of frequency components in the base matrix include the first column with a frequency of 100Hz and an amplitude of the fundamental frequency component of 0.8, the second column with a frequency of 500Hz and an amplitude of the fundamental frequency component of 0.5, the third column with a frequency of 1000Hz and an amplitude of the fundamental frequency component of 0.3, the fourth column with a frequency of 1500Hz and an amplitude of the fundamental frequency component of 0.2, the cross-band is 1000Hz, and the overlap coefficient is 0.8. Based on the above characteristics, a least mean square adaptive filter is designed using the least mean square algorithm to remove noise signals in real time. The initial parameters of the filter include: the filter length is set to 64 sampling points, which means that the filter processes each frame of the signal based on the historical data of 64 sampling points; the bandwidth is set to 200Hz to adapt to the current noise frequency range; and the gain coefficient is initially set to 1.2, which represents the filter's denoising strength. If a short-term frame of transient noise is identified in the noise signal, an energy adjustment factor is calculated based on the energy of the noise signal and the effective audio in that frame. The denoising strength of the filter is adjusted based on the energy ratio of the noise to the effective audio. If the noise energy of the current frame is 2.0 and the effective audio energy is 6.0, the energy adjustment factor formula can be used to calculate the energy adjustment factor α(t)≈0.474. Since α(t)≈0.474. The filter's gain and bandwidth are dynamically adjusted based on the calculated energy adjustment factor α(t). Since α(t) ≈ 0.474, which is less than the preset factor threshold of 0.7, the filter gain remains at 1.2, eliminating the need for further denoising. The 200Hz bandwidth is suitable for the current noise frequency range, so the filter bandwidth remains unchanged. The filter parameters are dynamically adjusted based on the time-domain envelope characteristics, spectral flatness, statistical properties, basis matrix, crossover frequency bands, and overlap coefficients of the noise signal in the karaoke audio signal acquired in real time. For example, as the audio signal changes, the zero-crossing rate in some frames may increase significantly, leading to an increase in the filter gain. In this case, the filter gain may be increased to 1.5, and the bandwidth may be reduced to 150Hz to more accurately remove transient noise. The adaptive filter filters the noise in each frame of the audio signal in real time. If the noise energy is high in the 10th frame, the filter gain is increased to 1.5, and the bandwidth is reduced. After filtering, the significant background noise is successfully removed, while preserving the effective audio components. Finally, the time-frequency matrix of the denoised audio signal shows that the noise components of the amplitude spectrum and phase spectrum are significantly reduced, while the vocals and accompaniment are retained.

[0043] If a short-term frame of transient noise is identified, an energy adjustment factor is calculated based on the energy of the noise signal and valid audio of the frame, and the denoising strength of each frame of the audio signal is adjusted based on the energy adjustment factor.

[0044] If a short-term frame of transient noise is identified, the energy difference ratio between the noise source and the effective audio source is calculated based on the energy of the noise signal and the effective audio in the frame. Among them, E n (t) is the energy of the noise signal, E a (t) is the energy of the effective audio. Based on the energy difference ratio between the noise source and the effective audio source, the energy adjustment factor formula is used. Calculate the energy adjustment factor α(t), where β and γ are adjustment parameters used to control the sensitivity of the noise-audio balance, obtained by fitting historical data. Based on the energy adjustment factor, dynamically adjust the denoising strength of each audio frame.

[0045] For example, there is transient noise in the 50th frame of a karaoke audio signal, the noise energy is 3.2, and the effective audio energy is 8.0. By calculating the ratio of the two, the energy difference ratio is obtained. =0.4, based on the energy difference ratio, use the energy adjustment factor formula Calculate the energy adjustment factor α(t), where β and γ are adjustment parameters used to control the sensitivity of the noise and audio balance, which are obtained by fitting historical data. If β = 2 and γ = 1.5, then the energy adjustment factor α(50) ≈ 0.418. According to the energy adjustment factor α(50) ≈ 0.418, it can be concluded that the denoising strength of the frame should be moderate. At this time, the denoising strength of the filter can be adjusted according to this factor. Because α(50) is small, it means that the noise is relatively weak, so the denoising strength does not need to be very strong. The gain of the filter can be kept at around 1.2, and the bandwidth can also be kept moderate. In other frames, according to the changes in noise energy and effective audio energy, the energy difference ratio R(t) will change, resulting in a change in the energy adjustment factor. If the noise energy of a frame is relatively large, resulting in a large R(t), such as R(t) = 1.5, then α(t) will increase, the gain coefficient of the filter will increase accordingly, and the bandwidth may also be reduced to strengthen the denoising process and ensure that the noise component is effectively removed.

[0046] Step S105: Determine the lost phase information by comparing the amplitude spectrum and phase spectrum of the denoised karaoke audio signal, calculate the envelope based on the amplitude spectrum, use linear interpolation to interpolate the phase of the missing area, and compensate for the time delay through time-frequency matrix shifting.

[0047] By comparing the amplitude spectrum and phase spectrum of the denoised karaoke audio signal, it is determined whether any phase information is missing. If the denoised karaoke audio signal has lost phase information, the frequency range of the missing phase is determined by examining the missing region in the phase spectrum. At each frequency point, the instantaneous phase information of the signal is obtained by combining the logarithm of the amplitude spectrum with the phase spectrum, and the signal envelope is calculated based on the amplitude spectrum. Based on the instantaneous phase and envelope information of known adjacent frequency points, the missing phase region is interpolated using a linear interpolation method to restore the lost phase information. Based on the recovered lost phase information and the amplitude spectrum of the denoised karaoke audio signal, the time-frequency matrix of the repaired karaoke audio signal is reconstructed. Based on the repaired time-frequency matrix of the karaoke audio signal, the time axis of the audio signal is aligned using a dynamic time warping algorithm to calculate the time delay between the input and output of the karaoke audio signal. Based on the delay calculation results, the time columns in the time-frequency matrix of the repaired karaoke audio signal are shifted to obtain the time-frequency matrix of the delay-compensated karaoke audio signal. The inverse short-time Fourier transform is used to convert the time-frequency matrix back to the time domain signal to obtain the output signal of the karaoke audio.

[0048] For example, in the denoised audio signal, it is identified that some frequency components of the phase spectrum are missing. This may be because the phase information is destroyed during the denoising process, or the noise removal algorithm loses the phase information of some frequency points, such as the phase spectrum of the 1000Hz and 1500Hz frequency bands. Therefore, this missing phase information needs to be repaired. By examining the denoised phase spectrum, it is determined that the frequency components between 1000Hz and 1500Hz are missing in the phase spectrum. In order to restore the phase information in these frequency ranges, it is necessary to combine the amplitude spectrum and phase spectrum to calculate the instantaneous phase and perform interpolation. The logarithm of the magnitude spectrum is combined with the phase spectrum to obtain the instantaneous phase at each frequency point. For example, for the magnitude spectra at 1000Hz and 1500Hz, if their amplitude values are 2.0 and 1.8, respectively, their logarithmic values are log(2.0)≈0.3010 and log(1.8)≈0.2553. Calculating the instantaneous phase at these frequency points yields a value of 2.5 rad for 1000Hz and 1.8 rad for 1500Hz. Calculating the signal envelope based on the magnitude spectrum yields envelope values of 1.2 and 1.0 for the 1000Hz and 1500Hz frequency bands, respectively. Since the phase information between 1000Hz and 1500Hz is missing, the phase values between them are obtained through linear interpolation based on the instantaneous phase of 2.5rad at 1000Hz and 1.8rad at 1500Hz. For the 1200Hz band between 1000Hz and 1500Hz, the phase value of the 1200Hz band is obtained as 2.22rad using a linear interpolation formula. The phase information of the other missing frequency bands is restored in a similar manner. After recovering the lost phase information, the repaired phase spectrum is combined with the amplitude spectrum to reconstruct the repaired time-frequency matrix. The dynamic time warping algorithm is used to calculate the time delay between the denoised audio signal and the original audio signal, and the delay is calculated to be 20ms. Based on the calculated delay, the time column of the repaired time-frequency matrix is shifted backward by 20 milliseconds to eliminate the delay caused by the denoising process. By inverse short-time Fourier transform, the delay-compensated time-frequency matrix is converted back to a time domain signal, and a karaoke audio output signal with noise removed and phase information restored can be obtained.

[0049] Step S106: judging the denoising effect and audio restoration effect of the karaoke audio signal based on the user feedback data, and formulating and executing a karaoke audio signal processing improvement plan.

[0050] By setting up a feedback mechanism, we can obtain user feedback data on sound quality, audio delay, vocal clarity, and accompaniment synchronization after the user's karaoke experience, and judge the denoising effect and audio repair effect of the karaoke audio signal. If the denoising effect or audio repair effect of the karaoke audio signal does not meet the preset effect standards, we will formulate and implement a karaoke audio signal processing improvement plan based on user feedback data until the denoising effect and audio repair effect of the karaoke audio signal meet the preset effect standards. The processing improvement plan includes but is not limited to adjusting filter parameters, optimizing phase interpolation methods to fill in lost phase information, and readjusting delay compensation through dynamic time warping.

[0051] For example, by setting up a feedback mechanism, user feedback data after the karaoke experience is collected to evaluate the denoising effect and audio restoration effect. The feedback data of the user after the experience is obtained, including sound quality of 3 stars, audio delay of 50ms, vocal clarity of 4 stars, and synchronization of accompaniment and vocals of 2 stars. Based on these feedback data, it can be preliminarily judged that although the denoising effect of the audio signal has improved, the audio delay and accompaniment synchronization still do not meet the preset standard of 3 stars and need further improvement. Based on user feedback, a processing improvement plan for karaoke audio signals is formulated. Among them, for the denoising effect, users reflect that the noise still exists, so the filter parameters need to be adjusted to further improve the denoising effect. The bandwidth of the currently used filter may be too wide, resulting in some effective audio components, such as the high-frequency part of the human voice, being mistakenly removed. Based on user feedback, the filter bandwidth is narrowed to 150Hz to focus more on removing low-frequency noise while maintaining high-frequency sound quality. In addition, the gain coefficient needs to be dynamically adjusted according to the energy adjustment factor. When the noise is more significant, the gain coefficient is increased to enhance the denoising strength. When noise is weak, the gain factor is reduced to ensure uncompromised audio quality. Regarding audio restoration, user feedback indicates that vocals occasionally blur in the high-pitched portion, indicating a loss of phase information, particularly during the noise removal process. To recover this lost phase information, the phase interpolation method has been optimized. By analyzing the denoised amplitude and phase spectra, missing phase information between 1000Hz and 1500Hz is identified. Linear interpolation is used to recover the lost phase, and combined with the logarithm of the amplitude spectrum, the phase in the missing region is restored to an appropriate range, ensuring clarity in the high-pitched portion. Due to user feedback regarding long audio latency, a dynamic time warping algorithm is used to re-optimize the audio signal's delay compensation. The calculated delay between the input and output of the audio signal is 50ms, indicating a significant time difference between the denoised and original audio signals. Furthermore, a time column shift is performed on the time-frequency matrix of the denoised audio signal, shifting the signal's time axis forward by 50ms to ensure synchronization between the audio and the vocals. After implementing these improvements, the audio signal was reprocessed and user feedback was collected again. This new feedback confirmed that both the denoising and audio restoration effects met the pre-set performance standards. By adjusting filter parameters, optimizing phase interpolation methods, and adjusting delay compensation using dynamic time warping, we successfully addressed noise, phase loss, audio delay, and accompaniment synchronization issues in the audio signal, significantly improving the quality of karaoke audio signals.

[0052] The above description is merely a preferred embodiment of the present application and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this application is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also encompasses other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the concept of this application. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this application.

Claims

1. A karaoke audio processing method based on voice separation and restoration, characterized in that: The method comprises: The karaoke audio signal is acquired through the audio acquisition module, and the short-time Fourier transform algorithm is used to convert the karaoke audio signal from the time domain to the frequency domain to obtain the time-frequency matrix of the karaoke audio signal, and determine whether there is noise component in the karaoke audio signal; Based on the time-frequency matrix of the karaoke audio signal, non-negative matrix decomposition is used to obtain the basis matrix and activation matrix, identify the noise basis matrix and the effective audio basis matrix, calculate the overlap coefficient of noise and effective audio in different frequency bands, and determine the cross-band; According to the noise basis matrix and the corresponding activation matrix, the energy value, zero-crossing rate, spectrum flatness and statistical characteristics of each frame of the noise signal are determined, and transient noise and steady-state noise are identified; Based on the time domain envelope characteristics, spectral flatness, statistical characteristics, basis matrix, cross-band and overlap coefficient of the noise signal, an adaptive filtering algorithm is used to design a filter. The noise of each frame is filtered in real time to obtain the denoised karaoke audio signal. By comparing the amplitude spectrum and phase spectrum of the denoised karaoke audio signal, the lost phase information is determined, the envelope is calculated based on the amplitude spectrum, the phase of the missing area is interpolated using linear interpolation, and the time delay is compensated by time-frequency matrix translation; Based on user feedback data, determine the denoising and audio restoration effects of karaoke audio signals, and formulate and implement a karaoke audio signal processing improvement plan.

2. The method according to claim 1, wherein The method includes obtaining a karaoke audio signal through an audio acquisition module, converting the karaoke audio signal from a time domain to a frequency domain using a short-time Fourier transform algorithm, obtaining a time-frequency matrix of the karaoke audio signal, and determining whether there is a noise component in the karaoke audio signal. The method includes: The karaoke audio signal is obtained through the audio acquisition module, and the audio signal is divided into several overlapping short-time frame sequences based on the preset frame length and preset frame shift; a Hamming window is applied to each frame, and the short-time Fourier transform algorithm is used to convert the karaoke audio signal from the time domain to the frequency domain, and the spectrum of each frame is calculated to obtain the time-frequency matrix of the karaoke audio signal, including the amplitude spectrum and phase spectrum; based on the amplitude spectrum of the karaoke audio signal, the power spectrum and spectrum flatness of each frame are calculated to determine the energy distribution of each frequency band; based on the energy distribution and spectrum flatness of each frequency band, the support vector machine algorithm is used for model training to construct a noise component discrimination model to determine whether there is noise component in the karaoke audio signal.

3. The method according to claim 1, wherein The method uses non-negative matrix decomposition to obtain a basis matrix and an activation matrix based on the time-frequency matrix of the karaoke audio signal, identifies the noise basis matrix and the effective audio basis matrix, calculates the overlap coefficient of noise and effective audio in different frequency bands, and determines the cross-frequency band, including: According to the time-frequency matrix of the karaoke audio signal, the amplitude spectrum is decomposed by non-negative matrix decomposition, the number of basis functions and the number of iterations are set, and the KL divergence is used as the loss function to decompose the amplitude spectrum of the karaoke audio signal. The basis matrix and the activation matrix are alternately optimized through the multiplication update rule, wherein the basis matrix represents the basic frequency components of the audio signal, and the activation matrix represents the activation degree of the basic frequency components in each frame; according to the column vector of the basis matrix, the K-maens clustering algorithm is used for clustering analysis to identify the noise basis matrix and the effective audio basis matrix, and the effective audio includes human voice and accompaniment; according to the activation matrix corresponding to the noise basis matrix and the activation matrix corresponding to the effective audio basis matrix, the cosine similarity calculation method is used to calculate the similarity between the activation value of the noise part and the activation value of the effective audio part in each frequency band, and determine the overlapping coefficient of the noise and effective audio in different frequency bands; if the overlapping coefficient is greater than the preset coefficient threshold, the frequency band is judged to be a cross-band.

4. The method according to claim 1, wherein The method of determining the energy value, zero-crossing rate, spectrum flatness, and statistical characteristics of each frame of the noise signal based on the noise basis matrix and the corresponding activation matrix, and identifying transient noise and steady-state noise, includes: According to the noise basis matrix and the corresponding activation matrix, according to the short-time energy calculation formula Obtain the energy value of each frame of the noise signal and determine the time domain envelope characteristics of the noise signal, where E(t) is the audio signal energy at time t, X(t,f) is the amplitude spectrum value at time t and frequency f, F is the frequency range, and the time domain envelope characteristics are the energy fluctuation amplitude of the noise signal; based on the noise basis matrix, calculate the zero-crossing rate and spectrum flatness of each frame of the noise signal, where the zero-crossing rate is the number of times the signal passes through the zero point per unit time; calculate the energy value difference of each frame of the noise signal, and if there is a short-time frame sequence with an energy value difference greater than a preset difference threshold or a zero-crossing rate greater than a preset zero-crossing rate threshold, mark the noise of the frame as transient noise, otherwise mark it as steady-state noise; by comparing the energy value of each frame with the energy values of the previous and next frames, judge the duration and intensity of the transient noise; use statistical methods to calculate the mean, variance, kurtosis, and skewness of the noise signal energy to determine the statistical characteristics of the noise signal.

5. The method according to claim 1, wherein The method designs a filter using an adaptive filtering algorithm based on the time domain envelope characteristics, spectrum flatness, statistical characteristics, base matrix, cross-band and overlap coefficient of the noise signal, performs real-time filtering on the noise of each frame, and obtains a denoised karaoke audio signal, including: According to the time domain envelope characteristics, spectral flatness, statistical characteristics, basis matrix, cross-frequency band and overlapping coefficient of the noise signal, an adaptive filtering algorithm is used to design the filter, and the initial parameters of the filter are determined to perform real-time filtering on the noise of each frame. The initial parameters of the filter include but are not limited to the length, bandwidth and gain coefficient of the filter; if a short-time frame of transient noise is identified in the noise signal, an energy adjustment factor is calculated according to the energy of the noise signal and the effective audio of the frame, and the denoising strength of each frame of the audio signal is adjusted based on the energy adjustment factor; if the energy adjustment factor is greater than the preset factor threshold, the gain coefficient of the filter is increased or the bandwidth is narrowed based on the energy adjustment factor; the parameters of the filter are dynamically adjusted by the time domain envelope characteristics, spectral flatness, statistical characteristics, basis matrix, cross-frequency band and overlapping coefficient of the noise signal in the karaoke audio signal obtained in real time; the noise component in the original karaoke audio signal is removed by the filter to obtain the denoised karaoke audio signal, and the time-frequency matrix of the denoised karaoke audio signal is generated.

6. The method according to claim 5, wherein: If a short-time frame of transient noise is identified, an energy adjustment factor is calculated according to the energy of the noise signal and the effective audio of the frame, and the denoising strength of each frame of the audio signal is adjusted based on the energy adjustment factor, including: If a short-term frame of transient noise is identified, the energy difference ratio between the noise source and the effective audio source is calculated based on the energy of the noise signal and the effective audio in the frame. Among them, E n (t) is the energy of the noise signal, E a (t) is the energy of the effective audio; according to the energy difference ratio between the noise source and the effective audio source, the energy adjustment factor formula is used Calculate the energy adjustment factor α(t ) , where β and γ are adjustment parameters used to control the sensitivity of noise and audio balance, which are obtained by fitting historical data; according to the energy adjustment factor, the denoising strength of each frame of audio signal is dynamically adjusted.

7. The method according to claim 1, wherein The method comprises the following steps: comparing the amplitude spectrum and phase spectrum of the denoised karaoke audio signal to determine the lost phase information, calculating the envelope in combination with the amplitude spectrum, interpolating the phase of the missing region using linear difference, and compensating for the time delay by time-frequency matrix shifting. By comparing the amplitude spectrum and phase spectrum of the denoised karaoke audio signal, it is determined whether there is any lost phase information; if there is any lost phase information in the denoised karaoke audio signal, the frequency range where the lost phase is located is determined by checking the missing area in the phase spectrum; at each frequency point, the instantaneous phase information of the signal is obtained by combining the logarithm of the amplitude spectrum with the phase spectrum, and the envelope of the signal is calculated based on the amplitude spectrum; based on the instantaneous phase and envelope information of the known adjacent frequency points, the linear difference method is used to interpolate the missing phase area to restore the lost phase information; based on the restored lost Phase information is combined with the amplitude spectrum of the denoised karaoke audio signal to reconstruct the time-frequency matrix of the repaired karaoke audio signal; according to the time-frequency matrix of the repaired karaoke audio signal, the time axis of the audio signal is matched using a dynamic time warping algorithm to calculate the time delay between the input and output of the karaoke audio signal; based on the delay calculation result, the time column in the time-frequency matrix of the repaired karaoke audio signal is shifted to obtain the time-frequency matrix of the karaoke audio signal after delay compensation; the inverse short-time Fourier transform is used to convert the time-frequency matrix back to the time domain signal to obtain the output signal of the karaoke audio.

8. The method according to claim 1, wherein The method of determining the denoising effect and audio restoration effect of the karaoke audio signal based on user feedback data and formulating and implementing a karaoke audio signal processing improvement plan includes: By setting up a feedback mechanism, we obtain user feedback data on sound quality, audio delay, vocal clarity, and accompaniment synchronization after their karaoke experience, and judge the denoising effect and audio repair effect of the karaoke audio signal; if the denoising effect or audio repair effect of the karaoke audio signal does not meet the preset effect standards, we will formulate and implement a karaoke audio signal processing improvement plan based on user feedback data until the denoising effect and audio repair effect of the karaoke audio signal meet the preset effect standards. The processing improvement plan includes but is not limited to adjusting filter parameters, optimizing phase interpolation methods to fill in lost phase information, and readjusting delay compensation through dynamic time warping.

Citation Information

Patent Citations

  • Speech enhancement method based on non-negative low-rank and sparse matrix decomposition principle

    CN103559888A

  • Method and system for suppressing metronome noise in audio

    CN113823305A

  • Bluetooth receiving end monaural upmixing method and device based on non-negative matrix factorization, and medium

    CN118782053A

  • Sound Processing Apparatus

    US20130010968A1