Method, device, speech transmission system, equipment and medium for estimating comfort noise

By estimating the power spectrum of the voice frame signal at the receiving end, generating a comfortable noise signal to fill the silent frame, solving the problem of insufficient bandwidth utilization in network live broadcast and achieving more efficient voice transmission.

CN115273868BActive Publication Date: 2025-09-02GUANGZHOU FANGGUI INFORMATION TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210894228.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-27
Publication Date
2025-09-02
Estimated Expiration
2042-07-27

AI Technical Summary

Technical Problem

In the prior art, in voice data transmission scenarios such as live broadcasts in networks, the bandwidth utilization rate of voice data transmission is insufficient, especially when processing silent frame signals, the transmission efficiency of comfortable noise signals needs to be improved.

Method used

By obtaining the power spectrum of the speech frame signal, the receiver determines the part of the amplitude value that does not reach the amplitude threshold by using the historical amplitude or the preset initial amplitude value, and determines the part of the amplitude value that reaches the threshold through the smaller amplitude value and the lift coefficient, generating a comfortable noise signal and filling it into the silent frame to maintain speech continuity.

Benefits of technology

Without losing voice continuity at the receiver, the utilization of transmission bandwidth is improved and the voice transmission performance in network live broadcast scenarios is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115273868B_ABST
    Figure CN115273868B_ABST
Patent Text Reader

Abstract

The present application relates to the field of audio and live broadcast technology, and provides a method, apparatus, system, device and medium for estimating comfort noise. The present application improves the transmission bandwidth utilization rate without losing the voice continuity of the receiving end. It includes: the receiving end receives the current voice frame signal of the transmitting end and obtains its first power spectrum, if there is a first part in the first power spectrum that does not reach the amplitude threshold, the current amplitude of the first part in the second power spectrum is determined according to the historical or preset amplitude of the first part in the second power spectrum, if there is a second part in the first power spectrum that reaches the amplitude threshold, the current amplitude of the second part in the second power spectrum is determined according to the smaller amplitude and the lifting coefficient, the smaller amplitude being the smaller of the amplitude of the second part in the first power spectrum and the historical amplitude of the second part in the second power spectrum; the current comfort noise signal is obtained according to the second power spectrum determined by the current amplitude, and the comfort noise signal is filled into the silent frame output when a silent frame is detected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical fields of audio processing and network live broadcasting, and in particular to a method for estimating comfort noise, an apparatus for estimating comfort noise, a voice transmission system for network live broadcasting, an electronic device, and a computer-readable storage medium. Background Art

[0002] In scenarios involving voice data transmission, such as live streaming, current technologies require that the transmitter first perform VAD (Voice Activity Detection) to distinguish between voice frame signals and silence frame signals. The transmitter then performs noise estimation on the silence frame signals to obtain a comfort noise signal. During transmission, empty packets carrying the comfort noise signal are sent to the silence frame signals, from which the receiver decodes the comfort noise signal to fill the silence frame. However, this technology still needs to improve the transmission bandwidth utilization. Summary of the Invention

[0003] Based on this, it is necessary to provide a method for estimating comfort noise, an apparatus for estimating comfort noise, a voice transmission system for live webcasting, an electronic device, and a computer-readable storage medium to address the above technical issues.

[0004] In a first aspect, the present application provides a method for estimating comfort noise, which is applied to a receiving end. The method comprises:

[0005] Upon receiving a current voice frame signal from a transmitting end, obtaining a first power spectrum of the current voice frame signal;

[0006] If there is a first portion in the first power spectrum that does not reach the amplitude threshold, determining the current amplitude of the first portion in the second power spectrum according to a historical amplitude or a preset initial amplitude of the first portion in the second power spectrum whose current amplitude is to be determined;

[0007] If there is a second part in the first power spectrum that reaches the amplitude threshold, determining a current amplitude of the second part in the second power spectrum according to the smaller amplitude and the lift coefficient; wherein the smaller amplitude is the smaller of the amplitude of the second part in the first power spectrum and the historical amplitude of the second part in the second power spectrum;

[0008] obtaining a current comfort noise signal according to a second power spectrum determined according to the current amplitude;

[0009] When a silence frame is detected, the current comfort noise signal is filled into the silence frame output.

[0010] In one embodiment, determining the current amplitude of the first portion in the second power spectrum according to a historical amplitude or a preset initial amplitude of the first portion in the second power spectrum whose current amplitude is to be determined includes:

[0011] When the comfort noise signal is acquired for the first time, the preset initial amplitude of the first part in the second power spectrum is used as the current amplitude of the first part in the second power spectrum; when the comfort noise signal is not acquired for the first time, the current amplitude of the first part in the second power spectrum is determined based on the historical amplitude of the first part in the second power spectrum.

[0012] In one embodiment, determining the current amplitude of the first portion in the second power spectrum based on the historical amplitude of the first portion in the second power spectrum includes:

[0013] The amplitude of the first part in the second power spectrum of the previous comfort noise signal is used as the current amplitude of the first part in the second power spectrum.

[0014] In one embodiment, before determining the current amplitude of the second part of the second power spectrum according to the smaller amplitude and the lift coefficient, the method further includes:

[0015] Determine an amplitude raising stage of the second part in the second power spectrum; and determine a corresponding raising coefficient according to the amplitude raising stage.

[0016] In one embodiment, determining the amplitude rising phase of the second portion in the second power spectrum includes:

[0017] If the number of times the second part in the second power spectrum rises does not reach a preset number of times, the amplitude raising stage is determined to be the first stage; if the number of times the second part in the second power spectrum rises reaches the preset number of times, and the historical amplitude of the second part in the second power spectrum does not reach the amplitude threshold, the amplitude raising stage is determined to be the second stage; if the historical amplitude of the second part in the second power spectrum reaches the amplitude threshold, the amplitude raising stage is determined to be the third stage;

[0018] The determining of the corresponding lifting coefficient according to the amplitude raising stage includes: when the amplitude raising stage is the first stage, using the lifting coefficient with a first value as the corresponding lifting coefficient; when the amplitude raising stage is the second stage, using the lifting coefficient with a second value as the corresponding lifting coefficient; the second value is smaller than the first value; when the amplitude raising stage is the third stage, using the lifting coefficient with a third value as the corresponding lifting coefficient; the third value is smaller than the second value.

[0019] In one embodiment, determining the current amplitude of the second portion of the second power spectrum according to the smaller amplitude and the lift coefficient includes:

[0020] Obtain an expected value of the smaller amplitude; and obtain a current amplitude of the second part according to the product of the expected value of the smaller amplitude and the lifting coefficient.

[0021] In one embodiment, the obtaining of the current comfort noise signal according to the second power spectrum determined based on the current amplitude includes:

[0022] A comfort noise spectrum after spectrum gain compensation is obtained based on a second power spectrum determined by the current amplitude and a spectrum gain compensation factor; and the current comfort noise signal is obtained according to the comfort noise spectrum after spectrum gain compensation.

[0023] In a second aspect, the present application provides a voice transmission system for live streaming. The system comprises: a voice transmitter and a voice receiver for live streaming; wherein the voice transmitter is configured to clear silence frame signals and transmit the voice frame signals to the voice receiver; and the voice receiver is configured to output silence frames according to the method for estimating comfort noise as described in any of the above embodiments.

[0024] In a third aspect, the present application provides a device for estimating comfort noise, which is applied to a receiving end. The device includes:

[0025] A voice frame signal receiving module, configured to obtain a first power spectrum of a current voice frame signal upon receiving the current voice frame signal from a transmitting end;

[0026] a first amplitude determination module, configured to determine, if there is a first portion in the first power spectrum that does not reach the amplitude threshold, a current amplitude of the first portion in the second power spectrum according to a historical amplitude or a preset initial amplitude of the first portion in the second power spectrum whose current amplitude is to be determined;

[0027] a second amplitude determination module configured to determine, if there is a second portion in the first power spectrum that reaches an amplitude threshold, a current amplitude of the second portion in the second power spectrum based on a smaller amplitude and a lift coefficient; wherein the smaller amplitude is the smaller of the amplitude of the second portion in the first power spectrum and a historical amplitude of the second portion in the second power spectrum;

[0028] a comfort noise acquisition module, configured to acquire a current comfort noise signal according to a second power spectrum determined by a current amplitude;

[0029] The comfort noise output module is configured to fill the current comfort noise signal into the silence frame output when detecting the occurrence of the silence frame.

[0030] In a fourth aspect, the present application further provides an electronic device. The electronic device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0031] When a current voice frame signal is received from the transmitting end, a first power spectrum of the current voice frame signal is obtained; if there is a first part in the first power spectrum that does not reach the amplitude threshold, the current amplitude of the first part in the second power spectrum is determined based on the historical amplitude or the preset initial amplitude of the first part in the second power spectrum whose current amplitude is to be determined; if there is a second part in the first power spectrum that reaches the amplitude threshold, the current amplitude of the second part in the second power spectrum is determined based on the smaller amplitude and the lifting coefficient; wherein the smaller amplitude is the smaller of the amplitude of the second part in the first power spectrum and the historical amplitude of the second part in the second power spectrum; based on the second power spectrum determined by the current amplitude, a current comfort noise signal is obtained; when a silence frame is detected, the current comfort noise signal is filled into the silence frame output.

[0032] In a fifth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the following steps:

[0033] When a current voice frame signal is received from the transmitting end, a first power spectrum of the current voice frame signal is obtained; if there is a first part in the first power spectrum that does not reach the amplitude threshold, the current amplitude of the first part in the second power spectrum is determined based on the historical amplitude or the preset initial amplitude of the first part in the second power spectrum whose current amplitude is to be determined; if there is a second part in the first power spectrum that reaches the amplitude threshold, the current amplitude of the second part in the second power spectrum is determined based on the smaller amplitude and the lifting coefficient; wherein the smaller amplitude is the smaller of the amplitude of the second part in the first power spectrum and the historical amplitude of the second part in the second power spectrum; based on the second power spectrum determined by the current amplitude, a current comfort noise signal is obtained; when a silence frame is detected, the current comfort noise signal is filled into the silence frame output.

[0034] The above-mentioned method, apparatus, and voice transmission system, device, and medium for estimating comfort noise are as follows: when the receiving end receives the current voice frame signal from the transmitting end, it obtains its first power spectrum; if there is a first part in the first power spectrum that does not reach the amplitude threshold, the current amplitude of the first part in the second power spectrum is determined based on the historical amplitude or preset initial amplitude of the first part in the second power spectrum; if there is a second part in the first power spectrum that reaches the amplitude threshold, the current amplitude of the second part in the second power spectrum is determined based on the smaller amplitude and the lift coefficient, where the smaller amplitude is the smaller of the amplitude of the second part in the first power spectrum and the historical amplitude of the second part in the second power spectrum; the current comfort noise signal is obtained based on the second power spectrum determined by the current amplitude, and when a silence frame is detected, the current comfort noise signal is filled into the silence frame output. This solution can achieve the estimation of the comfort noise signal of the silence frame using only the voice frame signal at the receiving end, thereby supporting the transmitting end to transmit only the voice frame signal without sending the silence frame signal, improving the transmission bandwidth utilization without losing the voice continuity of the receiving end, and improving the voice transmission performance in scenarios such as live broadcast. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 This is a diagram of an application scenario of the method for estimating comfort noise in an embodiment of the present application;

[0036] Figure 2 Schematic diagram of a flow chart of a method for estimating comfort noise according to an embodiment of the present application;

[0037] Figure 3 This is a flowchart of a method for estimating comfort noise in a specific embodiment of the present application;

[0038] Figure 4 This is a structural block diagram of the voice transmission system for live broadcasting in an embodiment of the present application;

[0039] Figure 5 This is a flow chart of the voice data processing process of the voice transmission system in the embodiment of the present application;

[0040] Figure 6 This is a structural block diagram of an apparatus for estimating comfort noise according to an embodiment of the present application;

[0041] Figure 7 FIG. 1 is a diagram showing the internal structure of an electronic device in one embodiment. DETAILED DESCRIPTION

[0042] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0043] The method for estimating comfort noise provided in the embodiment of the present application can be applied to Figure 1 The application scenario shown may include a transmitter and a receiver, which communicate with the receiver via a network. Both the transmitter and the receiver may be, but are not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices may include smart speakers, smart TVs, and smart car devices, while portable wearable devices may include smart watches, smart bracelets, and head-mounted devices.

[0044] In this application, the transmitter only needs to transmit voice frame signals to the receiver, which then uses the voice frame signals to estimate the comfort noise signal of the silence frame. This improves transmission bandwidth utilization without sacrificing voice continuity at the receiver, and enhances voice transmission performance in scenarios such as live streaming. The following describes the comfort noise estimation method provided by this application in conjunction with the following embodiments and accompanying figures.

[0045] In one embodiment, Figure 2 As shown, a method for estimating comfort noise is provided, which is performed by a receiving end and includes the following steps:

[0046] Step S201: upon receiving a current voice frame signal from a transmitting end, obtaining a first power spectrum of the current voice frame signal.

[0047] In this application, when the transmitting end transmits voice data to the receiving end, the transmitting end performs VAD detection and sends the detected voice frame signal to the receiving end. If a silent frame signal is detected, it is not sent. In this step, the receiving end receives the voice frame signal from the transmitting end and uses the voice frame signal as the current voice frame signal. Then, the power spectrum of the current voice frame signal is obtained to distinguish it from the power spectrum related to the comfort noise signal. The power spectrum of the current voice frame signal is recorded as the first power spectrum.

[0048] Specifically, the receiving end decodes the voice data from the transmitting end to obtain the current voice frame signal x(t), and transforms the current voice frame signal x(t) into a transform domain to obtain Xf(t). Xf(t) is a specific transform domain output signal obtained by performing signal transformations such as FFT (fast Fourier transform) on the time domain input signal. The multi-band output signal can also be extracted by frequency division filtering to achieve conversion from the time domain to the transform domain. k represents the transform domain coordinate, such as the frequency point number. The receiving end can then calculate the first power spectrum Psd(n,k) of the current voice frame signal x(t) according to the following formula:

[0049] Psd(n,k)=E{Xf 2 (n,k)}.

[0050] Where n represents the frame sequence, k represents the transform domain coordinates such as the frequency point number, and E{·} represents the expected operation.

[0051] For the expected operation E{·}, illustratively, let E{A} be the expected operation of signal A, and the expected value of signal A can be calculated using N frames of signal:

[0052]

[0053] Among them, β i is the i-th frame signal A i The weight coefficient of .

[0054] Step S202: If there is a first part in the first power spectrum that does not reach the amplitude threshold, the current amplitude of the first part in the second power spectrum is determined based on the historical amplitude or the preset initial amplitude of the first part in the second power spectrum whose current amplitude is to be determined.

[0055] Step S203: If there is a second portion in the first power spectrum that reaches the amplitude threshold, a current amplitude of the second portion in the second power spectrum is determined based on the smaller amplitude and the lift coefficient, where the smaller amplitude is the smaller of the amplitude of the second portion in the first power spectrum and a historical amplitude of the second portion in the second power spectrum.

[0056] The above-mentioned steps S202 and S203 are explained in a unified manner. In the present application, the second power spectrum is used to obtain the power spectrum of the current comfort noise signal, that is, after determining the current amplitude in the second power spectrum, it can be used to obtain the current comfort noise signal through the transformation from the transform domain to the time domain. Among them, steps S202 and S203 are respectively used to obtain the current amplitudes of the first part and the second part in the second power spectrum. Among them, the parts referred to in the first part and the second part are the parts that constitute the power spectrum, and specifically can be the various frequency points and frequency bands of the power spectrum. This application uses the frequency point as the part referred to in the first part and the second part for explanation.

[0057] Specifically, after obtaining the first power spectrum of the current voice frame signal x(t), the receiving end may detect whether the amplitude Psd(n,k) of each frequency point in the first power spectrum reaches a preset amplitude threshold Th.

[0058] Among them, step S202, if it is detected that there is a first frequency point Psd(n,k) in the first power spectrum that does not reach the amplitude threshold Th, the receiving end determines the current amplitude BN(n,k) of the first frequency point in the second power spectrum based on the historical amplitude or preset initial amplitude of the first frequency point in the second power spectrum. That is, in this case, the receiving end refers to the historical amplitude of the second power spectrum itself with respect to the first frequency point (such as the amplitude determined when estimating the amplitude of the frequency point related to the second power spectrum based on the previous voice frame signal as the historical amplitude) or refers to the preset initial amplitude of the second power spectrum itself with respect to the first frequency point (such as the initial amplitude preset for each frequency point based on white noise) to determine the current amplitude BN(n,k) of the first frequency point in the second power spectrum. The method of estimating the current amplitude of the first frequency point in the second power spectrum in this step can avoid the amplitude BN(n,k) of the corresponding frequency point in the second power spectrum estimated to be 0 when the signal of some frequency points in the voice is lost or disconnected during transmission.

[0059] In step S203, if it is detected that there is a second frequency point Psd(n,k) that reaches the amplitude threshold Th in the first power spectrum, the receiving end determines the current amplitude BN(n,k) of the second part of the second power spectrum based on the smaller amplitude MinPsd(n,k) and the lift coefficient β(n,k). The smaller amplitude MinPsd(n,k) refers to the smaller of the amplitude Psd(n,k) of the second frequency point in the first power spectrum and the historical amplitude of the second part in the second power spectrum. The historical amplitude of the second part in the second power spectrum can be the amplitude BN(n-1,k) determined when estimating the amplitude of the frequency point related to the second power spectrum based on the previous voice frame signal. That is, the smaller amplitude MinPsd(n,k) can be expressed as:

[0060] MinPsd(n,k)=MIN{Psd(n,k),BN(n-1,k)}.

[0061] After obtaining the minimum amplitude MinPsd(n,k), the receiver determines the current amplitude BN(n,k) of the second portion of the second power spectrum based on the minimum amplitude MinPsd(n,k) and the boost coefficient β(n,k). Specifically, the receiver first calculates the minimum amplitude MinPsd(n,k), and then uses this amplitude in combination with the boost coefficient β(n,k) to update the current amplitude BN(n,k) of the second portion of the second power spectrum. The boost coefficient β(n,k) is used to boost the noise signal, facilitating real-time tracking of noise changes in the voice frame signal and preventing MinPsd(n,k) from falling into a global minimum. In practical applications, the boost coefficient β(n,k) can vary with the voice transmission process. Each frame can have a corresponding boost coefficient, and each frequency point can also have a corresponding boost coefficient. This means that different frame sequences and different frequencies can have different or identical boost coefficients.

[0062] Regarding the calculation of the current amplitude BN(n,k) of the second part in the second power spectrum in step S203, as an embodiment, determining the current amplitude of the second part in the second power spectrum according to the smaller amplitude and the lift coefficient in step S203 specifically includes:

[0063] The expected value of the smaller amplitude is obtained; and the current amplitude of the second part is obtained according to the product of the expected value of the smaller amplitude and the lifting coefficient.

[0064] Specifically, after the receiving end obtains the smaller amplitude MinPsd(n,k), it obtains the expected value E{MinPsd(n,k)} of the smaller amplitude MinPsd(n,k), and then calculates the product of the expected value E{MinPsd(n,k)} of the smaller amplitude and the lifting coefficient β(n,k), and obtains the current amplitude of the second part of the second power spectrum as BN(n,k):

[0065] BN(n,k)=E{MinPsd(n,k)}*β(n,k).

[0066] Combining the above steps S202 and S203, taking the amplitude BN(n-1, k) determined when estimating the amplitude of the relevant frequency point of the second power spectrum based on the previous voice frame signal as the historical amplitude of the relevant frequency point as an example, the method of determining the current amplitude BN(n, k) of the corresponding frequency point in the second power spectrum represented by steps S202 and S203 can be expressed as follows:

[0067]

[0068] MinPsd(n,k)=MIN{Psd(n,k),BN(n-1,k)}.

[0069] Step S204: Acquire a current comfort noise signal according to the second power spectrum determined by the current amplitude.

[0070] The second power spectrum determined by the current amplitude refers to the second power spectrum determined by the current amplitude of the corresponding portion (corresponding frequency point) through steps S202 and S203. In this step, after the receiving end obtains the second power spectrum determined by the current amplitude according to steps S202 and S203, it can obtain the current comfort noise signal through a transform from the transform domain to the time domain. For example, the transform from the time domain to the transform domain is an FFT. In this step, the receiving end can use an IFFT (Inverse Fast Fourier Transform) to inversely transform the second power spectrum determined by the current amplitude to obtain the current comfort noise signal.

[0071] Specifically, in some embodiments, obtaining the current comfort noise signal by determining the second power spectrum according to the current amplitude in step S204 further includes:

[0072] A second power spectrum determined based on the current amplitude and a spectrum gain compensation factor is obtained, and a spectrum gain compensated comfort noise spectrum is obtained. A current comfort noise signal is obtained according to the spectrum gain compensated comfort noise spectrum.

[0073] In this embodiment, as shown in the following formula, the receiving end generates a spectrum gain compensated comfort noise spectrum CNG (CNG(n,2*k), CNG(n,2*k+1)) based on the second power spectrum BN(n,k) and the spectrum gain compensation factor SpecFactor(k):

[0074]

[0075]

[0076] Taking IFFT as an example, the inverse transform is performed. The process from the second power spectrum BN(n,k) to the comfort noise spectrum CNG is recorded as the comfort noise spectrum estimation process. The random phase generated in the comfort noise spectrum estimation process is There is no guarantee that the conjugate cancellation of the IFFT rotation factor is achieved, so the amplitude of the time domain signal cng(t) obtained after the IFFT accumulation calculation will be smaller than the spectrum amplitude:

[0077]

[0078] Therefore, the receiver performs spectral gain compensation during the comfort noise spectrum estimation process. SpecFactor(k) is the spectral gain compensation factor, which can be set empirically. The receiver then obtains the spectral gain-compensated CNG spectrum, also known as the comfort noise spectrum CNG. The receiver then performs an IFFT on the comfort noise spectrum CNG to obtain the time-domain comfort noise signal cng(t), which is stored as the current comfort noise signal in the CNG buffer (comfort noise generator buffer).

[0079] Step S205: When a silence frame is detected, the current comfort noise signal is filled into the silence frame and output.

[0080] In this step, during the voice communication between the receiving end and the sending end, the receiving end can determine whether a silent frame occurs based on whether the sending end sends data. When a silent frame is detected, the receiving end obtains the current comfort noise signal from the CNG buffer and fills it into the silent frame output to maintain voice continuity.

[0081] Combine Figure 3 The comfort noise estimation process implemented at the receiving end as shown in the above steps is generally described. First, when the receiving end receives the current voice frame signal from the transmitting end, it performs a spectrum transformation on the current voice frame signal to obtain a first power spectrum. The receiving end detects whether the amplitude of each frequency point in the first power spectrum is greater than a preset amplitude threshold. If so, it outputs the iterative amplitude; otherwise, it outputs the historical amplitude. The output iterative amplitude corresponds to the method described in step S203 above to determine the current amplitude of the corresponding frequency point in the second power spectrum, and the output historical amplitude corresponds to the method described in step S202 above to determine the current amplitude of the corresponding frequency point in the second power spectrum. After obtaining the second power spectrum determined by the current amplitude, spectrum compensation and random phase addition are performed, and then an inverse transformation is performed to obtain the current comfort noise signal and store it in the CNGbuffer (comfort noise generation buffer). When a silent frame is detected, the current comfort noise signal is filled into the silent frame output.

[0082] The comfort noise estimation method of this embodiment can estimate the comfort noise signal of the silence frame using only the voice frame signal at the receiving end, thereby supporting the transmitting end to transmit only the voice frame signal without sending the silence frame signal, thereby improving the transmission bandwidth utilization without losing the voice continuity at the receiving end, and enhancing the voice transmission performance in scenarios such as live broadcasting.

[0083] In some embodiments, determining the current amplitude of the first part of the second power spectrum according to the historical amplitude or the preset initial amplitude of the first part of the second power spectrum whose current amplitude is to be determined in step S202 specifically includes:

[0084] When the comfort noise signal is acquired for the first time, the preset initial amplitude of the first part in the second power spectrum is used as the current amplitude of the first part in the second power spectrum; when the comfort noise signal is not acquired for the first time, the current amplitude of the first part in the second power spectrum is determined based on the historical amplitude of the first part in the second power spectrum.

[0085] In this embodiment, during voice transmission between a receiving end and a transmitting end, if the receiving end is estimating a comfort noise signal for the first time, i.e., acquiring the comfort noise signal for the first time, then, because there is no historical value for reference, the receiving end uses a preset initial amplitude of a first frequency point in the second power spectrum as the current amplitude of the first frequency point in the second power spectrum. As described above, the preset initial amplitude may be an initial amplitude preset for each frequency point based on white noise. If the receiving end is not estimating a comfort noise signal for the first time, i.e., acquiring the comfort noise signal for the first time, the receiving end may determine the current amplitude of the first frequency point in the second power spectrum based on the historical amplitude of the first frequency point in the second power spectrum.

[0086] As an embodiment, the above-mentioned determining the current amplitude of the first part in the second power spectrum based on the historical amplitude of the first part in the second power spectrum may specifically include: taking the amplitude of the first part in the second power spectrum of the previous comfort noise signal as the current amplitude of the first part in the second power spectrum.

[0087] In this embodiment, the receiving end obtains the amplitude of the first frequency point in the second power spectrum obtained when the comfort noise signal was last estimated (that is, the second power spectrum of the previous comfort noise signal), and uses the amplitude as the current amplitude of the first frequency point in the second power spectrum estimated for the comfort noise signal this time.

[0088] Regarding the determination of the lifting coefficient, in some embodiments, before determining the current amplitude of the second portion of the second power spectrum according to the smaller amplitude and the lifting coefficient in step S203, the lifting coefficient may be determined by the following steps, including:

[0089] Determine the amplitude raising stage of the second part in the second power spectrum; and determine the corresponding raising coefficient according to the amplitude raising stage.

[0090] This embodiment is a process for determining the lifting coefficient β(n,k). Specifically, when there is a second frequency point that reaches the amplitude threshold in the first power spectrum, the receiving end needs to determine the current amplitude of the second frequency point in the second power spectrum based on the smaller amplitude and the lifting coefficient β(n,k). As mentioned above, the lifting coefficient β(n,k) can change with the voice transmission process, and each frequency point of each frame can have a lifting coefficient of a corresponding value. In this embodiment, the receiving end determines the amplitude lifting stage of the second frequency point in the second power spectrum. The amplitude lifting stage can include at least two stages. Exemplarily, the amplitude lifting stage can include a first stage, a second stage, and a third stage, wherein the first stage can be called the initial stage of the comfort noise iteration, the second stage can be called the convergence stage of the comfort noise iteration, and the third stage can be called the limiting stage of the comfort noise iteration. Different amplitude lifting stages can correspond to lifting coefficients of different values. The receiving end determines the lifting coefficient of the corresponding value according to the amplitude lifting stage, thereby effectively tracking the noise by setting the lifting coefficient in stages.

[0091] Furthermore, in one embodiment, determining the amplitude rising phase of the second portion in the second power spectrum specifically includes:

[0092] If the number of times the second part of the second power spectrum rises does not reach the preset number of times, the amplitude raising stage is determined to be the first stage; if the number of times the second part of the second power spectrum rises reaches the preset number of times, and the historical amplitude of the second part of the second power spectrum does not reach the amplitude threshold, the amplitude raising stage is determined to be the second stage; if the historical amplitude of the second part of the second power spectrum reaches the amplitude threshold, the amplitude raising stage is determined to be the third stage.

[0093] In this embodiment, combined with Figure 3 The receiving end determines the amplitude raising phase it is in based on the number of times the second portion of the second power spectrum is raised and the historical amplitude of the second portion of the second power spectrum. The amplitude raising phase includes a first phase, a second phase, and a third phase. The first phase is the initial phase of the comfort noise iteration, the second phase is the convergence phase of the comfort noise iteration, and the third phase is the limiting phase of the comfort noise iteration. First, for the first stage, i.e., the initial stage of the comfort noise iteration, the receiving end obtains the number of times the second frequency point in the second power spectrum is raised. If the receiving end determines that the number of times the second frequency point is raised does not reach the preset number of times, the receiving end determines that its amplitude raising stage is the first stage; for the second stage, i.e., the convergence stage of the comfort noise iteration, if the receiving end determines that the number of times the second frequency point is raised reaches the preset number of times, the receiving end obtains the historical amplitude of the second frequency point in the second power spectrum (such as the amplitude determined when estimating the amplitude of the corresponding frequency point in the second power spectrum based on the previous voice frame signal). If the receiving end determines that the historical amplitude does not reach the preset amplitude threshold, the receiving end determines that its amplitude raising stage is the second stage; for the third stage, i.e., the limiting stage of the comfort noise iteration, if the receiving end determines that the aforementioned historical amplitude reaches the amplitude threshold, the receiving end determines that its amplitude raising stage is the third stage. Thus, this embodiment accurately determines the amplitude raising stage of the relevant part in the second power spectrum based on the number of times the amplitude is raised and the historical amplitude.

[0094] Based on this, the above-mentioned determination of the corresponding lifting coefficient according to the amplitude lifting stage further includes:

[0095] When the amplitude raising stage is the first stage, the lifting coefficient with the first value is used as the corresponding lifting coefficient; when the amplitude raising stage is the second stage, the lifting coefficient with the second value is used as the corresponding lifting coefficient; the second value is smaller than the first value; when the amplitude raising stage is the third stage, the lifting coefficient with the third value is used as the corresponding lifting coefficient; the third value is smaller than the second value.

[0096] In this embodiment, when the receiving end determines that the amplitude raising stage of the second frequency point in the second power spectrum is the first stage, that is, the initial stage of the comfort noise iteration, the receiving end will use the raising coefficient with the first value as its corresponding raising coefficient; when the receiving end determines that the amplitude raising stage of the second frequency point is the second stage, that is, the convergence stage of the comfort noise iteration, the receiving end will use the raising coefficient with a second value smaller than the first value as its corresponding raising coefficient; when the receiving end determines that the amplitude raising stage of the second frequency point is the third stage, that is, the limiting stage of the comfort noise iteration, the receiving end will use the raising coefficient with a third value smaller than the second value as its corresponding raising coefficient, for example, the third value is 1.0, the second value is 1.01, the first value is 1.02, etc.

[0097] The lifting coefficient β(n,k) of this embodiment can be expressed as:

[0098]

[0099] This embodiment can use a larger lifting coefficient to quickly track the background noise in the initial stage of the comfort noise iteration. After the iteration is stable, a medium-sized lifting coefficient can be used to keep the tracking stable. When the estimated noise is detected to exceed the threshold, the lifting coefficient is limited. In this way, the estimated background noise amplitude is protected by the lifting coefficient to prevent it from tracking the speech characteristics, that is, to avoid the estimated noise from tracking the speech characteristics.

[0100] In one embodiment, Figure 4 As shown, a voice transmission system for live network broadcasting is provided. The system may include: a voice transmitter for live network broadcasting and a voice receiver for live network broadcasting; the voice transmitter is communicatively connected to the voice receiver via a network. The voice transmitter is configured to clear silence frame signals and transmit the voice frame signals to the voice receiver; the voice receiver is configured to output silence frames according to the method for estimating comfort noise described in any of the above embodiments.

[0101] Specific, combined Figure 5 The voice transmission between the voice transmitter and the voice receiver in this embodiment is described. The voice transmitter uses DTX (Discontinuous Transmission) to transmit voice to the voice receiver. DTX transmission is a discontinuous transmission of voice data to save transmission bandwidth. For example, in this embodiment, the voice transmitter only transmits voice frame signals. The voice transmitter performs VAD detection on the voice to determine whether the detected signal is a voice frame signal. If a silence frame signal is detected, the voice transmitter clears the silence frame signal, i.e., does not transmit the silence frame signal. If a voice frame signal is detected, it is encoded and transmitted to the voice receiver. The voice receiving end decodes the voice data from the voice transmitting end and determines whether it is a frame signal that needs to be filled with noise. If so, that is, a silence frame is detected, the voice receiving end loads comfort noise from the CNG buffer and fills it into the playback buffer, that is, filling the silence frame. The playback device at the voice receiving end can then output comfortable noise that sounds continuous between voice frames. If not, that is, the current voice frame signal is detected, the voice receiving end estimates the current comfort noise signal according to the comfort noise estimation method described in the above embodiment and stores it in the CNG buffer. When a silence frame is detected, the comfort noise signal is filled into the silence frame for output, and the voice receiving end plays and outputs the voice frame signal.

[0102] This embodiment can estimate the comfort noise used to fill the silent frame by using only the voice frame signal at the voice receiving end, thereby supporting the voice sending end to transmit only the voice frame signal, while improving the transmission bandwidth utilization without losing the voice continuity at the voice receiving end, thereby improving the DTX performance of the voice room service in the live broadcast scenario, increasing the applicability of the comfort noise, and improving the compatibility of related services.

[0103] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0104] Based on the same inventive concept, embodiments of the present application further provide a device for estimating comfort noise for implementing the aforementioned method for estimating comfort noise. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations in one or more embodiments of the device for estimating comfort noise provided below can be found in the aforementioned method for estimating comfort noise, and are not further elaborated here.

[0105] In one embodiment, Figure 6 As shown, a device for estimating comfort noise is provided, which is applied to a receiving end. The device 600 includes:

[0106] The voice frame signal receiving module 601 is configured to obtain a first power spectrum of a current voice frame signal upon receiving the current voice frame signal from the transmitting end;

[0107] a first amplitude determination module 602 configured to determine, if there is a first portion in the first power spectrum that does not reach the amplitude threshold, a current amplitude of the first portion in the second power spectrum according to a historical amplitude or a preset initial amplitude of the first portion in the second power spectrum whose current amplitude is to be determined;

[0108] a second amplitude determination module 603 configured to determine, if there is a second portion in the first power spectrum that reaches the amplitude threshold, a current amplitude of the second portion in the second power spectrum based on a smaller amplitude and a lift coefficient; wherein the smaller amplitude is the smaller of the amplitude of the second portion in the first power spectrum and a historical amplitude of the second portion in the second power spectrum;

[0109] The comfort noise acquisition module 604 is configured to acquire a current comfort noise signal according to the second power spectrum determined by the current amplitude;

[0110] The comfort noise output module 605 is configured to fill the current comfort noise signal into the silence frame output when a silence frame is detected.

[0111] In one embodiment, the first amplitude determination module 602 is configured to use a preset initial amplitude of the first part in the second power spectrum as the current amplitude of the first part in the second power spectrum when the comfort noise signal is acquired for the first time; and to determine the current amplitude of the first part in the second power spectrum based on a historical amplitude of the first part in the second power spectrum when the comfort noise signal is not acquired for the first time.

[0112] In one embodiment, the first amplitude determining module 602 is configured to use the amplitude of the first part in the second power spectrum of the previous comfort noise signal as the current amplitude of the first part in the second power spectrum.

[0113] In one embodiment, the apparatus 600 further includes: a lifting coefficient determination module, configured to determine an amplitude lifting stage of the second portion of the second power spectrum; and determine a corresponding lifting coefficient according to the amplitude lifting stage.

[0114] In one embodiment, a lifting coefficient determination module is used to determine that the amplitude lifting stage is the first stage if the number of lifting times for the second part in the second power spectrum does not reach the preset number of lifting times; if the number of lifting times for the second part in the second power spectrum reaches the preset number of lifting times, and the historical amplitude of the second part in the second power spectrum does not reach the amplitude threshold, then the amplitude lifting stage is determined to be the second stage; if the historical amplitude of the second part in the second power spectrum reaches the amplitude threshold, then the amplitude lifting stage is determined to be the third stage; and, when the amplitude lifting stage is the first stage, the lifting coefficient with a first value is used as the corresponding lifting coefficient; when the amplitude lifting stage is the second stage, the lifting coefficient with a second value is used as the corresponding lifting coefficient; the second value is less than the first value; when the amplitude lifting stage is the third stage, the lifting coefficient with a third value is used as the corresponding lifting coefficient; the third value is less than the second value.

[0115] In one embodiment, the second amplitude determination module 603 is configured to obtain an expected value of the smaller amplitude; and obtain the current amplitude of the second part according to the product of the expected value of the smaller amplitude and the lifting coefficient.

[0116] In one embodiment, the comfort noise acquisition module 604 is configured to acquire a spectral gain compensated comfort noise spectrum based on the second power spectrum determined by the current amplitude and a spectral gain compensation factor; and acquire the current comfort noise signal according to the spectral gain compensated comfort noise spectrum.

[0117] Each module in the above-mentioned apparatus for estimating comfort noise may be implemented in whole or in part via software, hardware, or a combination thereof. Each module may be embedded in or independent of a processor in an electronic device in the form of hardware, or may be stored in a memory in the electronic device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0118] In one embodiment, an electronic device is provided. The electronic device may be a terminal, and its internal structure diagram may be as follows: Figure 7 As shown. The electronic device includes a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the electronic device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a method for estimating comfort noise is implemented. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the electronic device can be a touch layer covering the display screen, or a button, trackball or touchpad provided on the housing of the electronic device, or an external keyboard, touchpad or mouse.

[0119] Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the electronic device to which the solution of the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0120] In one embodiment, an electronic device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0121] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0122] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.

[0123] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0124] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A method for estimating comfort noise, characterized in that Applied to a receiving end, the method includes: Upon receiving a current voice frame signal from a transmitting end, obtaining a first power spectrum of the current voice frame signal; If there is a first portion in the first power spectrum that does not reach the amplitude threshold, determining the current amplitude of the first portion in the second power spectrum according to a historical amplitude or a preset initial amplitude of the first portion in the second power spectrum whose current amplitude is to be determined; If there is a second part in the first power spectrum that reaches the amplitude threshold, determining a current amplitude of the second part in the second power spectrum according to the smaller amplitude and the lift coefficient; wherein the smaller amplitude is the smaller of the amplitude of the second part in the first power spectrum and the historical amplitude of the second part in the second power spectrum; Acquiring a current comfort noise signal according to a second power spectrum determined by the current amplitude; comprising: acquiring a comfort noise spectrum compensated by spectrum gain based on the second power spectrum determined by the current amplitude and a spectrum gain compensation factor; acquiring the current comfort noise signal according to the comfort noise spectrum compensated by spectrum gain; When a silence frame is detected, the current comfort noise signal is filled into the silence frame output.

2. The method according to claim 1, characterized in that The determining the current amplitude of the first part in the second power spectrum according to the historical amplitude or the preset initial amplitude of the first part in the second power spectrum whose current amplitude is to be determined includes: When acquiring the comfort noise signal for the first time, using the preset initial amplitude of the first part in the second power spectrum as the current amplitude of the first part in the second power spectrum; When the comfort noise signal is not acquired for the first time, the current amplitude of the first part in the second power spectrum is determined according to the historical amplitude of the first part in the second power spectrum.

3. The method according to claim 2, characterized in that The determining the current amplitude of the first part in the second power spectrum according to the historical amplitude of the first part in the second power spectrum includes: The amplitude of the first part in the second power spectrum of the previous comfort noise signal is used as the current amplitude of the first part in the second power spectrum.

4. The method according to claim 1, wherein Before determining the current amplitude of the second part in the second power spectrum according to the smaller amplitude and the lift coefficient, the method further includes: determining an amplitude rise phase of the second portion in the second power spectrum; According to the amplitude raising stage, a corresponding raising coefficient is determined.

5. The method according to claim 4, characterized in that The determining the amplitude raising stage of the second part in the second power spectrum includes: If the number of times the second part of the second power spectrum is raised does not reach a preset number of times, determining that the amplitude raising stage is the first stage; If the number of times the second part of the second power spectrum is raised reaches the preset number of times, and the historical amplitude of the second part of the second power spectrum does not reach the amplitude threshold, determining that the amplitude raising stage is the second stage; If the historical amplitude of the second part in the second power spectrum reaches an amplitude threshold, determining that the amplitude raising stage is the third stage; Determining the corresponding raising coefficient according to the amplitude raising stage includes: When the amplitude raising stage is the first stage, the raising coefficient having the first value is used as the corresponding raising coefficient; When the amplitude raising stage is the second stage, a raising coefficient having a second value is used as the corresponding raising coefficient; the second value is smaller than the first value; When the amplitude raising stage is the third stage, a raising coefficient having a third value is used as the corresponding raising coefficient; the third value is smaller than the second value.

6. The method according to claim 1, characterized in that The determining, according to the smaller amplitude and the lift coefficient, a current amplitude of the second part in the second power spectrum, comprises: obtaining an expected value of the smaller amplitude; The current amplitude of the second part is obtained according to the product of the expected value of the smaller amplitude and the lifting coefficient.

7. A voice transmission system for live webcasting, characterized in that: include: The voice sending end and voice receiving end of the live broadcast; The voice transmitting end is used to clear the silence frame signal and send the voice frame signal to the voice receiving end; The speech receiving end is configured to output silence frames according to the method for estimating comfort noise according to any one of claims 1 to 6.

8. A device for estimating comfort noise, characterized in that Applied to a receiving end, the device includes: A voice frame signal receiving module, configured to obtain a first power spectrum of a current voice frame signal upon receiving the current voice frame signal from a transmitting end; a first amplitude determination module, configured to determine, if there is a first portion in the first power spectrum that does not reach the amplitude threshold, a current amplitude of the first portion in the second power spectrum according to a historical amplitude or a preset initial amplitude of the first portion in the second power spectrum whose current amplitude is to be determined; a second amplitude determination module configured to determine, if there is a second portion in the first power spectrum that reaches an amplitude threshold, a current amplitude of the second portion in the second power spectrum based on a smaller amplitude and a lift coefficient; wherein the smaller amplitude is the smaller of the amplitude of the second portion in the first power spectrum and a historical amplitude of the second portion in the second power spectrum; A comfort noise acquisition module is configured to acquire a current comfort noise signal based on a second power spectrum determined by a current amplitude; the module comprises: acquiring a comfort noise spectrum compensated by spectrum gain based on the second power spectrum determined by the current amplitude and a spectrum gain compensation factor; and acquiring the current comfort noise signal based on the comfort noise spectrum compensated by spectrum gain. The comfort noise output module is configured to fill the current comfort noise signal into the silence frame output when detecting the occurrence of the silence frame.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Background noise generation method and device

    CN105721656A