Systems and methods for providing high-quality audio communication in low-bitrate network connections

By using machine learning-based noise suppression and new encoder in real-time communication, decomposing audio data into subbands, extracting and quantizing features, and combining deep learning to predict residual signal, the problem of audio signals discontinuity under low-code rate networks is solved, and high-quality audio playback is achieved.

CN116137151BActive Publication Date: 2025-07-22AGORA LAB INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210666398.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-11-17
Filing Date
2022-06-13
Publication Date
2025-07-22
Estimated Expiration
2042-06-13

AI Technical Summary

Technical Problem

In real-time communication, under low-code network connections, existing audio codecs lead to high computing costs and cannot be effectively decoded on portable devices, resulting in discontinuity of audio signals and affecting the listening experience.

Method used

The noise suppression module based on machine learning and a new encoder are used to decompose the audio data into low-frequency and high-frequency subbands, extract audio characteristics and quantify and compress them, and combine deep learning methods to predict and de-emphasize residual signals at the receiving end to generate high-quality audio signals.

Benefits of technology

High-quality audio playback is achieved under low-code rate networks, reducing network bandwidth requirements, and maintaining the continuity of audio signals and listening experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116137151B_ABST
    Figure CN116137151B_ABST
Patent Text Reader

Abstract

The present invention provides a novel system and method for providing high-quality audio under low-bitrate network connections in real-time communication. The system includes a real-time communication software application equipped with an improved encoder and an improved decoder. The encoder divides audio data corresponding to two frequency ranges of an ultra-wideband mode and a wideband mode into low-frequency subband and high-frequency subband audio data. Audio features are extracted from the low-frequency subband and high-frequency subband audio data. The audio features are quantized and packed. The decoder reconstructs the audio data based on the compressed audio features in the ultra-wideband mode and the wideband mode for playback on a receiving device.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims priority to a prior U.S. application with an application date of November 17, 2021 and application number 17,528 / 217. Technical field

[0003] The present invention generally relates to the field of real - time communication with audio data capture and remote playback capabilities. Specifically, the present invention relates to a real - time communication system that provides high - quality audio playback in a low - bitrate network connection. More specifically, the present invention relates to a real - time communication software application equipped with a codec that includes a low - bitrate audio encoder and a high - quality decoder. Background art

[0004] In real - time communication (RTC), the network bandwidth (also known as bitrate) is usually limited. The audio signal in RTC is encoded by a sending - end electronic device (such as a smartphone, tablet, laptop computer, or desktop computer) at the sending end and decoded at the receiving end. When the bitrate is low, compared to when the bitrate is high, the audio signal in RTC needs to pack data packets into smaller ones for transmission over the Internet. Therefore, an audio codec is used to compress the audio data packets as small as possible while trying to ensure the decoded audio quality.

[0005] Audio codecs based on deep learning usually result in too high a computational cost on the computer running the deep learning. The high computational cost makes the codec not suitable for portable devices such as smartphones and laptops. This is especially true when multiple audio signals need to be decoded simultaneously on the same computer, such as in a multi - user online meeting. If the audio data packets cannot be decoded in time, there will be discontinuities in playback on the receiving device, resulting in a significantly reduced listening experience.

[0006] Therefore, in RTC communication, a new type of low - bitrate audio codec with a high - quality decoder is needed to achieve the purpose of saving network bandwidth costs and maintaining the quality of the RTC experience in a weak network situation. The network bandwidth may vary over time. For example, when the network signal is weak or when there are too many devices sharing the same network, the available network bandwidth may drop to a very low level or range. In this case, the audio packet loss rate will increase, resulting in discontinuous audio signals. The reason is that the poor network bandwidth causes some audio data packets (also called audio signals in the present invention) to be discarded or blocked. Therefore, in the case of limited network bandwidth, only a low - bitrate audio codec can provide continuous audio stream playback at the receiving end. Summary of the invention

[0007] Generally speaking, based on various embodiments, the present invention provides a computer operation method for providing high-quality audio playback through a low-bitrate network connection in real-time communication. This method is run by a real-time communication software application program, and includes: receiving an audio input data stream on a sending device; suppressing noise in the audio input data stream on the sending device to generate clean audio input data; splitting the clean audio input data into a set of audio data frames on the sending device; normalizing each frame in the set of audio data frames on the sending device to generate a set of normalized audio data frames, where the audio data of the frame is resampled according to two frequency ranges corresponding to the wideband mode and the ultra-wideband mode, so as to form low-frequency subband audio data and high-frequency subband audio data; extracting a set of audio features from each frame in the set of normalized audio data frames on the sending device, thereby forming a set of audio features; quantifying the set of audio features of each frame in the set of normalized audio data frames on the sending device into a set of compressed audio features; packing a set of compressed audio features into an audio data packet on the sending device; sending the audio data packet from the sending device to the receiving device; receiving the audio data packet in the ultra-wideband mode on the receiving device; obtaining the set of audio features of each frame in the set of normalized audio data frames from the audio data packet on the receiving device; in the low-frequency subband and high-frequency subband of the ultra-wideband mode, on the receiving device, determining the linear prediction value of the next sample for each audio data sample of the frame according to the audio feature set corresponding to the data frame; using a deep learning method on the receiving device to extract a context vector for residual signal prediction from the acoustic feature vector of the low-frequency subband samples; determining the first residual prediction value of the samples in the low-frequency subband on the receiving device; combining the linear prediction value with the first residual prediction value on the receiving device to generate a subband audio signal for the samples in the low-frequency subband; performing de-emphasis processing on the subband audio signal on the receiving device to form a de-emphasized low-frequency subband audio signal; determining the second residual prediction value of the samples in the high-frequency subband on the receiving device; combining the linear prediction value and the second residual prediction value, and generating a subband audio signal for the samples in the high-frequency subband on the receiving device; merging the de-emphasized low-frequency subband audio signal and the subband audio signal of the samples in the high-frequency subband on the receiving device to form a merged audio sample; and then converting the merged audio sample into audio data for playback on the receiving device.

[0008] Extract an audio feature set from each frame within a set of frames of standardized audio data in ultra-wideband mode, including: pre-emphasizing the low-frequency subband audio data using a high-pass filter to form pre-emphasized low-frequency subband audio data; running Bark-Frequency Cepstrum Coefficients (BFCC) calculations on the pre-emphasized low-frequency subband audio data to extract the BFCC features of the audio, and performing pitch prediction processing on the pre-emphasized low-frequency subband audio data to extract the pitch features of the audio, including information such as fundamental period and pitch correlation; calculating the Linear Prediction Coding (LPC) coefficients of the audio based on the high-frequency subband audio data; converting the LPC coefficients to Line Spectral Frequencies (LPF) coefficients; determining the ratio of the sum of energies between the low-frequency subband data and the high-frequency subband audio data, where the ratio of the sum of energies, the LPF coefficients, the pitch features of the audio, and the BFCC features of the audio form part of the audio feature set.

[0009] Extract an audio feature set from each frame in a set of frames of standardized audio data in wideband mode, including: pre-emphasizing the standardized audio data of each frame using a high-pass filter to form pre-emphasized standardized audio data; running Bark-Frequency Cepstrum Coefficients (BFCC) calculations on the pre-emphasized standardized audio data to extract the BFCC features of the audio, and performing pitch prediction processing on the pre-emphasized standardized audio data to extract the audio pitch features including information such as fundamental period and pitch correlation, where the audio pitch features and the audio BFCC features form part of the audio feature set.

[0010] Obtain the audio feature set of each frame within a set of frames of standardized audio data from an audio data packet on a receiving device, including: performing inverse quantization processing on the compressed audio feature set to obtain the audio feature set; determining the LPC (Linear Prediction Coding) coefficients of the high-frequency subband based on the LPF coefficients; determining the LPC coefficients of the low-frequency subband based on the BFCC coefficients.

[0011] In one implementation, the inverse quantization processing uses the inverse differential vector quantization (DVQ) method, the inverse residual vector quantization (RVQ) method, or the inverse interpolation method.

[0012] The method for quantizing an audio feature set includes: compressing the audio feature set of each I-frame (key frame) within the set of frames using the residual vector quantization (RVQ) method or the differential vector quantization (DVQ) method, where there is at least one I-frame in the set of frames; compressing the audio feature set of each non-I frame within the set of frames using the interpolation method.

[0013] In one embodiment, the two frequency ranges are 0 to 16 kHz and 16 kHz to 32 kHz respectively, and a machine learning-based method is used to suppress noise.

[0014] In addition, according to the present invention, a computer operation method is also provided for providing high-quality audio playback through a low-bitrate network connection in real-time communication. This method is carried out through a real-time communication software application, and includes: receiving an audio input data stream on the sending device; suppressing noise in the audio input data stream on the sending device to generate clean audio input data; splitting the clean audio input data into a set of audio data frames on the sending device; normalizing each frame in the set of frames, so as to generate a set of normalized audio data frames on the sending device, wherein the audio data of the frame is resampled according to two frequency ranges corresponding to the wideband mode and the ultra-wideband mode, so as to form low-frequency sub-band audio data and high-frequency sub-band audio data; extracting an audio feature set from each frame in the set of normalized audio data frames, so as to form a set of audio feature sets on the sending device; quantifying the audio feature set of each frame in the set of normalized audio data frames into a compressed audio feature set on the sending device; packing a set of compressed audio feature sets into an audio data packet on the sending device; sending the audio data packet from the sending device to the receiving device; receiving the audio data packet in the wideband mode on the receiving device; obtaining the audio feature set of each frame in the set of frames by performing inverse quantization processing on the receiving device, wherein the audio feature set includes a set of Bark-Frequency Cepstrum Coefficients (BFCC) on the receiving device; determining a set of Linear Prediction Coding (LPC) coefficients on the receiving device according to the set of BFCC coefficients; determining the linear prediction value of the next sample for each sample of the audio data of each frame in the set of frames according to the audio feature set on the receiving device; using a deep learning method to extract a context vector for residual signal prediction from the acoustic feature vector of the sample on the receiving device; determining the residual signal prediction value of the sample based on the context vector, deep learning network, linear prediction value, final output signal value and final predicted residual signal; combining the linear prediction value and the residual signal prediction value to generate the audio signal of the sample; performing de-emphasis processing on the audio signal of the sample to generate a de-emphasized audio signal for playback on the receiving device.

[0015] Extract an audio feature set for each frame within a set of standardized audio data frames in ultra-wideband mode, including: pre-emphasizing the low-frequency subband audio data using a high-pass filter to form pre-emphasized low-frequency subband audio data; running Bark-Frequency Cepstrum Coefficients (BFCC) calculations on the pre-emphasized low-frequency subband audio data to extract audio BFCC features, and performing pitch prediction processing on the pre-emphasized low-frequency subband audio data to extract audio pitch features, where the audio pitch features include information such as fundamental period and pitch correlation; calculating audio Linear Prediction Coding (LPC) coefficients based on the high-frequency subband audio data; converting the LPC coefficients to Line Spectral Frequencies (LPF) coefficients; determining the ratio of the sum of energies between the low-frequency subband data and the high-frequency subband audio data, where the ratio of the sum of energies, LPF coefficients, audio pitch features, and audio BFCC features form part of the audio feature set.

[0016] Extract an audio feature set for each frame in a set of frames of standardized audio data in wideband mode, including: pre-emphasizing the standardized audio data of each frame using a high-pass filter to form pre-emphasized standardized audio data; running Bark-Frequency Cepstrum Coefficients (BFCC) calculations on the pre-emphasized standardized audio data to extract audio BFCC features, and performing pitch prediction processing on the pre-emphasized standardized audio data to extract audio pitch features including information such as fundamental period and pitch correlation, where the audio pitch features and audio BFCC features form part of the audio feature set.

[0017] In one embodiment, the inverse quantization process employs an inverse Differential Vector Quantization (DVQ) method, an inverse Residual Vector Quantization (RVQ) method, or an inverse interpolation method.

[0018] The method for quantizing the audio feature set includes: compressing the audio feature set of each I-frame within the set of frames using a Residual Vector Quantization (RVQ) method or a Differential Vector Quantization (DVQ) method, where there is at least one I-frame in the set of frames; compressing the audio feature set of each non-I-frame within the set of frames using an interpolation method.

[0019] In one embodiment, the two frequency ranges are 0 to 16 kHz and 16 kHz to 32 kHz respectively, and noise suppression employs a machine learning-based method. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] This patent or application document contains at least one color drawing. The Patent Office will provide a copy of the disclosure of this patent or patent application with color drawings upon request and payment of the relevant fees.

[0021] The technical features of the present invention will be specifically pointed out in the claims. Meanwhile, the present invention itself, as well as its constitution and usage method, can be better understood by referring to the following description and the accompanying drawings which are part of the description. All the accompanying drawings of the present invention are also part of the present invention, where the same reference numerals represent the same components:

[0022] Figure 1 is a schematic block diagram of a real-time communication system drawn according to the present invention.

[0023] Figure 2 is a schematic block diagram of a real-time communication device installed with an improved real-time communication application drawn according to the present invention.

[0024] Figure 3 is a flowchart of the process in which an improved real-time communication application provides high-quality audio to a remote listener when connected to a low-bitrate network, drawn according to the present invention.

[0025] Figure 4A is a flowchart of the process in which an improved encoder in an improved real-time communication application extracts ultra-wideband audio features, drawn according to the present invention.

[0026] Figure 4B is a flowchart of the process in which an improved encoder in an improved real-time communication application extracts wideband audio features, drawn according to the present invention.

[0027] Figure 5 is a flowchart of the process in which an improved encoder in an improved real-time communication application compresses audio features, drawn according to the present invention.

[0028] Figure 6 is a flowchart of the process in which an improved decoder in an improved real-time communication application decodes the received ultra-wideband data packets and obtains audio data for playback, drawn according to the present invention.

[0029] Figure 7 is a flowchart of the process in which an improved decoder in an improved real-time communication application decodes the received wideband data packets and obtains audio data for playback, drawn according to the present invention.

[0030] Figure 8 is a flowchart of the process in which an improved decoder in an improved real-time communication application dequantizes and decodes a set of ultra-wideband compressed audio features of a frame, drawn according to the present invention.

[0031] Those of ordinary skill in the art should understand that, for the sake of simply and clearly presenting the above-mentioned drawings, the components in the drawings are not necessarily drawn to scale, and the sizes of some components may be enlarged relative to other components to help understand the present invention. In addition, the specific order of certain elements, parts, components, modules, steps, operations, events, and / or processes described or illustrated in the present invention can also be changed in actual applications. Those of ordinary skill in the art should understand that, for the sake of simple and clear elaboration, useful and / or necessary elements that are well-known and easily understood in existing feasible implementation schemes may not be described in the present invention so as to clearly present various implementation schemes of the present invention. Detailed implementation manners

[0032] Figure 1 is a schematic block diagram of a real-time communication (RTC) system, which is generally denoted by 100. The RTC system includes a group of electronic communication devices, as shown by 102 and 104, which communicate with each other through a network (such as the Internet) 122. In one implementation manner, the network communication protocol adopted is the Transmission Control Protocol (TCP) and the Internet Protocol (IP) (collectively referred to as TCP / IP). Devices 102 - 104 are also referred to as participating devices in the present invention. Devices 102 - 104 are connected to the Internet 122 through a wireless or wired network (such as a Wi-Fi network and an Ethernet network).

[0033] Each of the communication devices 102 - 104 can be a laptop computer, a tablet computer, a smart phone, or other types of portable devices that can access the Internet 122 through a network connection. In Figure 2 Device 102 will be further described as an example for devices 102 - 104.

[0034] Figure 2 is a schematic block diagram of a wireless communication device 102. Device 102 includes a processor 202, a memory 204 with a certain capacity adapted to the processor 202, one or more user input interfaces 206 (such as a touchpad, a keyboard, a mouse, etc.) adapted to the processor 202, a voice input interface 208 (such as a microphone) adapted to the processor 202, a voice output interface 210 (such as a speaker) adapted to the processor 202, a video input interface 212 (such as a camera) adapted to the processor 202, a video output interface 214 (such as a display screen) adapted to the processor 202, and a network interface 216 (such as a Wi-Fi network interface) adapted to the processor 202 for connecting to the Internet 122. Device 102 also includes an operating system 220 running on the processor 202 (such as etc.). One or more computer software applications 222-224 are loaded and run on the device 102. The computer software applications 222-224 are compiled using one or more computer software programming languages (such as C, C++, C#, Java, etc.).

[0035] In one embodiment, the computer software application 222 is a real-time communication software application. For example, two or more people can use the application 222 to conduct an online meeting via the Internet 122. This real-time communication involves audio and / or video communication.

[0036] Back to Figure 1 In, the RTC devices 102-104 can be used to participate in an RTC session. Each of the RTC devices 222-224 runs an improved RTC application software 222, which includes a machine learning-based noise suppression module 112, an encoder 114, and a decoder 116. The voice input interface 208 of the device 102 captures the audio data 132 and sends it to other participating devices in the RTC session, such as the device 104. For a specific audio data 132, the device 102 is the sending device, that is, the sender; the device 104 is the receiving device, that is, the receiver. For the audio data captured by the device 104 and sent to the device 102, the device 104 is the sender, and the device 102 is the receiver. The encoder 114 and the decoder 116 are also collectively referred to as codecs in the present invention.

[0037] First, the audio data 132 is processed using the machine learning-based noise reduction module 112, and then the processed audio data is encoded by the new encoder 114. The encoded audio data is then sent to the device 104. The new decoder 116 processes the received audio data, and then the decoded audio data 134 is played on the voice output interface 210 of the device 104.

[0038] When the network connection between the devices 102-104 becomes slow due to various reasons (such as network congestion or packet loss, etc.) and is for low-bandwidth (i.e., low bitrate) transmission, the encoder 114 will run with a low-bitrate audio codec, while the decoder 116 will run with a high-quality decoder, thereby reducing the demand and requirement for network bandwidth while maintaining the quality of the audio data 134 received by the listener. Figure 3 It further illustrates the process of the improved RTC application 222 providing high-quality audio communication in a weak network situation.

[0039] Figure 3A flowchart showing the process by which the improved RTC application 222 provides high-quality audio using the new low-bitrate audio encoder 114 and the new high-quality decoder 116 when the network connection is at a low bitrate, with the overall process represented by 300. At 302, the RTC application 222 receives a stream of audio data 132. At 304, the machine learning-based noise suppression module 112 in the RTC application 222 processes the audio data 132 to suppress and reduce noise.

[0040] When there is noise in the audio data, the performance of traditional neural network-based vocoders will degrade. In particular, transitional noise will significantly reduce the clarity of the synthesized speech. Therefore, it is preferable to reduce or even eliminate the noise in the audio data before the encoding stage. Traditional noise suppression (NS) algorithms based on statistical methods are only effective when there is stable background noise. The improved RTC application 222 is equipped with a machine learning-based noise suppression (ML-NS) module 112 to reduce the noise in the audio data 132. The ML-NS module uses methods such as recurrent neural network (RNN) and / or convolutional neural network (CNN) algorithms to reduce the noise in the audio data 132.

[0041] The output of step 304 is also referred to as clean audio data in the present invention. Without performing step 304, the audio data 132 is also referred to as clean audio data in the present invention. At 306, the improved encoder 114 divides the clean audio data into a set of audio data frames. For example, the length of each frame in the set can be 5 milliseconds (ms) or 10 milliseconds.

[0042] At 308, the improved encoder 114 normalizes each frame within the set of data frames. The audio data in each frame is Pulse-code Modulation (PCM) data. The improved encoder 114 and decoder 116 operate in two modes: wideband mode and ultra-wideband mode. In one embodiment, at 308, the clean audio data is resampled to 16 kHz and 32 kHz for wideband mode and ultra-wideband mode respectively. Their bit rates are 2.1 kbps and 3.5 kbps respectively. Thus at 308, the improved encoder 114 decomposes the normalized PCM data of each frame into audio data of two subbands. In one embodiment, the lower-frequency subband of the audio data (also referred to as the low-frequency subband in the present invention) contains audio data with a sampling rate from 0 kHz to 16 kHz, while the higher-frequency subband (also referred to as the high-frequency subband in the present invention) contains audio data with a sampling rate from 16 kHz to 32 kHz. Thus, if divided into two subbands, each frame contains decomposed low-frequency subband audio data and decomposed high-frequency subband audio data. After running step 308, each frame is also referred to as a decomposed frame or a decomposed frame of audio data in the present invention. In one embodiment, the decomposition process is performed using Quadrature Mirror Filter (QMF). The QMF filter can also avoid spectral aliasing.

[0043] At 310, the improved encoder 114 extracts an audio feature set for each frame of the audio data. In ultra-wideband mode, the feature set includes 18 Bark Frequency Cepstral Coefficients (BFCCs), fundamental period, fundamental correlation of the low-frequency subband, Line Spectral Frequencies (LSFs) of the high-frequency subband, and the ratio of the sum of energies between the low-frequency subband audio data and the high-frequency subband audio data of each frame. In wideband mode, the feature set includes 18 BFCCs, fundamental period, and fundamental correlation. The feature vector retains the original waveform information with a smaller amount of data. Performing vector quantization methods can further reduce the amount of data of the feature vector. The present invention compresses the original PCM data by more than 95% while only losing a small part of the audio quality.

[0044] Figure 4A The audio feature extraction process in ultra-wideband mode at 310 is further illustrated. Figure 4AThe flowchart shows the process by which the encoder 114 extracts the audio features of each frame in the ultra-wideband mode, and the overall process is represented by 400. At 404, the improved encoder 114 pre-emphasizes the PCM data using a high-pass filter (such as an infinite impulse response (IIR) filter) to form pre-emphasized low-frequency subband audio data. At 406, the improved encoder 114 performs BFCC operations on the pre-emphasized low-frequency subband audio data. Additionally, at 406, the improved encoder 114 extracts pitch features such as the pitch period and pitch correlation from the low-frequency subband audio data. Since the LPC coefficient α can be predicted based on BFCC, only BFCC, the pitch period, and pitch correlation are explicitly represented in the feature vector. LPC refers to linear predictive coding.

[0045] At steps 408, 410, and 412, for each frame of the audio data, the improved encoder 114 operates on the higher-frequency subband audio data. At 408, the encoder 114 uses an algorithm such as the Burg's algorithm to calculate the LPC coefficient (such as α_h). At 410, the encoder 114 converts the LPC coefficient to line spectral frequencies (LSFs). At 412, the improved encoder 114 determines the ratio of the total energy between the low-frequency subband audio data and the high-frequency subband audio data for each frame. In one embodiment, the feature set includes the energy ratio of the two subbands. Thus, the audio feature vector for each frame includes BFCC, pitch, LSF, and the energy ratio between the two subbands. Steps 402 - 406 are collectively referred to in the present invention as extracting an audio feature set for one frame in the low-frequency subband of the audio data, and steps 408 - 412 are collectively referred to in the present invention as extracting an audio feature set for one frame in the high-frequency subband of the audio data. The audio features include the energy sum ratio and line spectral frequencies (LSFs), which are referred to as audio energy features and audio LPC features respectively in the present invention.

[0046] Figure 4B Further illustrates the audio feature extraction process at 310 in the wideband mode. At 422, the improved encoder 114 pre-emphasizes the PCM data using a high-pass filter (such as an infinite impulse response (IIR) filter) to form pre-emphasized audio data. At 424, the improved encoder 114 performs BFCC and pitch prediction operations including the calculation of the pitch period and pitch correlation on the pre-emphasized audio data.

[0047] Back to Figure 3, at 312, the improved encoder 114 compresses the set of audio features extracted for each frame using signal compression methods such as vector quantization and frame correlation methods. In one embodiment, the signal compression method employed is the differential vector quantization (DVQ) method. Alternatively, the residual vector quantization (RVQ) method can also be used as the signal compression method. In a further embodiment, an appropriate interpolation strategy is adopted for the compression operation. Figure 5 A further description of the compression process is provided.

[0048] Figure 5 The flowchart showing the process of the improved encoder 114 compressing a set of audio feature sets of a frame set is presented, and the overall process is denoted by 500. At 502, the improved encoder 114 compresses the audio feature sets of each key frame within the frame set using methods such as residual vector quantization (RVQ). In one embodiment, in each data packet, at least one frame is encoded using the RVQ method. Such a frame is referred to as a key frame (I-frame) in the present invention. Other frames are referred to as non-I frames, non-key frames, or other frames in the present invention. At 504, the improved encoder 114 uses methods such as interpolation to compress the audio feature sets of each non-I frame within the frame set.

[0049] The acoustic features of adjacent audio frames have strong local correlations. For example, the pronunciation of a phoneme typically spans several frames. Therefore, the feature vector of a non-I frame can be obtained from the feature vectors of its adjacent frames through interpolation. This operation can be achieved using interpolation methods such as differential vector quantization (DVQ) or polynomial interpolation. For example, in a data packet with 4 frames (i.e., 4 audio feature sets of 4 frames of audio data in the same data packet), only the 2nd and 4th frames are subjected to RVQ quantization operations. The 1st frame is inserted using the 2nd and 4th frames of the previous data packet, and the 3rd frame is inserted using the 2nd and 4th frames using the DVQ method. The encoding interpolation parameters require fewer data bits than the RVQ method. However, the interpolation method may not be as accurate as the RVQ method.

[0050] Looking back Figure 3 , at 314, the improved encoder 114 packs a set of compressed audio feature sets of the frame set into audio data packets. In one embodiment, each data packet contains 4 compressed audio feature sets corresponding to 4 frames of audio data. The following table shows an example of a data packet:

[0051] Example of a 40-millisecond (4-frame) data packet with bit field allocation

[0052]

[0053] In this example, the total number of bits in the data payload of a 40 - ms data packet is 140, corresponding to bit rates of 2.1 kbps and 3.5 kbps in broadband and ultra - wideband modes respectively. At 316, the RTC application 222 sends the data packet to the device 104 via the Internet 122. For example, the UDP protocol can be used to implement the transmission. The RTC application 222 running on the device 104 receives the data packet and processes it.

[0054] Figure 6 A flowchart showing the process of the improved decoder 116 decoding the received data packet in ultra - wideband mode and obtaining audio data for playback on the receiving device 116 is presented, and the overall process is denoted by 600. At 602, the improved decoder 116 receives the audio data packet sent by the sending device 102 at 316. After receiving the data packet, at 604, the improved decoder 116 obtains the audio feature set for each frame from the data packet. When the sub - bands are 0 kHz - 16 kHz and 16 kHz - 32 kHz, the sampling frequency range of the high - frequency sub - band is 16 kHz - 32 kHz, and the sampling frequency range of the low - frequency sub - band is the other frequency bands except this. In the high - frequency sub - band, the LPC coefficients and energy features (such as the ratio of energy sum between low - frequency and high - frequency sub - bands) can be directly obtained from the data packet.

[0055] Figure 8 The process of obtaining the audio feature set for each frame is further illustrated. Figure 8 A flowchart showing the process of the improved decoder 116 de - quantizing the compressed audio feature set of the data frame in ultra - wideband mode is presented. At 802, the improved decoder 116 obtains the audio features of the data frame from the data packet by performing the reverse operation of step 312, i.e., the de - quantization process, such as feature information like BFCC, pitch period and correlation, LSF, and energy ratio. At 804, the improved decoder 116 determines the LPC coefficients of the high - frequency sub - band of the audio data in the frame. At 806, the improved decoder 116 determines the LPC coefficients of the low - frequency sub - band according to the BFCC feature. In the description of the present invention, the audio features obtained at 802 are also referred to as the first subset of audio features; the audio features obtained at 804 are also referred to as the second subset of audio features; and the audio features obtained at 806 are also referred to as the third subset of audio features.

[0056] The total speech signal of each sub - band is decomposed into linear and non - linear parts. In one implementation, an LPC model is used to determine the linear prediction value, which takes the LPC coefficients as the audio feature input and generates the value in an autoregressive manner. The total speech signal of each sub - band at time t can be expressed as:

[0057]

[0058] where k is the order of the LPC model, and α i is the i-th LPC coefficient, s t-i is the previous i-th sample, and e t is the residual signal. The LPC coefficients are optimized by minimizing the excitation e t . As shown below, the first term represents the LPC prediction value:

[0059]

[0060] The above equation is used to predict the LPC prediction value in each sub-band at 606. While the neural network model can only predict the non-linear residual signal for the low-frequency sub-band at 612 and 614. In this way, the computational complexity can be significantly reduced while achieving high-quality speech generation.

[0061] Next, Figure 6 , at 606, within each sub-band, the linear prediction value of the next sample is determined for each sample of the audio data per frame based on the audio features. For example, the audio sample can be a PCM sample. In one implementation, the linear prediction value of each audio data sample is determined at 606. At 612, the improved decoder 116 extracts the context vector for the residual signal prediction at 614 from the acoustic feature vector.

[0062] Taking the audio features BFCC, pitch period, and correlation as inputs, step 612 is performed for each frame. Since the pitch period is an important feature for residual prediction, the pitch periods are first combined and then mapped to a larger feature space to enrich their representation. Then the pitch feature is concatenated with other acoustic features and input into a 1D convolutional layer. The convolutional layer has a larger perception threshold in the time dimension. After that, the output of the CNN layer is connected through a fully connected layer, and the fully connected layer is used as the output layer to obtain the final context vector c f (also referred to as c in the present invention l,f ). The context vector c f is an input to the residual prediction network and remains unchanged during the data generation process of the f-th frame.

[0063] At 614, the improved decoder 116 determines the prediction error (also referred to as the residual signal prediction value in the present invention). In other words, at 614, the improved decoder 116 performs the residual signal prediction. The residual signal e t is modeled and predicted through a neural network (also referred to as the residual prediction network) algorithm. The input features include the conditional network output vector c f , the current LPC prediction signal p t , and the non-linear residual signal e t and the full signal s tFinal prediction. To enrich the embedding of the signal, the signal is first converted to the mu-law domain and then mapped to a high-dimensional vector using a shared embedding matrix. The concatenated features are input into the RNN layer, followed by a fully connected layer. Thereafter, the softmax activation method is used to calculate the t probability distribution, restricting the value range of the signal in the asymmetric quantization pulse code modulation (PCM) domain, such as the μ-law or A-law domain. A sampling strategy is used instead of choosing the value with the maximum probability to select the t final value.

[0064] At 616, the improved decoder 116 combines the linear prediction value and the non-linear prediction error to generate the sub-band audio signal for each sample. The generated sub-band audio signal (s t ) is the sum of p t and e t . Since the low-frequency sub-band signal is emphasized during the encoding process, the output signal s t needs to be de-emphasized to obtain the original signal. Therefore, at 618, the improved decoder 116 de-emphasizes the generated low-frequency sub-band signal to restore the un-emphasized low-frequency sub-band audio signal. For example, if a high-pass filter is used to emphasize the PCM samples during encoding, a low-pass filter is used for the output signal to perform de-emphasis, which is also called de-emphasis in the present invention.

[0065] At 622, for the higher frequency sub-band signal, the residual signal is predicted using the following equation:

[0066]

[0067] where e h,t and e l,t are the residual signals in the high-frequency and low-frequency bands at time t. E h and E l are the energies of the current frame in the high-frequency and low-frequency bands respectively.

[0068] At 624, the improved decoder 116 combines the linear prediction value and the residual prediction value to generate a subband audio signal for each sample in the high-frequency subband. At 632, the improved decoder 116 combines the de-emphasized low-frequency subband audio signal generated at 618 and the subband audio signal of the high-frequency subband generated at 624, and uses an inverse Quadrature Mirror Filter (QMF) to generate audio data. Steps 622-624 are performed for each frame of audio feature set in the high-frequency subband audio data. For example, if a high-pass filter is used to emphasize the PCM samples during encoding, a low-pass filter is used to de-emphasize the output signal, which is also referred to as de-emphasis in the present invention. The generated audio data is also referred to as de-emphasized audio data or samples in the present invention, such as a waveform signal of 32 kHz. If the combined audio samples do not match the correct playback format, for example, if the format of the combined audio samples is 8-bit μ-law, it needs to be converted to 16-bit linear PCM format for playback on device 104. In this case, at 634, the improved decoder 116 converts the combined audio samples into audio data 134 for playback on device 104.

[0069] Figure 7 The flowchart shows the process of the improved decoder 116 decoding the received data packet in the wideband mode, and the overall process is represented by 700. At 702, the improved decoder 116 receives the audio data packet sent by the transmitting device 102 at 316. At 704, the improved decoder 116 performs the reverse process of step 312, that is, the inverse vector quantization process to obtain audio features, such as features like BFCC, pitch period, and pitch-related vector of the wideband audio data. At 706, the improved decoder 116 determines the LPC coefficients according to the BFCC features. Then the improved decoder 116 reconstructs the signal in an autoregressive manner. At 708, the improved decoder 116 calculates the prediction value of the current sample using the LPC coefficients and the previous 16 output signals. In one embodiment, the prediction value is a linear prediction value. At 710, a context vector is extracted using the BFCC and pitch features. At 712, the non-linear residual signal prediction is performed based on information such as the context vector, the current linear prediction value, the final output signal value, and the final prediction residual signal. At 714, the current signal is determined by summing the linear and non-linear residual prediction values. At 716, since the corresponding original signal was emphasized at 404 before, the output signal is de-emphasized here.

[0070] In view of the above description, it is obvious that the present invention can have many other modifications and variations. Therefore, it should be noted that within the scope of the appended claims, the present invention can be implemented in a manner different from the above specific description. For example, the residual prediction network can be implemented using some different designs. First, there are many variants of RNN, such as GRU, LSTM, SRU cells, etc. Second, directly predicting s t instead of predicting the residual signal e t , which is also an alternative. Third, batch sampling enables multiple samples to be predicted in a single time step. This method usually improves the decoding efficiency at the cost of reducing the audio quality. The residual signal e l,t is predicted using the above network, where the subscript l represents the low-frequency subband (h represents the high-frequency subband), and t is the time step. Therefore, the full signal s′ l,t is the sum of the LPC prediction p l,t and the residual signal e l,t . Then this value is input into the LPC module to predict p l,t+1 .

[0071] The above description of the present invention is for better illustration and explanation, and is not intended to be exclusive or to limit the present invention to the above specific form. The above description is for better explaining the principles of the present invention and the practical applications of these principles, so that those skilled in the relevant art can best utilize the present invention to implement various embodiments and make various modifications in the intended appropriate uses. It should be recognized that the words "a" or "an" in the present invention include both the singular and plural forms. At the same time, in appropriate cases, the situation of multiple elements mentioned in the present invention should also include its singular form.

[0072] The scope of the present invention is not limited to the content of the above specification alone, but is determined by the claims. In addition, although the claims proposed may have a relatively narrow scope, it should be recognized that the scope provided by the present invention is much broader than the scope proposed by the claims. We will propose claims with a larger scope in one or more applications claiming the priority of this application. If the parts disclosed in the above specification and drawings are not included within the scope of the claims, then these inventive contents are not publicly disclosed, and we reserve the right to file one or more patent applications for the above inventive contents in the future.

Claims

1. A computer operation method for providing high-quality audio playback in a low-bitrate network connection for real-time communication. The method is run by a real-time communication software application and includes: 1) Receiving an audio input data stream on a sending device; 2) Suppressing noise in the audio input data stream on the sending device to generate clean audio input data; 3) Splitting the clean audio input data into a set of audio data frames on the sending device; 4) Normalizing each frame in the set of frames on the sending device to generate a set of normalized audio data frames, where the audio data of the frames is resampled according to two frequency ranges corresponding to the wideband mode and the ultra-wideband mode, so as to form low-frequency subband audio data and high-frequency subband audio data; 5) Extracting an audio feature set from each frame in the set of normalized audio data frames on the sending device to form a set of audio feature sets; 6) Quantifying the audio feature set of each frame in the set of normalized audio data frames into a compressed audio feature set on the sending device; 7) Packing the compressed audio feature set into an audio data packet on the sending device; 8) Sending the audio data packet from the sending device to a receiving device; 9) Receiving the audio data packet in the ultra-wideband mode on the receiving device; 10) Obtaining the audio feature set of each frame in the set of normalized audio data frames from the audio data packet on the receiving device; 11) On the receiving device, within the low-frequency subband and the high-frequency subband in the ultra-wideband mode, determining the linear prediction value of the next sample for the audio data samples of each frame according to the audio feature set corresponding to the data frame; 12) On the receiving device, extracting a context vector for residual signal prediction from the acoustic feature vectors of the samples in the low-frequency subband; 13) On the receiving device, using a deep learning method to determine the first residual signal prediction value for the samples in the low-frequency subband; 14) Combining the linear prediction value and the first residual prediction value on the receiving device to generate a subband audio signal for the samples in the low-frequency subband; 15) Performing de-emphasis processing on the subband audio signal on the receiving device to form a de-emphasized low-frequency subband audio signal; 16) On the receiving device, determining the second residual prediction value of the samples in the high-frequency subband; 17) On the receiving device, combining the linear prediction value and the second residual prediction value to generate a subband audio signal for the samples in the high-frequency subband; 18) On the receiving device, merging the de-emphasized low-frequency subband audio signal and the subband audio signal of the samples in the high-frequency subband to form a merged audio sample; And 19) Converting the merged audio sample into audio data for playback on the receiving device.

2. The method according to claim 1, wherein Extracting an audio feature set for each frame in the set of normalized audio data frames in the ultra-wideband mode includes: 1) Pre-emphasize the low-frequency subband audio data using a high-pass filter to form pre-emphasized low-frequency subband audio data; 2) Perform Bark frequency cepstral coefficient operations on the pre-emphasized low-frequency subband audio data to extract the Bark frequency cepstral coefficient features of the audio, and perform pitch prediction processing on the pre-emphasized low-frequency subband audio data to extract audio pitch features, including fundamental period and fundamental correlation; 3) Calculate the audio linear prediction coding coefficients according to the high-frequency subband audio data; 4) Convert the linear prediction coding coefficients into line spectral frequency coefficients; and 5) Determine the ratio of the sum of energies between the low-frequency subband audio data and the high-frequency subband audio data, where the ratio of the sum of energies, line spectral frequency coefficients, audio pitch features, and audio Bark frequency cepstral coefficient features form part of the audio feature set.

3. The method according to claim 1, wherein, In the wideband mode, extracting an audio feature set for each frame in the set of normalized audio data frames includes: 1) Pre-emphasize the normalized audio data of each frame using a high-pass filter to form pre-emphasized normalized audio data; and 2) Perform Bark frequency cepstral coefficient operations on the pre-emphasized normalized audio data to extract the Bark frequency cepstral coefficient features of the audio, and perform pitch prediction processing on the pre-emphasized normalized audio data to extract audio pitch features including fundamental period and fundamental correlation, where the audio pitch features and the audio Bark frequency cepstral coefficient features form part of the audio feature set.

4. The method according to claim 1, wherein On the receiving device, obtaining the audio feature set for each frame in the set of normalized audio data frames in the audio data packet includes: 1) Perform an inverse quantization process on the compressed audio feature set to obtain the audio feature set; 2) Determine the linear prediction coding coefficients of the high-frequency subband according to the line spectral frequency coefficients; and 3) Determine the linear prediction coding coefficients of the low-frequency subband according to the Bark frequency cepstral coefficients.

5. The method according to claim 4, wherein the inverse quantization process uses an inverse differential vector quantization method, an inverse residual vector quantization method, or an inverse interpolation method.

6. The method according to claim 1, wherein the method of quantifying the audio feature set includes: 1) Compress the audio feature set of each key frame in the set of frames using a residual vector quantization method or a differential vector quantization method, where there is at least one key frame in the set of frames; and 2) Compress the audio feature set of each non-key frame in the set of frames using an interpolation method.

7. The method according to claim 1, wherein the two frequency ranges are 0 to 16 kHz and 16 kHz to 32 kHz respectively.

8. The method according to claim 1, wherein the noise is suppressed by a machine learning-based method.

9. A computer operation method for providing high-quality audio playback in a low-bitrate network connection for real-time communication, the method being run by a real-time communication software application, including: 1) Receive an audio input data stream on the sending device; 2) Suppress the noise in the audio input data stream on the sending device to generate clean audio input data; 3) Split the clean audio input data into a set of audio data frames on the sending device; 4) Normalize each frame in the set of frames on the sending device to generate a set of normalized audio data frames, where the audio data of the frames is resampled according to two frequency ranges corresponding to the wideband mode and the ultra-wideband mode, so as to form low-frequency subband audio data and high-frequency subband audio data; 5) On the sending device, extract an audio feature from each frame in the set of normalized audio data frames, thereby forming a set of audio feature sets; 6) Quantize the audio feature set of each frame in the set of normalized audio data frames on the sending device into a compressed audio feature set; 7) Pack the compressed audio feature set into an audio data packet on the sending device; 8) Send the audio data packet from the sending device to the receiving device; 9) Receive the audio data packet in the wideband mode on the receiving device; 10) Obtain the audio feature set of each frame in the set of frames by performing an inverse quantization process on the receiving device, where the audio feature set includes a set of Bark frequency cepstral coefficients; 11) On the receiving device, determine a set of linear prediction coding coefficients according to the set of Bark frequency cepstral coefficients; 12) On the receiving device, determine the linear prediction value of the next sample for each sample of the audio data of each frame in the set of frames according to the audio feature set; 13) On the receiving device, use a deep learning method to extract a context vector for residual signal prediction from the acoustic feature vector of the sample; 14) Determine the residual signal prediction value of the sample based on the context vector, the deep learning network, the linear prediction value, the final output signal value, and the final predicted residual signal; 15) Combine the linear prediction value and the residual signal prediction value to generate the audio signal of the sample; and 16) Perform de-emphasis processing on the audio signal of the sample to generate a de-emphasized audio signal for playing on the receiving device.

10. The method according to claim 9, wherein In the ultra-wideband mode, extracting an audio feature set from each frame in the set of normalized audio data frames includes: 1) Perform pre-emphasis processing on the low-frequency subband audio data using a high-pass filter to form pre-emphasized low-frequency subband audio data; 2) Run Bark frequency cepstral coefficient calculation on the pre-emphasized low-frequency subband audio data to extract audio Bark frequency cepstral coefficient features, and perform pitch prediction processing on the pre-emphasized low-frequency subband audio data to extract audio pitch features, where the audio pitch features include fundamental period and pitch correlation; 3) Calculate audio linear prediction coding coefficients according to the high-frequency subband audio data; 4) Convert the linear prediction coding coefficients to line spectral frequency coefficients; and 5) Determine the ratio of the sum of energies between the low-frequency subband audio data and the high-frequency subband audio data, where the ratio of the sum of energies, the line spectral frequency coefficients, the audio pitch characteristics, and the audio Bark frequency cepstral coefficient characteristics form part of the audio feature set.

11. The method according to claim 9, wherein In the wideband mode, extracting an audio feature set from each frame in the set of normalized audio data frames includes: 1) Pre-emphasize the normalized audio data of each frame using a high-pass filter to form pre-emphasized normalized audio data; and 2) Run Bark frequency cepstral coefficient calculation on the pre-emphasized normalized audio data to extract the audio Bark frequency cepstral coefficient characteristics, and perform pitch prediction processing on the pre-emphasized normalized audio data to extract audio pitch characteristics including the fundamental period and fundamental correlation, where the audio pitch characteristics and the audio Bark frequency cepstral coefficient characteristics form part of the audio feature set.

12. The method according to claim 9, wherein the inverse quantization process employs an inverse differential vector quantization method, an inverse residual vector quantization method, or an inverse interpolation method.

13. The method according to claim 9, wherein the method of quantifying the audio feature set includes: 1) Compress the audio feature set of each key frame in the set of frames using a residual vector quantization method or a differential vector quantization method, where there is at least one key frame in the set of frames; and 2) Compress the audio feature set of each non-key frame in the set of frames using an interpolation method.

14. The method according to claim 9, wherein the two frequency ranges are 0 to 16 kHz and 16 kHz to 32 kHz, respectively.

15. The method according to claim 9, wherein noise suppression employs a machine learning-based method.

Citation Information

Patent Citations

  • Coding and decoding device and method of ultralow-bit-rate speech

    CN103325375A

  • Signal encoding method and apparatus and signal decoding method and apparatus

    CN107077855A