Bluetooth low-bit-rate audio encoding and decoding method and related device

Through dynamic wavelet processing, FSQ, lightweight Transformer model and forward error correction code, traditional Bluetooth audio codec has solved the problem of deterioration of sound quality and insufficient anti-interference capabilities at low code rates, and achieved efficient, low latency and low power audio transmission.

CN120199259APending Publication Date: 2025-06-24GUANGDONG UNIV OF TECH +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510461353.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Traditional Bluetooth audio encoding and decoding methods have problems such as deterioration of sound quality, insufficient anti-interference capability, high latency and large power consumption in low-code rate scenarios, which are difficult to meet the needs of modern audio transmission.

Method used

Dynamic wavelet processing, layered finite state quantization (FSQ) and lightweight Transformer model are adopted, combining forward error correction code and psychoacoustic postprocessing, and dynamically adjust the bit rate to adapt to Bluetooth real-time bandwidth.

Benefits of technology

It realizes high-quality audio transmission at low bitrates, reduces latency and power consumption, improves anti-interference ability and sound quality performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199259A_ABST
    Figure CN120199259A_ABST
Patent Text Reader

Abstract

The invention discloses a Bluetooth low-bit-rate audio encoding and decoding method and a related device, and the method comprises the steps: obtaining a Bluetooth real-time bandwidth, carrying out the dynamic wavelet processing of an input audio signal based on the Bluetooth real-time bandwidth, and decomposing to obtain a plurality of initial sub-band signals; quantizing the initial sub-band signal, and performing layered compression on the quantized initial sub-band signal through FSQ to obtain a compressed sub-band signal; performing differential coding processing on the compressed sub-band signal to obtain coded data; inputting the coded data into a lightweight Transform model to obtain a residual error correction item, and adding a forward error correction code to the coded data to obtain an audio data packet; packaging the residual error correction item and the audio data packet into a bit stream, and transmitting the bit stream to a decoding end through Bluetooth; and at a decoding end, performing decoding processing according to the bit stream, and recovering to obtain the target audio signal. The invention provides a high-efficiency, low-delay and low-power-consumption solution for Bluetooth audio transmission.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio processing, and in particular, to a Bluetooth low-bitrate audio encoding and decoding method and related devices. Background Art

[0002] With the wide application of Bluetooth audio technology, especially in portable audio devices such as Bluetooth headsets and Bluetooth speakers, users have put forward higher requirements for the sound quality, latency, and power consumption of audio transmission. However, traditional Bluetooth audio encoding and decoding methods have many problems in low-bitrate scenarios and are difficult to meet the needs of modern audio transmission.

[0003] Traditional Bluetooth audio encoding and decoding methods usually adopt fixed encoding strategies, such as SBC (Sub-Band Coding) or AAC (Advanced Audio Coding), etc. These methods have serious sound quality degradation at high compression rates, mainly manifested as the loss of high-frequency details, blurred low frequencies, and limited dynamic range of the overall audio. For example, SBC is prone to obvious distortion at low bitrates, especially in high-dynamic-range music or voice signals, and the sound quality drops significantly. In addition, the anti-interference ability of traditional encoding and decoding methods is insufficient and is easily affected by Bluetooth channel noise and packet loss, resulting in audio transmission interruption or increased distortion. In terms of latency, the end-to-end latency of traditional Bluetooth audio encoding and decoding methods is usually high and difficult to meet the needs of real-time audio interaction. For example, in scenarios such as games and video calls on Bluetooth audio devices, excessive latency will cause audio-video out-of-sync, seriously affecting the user experience. In addition, traditional methods have high power consumption, especially during high-resolution audio transmission, which will significantly shorten the battery life of the device and limit its application in portable devices. Summary of the Invention

[0004] The present invention provides a Bluetooth low-bitrate audio encoding and decoding method and related devices, which are used to solve the technical problems that existing audio encoding and decoding methods have serious sound quality degradation and insufficient anti-interference ability at high compression rates and are difficult to meet the needs of modern audio transmission.

[0005] A Bluetooth low-bitrate audio encoding and decoding method provided by the present invention, the method includes:

[0006] Obtain the Bluetooth real-time bandwidth, and perform dynamic wavelet processing on the input audio signal based on the Bluetooth real-time bandwidth to decompose it into multiple initial sub-band signals;

[0007] Quantize the initial sub-band signals, and perform hierarchical compression on the quantized initial sub-band signals through FSQ to obtain compressed sub-band signals;

[0008] Perform differential encoding processing on the compressed subband signal to obtain encoded data; input the encoded data into a lightweight Transformer model to obtain a residual correction term, and add a forward error correction code to the encoded data to obtain an audio data packet;

[0009] Encapsulate the residual correction term and the audio data packet into a bitstream, and transmit the bitstream to the decoding end via Bluetooth;

[0010] At the decoding end, perform decoding processing according to the bitstream to recover the target audio signal.

[0011] Further, the step of obtaining the Bluetooth real-time bandwidth and performing dynamic wavelet processing on the input audio signal based on the Bluetooth real-time bandwidth to decompose into multiple initial subband signals includes:

[0012] Perform pre-emphasis filtering processing on the input audio signal through a pre-emphasis filter to obtain a filtered audio signal;

[0013] Obtain the Bluetooth real-time bandwidth. If the Bluetooth real-time bandwidth exceeds the first preset bandwidth threshold, perform three-level wavelet decomposition on the filtered audio signal through a Le Gall filter; otherwise, perform two-level wavelet decomposition on the filtered audio signal through a Le Gall filter to decompose into multiple initial subband signals;

[0014] The decomposition process of the Le Gall filter is as follows:

[0015]

[0016] Where: represents the high-frequency subband coefficient, represents the low-frequency subband coefficient, and represent the sample points separated from the input audio signal according to odd and even indices.

[0017] Further, the step of quantifying the initial subband signal and performing hierarchical compression on the quantified initial subband signal through FSQ to obtain a compressed subband signal includes:

[0018] Calculate the energy value of each initial subband signal, and set a differential threshold based on the maximum energy value;

[0019] Judge whether the energy value of each initial subband signal is less than the differential threshold. If it is less, perform coarse quantization of 4 bit / sample on the initial subband signal; otherwise, perform high-precision quantization of 8 bit / sample on the initial subband signal; during the quantization process, if the Bluetooth real-time bandwidth is less than the second preset bandwidth threshold, discard the initial subband signals exceeding the preset frequency band;

[0020] The initial subband signal after quantization is hierarchically compressed by FSQ to obtain a compressed subband signal.

[0021] Further, the step of hierarchically compressing the quantized initial subband signal by FSQ to obtain a compressed subband signal includes: hierarchically compressing the quantized initial subband signal by FSQ based on a preset hierarchical finite state quantization codebook strategy to obtain a compressed subband signal; under the preset hierarchical finite state quantization codebook strategy, designing corresponding codebook sizes and quantization step sizes for the quantized initial subband signals in different frequency bands, and dynamically adjusting the quantization codebook and quantization step sizes in combination with the state transition rules.

[0022] Further, the steps of performing differential encoding processing on the compressed subband signal to obtain encoded data, inputting the encoded data into a lightweight Transformer model to obtain a residual correction term, and adding a forward error correction code to the encoded data to obtain an audio data packet include:

[0023] Performing differential encoding processing on the compressed subband signal to calculate the inter-frame difference of the compressed subband signal and generating encoded data;

[0024] Inputting the encoded data into a lightweight Transformer model, and using a sparse self-attention mechanism in the lightweight Transformer model to calculate the inter-frame correlation of the encoded data to obtain a residual correction term;

[0025] Adding a forward error correction code to the encoded data to obtain an audio data packet.

[0026] Further, the step of, at the decoding end, performing decoding processing according to the bitstream to recover the target audio signal includes:

[0027] At the decoding end, performing decoding error correction on the bitstream to obtain an error-corrected subband signal and a residual correction term; performing inverse differential calculation on the error-corrected subband signal to obtain a reconstructed signal;

[0028] Superposing the reconstructed signal and the residual correction term obtained by decoding error correction to generate a restored subband signal;

[0029] Performing reconstruction processing through the restored subband signal to recover the target audio signal.

[0030] Further, the step of performing reconstruction processing through the restored subband signal to recover the target audio signal includes:

[0031] Performing inverse quantization operation on the restored subband signal by FSQ to restore the audio subband coefficients;

[0032] Perform inverse wavelet transform on the audio subband coefficients to reconstruct a time-domain signal;

[0033] Perform error propagation suppression operation on the reconstructed time-domain signal, and perform psychoacoustic post-processing on the time-domain signal after the error propagation suppression operation to recover the target audio signal.

[0034] The present invention also provides a Bluetooth low-bitrate audio codec system, and the system includes:

[0035] A decomposition unit, configured to obtain the Bluetooth real-time bandwidth, and perform dynamic wavelet processing on the input audio signal based on the Bluetooth real-time bandwidth to decompose a plurality of initial subband signals;

[0036] A quantization unit, configured to quantize the initial subband signals, and perform hierarchical compression on the quantized initial subband signals through FSQ to obtain compressed subband signals;

[0037] An encoding unit, configured to perform differential encoding processing on the compressed subband signals to obtain encoded data; input the encoded data into a lightweight Transformer model to obtain a residual correction term, and add a forward error correction code to the encoded data to obtain an audio data packet;

[0038] An encapsulation unit, configured to encapsulate the residual correction term and the audio data packet into a bitstream, and transmit the bitstream to the decoding end through Bluetooth;

[0039] A decoding unit, configured to perform decoding processing on the bitstream at the decoding end to recover the target audio signal.

[0040] Further, the system further includes: a heterogeneous computing platform having an STM32H743 microcontroller and an FPGA chip.

[0041] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of any one of the above-mentioned Bluetooth low-bitrate audio codec methods are implemented.

[0042] It can be seen from the above technical solutions that the present invention has the following advantages:

[0043] The present invention provides a Bluetooth low bitrate audio encoding and decoding method and related devices. The method includes: obtaining the Bluetooth real-time bandwidth, performing dynamic wavelet processing on the input audio signal based on the Bluetooth real-time bandwidth, and decomposing it to obtain multiple initial subband signals; quantifying the initial subband signals, and performing hierarchical compression on the quantified initial subband signals through FSQ to obtain compressed subband signals; performing differential encoding processing on the compressed subband signals to obtain encoded data; inputting the encoded data into a lightweight Transformer model to obtain a residual correction term, and adding a forward error correction code to the encoded data to obtain an audio data packet; encapsulating the residual correction term and the audio data packet into a bitstream, and transmitting the bitstream to the decoding end through Bluetooth; at the decoding end, performing decoding processing according to the bitstream to recover the target audio signal. The present invention provides an efficient, low-latency, and low-power solution for Bluetooth audio transmission, solving the technical problems that existing audio encoding and decoding methods have serious deterioration of sound quality and insufficient anti-interference ability at high compression rates, and are difficult to meet the requirements of modern audio transmission. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0045] Figure 1 It is a flowchart of the steps of a Bluetooth low bitrate audio encoding and decoding method provided by an embodiment of the present invention;

[0046] Figure 2 It is an encoding and decoding schematic diagram of a Bluetooth low bitrate audio encoding and decoding method provided by an embodiment of the present invention;

[0047] Figure 3 It is a schematic diagram of the level selection of dynamic wavelet decomposition provided by an embodiment of the present invention;

[0048] Figure 4 It is a structural block diagram of a Bluetooth low bitrate audio encoding and decoding system provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] Embodiments of the present invention provide a Bluetooth low bitrate audio encoding and decoding method and related devices, which are used to solve the technical problems that existing audio encoding and decoding methods have serious deterioration of sound quality and insufficient anti-interference ability at high compression rates, and are difficult to meet the requirements of modern audio transmission.

[0050] To make the objectives, features, and advantages of the present invention more apparent and understandable, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the embodiments described below are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0051] Please refer to Figure 1 and Figure 2 , the present invention provides a Bluetooth low bitrate audio encoding and decoding method, which includes:

[0052] Step 101, obtain the Bluetooth real-time bandwidth, and perform dynamic wavelet processing on the input audio signal based on the Bluetooth real-time bandwidth to decompose it into multiple initial subband signals.

[0053] In this step, first perform pre-emphasis filtering on the input audio signal through a pre-emphasis filter to obtain a filtered audio signal; then obtain the Bluetooth real-time bandwidth. If the Bluetooth real-time bandwidth exceeds the first preset bandwidth threshold, perform three-level wavelet decomposition on the filtered audio signal through a LeGall filter; otherwise, perform two-level wavelet decomposition on the filtered audio signal through a Le Gall filter to decompose it into multiple initial subband signals.

[0054] Among them, the pre-emphasis filter specifically uses a first-order high-pass filter, and its transfer function is:

[0055]

[0056] where is the filter coefficient, which is used to control the high-frequency enhancement amplitude; is the unit delay operator, which represents that the signal is delayed by one sampling period; this transfer function is used to enhance the high-frequency components above 2 kHz in the signal and compensate for the possible high-frequency loss caused by subsequent processing (such as wavelet decomposition and quantization).

[0057] Then, the relationship between the input signal and the output signal of the transfer function is expressed as:

[0058]

[0059] In the formula: represents the output signal, represents the input signal of the current frame, represents the input signal of the previous frame.

[0060] Exemplarily, the input audio signal is an audio signal with a sampling rate of 16 kHz and a depth of 24 bits, and the dynamic range is 96 dB. The signal is framed, and the input audio signal is pre-emphasized and filtered, and the signal is cut into 20 ms frames (320 samples / frame), with 10 ms overlap between frames to reduce spectral leakage.

[0061] Specifically, the processing process of wavelet decomposition includes: dynamic level selection, integer arithmetic implementation, and sub-band division.

[0062] Please refer to Figure 3 , the dynamic level selection is to select the wavelet decomposition level according to the Bluetooth real-time bandwidth. Taking the first preset bandwidth threshold set to 256 kbps as an example, when the Bluetooth real-time bandwidth is used, 3-level decomposition is adopted, otherwise 2-level decomposition is adopted;

[0063] The integer arithmetic uses the 5 / 3 lifting wavelet, that is, the Le Gall filter (Le Gall Filter) for decomposition. The LeGall filter is a digital filter, and the decomposition process of the Le Gall filter is as follows:

[0064]

[0065] In the formula: represents the high-frequency sub-band coefficient, represents the low-frequency sub-band coefficient, and represent the sample points separated from the input audio signal according to odd and even indexes.

[0066] It can be understood that the 5 / 3 lifting wavelet (Le Gall filter) is used for integer arithmetic to implement wavelet lifting, avoiding the complexity and increased latency caused by floating-point calculations; if the input signal is , then when n = 0, we get:

[0067]

[0068] The sub-band division is specifically as follows: when performing 3-level decomposition, the output low-frequency (LL) sub-band is 0 - 2 kHz (fundamental frequency and timbre), the middle-frequency (LH) sub-band is 2 - 4 kHz (human voice and harmonics), and the high-frequency (HL / HH) sub-band is 4 - 8 kHz (transient details). When performing 2-level decomposition, the output low-frequency sub-band is 0 - 4 kHz and the high-frequency sub-band is 4 - 8 kHz, that is, the middle frequency is combined with the high frequency above 4 kHz.

[0069] It should be noted that traditional wavelet transforms (such as the discrete wavelet transform DWT) use a fixed decomposition level, unable to dynamically adjust the frequency domain resolution according to the channel bandwidth, and with high floating-point operation complexity, resulting in increased latency and power consumption. In contrast, the present invention dynamically selects the wavelet decomposition level (3 levels or 2 levels) according to the Bluetooth real-time bandwidth (such as 256 kbps / 128 kbps), preferentially retains the key low-frequency subbands (0 - 2 kHz), and trims the high-frequency subbands when the bandwidth is insufficient for dynamic level adjustment. At the same time, the present invention uses integer lifting wavelets and implements integer operations using the 5 / 3 lifting wavelet (Le Gall wavelet), avoiding floating-point calculations and increasing the operation speed by 40%, which is suitable for embedded devices.

[0070] Step 102: Quantize the initial subband signals and perform hierarchical compression on the quantized initial subband signals through FSQ to obtain compressed subband signals.

[0071] This step specifically includes:

[0072] Sub-step 1021: Calculate the energy value of each initial subband signal and set a differential threshold based on the maximum energy value.

[0073] Among them, the process of the energy value of the initial subband signal is as follows:

[0074]

[0075] In the formula: represents the energy value of the b-th initial subband signal, represents the i-th sample value of the b-th initial subband signal.

[0076] In this embodiment, 0.1 × the maximum subband energy is used as the differential threshold, and the energy values of each initial subband signal are compared with the differential threshold to select the subsequent signal quantization process.

[0077] Sub-step 1022: Determine whether the energy value of each initial subband signal is less than the differential threshold. If it is less, perform coarse quantization of 4 bits / sample on the initial subband signal; otherwise, perform high-precision quantization of 8 bits / sample on the initial subband signal. During the quantization process, if the Bluetooth real-time bandwidth is less than the second preset bandwidth threshold, discard the initial subband signals exceeding the preset frequency band. Among them, the preset frequency band can be set to 8 kHz.

[0078] Sub-step 1023: Perform hierarchical compression on the quantized initial subband signals through FSQ to obtain compressed subband signals.

[0079] In this sub-step, based on the preset hierarchical finite-state quantization codebook strategy, the FSQ is used to perform hierarchical compression on the quantized initial subband signal to obtain the compressed subband signal. Under the preset hierarchical finite-state quantization codebook strategy, corresponding codebook sizes and quantization steps are designed for the quantized initial subband signals in different frequency bands, and the quantization codebook and quantization step are dynamically adjusted in combination with the state transition rules. For details, please refer to Table 1.

[0080] Table 1 Preset Hierarchical Finite-State Quantization Codebook Strategy

[0081]

[0082] Specifically, a psychoacoustic model is used to design the state transition rules to optimize the quantization strategy. The state transition rules include: if the quantization error of three consecutive samples in the middle-frequency subband > 30 dB, the step size is reduced to ; if a high-frequency transient signal is detected in the high-frequency subband, the codebook is temporarily extended to 128 states. Among them, the above transient detection algorithm is to calculate the energy difference between adjacent frames in the high-frequency subband , if , it is determined as a transient signal; during the compression process, the sample x is uniformly quantized: , where q represents the quantization value obtained after the sample x undergoes the quantization operation; round( ) represents the rounding function.

[0083] Step 103: Perform differential encoding on the compressed subband signal to obtain the encoded data; input the encoded data into the lightweight Transformer model to obtain the residual correction term, and add a forward error correction code to the encoded data to obtain the audio data packet.

[0084] This step specifically includes:

[0085] Sub-step 1031: Perform differential encoding on the compressed subband signal to calculate the inter-frame difference of the compressed subband signal and generate the encoded data;

[0086] It can be understood that calculating the inter-frame difference can reduce data redundancy, and the quantity can be reduced by 30%. Among them, the calculation process of the inter-frame difference is:

[0087]

[0088] Among them, if the reconstructed value of the previous frame = 50 and the previous frame = 55, then = 5.

[0089] Sub-step 1032: Input the encoded data into the lightweight Transformer model, and use the sparse self-attention mechanism in the lightweight Transformer model to calculate the inter-frame correlation of the encoded data to obtain the residual correction term.

[0090] Specifically, use the lightweight Transformer for model training. Its sparse self-attention only calculates the local correlation between the current frame and the previous 5 frames. The attention matrix is a banded structure (bandwidth = 5). The corresponding encoding method specifically adopts binary encoding, using 1 bit to represent the parity frame position (0 = odd frame, 1 = even frame), and the computational complexity is reduced by 80%. The residual correction term is output through 4 layers of Transformer. The residual correction term compensates for the distortion or loss caused by quantization or transmission, thereby suppressing long-term error propagation.

[0091] Sub-step 1033: Add a forward error correction code to the encoded data to obtain an audio data packet.

[0092] Among them, forward error correction adds a Reed-Solomon (15, 9) error correction code to the key low-frequency subbands, that is, each codeword contains 15 symbols (9 data symbols + 6 parity symbols), and can correct up to 3 symbol errors within each codeword. At a 10% packet loss rate, 98% of the low-frequency data can be recovered.

[0093] Step 104: Package the residual correction term and the audio data packet into a bitstream, and transmit the bitstream to the decoding end via Bluetooth.

[0094] Specifically, when packaging, dual-mode compatibility needs to be satisfied, that is, package according to the format that supports the Bluetooth classic mode (A2DP) and the LE Audio LC3 protocol. Among them, the A2DP mode is packaged according to the SBC specification (4 subbands / 8 block groups); the LE Audio LC3 mode uses a 10ms frame length, and the dynamic bit rate field occupies 2 bits.

[0095] In addition, the encoding end and the decoding end transmit structured data streams through the Bluetooth transmission protocol to ensure synchronization when dynamically adjusting parameters and signal reconstruction. The two parties exchange channel quality information (such as received signal strength RSSI, bit error rate BER, and channel occupancy rate, etc.) through the Bluetooth HCI layer. The encoding end dynamically adjusts the bit rate, and the decoding end pre-allocates resources accordingly.

[0096] Among them, the decoding end obtains the RSSI and BER through the Bluetooth HCI layer for the received bitstream to monitor the channel quality, thereby realizing the adjustment of the dynamic bit rate (64kbps / 128kbps / 256kbps), so as to smoothly switch the bit rate to avoid sudden changes in sound quality. The bit rate strategy corresponding to the channel quality monitored by the decoding end is shown in Table 2.

[0097] Table 2 Dynamic Bit Rate Adjustment Strategy

[0098]

[0099] Step 105. At the decoding end, perform decoding processing according to the bitstream to recover the target audio signal.

[0100] This step specifically includes:

[0101] Sub-step 1051. At the decoding end, perform decoding error correction on the bitstream to obtain an error-corrected subband signal and a residual correction term; perform inverse difference calculation on the error-corrected subband signal to obtain a reconstructed signal;

[0102] Among them, for decoding error correction, forward error correction (FEC) is specifically used to add Reed-Solomon (15, 9) error correction codes to the low-frequency subband data in the bitstream, that is, 6 check blocks are generated for every 9 data blocks, and the error correction ability is 3 symbols / codes, so as to protect the low-frequency key data of the bitstream.

[0103] Sub-step 1052. Superimpose the reconstructed signal and the residual correction term to generate a restored subband signal;

[0104] Among them, the superimposing process is shown in the following formula:

[0105]

[0106] At the decoding end, by superimposing the residual correction term on the reconstructed signal, these error correction terms can be used to compensate for the distortion or loss caused by quantization or transmission.

[0107] Sub-step 1053. Perform reconstruction processing through the restored subband signal to recover the target audio signal.

[0108] In this sub-step, first perform inverse quantization operation on the restored subband signal through FSQ, so as to restore the restored subband signal from the quantization state to the subband coefficient state, and restore the audio subband coefficients; then perform inverse wavelet transform on the audio subband coefficients to reconstruct the time-domain signal; finally, perform error propagation suppression operation on the reconstructed time-domain signal, and perform psychoacoustic post-processing on the time-domain signal after the error propagation suppression operation to recover the target audio signal.

[0109] In the error propagation suppression operation, when the signal continuously loses frames , reconstruct the signal through median filtering. For example, when the historical frame values are [0.5, 0.7, 0.6], take the median value, that is = 0.7.

[0110] During the psychoacoustic post - processing, the mid - high - frequency harmonics are enhanced by an IIR filter, that is, a second - order band - pass filter is designed in the range of 2 - 4 kHz with a gain of +3 dB to compensate for the high - frequency details (such as vocal clarity) lost due to quantization. At the same time, based on the Zwicker model, noise below the masking threshold is added in the low - frequency strong - sound region to mask the quantization distortion.

[0111] Specifically, calculate the masking threshold in the low - frequency strong - sound region (such as 50 - 200 Hz). , and inject Gaussian noise. The corresponding power spectral density is expressed as:

[0112]

[0113] In the formula: represents the power spectral density of the injected Gaussian noise at frequency , is the power spectral density of the original audio signal at frequency . - 10 is an offset, indicating that in order to ensure the masking effect, the injected noise should be 10 dB lower than the theoretical masking threshold, providing a certain safety margin for the subsequently generated target audio signal to ensure that the masking noise itself cannot be detected by the human ear.

[0114] At a bitrate of 128 kbps, the present invention can achieve end - to - end delay , high - fidelity audio transmission of Perceptual Evaluation of Speech Quality (PESQ) , providing an efficient, low - delay and low - power solution for Bluetooth audio transmission.

[0115] The present invention provides a Bluetooth low - bitrate audio encoding and decoding method, aiming to solve the problems of deteriorated sound quality, insufficient anti - interference ability, high delay and high power consumption in the traditional Bluetooth audio encoding and decoding in low - bitrate scenarios. The present invention has the following advantages:

[0116] 1. By dynamically adjusting the wavelet decomposition level to adapt to the real - time bandwidth, preferentially retaining the key low - frequency sub - bands and cropping the high - frequency sub - bands according to the bandwidth, and at the same time using integer - lifting wavelets to avoid floating - point calculations, the operation speed is significantly improved and the power consumption is reduced, thus solving the problems of the traditional wavelet transform (such as the discrete wavelet transform DWT) using a fixed decomposition level, being unable to dynamically adjust the frequency - domain resolution according to the channel bandwidth, and having a high floating - point operation complexity, resulting in increased delay and power consumption.

[0117] 2. In combination with a hierarchical finite-state quantizer (FSQ), the quantization step is designed according to different subbands, and the quantization strategy is optimized using a psychoacoustic model to make the noise power spectrum negatively correlated with the auditory sensitivity, thereby achieving high-quality audio transmission at a low bit rate. This solves the problems of significant noise at low bit rates in traditional uniform quantization, the lack of combination with a psychoacoustic model, and the ineffective elimination of inter-frame redundancy.

[0118] 3. Inter-frame redundancy is eliminated through differential coding. Combining the sparse self-attention mechanism of a lightweight Transformer, only the local correlation between the current frame and the previous 5 frames is calculated, significantly reducing the computational amount and generating an error correction term to suppress long-term error propagation. This solves the problems of high computational and memory overhead caused by the global self-attention mechanism in traditional Transformers and the easy introduction of attention noise.

[0119] 4. The encoder and decoder are coordinated and adapted. Reed-Solomon error-correcting codes are used to protect key low-frequency data, and the bit rate is dynamically adjusted according to the channel quality to ensure transmission stability. This solves the problems of serious error propagation and a significant drop in the MOS score in traditional encoders (such as SBC) in case of packet loss or channel fluctuations.

[0120] 5. At the decoder side, through psychoacoustic post-processing, an IIR filter is used to compensate for the loss of high-frequency details and inject masking noise to cover quantization distortion, further optimizing the sound quality performance. This solves the problems that traditional denoising algorithms (such as spectral subtraction) are prone to introducing "musical noise" and do not specifically enhance key frequency bands.

[0121] Please refer to Figure 4 For this, the present invention also provides a Bluetooth low-bit-rate audio encoding and decoding system, which includes:

[0122] A decomposition unit 201, configured to obtain the Bluetooth real-time bandwidth, perform dynamic wavelet processing on the input audio signal based on the Bluetooth real-time bandwidth, and decompose it into multiple initial subband signals;

[0123] A quantization unit 202, configured to quantize the initial subband signals, and perform hierarchical compression on the quantized initial subband signals through FSQ to obtain compressed subband signals;

[0124] An encoding unit 203, configured to perform differential coding processing on the compressed subband signals to obtain encoded data; input the encoded data into a lightweight Transformer model to obtain a residual correction term, and add a forward error correction code to the encoded data to obtain an audio data packet;

[0125] An encapsulation unit 204, configured to encapsulate the residual correction term and the audio data packet into a bit stream, and transmit the bit stream to the decoder side through Bluetooth;

[0126] A decoding unit 205 is configured to perform decoding processing according to a bitstream at a decoding end to recover a target audio signal.

[0127] Furthermore, the system further includes a heterogeneous computing platform having an STM32H743 microcontroller and an FPGA chip. Among them, the FPGA chip is used to implement the acceleration control of wavelet integer operation and FSQ state machine (clock synchronization control). Using this system can achieve end-to-end delay: (Encoding 10ms + Transmission 12ms); Power consumption: Total power consumption of encoding and decoding .

[0128] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above Bluetooth low-bitrate audio encoding and decoding methods are implemented.

[0129] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described system, device, and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0130] In several embodiments provided by the present application, it should be understood that the disclosed system, device, and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.

[0131] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0132] In addition, each functional unit in various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0133] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0134] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A Bluetooth low bit rate audio encoding and decoding method, characterized in that: The method comprises: Acquire a Bluetooth real-time bandwidth, perform dynamic wavelet processing on an input audio signal based on the Bluetooth real-time bandwidth, and decompose the input audio signal to obtain a plurality of initial sub-band signals; quantizing the initial sub-band signal, and performing hierarchical compression on the quantized initial sub-band signal through FSQ to obtain a compressed sub-band signal; Performing differential coding processing on the compressed subband signal to obtain coded data; inputting the coded data into a lightweight Transformer model to obtain a residual correction term, and adding a forward error correction code to the coded data to obtain an audio data packet; Encapsulating the residual correction term and the audio data packet into a bit stream, and transmitting the bit stream to a decoding end via Bluetooth; At the decoding end, decoding processing is performed according to the bit stream to restore the target audio signal.

2. The Bluetooth low bit rate audio encoding and decoding method according to claim 1, characterized in that: The step of obtaining the Bluetooth real-time bandwidth, performing dynamic wavelet processing on the input audio signal based on the Bluetooth real-time bandwidth, and decomposing to obtain a plurality of initial sub-band signals comprises: Performing pre-emphasis filtering on the input audio signal through a pre-emphasis filter to obtain a filtered audio signal; Acquire the Bluetooth real-time bandwidth, and if the Bluetooth real-time bandwidth exceeds a first preset bandwidth threshold, perform a three-level wavelet decomposition on the filtered audio signal through a Le Gall filter; otherwise, perform a two-level wavelet decomposition on the filtered audio signal through a Le Gall filter to obtain a plurality of initial sub-band signals; The decomposition process of the Le Gall filter is as follows: Where: represents the high frequency subband coefficient, represents the low-frequency subband coefficient, and Indicates the sample points of the input audio signal separated according to the odd and even indexes.

3. The Bluetooth low bit rate audio encoding and decoding method according to claim 1, characterized in that: The step of quantizing the initial sub-band signal and performing hierarchical compression on the quantized initial sub-band signal through FSQ to obtain a compressed sub-band signal comprises: Calculating the energy value of each of the initial sub-band signals, and setting a differentiation threshold based on a maximum energy value; Determine whether the energy value of each of the initial sub-band signals is less than the differentiation threshold, if so, perform coarse quantization of the initial sub-band signal at 4 bits / sample; otherwise, perform high-precision quantization of the initial sub-band signal at 8 bits / sample; during the quantization process, if the Bluetooth real-time bandwidth is less than the second preset bandwidth threshold, discard the initial sub-band signal that exceeds the preset frequency band; The quantized initial sub-band signal is hierarchically compressed through FSQ to obtain a compressed sub-band signal.

4. The Bluetooth low bit rate audio encoding and decoding method according to claim 3, characterized in that: The step of performing hierarchical compression on the quantized initial subband signal through FSQ to obtain a compressed subband signal includes: based on a preset hierarchical finite state quantization codebook strategy, performing hierarchical compression on the quantized initial subband signal through FSQ to obtain a compressed subband signal; under the preset hierarchical finite state quantization codebook strategy, designing corresponding codebook sizes and quantization step sizes for the quantized initial subband signals of different frequency bands, and dynamically adjusting the quantization codebook and quantization step size in combination with a state transition rule.

5. The Bluetooth low bit rate audio encoding and decoding method according to claim 1, characterized in that: The steps of performing differential coding processing on the compressed subband signal to obtain coded data; inputting the coded data into a lightweight Transformer model to obtain a residual correction term, and adding a forward error correction code to the coded data to obtain an audio data packet include: Performing differential encoding processing on the compressed sub-band signal to calculate the inter-frame difference of the compressed sub-band signal to generate encoded data; The encoded data is input into a lightweight Transformer model, and a sparse self-attention mechanism is used in the lightweight Transformer model to calculate the inter-frame correlation of the encoded data to obtain a residual correction term; A forward error correction code is added to the encoded data to obtain an audio data packet.

6. The Bluetooth low bit rate audio encoding and decoding method according to claim 1, characterized in that: The step of performing decoding processing on the bit stream at the decoding end to recover the target audio signal comprises: At the decoding end, the bit stream is decoded and error corrected to obtain an error correction subband signal and a residual correction term; the error correction subband signal is reversely differentially calculated to obtain a reconstructed signal; Superimposing the reconstructed signal and the residual correction term obtained by decoding and error correction to generate a restored subband signal; The target audio signal is restored by performing reconstruction processing on the restored subband signal.

7. The Bluetooth low bit rate audio encoding and decoding method according to claim 6, characterized in that: The step of performing reconstruction processing on the restored sub-band signal to restore the target audio signal comprises: Performing an inverse quantization operation on the restored subband signal through FSQ to restore the audio subband coefficients; Performing inverse wavelet transform on the audio subband coefficients to reconstruct a time domain signal; The reconstructed time domain signal is subjected to an error propagation suppression operation, and the time domain signal after the error propagation suppression operation is subjected to psychoacoustic post-processing to restore the target audio signal.

8. A Bluetooth low bit rate audio codec system, characterized in that: The system comprises: A decomposition unit, used to obtain a Bluetooth real-time bandwidth, perform dynamic wavelet processing on the input audio signal based on the Bluetooth real-time bandwidth, and decompose to obtain a plurality of initial sub-band signals; A quantization unit, used for quantizing the initial sub-band signal, and performing hierarchical compression on the quantized initial sub-band signal through FSQ to obtain a compressed sub-band signal; An encoding unit, configured to perform differential encoding processing on the compressed subband signal to obtain encoded data; input the encoded data into a lightweight Transformer model to obtain a residual correction term, and add a forward error correction code to the encoded data to obtain an audio data packet; An encapsulation unit, used for encapsulating the residual correction term and the audio data packet into a bit stream, and transmitting the bit stream to a decoding end via Bluetooth; The decoding unit is used to perform decoding processing according to the bit stream at the decoding end to restore the target audio signal.

9. The Bluetooth low bit rate audio codec system according to claim 8, characterized in that: Also includes: Heterogeneous computing platform with STM32H743 microcontroller and FPGA chip.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the Bluetooth low bit rate audio encoding and decoding method as described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Audio signal processing method and device, equipment and storage medium

    CN121054011A