An audio decoding method, apparatus, electronic device, and storage medium
By adding dedicated fields and control fields for speech decoding to the L2HC bitstream data, and independently decoding speech data segments, the problem of poor decoding quality at low rates in speech scenarios in L2HC is solved, and high-quality speech and audio decoding is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-29
- Publication Date
- 2026-04-03
AI Technical Summary
L2HC encoding and decoding technology suffers from poor decoding quality and incomplete parsing of side information in low-speed speech decoding scenarios, making it unsuitable for speech decoding.
A dedicated field for voice decoding is added to the side information segment of the L2HC bitstream data, and the voice data segment is carried by the payload segment. High-quality voice and audio decoding is achieved by extracting the frequency division control mark and the voice coding control field for independent decoding.
High-quality speech and audio decoding was achieved within the L2HC framework, solving the problem of poor quality at low decoding rates and improving the integrity and quality of speech decoding.
Smart Images

Figure CN121418495B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio decoding, and in particular to an audio decoding method, apparatus, electronic device, and storage medium. Background Technology
[0002] L2HC (Low Latency High Clarity) is a codec technology suitable for high-quality audio. However, when dealing with low-rate decoding in speech scenarios, it is prone to defects such as poor low-rate decoding quality and incomplete parsing of side information, making it unsuitable for speech decoding. Summary of the Invention
[0003] The purpose of this invention is to provide an audio decoding method, apparatus, electronic device, and storage medium that can add a dedicated voice decoding field to the side information segment of L2HC bitstream data, and use the payload segment of L2HC bitstream data to carry voice data segments and perform independent decoding on the voice data segments, thereby achieving high-quality voice and audio decoding under the L2HC framework.
[0004] To solve the above-mentioned technical problems, the present invention provides an audio decoding method, comprising:
[0005] Receive L2HC bitstream data, extract the packet header segment from the L2HC bitstream data, and parse the mixed coding marker in the packet header segment;
[0006] The side information segment and payload segment are extracted from the L2HC bitstream data. When determining the hybrid coding mark to represent speech coding, the frequency division control mark and speech coding control field are extracted from the side information segment. The speech data segment of the specified frequency band is extracted from the payload segment according to the frequency division control mark.
[0007] The voice data segment is decoded according to the voice encoding control field to obtain the voice audio.
[0008] Optionally, the frequency division control flag and the speech coding control field are set in the reserved field or extended field of the side information segment;
[0009] The method further includes:
[0010] When determining that the hybrid coding mark represents music coding, the frequency division control mark and the speech coding control field are masked, and the L2HC decoding process is performed based on the remaining content in the side information segment and the payload segment.
[0011] Optionally, the frequency division control flag is a preset frequency division cutoff frequency;
[0012] The step of extracting a speech data segment of a specified frequency band from the payload segment according to the frequency division control flag includes:
[0013] Extract the data segment with a frequency band from zero to the preset frequency division cutoff frequency from the payload segment, and use it as the voice data segment.
[0014] Optionally, the speech coding control field includes residual coding mode, pitch period range, and LPC order;
[0015] The speech data segment includes pitch period quantization value, LSP parameter quantization value, and pulse sequence or codebook data;
[0016] Decoding the voice data segment according to the voice encoding control field includes:
[0017] Based on the fundamental period range, the fundamental period quantization value is inversely quantized to obtain the fundamental period;
[0018] The target decoding mode is determined based on the residual coding mode;
[0019] If the target decoding mode is CELP mode, then the excitation signal is obtained by reconstructing the codebook using the pitch period and the codebook data;
[0020] If the target decoding mode is MPE mode, then the excitation signal is obtained by multi-pulse excitation reconstruction using the pulse sequence;
[0021] Based on the LPC order, the LSP parameter quantization value is inversely quantized to obtain LPC coefficients, and the LPC coefficients are used to construct an LPC synthesis filter.
[0022] The excitation signal and the LPC synthesis filter are used to synthesize speech audio with a frequency band from zero to a preset cutoff frequency.
[0023] Optionally, it also includes:
[0024] When determining the hybrid coding marker to represent speech coding, the frame length in the packet header segment is extracted, and the buffer size of the speech audio is adjusted according to the frame length.
[0025] Optionally, before extracting the frequency division control flag and the speech coding control field from the side information segment, the method further includes:
[0026] Read the verification field in the L2HC bitstream data, and perform integrity verification on the edge information segment based on the verification field to determine whether the edge information segment is complete;
[0027] If it is determined that the side information segment is complete, then proceed to the step of extracting the frequency division control tag and the speech coding control field from the side information segment;
[0028] If it is determined that the side information segment is incomplete, a request is made to the encoding end to retransmit the L2HC bitstream data.
[0029] Optionally, it also includes:
[0030] The frame length of the speech audio is adjusted according to the L2HC frame length so that the speech audio frame is aligned with the L2HC audio frame.
[0031] The present invention also provides an audio decoding device, comprising:
[0032] The header parsing module is used to receive L2HC bit stream data, extract the header segment from the L2HC bit stream data, and parse the mixed coding markers in the header segment;
[0033] The data segment separation module is used to extract the side information segment and the payload segment from the L2HC bitstream data, and when determining that the hybrid coding mark represents speech coding, it extracts the frequency division control mark and the speech coding control field from the side information segment, and extracts the speech data segment of the specified frequency band from the payload segment according to the frequency division control mark.
[0034] The voice decoding module is used to decode the voice data segment according to the voice encoding control field to obtain the voice audio.
[0035] The present invention also provides an electronic device, comprising:
[0036] Memory, used to store computer programs;
[0037] A processor for implementing the audio decoding method described above when executing the computer program.
[0038] The present invention also provides a computer-readable storage medium storing computer-executable instructions, which, when loaded and executed by a processor, implement the audio decoding method described above.
[0039] This invention provides an audio decoding method, comprising: receiving L2HC bitstream data; extracting a header segment from the L2HC bitstream data and parsing a hybrid coding marker in the header segment; extracting a side information segment and a payload segment from the L2HC bitstream data; and, when determining that the hybrid coding marker represents speech coding, extracting a frequency division control marker and a speech coding control field from the side information segment, and extracting a speech data segment of a specified frequency band from the payload segment according to the frequency division control marker; and decoding the speech data segment according to the speech coding control field to obtain speech audio.
[0040] The beneficial effects of this invention are as follows: Upon receiving L2HC bitstream data, this invention extracts the packet header segment from the L2HC bitstream data and parses the hybrid coding marker in the packet header segment. The hybrid coding marker is a newly added field in the packet header segment used to distinguish between speech coding and traditional music coding. Subsequently, this invention can extract the side information segment and payload segment from the L2HC bitstream data. When determining that the hybrid coding marker represents speech coding, it extracts the frequency division control marker and speech coding control field from the side information segment, and extracts the speech data segment of the specified frequency band from the payload segment according to the frequency division control marker. That is, it can add a speech decoding-specific frequency division control marker and speech coding control field to the side information segment of the L2HC bitstream data, and utilize the payload segment of the L2HC bitstream data to carry the speech data segment. Finally, this invention can decode the speech data segment according to the speech coding control field to obtain the speech audio. Thus, by improving the L2HC bitstream data structure and independently decoding the speech data segment, this invention can achieve high-quality speech and audio decoding within the L2HC framework.
[0041] The present invention also provides an audio decoding device, an electronic device, and a computer-readable storage medium, which have the above-mentioned beneficial effects. Attached Figure Description
[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0043] Figure 1 A flowchart of an audio decoding method provided in an embodiment of the present invention;
[0044] Figure 2 A structural block diagram of a decoder provided in an embodiment of the present invention;
[0045] Figure 3A flowchart for decoding a side information segment is provided as an embodiment of the present invention;
[0046] Figure 4 A flowchart of voice audio decoding is provided as an embodiment of the present invention;
[0047] Figure 5 This is a structural block diagram of an audio decoding device provided in an embodiment of the present invention;
[0048] Figure 6 This is a structural block diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0050] L2HC (Low Latency High Clarity) is a codec technology suitable for high-quality audio. However, when dealing with low-bitrate decoding in speech scenarios (mono below 96kbps), it is prone to defects such as poor low-bitrate decoding quality and incomplete side information parsing, and is not suitable for speech decoding.
[0051] In view of this, in order to address the technical problem of how to improve the quality of speech decoding in the L2HC codec framework, the present invention provides an audio decoding method, which can add a speech decoding-specific field to the side information segment of the L2HC bitstream data, and use the payload segment of the L2HC bitstream data to carry the speech data segment, and perform independent decoding on the speech data segment, thereby achieving high-quality speech and audio decoding under the L2HC framework.
[0052] It should be noted that this method can be executed by various types of devices, such as mobile devices (phones, tablets), personal computers, or other audio devices.
[0053] For easier understanding, please refer to Figure 1 , Figure 1 A flowchart of an audio decoding method provided in an embodiment of the present invention, the method may include:
[0054] S100: Receive L2HC bit stream data, extract the packet header segment from the L2HC bit stream data, and parse the mixed coding mark in the packet header segment.
[0055] L2HC bitstream data refers to audio encoded data based on the L2HC encoding and decoding framework and which has undergone binary bitstream conversion. It can consist of a header segment, side information segment, and payload segment. In this embodiment, the header segment can be parsed first to provide basic data for subsequent decoding.
[0056] Specifically, in addition to parsing the codec type, sample rate, channel number, and frame length in the packet header, a hybrid encoding marker (represented as "hybrid Mode") can be added to the packet header to integrate independent speech encoding and decoding into the L2HC codec framework and achieve hybrid encoding of speech and music. This hybrid encoding marker is then parsed. The hybrid encoding marker is at least 1 bit in size and has at least two values, corresponding to speech encoding and traditional music encoding respectively. For example, when the hybrid encoding marker is 1, it indicates that the current L2HC bitstream data is speech encoded; when the hybrid encoding marker is 0, it indicates that the current L2HC bitstream data is traditional music encoded. In this way, by parsing the hybrid encoding marker, the audio decoding method can be adjusted in a timely manner to adapt to speech decoding.
[0057] It should be noted that this embodiment does not limit the position of the hybrid encoding mark in the packet header segment. It can be set according to the actual application requirements. For example, it can be located at the 8th bit position of the packet header.
[0058] S200: Extract the side information segment and payload segment from the L2HC bitstream data. When determining the hybrid coding mark to represent speech coding, extract the frequency division control mark and speech coding control field from the side information segment, and extract the speech data segment of the specified frequency band from the payload segment according to the frequency division control mark.
[0059] After parsing the header segment, this embodiment can separate the data segments in the L2HC bitstream data into side information segments and payload segments, and continue parsing the side information segments and payload segments. The improvements to the side information segments and payload segments are described below.
[0060] In this embodiment, the side information segment can be equipped with a dedicated frequency division control flag and a speech coding control field for speech decoding. The frequency division control flag is used to determine the separation rules, which are used to extract speech data segments from the payload segment. The speech coding control field contains control parameters used to encode speech audio into speech data segments, and is used to decode speech data segments into speech audio. Since the frequency division control flag and the speech coding control field are fields added specifically for speech coding and decoding, they can be parsed when determining the hybrid coding flag in the packet header to represent speech coding. Furthermore, the frequency division control flag and the speech coding control field are set in the reserved or extended fields of the side information segment. This preserves the original structure of the L2HC side information segment while adding dedicated speech decoding fields to the L2HC side information segment, thus integrating the speech coding and decoding mechanism into the L2HC coding and decoding architecture.
[0061] Of course, when determining the hybrid coding marker to represent music coding, the frequency division control marker and the speech coding control field can be masked, and only the original content in the side information segment can be parsed. The original L2HC decoding process can be executed on the payload segment according to the original content in the side information segment, thus making it compatible with the original L2HC decoding process.
[0062] Based on this, the method may also include:
[0063] Step S211: When determining the hybrid coding mark to represent music coding, mask the frequency division control mark and the speech coding control field, and perform the L2HC decoding process based on the remaining content in the side information segment and the payload segment.
[0064] Furthermore, in this embodiment, the payload segment may include a speech data segment, which refers to the encoded data of the speech audio. Since the speech audio content is mainly distributed in the low-frequency band, when using the payload segment to carry the speech data segment, this embodiment can use a frequency division control flag to separate the speech data segment from other data segments, and can achieve separate extraction of the speech data segment based on the frequency division control flag.
[0065] In one implementation, the frequency division control flag can be a preset frequency division cutoff frequency, and the separation rule determined by the frequency division control flag can be: the data segment with the intermediate frequency band of the payload segment from zero to the preset frequency division cutoff frequency is taken as the voice data segment.
[0066] Based on this, the frequency division control marker is a preset frequency division cutoff frequency; extracting a voice data segment of a specified frequency band from the payload segment according to the frequency division control marker may include:
[0067] Step S221: Extract the data segment from the payload segment with a frequency band from zero to the preset frequency division cutoff frequency, and use it as the voice data segment.
[0068] Furthermore, since the integrity of the side information segments is crucial for decoding the voice data segments, this embodiment can set a check field (e.g., a CRC field, Cyclic Redundancy Check) in the L2HC bitstream data, and use this check field to verify the integrity of the side information segments. If the side information segment is determined to be complete, parsing of the side information segment can continue. If the side information segment is determined to be incomplete, a request can be made to the encoding end to retransmit the current L2HC bitstream data.
[0069] Based on this, before extracting the frequency division control flag and speech coding control field from the side information segment, it may also include:
[0070] Step S221: Read the verification field from the L2HC bitstream data and perform integrity verification on the edge information segment based on the verification field to determine whether the edge information segment is complete. If the edge information segment is determined to be complete, proceed to step S222; if the edge information segment is determined to be incomplete, proceed to step S223.
[0071] Step S222: Extract the frequency division control tag and speech coding control field from the side information segment;
[0072] Step S223: Request the encoding end to retransmit the L2HC bitstream data.
[0073] S300: Decode the voice data segment according to the voice encoding control field to obtain the voice audio.
[0074] In this embodiment, after parsing the side information segment and the payload segment, the speech data segment can be decoded separately using an independent speech decoding unit based on the speech coding control field to obtain the speech audio. Thus, this embodiment can solve the shortcomings of the original L2HC codec architecture, which is prone to speech distortion and decreased intelligibility due to quantization errors when decoding low-bitrate speech.
[0075] Furthermore, to ensure compatibility between the voice audio decoding and the original L2HC decoding, this embodiment can also adjust the voice audio frame length according to the L2HC frame length so that the voice audio frame is aligned with the L2HC audio frame.
[0076] Based on this, the method may also include:
[0077] Step S311: Adjust the frame length of the speech audio according to the L2HC frame length so that the speech audio frame is aligned with the L2HC audio frame.
[0078] Furthermore, to meet the requirements of low-bit audio decoding and avoid frame overflow at low bit rates, this embodiment can also dynamically adjust the size of the decoding frame buffer based on the frame length field (Frame Length, 5ms / 10ms) in the packet header.
[0079] Based on this, the method may also include:
[0080] Step S321: When determining the hybrid coding marker to represent speech coding, extract the frame length in the packet header segment and adjust the buffer size of the speech audio according to the frame length.
[0081] Furthermore, to improve the quality of speech decoding, this embodiment can also parse the low-rate mode field (lowBrFlag) in the side information segment, and when it is determined that the low-rate mode is enabled, enable residual signal smoothing filtering to process the speech audio, reduce the impact of quantization noise on speech formants, and improve the signal-to-noise ratio in the low-frequency band.
[0082] Based on the above embodiments, upon receiving L2HC bitstream data, this invention extracts the packet header segment from the L2HC bitstream data and parses the hybrid coding marker in the packet header segment. The hybrid coding marker is a newly added field in the packet header segment used to distinguish between speech coding and traditional music coding. Subsequently, this invention can extract the side information segment and payload segment from the L2HC bitstream data. When determining that the hybrid coding marker represents speech coding, it extracts the frequency division control marker and speech coding control field from the side information segment, and extracts the speech data segment of the specified frequency band from the payload segment according to the frequency division control marker. That is, it can add a speech decoding-specific frequency division control marker and speech coding control field to the side information segment of the L2HC bitstream data, and utilize the payload segment of the L2HC bitstream data to carry the speech data segment. Finally, this invention can decode the speech data segment according to the speech coding control field to obtain the speech audio. Thus, by improving the L2HC bitstream data structure and independently decoding the speech data segment, this invention can achieve high-quality speech and audio decoding within the L2HC framework.
[0083] Based on the above embodiments, the specific process of speech decoding is described below. In one embodiment, the speech coding control field includes a residual coding mode, a pitch period range, and an LPC order; the speech data segment includes a pitch period quantization value, an LSP parameter quantization value, and pulse sequence or codebook data. Decoding the speech data segment according to the speech coding control field may include:
[0084] S301. Based on the pitch period range, the pitch period quantization value is inversely quantized to obtain the pitch period.
[0085] The fundamental period refers to the time required for one vibration of the vocal cords, while the fundamental period range is the range of variation of the fundamental period. This step first involves inverse quantization of the fundamental period quantization value based on the fundamental period range to obtain the fundamental period.
[0086] S302. Determine the target decoding mode based on the residual coding mode; if the target decoding mode is CELP mode, proceed to S303; if the target decoding mode is MPE mode, proceed to S304.
[0087] S303. The excitation signal is obtained by reconstructing the codebook excitation using the pitch period and codebook data.
[0088] S304. The excitation signal is obtained by multi-pulse excitation reconstruction using pulse sequences.
[0089] The residual coding mode characterizes the residual coding mode selected by the encoder for audio encoding. Generally, there are two modes: Code Excited Linear Prediction (CELP) and Multi-Pulse Excited (MPE). When using CELP, the excitation signal can be obtained by reconstructing the codebook excitation using the pitch period and codebook data. When using MPE, the excitation signal can be obtained by reconstructing the multi-pulse excitation using a pulse sequence.
[0090] S305. Based on the LPC order, the LSP parameter quantization values are inversely quantized to obtain LPC coefficients, and LPC coefficients are used to construct an LPC synthesis filter.
[0091] The LPC order and LSP parameters (line spectral pairs) are key parameters used in Linear Predictive Coding (LPC) for modeling speech signals. This step involves inverse quantization of the LSP parameter values based on the LPC order to obtain the LPC coefficients. For example, when the LPC order is 10, this embodiment can inversely quantize to obtain 10th-order LPC coefficients. Subsequently, the LPC coefficients can be used to construct an LPC synthesis filter.
[0092] S306. Use the excitation signal and LPC synthesis filter to synthesize speech audio with a frequency band from zero to a preset cutoff frequency.
[0093] Based on the above embodiments, the audio decoding method described above will be fully introduced below with specific schematic diagrams. This invention aims to provide a low-bitrate optimized L2HC standalone decoder suitable for speech scenarios. By reusing the L2HC decoding structure and adding a dedicated speech decoding module and side information parsing logic, it achieves efficient decoding and high-quality wideband reconstruction of low-bitrate L2HC speech bitstreams, while ensuring compatibility with existing L2HC encoding formats.
[0094] The decoder of this invention includes the following core modules, which can reuse the core structure of L2HC technology (such as MDCT inverse transform and side information basic parsing). Hybrid decoding logic is designed for speech scenarios. Module connections and data flow are as follows: Figure 2 As shown.
[0095] The functions and specific implementations of each module are as follows:
[0096] 1. Low-rate bitstream parsing module:
[0097] 1) Core function: Receive L2HC encoded bit stream, reuse L2HC bit stream syntax rules, complete packet header parsing and data segment separation, and provide basic data for subsequent decoding.
[0098] 2) Specific implementation:
[0099] a. Packet Header Parsing: Based on the syntax of the packet header parsing function L2hcHeaderUnpack(), the codec type (encoder type, 2 bits), sample rate (sample rate, 2 bits), channel number (channel number, 1 bit), and frame length (frame length, 2 bits) are parsed. In this embodiment, a hybrid encoding identification flag (hybridMode, 1 bit, located at the 8th bit of the packet header) can be added to the packet header to determine whether the current bitstream is "speech encoding (1)" or "traditional music encoding (0)".
[0100] b. Data segment separation: The bit stream is split into a side information segment (SideInfo) and a payload segment (Payload), where the payload segment includes the "LPC data segment". The separation rules are determined by the "frequency division control flag" in the side information.
[0101] c. Rate adaptation: Supports a specified low rate range (mono 2kbps-96kbps), and dynamically adjusts the decoding frame buffer size through the Frame Length field (5ms / 10ms) to avoid frame overflow at low rates.
[0102] 2. Improved side information decoding module:
[0103] 1) Core function: Reuse the basic logic of L2HC side information decoding, extend the parsing of voice-specific fields, and provide control parameters for decoding.
[0104] 2) Specific implementation (process as follows) Figure 3 (as shown)
[0105] a. Input the edge information segment.
[0106] b. Basic Field Parsing. In addition to parsing the lowBrFlag (low bit rate identifier), drQuater (auxiliary adjustment factor), sfld (subband division index), bandNum (number of coded subbands), and diffFlag (subband envelope differential coding enable flag), parsing of speech coding control fields has been added. Speech coding control fields include LPC order, residual coding mode, and pitch period range.
[0107] c. Output the complete edge information parameter set.
[0108] 3) Compatibility design: The voice-specific field is embedded in the "Reserved" bit or an extended field without destroying the original L2HC side information structure. When hybridMode=0, the parsing of the voice field is automatically blocked, which is compatible with the traditional L2HC side information format.
[0109] 4) Error verification: The L2HC CRC verification mechanism (8 bits) is reused to perform integrity verification on the edge information field. When the verification fails at low speed, a retransmission request is triggered to ensure the reliability of the parameters.
[0110] 4. LPC decoding module:
[0111] 1) Core function: Adapt the LPC compression strategy of the encoding end for the audio segment, and complete the LPC parameter decoding and reconstruction of the core speech signal.
[0112] 2) Specific implementation (process as follows) Figure 4 (as shown)
[0113] a. Based on the pitch period range, the 8~10 bit pitch period quantization value is dequantized to recover the 5~20ms pitch period.
[0114] b. Determine the target decoding mode based on the residual coding mode. If the target decoding mode is CELP mode, the excitation signal is obtained by reconstructing the codebook using the pitch period and codebook data, adapting to L2HC low-rate bit allocation. If the target decoding mode is MPE mode, the excitation signal is obtained by reconstructing the excitation signal using a pulse sequence, realizing the inverse decoding of 64~128 bits of data.
[0115] c. Based on the LPC order, inverse quantize the 18~24 bit LSP parameter quantization values to obtain LPC coefficients (e.g., 10th order LPC coefficients), and use the LPC coefficients to construct an LPC synthesis filter.
[0116] d. Use the excitation signal and LPC synthesis filter to synthesize speech audio with a frequency band from zero to a preset cutoff frequency.
[0117] 3) Compatible with L2HC: The frame length of LPC synthesis filtering is adaptively aligned with the L2HC frame length (5ms / 10ms).
[0118] 4) Low-rate optimization: When lowBrFlag=1 (low bit rate mode), "residual signal smoothing filter" is enabled to reduce the impact of quantization noise on speech formants and improve the signal-to-noise ratio in the low-frequency band.
[0119] The beneficial effects of this invention are as follows:
[0120] 1. Strong compatibility: It reuses the basic L2HC decoding structure (side information syntax, MDCT inverse transform) and achieves compatible decoding with traditional L2HC encoded bitstreams through the hybridMode flag, without the need to reconstruct existing L2HC devices;
[0121] 2. Low-rate speech coding quality: Addressing the issue of quality degradation in low-rate speech application scenarios using the L2HC algorithm.
[0122] The audio decoding device, electronic device, computer program product, and computer-readable storage medium provided in the embodiments of the present invention will be described below. The audio decoding device, electronic device, computer program product, and computer-readable storage medium described below can be referred to in correspondence with the audio decoding processing method described above.
[0123] Please refer to Figure 5 , Figure 5 This is a structural block diagram of an audio decoding device provided in an embodiment of the present invention. The device may include:
[0124] The header parsing module 501 is used to receive L2HC bit stream data, extract the header segment from the L2HC bit stream data, and parse the mixed coding mark in the header segment;
[0125] The data segment separation module 502 is used to extract the side information segment and the payload segment from the L2HC bit stream data, and when determining the hybrid coding mark to characterize the speech coding, it extracts the frequency division control mark and the speech coding control field from the side information segment, and extracts the speech data segment of the specified frequency band from the payload segment according to the frequency division control mark.
[0126] The voice decoding module 503 is used to decode the voice data segment according to the voice encoding control field to obtain the voice audio.
[0127] Optionally, the frequency division control flag and the voice coding control field are set in the reserved field or extended field of the side information segment;
[0128] This device may also include:
[0129] The masking module is used to mask the speech coding control field when determining the hybrid coding mark to represent music coding, and to perform the L2HC decoding process based on the remaining content in the side information segment and the payload segment.
[0130] Optionally, the frequency division control flag is set to a preset frequency division cutoff frequency;
[0131] The data segment separation module 502 may include:
[0132] The voice data segment separation submodule is used to extract data segments from the payload segment with frequency bands from zero to the preset frequency division cutoff frequency, and use them as voice data segments.
[0133] Optionally, the speech coding control field includes residual coding mode, pitch period range, and LPC order;
[0134] The speech data segment includes pitch period quantization values, LSP parameter quantization values, and pulse sequence or codebook data;
[0135] The voice decoding module 503 may include:
[0136] The pitch period decoding submodule is used to inverse quantize the pitch period quantization value according to the pitch period range to obtain the pitch period;
[0137] The excitation signal reconstruction submodule is used to determine the target decoding mode based on the residual coding mode. If the target decoding mode is CELP mode, the excitation signal is obtained by codebook excitation reconstruction using the pitch period and codebook data. If the target decoding mode is MPE mode, the excitation signal is obtained by multi-pulse excitation reconstruction using the pulse sequence.
[0138] The LPC decoding submodule is used to inverse quantize the LSP parameter quantization values according to the LPC order to obtain LPC coefficients, and to construct an LPC synthesis filter using the LPC coefficients.
[0139] The speech audio synthesis submodule is used to synthesize speech audio with a frequency band from zero to a preset cutoff frequency using an excitation signal and an LPC synthesis filter.
[0140] Optionally, the device may further include:
[0141] The buffer adjustment module is used to extract the frame length in the packet header segment when determining the mixed coding marker to represent the speech coding, and adjust the buffer size of the speech audio according to the frame length.
[0142] Optionally, the device may further include:
[0143] The integrity verification module is used to read the verification field in the L2HC bit stream data and perform integrity verification on the side information segment based on the verification field to determine whether the side information segment is complete. If the side information segment is determined to be complete, the module proceeds to the step of extracting the frequency division control flag and the speech coding control field from the side information segment. If the side information segment is determined to be incomplete, the module requests the encoding end to retransmit the L2HC bit stream data.
[0144] Optionally, the device may further include:
[0145] The frame length adjustment module is used to adjust the frame length of the speech audio according to the L2HC frame length so that the speech audio frame is aligned with the L2HC audio frame.
[0146] Please refer to Figure 6 , Figure 6 This is a structural block diagram of an electronic device provided in an embodiment of the present invention. The present invention provides an electronic device 10, including a processor 11 and a memory 12; wherein, the memory 12 is used to store a computer program; the processor 11 is used to execute the audio decoding method provided in the foregoing embodiment when executing the computer program.
[0147] For details regarding the audio decoding method described above, please refer to the relevant content provided in the foregoing embodiments; further details will not be repeated here.
[0148] Furthermore, the memory 12, as a carrier for resource storage, can be a read-only memory, random access memory, disk, or optical disk, and the storage method can be temporary storage or permanent storage.
[0149] In addition, the electronic device 10 also includes a power supply 13, a communication interface 14, an input / output interface 15, and a communication bus 16; wherein, the power supply 13 is used to provide operating voltage for the various hardware devices on the electronic device 10; the communication interface 14 can create a data transmission channel between the electronic device 10 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this invention, and is not specifically limited here; the input / output interface 15 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.
[0150] This invention also provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the audio decoding method described in the above embodiments.
[0151] Since the embodiments of the computer program product portion correspond to the embodiments of the audio decoding method portion, please refer to the description of the embodiments of the audio decoding method portion for the embodiments of the computer program product portion, and will not be repeated here.
[0152] This invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the audio decoding method described in the above embodiments.
[0153] Since the embodiments of the computer-readable storage medium portion correspond to the embodiments of the audio decoding method portion, the embodiments of the storage medium portion are described in the description of the embodiments of the audio decoding method portion, and will not be repeated here.
[0154] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0155] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0156] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0157] The above provides a detailed description of the audio decoding method, apparatus, electronic device, and storage medium provided by the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and core ideas of the present invention. It should be noted that those skilled in the art can make various improvements and modifications to the present invention without departing from its principles, and these improvements and modifications also fall within the protection scope of the present invention.
Claims
1. An audio decoding method, characterized in that, include: Receive L2HC bitstream data, extract the packet header segment from the L2HC bitstream data, and parse the mixed coding marker in the packet header segment; The side information segment and payload segment are extracted from the L2HC bitstream data. When determining the hybrid coding mark to represent speech coding, the frequency division control mark and speech coding control field are extracted from the side information segment. The speech data segment of the specified frequency band is extracted from the payload segment according to the frequency division control mark. The side information segment adds a frequency division control mark and speech coding control field dedicated to speech decoding. The frequency division control mark and speech coding control field are set in the reserved field or extended field of the side information segment. The voice data segment is decoded according to the voice encoding control field to obtain the voice audio.
2. The audio decoding method according to claim 1, characterized in that, The frequency division control flag and the voice coding control field are set in the reserved field or extended field of the side information segment; The method further includes: When determining that the hybrid coding mark represents music coding, the frequency division control mark and the speech coding control field are masked, and the L2HC decoding process is performed based on the remaining content in the side information segment and the payload segment.
3. The audio decoding method according to claim 1, characterized in that, The frequency division control flag is a preset frequency division cutoff frequency; The step of extracting a speech data segment of a specified frequency band from the payload segment according to the frequency division control flag includes: Extract the data segment with a frequency band from zero to the preset frequency division cutoff frequency from the payload segment, and use it as the voice data segment.
4. The audio decoding method according to claim 3, characterized in that, The speech coding control field includes residual coding mode, pitch period range, and LPC order; The speech data segment includes pitch period quantization value, LSP parameter quantization value, and pulse sequence or codebook data; Decoding the voice data segment according to the voice encoding control field includes: Based on the fundamental period range, the fundamental period quantization value is inversely quantized to obtain the fundamental period; The target decoding mode is determined based on the residual coding mode; If the target decoding mode is CELP mode, then the excitation signal is obtained by reconstructing the codebook using the pitch period and the codebook data; If the target decoding mode is MPE mode, then the excitation signal is obtained by multi-pulse excitation reconstruction using the pulse sequence; Based on the LPC order, the LSP parameter quantization value is inversely quantized to obtain LPC coefficients, and the LPC coefficients are used to construct an LPC synthesis filter. The excitation signal and the LPC synthesis filter are used to synthesize speech audio with a frequency band from zero to a preset cutoff frequency.
5. The audio decoding method according to claim 1, characterized in that, Also includes: When determining the hybrid coding marker to represent speech coding, the frame length in the packet header segment is extracted, and the buffer size of the speech audio is adjusted according to the frame length.
6. The audio decoding method according to claim 1, characterized in that, Before extracting the frequency division control flag and speech coding control field from the side information segment, the method further includes: Read the verification field in the L2HC bitstream data, and perform integrity verification on the edge information segment based on the verification field to determine whether the edge information segment is complete; If it is determined that the side information segment is complete, then proceed to the step of extracting the frequency division control tag and the speech coding control field from the side information segment; If it is determined that the side information segment is incomplete, a request is made to the encoding end to retransmit the L2HC bitstream data.
7. The audio decoding method according to claim 1, characterized in that, Also includes: The frame length of the speech audio is adjusted according to the L2HC frame length so that the speech audio frame is aligned with the L2HC audio frame.
8. An audio decoding device, characterized in that, include: The header parsing module is used to receive L2HC bit stream data, extract the header segment from the L2HC bit stream data, and parse the mixed coding markers in the header segment; The data segment separation module is used to extract the side information segment and the payload segment from the L2HC bitstream data. When determining that the hybrid coding mark represents speech coding, it extracts the frequency division control mark and the speech coding control field from the side information segment, and extracts the speech data segment of the specified frequency band from the payload segment according to the frequency division control mark. The side information segment adds a frequency division control mark and a speech coding control field dedicated to speech decoding. The frequency division control mark and the speech coding control field are set in the reserved field or extended field of the side information segment. The voice decoding module is used to decode the voice data segment according to the voice encoding control field to obtain the voice audio.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the audio decoding method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when loaded and executed by a processor, implement the audio decoding method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Decoding audio bitstreams with enhanced spectral band replication metadata in at least one fill element
CN107408391A
High resolution audio coding
CN113302684A