Method and device for joint time-domain / frequency-domain coding of an acoustic signal
The joint time-domain/frequency-domain coding model addresses inefficiencies in encoding general audio signals by dynamically allocating bits and applying specific coding sub-modes, improving synthesis quality and reducing artifacts for unclear signal types.
Patent Information
- Application Number
- JP2023541804
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-01-08
- Filing Date
- 2022-01-05
- Publication Date
- 2025-11-05
- Estimated Expiration
- 2042-01-05
AI Technical Summary
Existing speech codecs face challenges in encoding general audio signals like music and reverberant speech at low bit rates due to longer processing delays and inefficiencies in bit allocation between time and frequency domains, leading to degradation in synthesis quality.
A joint time-domain/frequency-domain coding model that dynamically allocates bits between domains based on signal characteristics, using a novel speech/music classifier to identify unclear signal types and apply specific coding sub-modes, including frequency band selection and bit allocation, to enhance synthesis quality without increasing processing delay or bit rate.
The model improves synthesis quality for general audio signals by efficiently allocating bits between time and frequency domains, reducing artifacts, and maintaining low processing delay, especially for unclear signal types, thus enhancing the performance of speech and music encoding.
Smart Images

Figure 0007764480000057 
Figure 0007764480000058 
Figure 0007764480000059
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to unified time-domain / frequency-domain coding devices and methods that use a mixed time-domain and frequency-domain coding mode for encoding an input sound signal, and corresponding decoder devices and decoding methods.
[0002] In this disclosure and the accompanying claims: The term "acoustic" may relate to generic audio signals such as speech, music and reverberant sounds, as well as any other acoustics. [Background technology]
[0003] Prior art speech codecs can represent clean speech signals with very good quality at bit rates of around 8 kbps, and approach transparency at bit rates of 16 kbps. However, at bit rates below 16 kbps, low-processing-delay speech codecs, which mostly encode input speech signals in the time domain, are not suitable for general audio signals, such as music and reverberant speech. To overcome this drawback, switched codecs have been introduced, which essentially use a time-domain approach to encode speech-dominated input acoustic signals and a frequency-domain approach to encode general audio signals. However, such switching solutions typically require longer processing delays, both for speech-music classification and for computing the transformation to the frequency domain.
[0004] To overcome the above drawbacks related to longer processing delays, a more integrated time-domain and frequency-domain coding model has been proposed in U.S. Patent No. 9,015,038 (see reference [1], the entire contents of which are incorporated herein by reference). This integrated time-domain and frequency-domain coding model is part of the EVS (Enhanced Voice Services) audio codec standardized by 3GPP (Third Generation Partnership Project) as described in reference [2], the entire contents of which are incorporated herein by reference. Recently, 3GPP has begun work on developing a 3D (three-dimensional) audio codec for immersive services called IVAS (Immersive Voice and Audio Services), based on the EVS codec (see reference [3], the entire contents of which are incorporated herein by reference).
[0005] To make the coding model more efficient for certain types of signals, additional coding modes have been added to efficiently allocate available bits between the time and frequency domains, and between low and high frequencies. The additional coding modes are triggered by a new speech / music classifier, whose output enables an indistinct category for signals that cannot be clearly classified as either music or speech (see [4], the entire contents of which are incorporated herein by reference). [Prior art documents] [Patent documents]
[0006] [Patent Document 1] U.S. Patent No. 9,015,038 Summary of the Invention [Means for solving the problem]
[0007] The present disclosure relates to a joint time-domain / frequency-domain coding method for encoding an input audio signal, the method including: classifying the input audio signal into one of a plurality of audio signal categories, the audio signal categories including an unclear signal type category indicating that the nature of the input audio signal is unclear; if the input audio signal is classified into the unclear signal type category, selecting one of a plurality of coding sub-modes for coding the input audio signal; and joint time-domain / frequency-domain coding the input audio signal using the selected coding sub-mode.
[0008] The present disclosure also relates to a joint time-domain / frequency-domain coding method for encoding an input audio signal, which includes classifying the input audio signal into one of a plurality of audio signal categories, the audio signal categories including an unclear signal type category indicating that the nature of the input audio signal is unclear, and joint time-domain / frequency-domain coding the input audio signal in response to classifying the input audio signal into the unclear signal type category. The joint time-domain / frequency-domain coding of the input audio signal includes frequency band selection and bit allocation for selecting frequency bands to quantize and distributing the bit budget available for quantization among the selected frequency bands.
[0009] According to the present disclosure, there is further provided a joint time domain / frequency domain encoding device for encoding an input acoustic signal, the joint time domain / frequency domain encoding device including: a classifier that classifies the input acoustic signal into one of a plurality of acoustic signal categories, the acoustic signal categories including an unclear signal type category indicating that the nature of the input acoustic signal is unclear; a selector that, if the input acoustic signal is classified into the unclear signal type category, selects one of a plurality of coding sub-modes for encoding the input acoustic signal; and a joint time domain / frequency domain encoder that encodes the input acoustic signal using the selected coding sub-mode.
[0010] The present disclosure still further relates to a joint time-domain / frequency-domain encoding device for encoding an input audio signal, including a classifier that classifies the input audio signal into one of a plurality of audio signal categories, the audio signal categories including an unclear signal type category indicating that the nature of the input audio signal is unclear, and a mixed time-domain / frequency-domain encoder that encodes the input audio signal in response to the classification of the input audio signal into the unclear signal type category. The mixed time-domain / frequency-domain encoder includes a frequency band selector and a bit allocator for selecting frequency bands to quantize and distributing a bit budget available for quantization among the selected frequency bands.
[0011] The present disclosure provides a method for decoding an acoustic signal, the method including: receiving a bitstream conveying information usable for reconstructing a mixed time domain / frequency domain excitation representing an acoustic signal classified in an unclear signal type category indicating that the nature of the acoustic signal is unclear, the information including one of a plurality of coding sub-modes used to encode an input acoustic signal classified in the unclear signal type category; reconstructing the mixed time domain / frequency domain excitation in response to the information conveyed in the bitstream including the coding sub-mode used to encode the input acoustic signal; transforming the mixed time domain / frequency domain excitation into a time domain; and filtering the mixed time domain / frequency domain excitation converted to the time domain through a synthesis filter to generate a synthesized version of the acoustic signal.
[0012] The present disclosure proposes an audio signal decoding method, which includes receiving a bitstream conveying information usable for reconstructing a mixed time-domain / frequency-domain excitation representing an audio signal (a) classified into an unclear signal type category indicating that the nature of the audio signal is unclear, and (b) encoded using (i) frequency bands selected for quantization and (ii) a bit budget available for quantization distributed among the frequency bands; reconstructing the mixed time-domain / frequency-domain excitation according to the information conveyed in the bitstream, where reconstructing the mixed time-domain / frequency-domain excitation includes selecting frequency bands to be used for quantization and a distribution of the bit budget available for quantization among the frequency bands; transforming the mixed time-domain / frequency-domain excitation into a time domain; and filtering the time-domain converted mixed time-domain / frequency-domain excitation through a synthesis filter to generate a synthesized version of the audio signal.
[0013] According to the present disclosure, an audio signal decoder is provided, which includes: a receiver that receives a bitstream conveying information usable to reconstruct a mixed time domain / frequency domain excitation representing an audio signal classified in an unclear signal type category indicating that the nature of the audio signal is unclear, the information including one of a plurality of coding sub-modes used to encode the audio signal classified in the unclear signal type category; a reconstructor that reconstructs the mixed time domain / frequency domain excitation in accordance with the information conveyed in the bitstream, the reconstructor including the coding sub-mode used to encode the input audio signal; a transformer that converts the mixed time domain / frequency domain excitation to a time domain; and a synthesis filter that filters the mixed time domain / frequency domain excitation converted to the time domain to generate a synthesized version of the audio signal.
[0014] The present disclosure still further relates to an audio signal decoder, including: a receiver that receives a bitstream that conveys information that can be used to reconstruct a mixed time-domain / frequency-domain excitation representing an audio signal that (a) has been classified into an unclear signal type category, indicating that the nature of the audio signal is unclear, and (b) has been coded using (i) frequency bands selected for quantization and (ii) a bit budget available for quantization distributed among the frequency bands; a reconstructor that reconstructs the mixed time-domain / frequency-domain excitation in accordance with the information conveyed in the bitstream, the reconstructor selecting the frequency bands to be used for quantization and a distribution of the bit budget available for quantization among the frequency bands; a converter that converts the mixed time-domain / frequency-domain excitation to a time domain; and a synthesis filter that filters the mixed time-domain / frequency-domain excitation converted to the time domain to generate a synthesized version of the audio signal.
[0015] The foregoing and other features will become more apparent on reading the following non-limiting description of exemplary embodiments of a joint time domain / frequency domain coding method, a joint time domain / frequency domain coding device, a decoding method and a decoder device, given by way of example only with reference to the accompanying drawings, in which:
[0016] The accompanying drawings are described below. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a schematic block diagram illustrating simultaneously an overview of a joint time domain / frequency domain CELP (Code Excited Linear Prediction) coding method and a corresponding joint time domain / frequency domain CELP coding device, e.g., ACELP (Algebraic Code Excited Linear Prediction) coding method and device; [Figure 2] FIG. 2 is a schematic block diagram illustrating a more detailed structure of the joint time domain / frequency domain coding method and device of FIG. 1, showing how the pre-processor performs a first level of analysis to classify the input acoustic signal. [Figure 3]FIG. 10 is a schematic block diagram illustrating simultaneously a calculator of the cut-off frequency of the time domain excitation contribution and an overview of the corresponding operation of estimating the cut-off frequency. [Figure 4] FIG. 4 is a schematic block diagram illustrating a more detailed structure of the cutoff frequency calculator of FIG. 3 and the corresponding operation of estimating the cutoff frequency. [Figure 5] 1 is a schematic block diagram illustrating simultaneously an overview of a frequency quantizer and the corresponding frequency quantization operation. [Figure 6] FIG. 6 is a schematic block diagram of the frequency quantizer of FIG. 5 and a more detailed structure of the frequency quantization operation. [Figure 7] 1 is a schematic block diagram illustrating simultaneously an alternative implementation of a joint time-domain / frequency-domain CELP coding method and a corresponding joint time-domain / frequency-domain CELP coding device; [Figure 8] FIG. 10 is a schematic block diagram illustrating simultaneously the operation of selecting a coding sub-mode and a corresponding sub-mode selector. [Figure 9] FIG. 9 is a schematic block diagram illustrating simultaneously a band selector and a bit allocator and corresponding operations of band selection and bit allocation for allocating the available bit budget to frequency domain coding modes when the input audio signal is classified as neither speech nor music in the alternative implementation of FIGS. 7 and 8. [Figure 10] 1 is a simplified block diagram of an exemplary configuration of hardware components forming a joint time-domain / frequency-domain coding device and method for encoding an input audio signal. [Figure 11] 11 is a schematic block diagram illustrating simultaneously a decoder device 1100 and a corresponding decoding method 1150 for decoding a bitstream from the joint time domain / frequency domain coding device and corresponding joint time domain / frequency domain coding method of FIG. 7. [Figure 12]1 is a schematic block diagram illustrating simultaneously an audio signal decoder and a corresponding audio signal decoding method for decoding a bitstream from a joint time domain / frequency domain coding device and a corresponding joint time domain / frequency domain coding method for an audio signal classified into an unclear signal type category. DETAILED DESCRIPTION OF THE INVENTION
[0018] This disclosure proposes a joint time-domain and frequency-domain coding model that improves synthesis quality for general audio signals, such as music and / or reverberant speech, without increasing processing delay and bit rate. The joint time-domain and frequency-domain coding model: - a time-domain coding mode that operates in the linear prediction (LP) residual domain and dynamically allocates available bits among an adaptive codebook, one or more fixed codebooks (e.g., algebraic codebooks, Gaussian codebooks, etc.), and a variable-length fixed codebook; and - includes a frequency domain coding mode, This depends on the characteristics of the input audio signal.
[0019] For example, to realize a low-processing-delay, low-bit-rate speech audio codec that improves the synthesis quality of general audio signals, such as music and / or reverberant speech, the frequency-domain coding mode is integrated as closely as possible to the CELP (Code Excited Linear Prediction) time-domain coding mode. To this end, the frequency-domain coding mode uses a frequency transform performed in the LP (Linear Prediction) residual domain. This allows for nearly artifact-free switching from one frame, e.g., a 20 ms frame, to another. As is well known in the audio codec art, the input audio signal is sampled at a given sampling rate and processed by groups of these samples, called "frames," which are typically divided into a number of "subframes." Here, the integration of the two time-domain and frequency-domain coding modes is sufficiently close to allow for dynamic reallocation of the bit budget to the other coding mode if it is determined that the current coding mode is not efficient enough.
[0020] One feature of the proposed joint time-domain and frequency-domain coding model is the variable temporal support of the time-domain components, which varies from a quarter frame (subframe) to a full frame per frame. As a non-limiting illustrative example, a frame may represent 20 ms of the input audio signal. Such a frame corresponds to 320 samples of the input audio signal if the audio codec's internal sampling rate is 16 kHz, or 256 samples per frame if the codec's internal sampling rate is 12.8 kHz. A subframe (a quarter of a frame in our example) represents 80 or 64 samples, depending on the audio codec's internal sampling rate. In our non-limiting illustrative embodiment, the audio codec's internal sampling rate is 12.8 kHz, giving a frame length of 256 samples and a subframe length of 64 samples of the input audio signal.
[0021] Variable temporal support allows capturing key temporal events with the minimum bit rate required to generate the basic time-domain excitation contribution. At very low bit rates, the temporal support is typically the entire frame. In that case, the time-domain contribution of the excitation consists only of the adaptive codebook, and the corresponding adaptive codebook (pitch) information and gain are transmitted once per frame. When more bit rate is available, it is possible to capture more temporal events by shortening the temporal support and increasing the bit rate allocated to the time-domain coding mode. Finally, when the temporal support is sufficiently short (shorter than one-quarter of a frame (subframe)) and the available bit rate is sufficiently high, the time-domain contribution of the excitation can include, for each subframe, an adaptive codebook contribution with the corresponding adaptive codebook gain, a fixed codebook contribution with the corresponding fixed codebook gain, or both an adaptive codebook contribution and a fixed codebook contribution with the corresponding gain. Alternatively, it is also possible to transport for each half of a frame (subframe) the adaptive codebook contribution with the corresponding adaptive codebook gain and the fixed codebook contribution with the corresponding fixed codebook gain, which has the advantage of not consuming too much bit rate while still being able to code temporal events as they are. Parameters describing the codebook index and gain are then transmitted for each subframe.
[0022] At low bit rates, speech codecs cannot adequately encode higher frequencies. This causes significant degradation in synthesis quality when the input audio signal contains music and / or reverberant speech. To solve this problem, a feature is added that calculates the efficiency of the time-domain excitation contribution. In some cases, the time-domain excitation contribution is worthless, regardless of the input bit rate and time frame support. In these cases, all bits are reallocated to the next step of frequency-domain coding. However, most of the time, the time-domain excitation contribution is worthless only up to a certain frequency (here, after the "cutoff frequency"). In these cases, the time-domain excitation contribution is filtered out above the cutoff frequency. The filtering operation allows us to retain the valuable information encoded in the time-domain excitation contribution and remove the worthless information above the cutoff frequency. In a non-limiting exemplary embodiment, filtering is performed in the frequency domain by setting frequency bins above a certain frequency (the cutoff frequency) to zero.
[0023] The variable time support combined with the variable cutoff frequency makes bit allocation inside the joint time-domain and frequency-domain coding model highly dynamic. The bitrate after LP filter quantization can be allocated entirely to the time domain, entirely to the frequency domain, or somewhere in between. The bitrate allocation between the time domain and the frequency domain is performed as a function of the number of subframes used for the time-domain excitation contribution, the available bit budget, and the calculated cutoff frequency. To make the joint time-domain and frequency-domain coding model even more efficient for certain types of input acoustic signals, specific coding submodes are added that efficiently allocate available bits between the time domain, the frequency domain, and the region between low and high frequencies. These additional specific coding submodes are determined using a novel speech / music audio classifier that generates output that allows for an unclear signal category (a signal that cannot be clearly classified as either music or speech).
[0024] To generate a total excitation that more efficiently matches the input LP residual, a frequency-domain coding mode is applied. Frequency-domain coding is performed on a vector containing the difference between the frequency representation of the input LP residual (frequency transform) and the frequency representation of the filtered time-domain excitation contribution (frequency transform) up to a cutoff frequency, and the frequency representation of the input LP residual itself (frequency transform) above that cutoff frequency. A smooth spectral transition is inserted between both segments just above the cutoff frequency. In other words, the high-frequency portion of the frequency representation of the time-domain excitation contribution is first zeroed out above the cutoff frequency. A transition region between the unchanged portion of the spectrum and the zeroed portion of the spectrum of the time-domain excitation contribution is inserted just above the cutoff frequency to ensure a smooth transition between both portions of the spectrum. This modified spectrum of the time-domain excitation contribution is then subtracted from the frequency representation of the input LP residual. The resulting spectrum thereby corresponds to the difference between both spectra below the cutoff frequency and the frequency representation of the LP residual above the cutoff frequency, with some transition region. The cutoff frequency can vary from frame to frame, as described above.
[0025] No matter what frequency quantization method (frequency-domain coding mode) is selected, there is always a possibility of pre-echoes, especially when long windows are used. In the technique disclosed herein, the window used is a square window, so the extra window length compared to the encoded input audio signal is zero, i.e., no overlap-add is used. This corresponds to the best window for reducing potential pre-echoes, but some pre-echoes may still be audible in temporal attacks. While many techniques exist for solving such pre-echo problems, this disclosure proposes a simple feature to counteract this pre-echo problem. This feature is based on a memoryless time-domain coding mode derived from "Transition Mode" in ITU-T Recommendation G.718, [5], Sections 6.8.1.4 and 6.8.4.2, the entire contents of which are incorporated herein by reference. The idea behind this feature is to take advantage of the fact that the proposed joint time-domain and frequency-domain coding model is integrated in the LP residual domain, which enables switching almost always without artifacts. When the input sound signal is considered to be general audio (music and / or reverberant speech) and a temporal attack is detected within a frame, only this frame is encoded in a memoryless time-domain coding mode, which prepares for the temporal attack and thereby avoids pre-echoes that may be introduced when using frequency-domain coding for that frame.
[0026] Non-limiting exemplary embodiments In the proposed joint time-domain and frequency-domain coding model, the aforementioned adaptive codebook, one or more fixed codebooks (e.g., algebraic codebook, Gaussian codebook, etc.), i.e., the so-called time-domain codebook, and frequency-domain quantization (frequency-domain coding mode) can be considered as a codebook library, and bits can be distributed among all available codebooks or a subset thereof. This means, for example, that if the input acoustic signal is clean speech, all bits are allocated to the time-domain coding mode, essentially reducing the coding to the legacy CELP scheme. On the other hand, for some music segments, all bits allocated to encoding the input LP residual are sometimes best spent in the frequency domain, e.g., the transform domain. Furthermore, specific cases can be added in which (a) the time domain uses a larger portion of the total available bitrate to code more time-domain events while still retaining bits to encode some of the frequency information, or (b) low-frequency content is prioritized higher than high-frequency content, and vice versa.
[0027] As shown in the previous discussion, the time support for the time-domain coding mode and the frequency-domain coding mode need not be the same. The bits spent on the different time-domain coding operations (adaptive and algebraic codebook search) are usually distributed on a subframe basis (typically one-quarter of a frame, or 5 ms of time support), while the bits allocated to the frequency-domain coding mode are distributed on a frame basis (typically 20 ms of time support) to improve frequency resolution.
[0028] The bit budget allocated to the time-domain CELP coding mode can also be dynamically controlled depending on the input audio signal. In some cases, the bit budget allocated to the time-domain CELP coding mode can be zero, which effectively means that the entire bit budget is attributed to the frequency-domain coding mode. The choice of operating in the LP residual domain for both the time-domain coding mode and the frequency-domain coding mode has two main advantages. First, it is compatible with the time-domain CELP coding mode, which has proven to be efficient in speech signal coding. As a result, no artifacts are introduced due to switching between the two types of coding modes (time-domain coding mode and frequency-domain coding mode). Second, the low and relatively flat dynamics of the LP residual relative to the original input audio signal facilitates the use of square windows for the frequency transform, enabling the use of non-overlapping windows.
[0029] In a non-limiting example where the codec's internal sampling rate is 12.8 kHz (meaning 256 samples per frame), similar to ITU-T Recommendation G.718 (reference [5]), the subframe length used in the time-domain CELP coding mode can vary from a typical quarter of the frame length (5 ms) to half a frame (10 ms) or full frame length (20 ms). The subframe length determination is based on the available bit rate and an analysis of the input audio signal, in particular its spectral dynamics. The subframe length determination can be performed in a closed-loop manner. To reduce complexity, it is also possible to base the subframe length determination on an open-loop manner. The subframe length determination can also be controlled by the characteristics of the input audio signal as detected by a signal classifier, e.g., a speech / music classifier. The subframe length can be changed from frame to frame.
[0030] After the subframe length is selected for the current frame, a standard closed-loop pitch analysis is performed, and a first contribution to the excitation signal is selected from an adaptive codebook. Then, depending on the available bit budget and the characteristics of the input acoustic signal (e.g., in the case of an input speech signal), a second contribution from one or more fixed codebooks may be added before transformation in the transform domain. The resulting excitation contribution is a time-domain excitation contribution. On the other hand, at very low bit rates and for general audio signals, it is often advantageous to skip the fixed codebook stage and use all remaining bits for transform-domain coding. Transform-domain coding can be, for example, a frequency-domain coding mode. As described above, the subframe length can be 1 / 4 of a frame, 1 / 2 of a frame, or full frame length. The fixed-codebook contribution is used only when the subframe length is equal to 1 / 4 of the frame length. If the subframe length is determined to be half the frame length or the entire frame length, only the adaptive codebook contribution is used to represent the time-domain excitation contribution, and all remaining bits are allocated to the frequency-domain coding mode. Alternatively, an additional coding mode is described in which a fixed codebook may be used when the subframe length is equal to half the frame length. This addition was made to improve the quality of certain types of input acoustic signals containing temporal events while maintaining an acceptable bit budget for coding the frequency-domain excitation contribution.
[0031] After the calculation of the time-domain excitation contribution is completed, its efficiency needs to be evaluated and quantized. If the gain of coding in the time domain is very low, it is more efficient to completely remove the time-domain excitation contribution and use all bits for the frequency-domain coding mode. On the other hand, for example, if the input speech signal is clean, the frequency-domain coding mode is not needed and all bits are allocated to the time-domain coding mode. However, in many cases, coding in the time domain is only efficient up to a certain frequency. This frequency corresponds to the above-mentioned cutoff frequency of the time-domain excitation contribution. Determining such a cutoff frequency ensures that the overall time-domain coding does not work against the frequency-domain coding, but rather helps to obtain a better final synthesis.
[0032] The cutoff frequency may be estimated in the frequency domain. To calculate the cutoff frequency, the spectra of both the LP residual and the time-domain excitation contribution are first divided into a predefined number of frequency bands, each of which defines a number of frequency bins. The number of frequency bands and the number of frequency bins covered by each frequency band may vary from implementation to implementation. For each frequency band, a normalized correlation between the frequency representation of the time-domain excitation contribution and the frequency representation of the LP residual is calculated, and the correlation is smoothed between adjacent frequency bands. As a non-limiting example, the correlation per band is limited to a low value of 0.5 and normalized between 0 and 1, and then the average correlation is calculated as the average of the correlations for all frequency bands. For the purpose of initial estimation of the cutoff frequency, the average correlation is then scaled between 0 and half the internal sampling rate (half the internal sampling rate corresponding to a normalized correlation value of 1). At very low bit rates or for additional coding submodes as described herein below, the average correlation is doubled before finding the cutoff frequency. This is done even if the correlation is not very high because the bit rate used is low, or if the type of input audio signal does not allow for high correlation and a time-domain excitation contribution is known to be necessary. An initial estimate of the cutoff frequency is found as the upper limit of the frequency band closest to the value of the scaled average correlation. In one example implementation, 16 fixed frequency bands with an internal sampling rate of 12.8 kHz are defined for the correlation calculation.
[0033] Taking advantage of the psychoacoustic properties of the human ear, the reliability of the cutoff frequency estimate can be improved by comparing the estimated position of the eighth harmonic frequency of the pitch with the cutoff frequency estimated by the correlation calculation. If this position is higher than the cutoff frequency estimated by the correlation calculation, the cutoff frequency is modified to correspond to the position of the eighth harmonic frequency of the pitch. If one of the additional coding submodes is used, the cutoff frequency has a minimum value, for example, above 2775 Hz (the seventh band). The final value of the cutoff frequency is then quantized and transmitted to the remote decoder. In one implementation, 3 or 4 bits are used for such quantization, thereby providing 8 or 16 possible cutoff frequencies depending on the bit rate.
[0034] After the cutoff frequency is known, frequency quantization of the frequency-domain excitation contribution is performed. First, the difference between the frequency representation of the input LP residual (frequency transform) and the frequency representation of the time-domain excitation contribution (frequency transform) is determined. A new vector is then created consisting of this difference up to the cutoff frequency and a smooth transition to the frequency representation of the input LP residual for the remaining spectrum. Frequency quantization is then applied to the entire new vector. In one implementation, the quantization consists of encoding the sign and position of the dominant (most energetic) spectral pulse. The number of pulses to be quantized per frequency band is related to the bit rate available for the frequency-domain coding mode. If the available bits are insufficient to cover all frequency bands, the remaining bands are filled with noise only.
[0035] Frequency quantization of a frequency band using the quantization method described in the previous paragraph does not guarantee that all frequency bins within this band are quantized. This is especially true at low bit rates, where the number of quantized spectral pulses per frequency band is relatively small. To prevent the appearance of audible artifacts due to these unquantized bins, some noise is added to fill these gaps. At low bit rates, the quantized spectral pulses should dominate the spectrum rather than the inserted noise, so the noise spectral amplitude corresponds to only a small fraction of the pulse amplitude. The amplitude of the added noise in the spectrum is high when the available bit budget is low (allowing for more noise) and low when the available bit budget is high.
[0036] In the frequency-domain coding mode, a gain is calculated for each frequency band to match the energy of the unquantized signal to the quantized signal. The gain is vector-quantized and applied to the quantized signal band by band. For example, when a joint time-domain and frequency-domain coding model changes bit allocation from a time-domain-only coding mode to a mixed time-domain / frequency-domain coding mode, the excitation spectral energy per band in the time-domain-only coding mode does not match the excitation spectral energy per band in the mixed time-domain / frequency-domain coding mode. This energy mismatch can cause some switching artifacts, especially at low bit rates. To reduce any audible degradation caused by this bit reallocation, a long-term gain can be calculated for each band and applied to correct the energy of each frequency band for a few frames after switching from the time-domain-only coding mode to the mixed time-domain / frequency-domain coding mode.
[0037] After the frequency-domain coding mode is completed, the total excitation is found by adding the frequency-domain excitation contribution to the frequency representation (frequency transform) of the time-domain excitation contribution, and then the sum of these two excitation contributions is transformed back to the time domain to form the total excitation. Finally, the synthesis signal is calculated by filtering the total excitation through an LP synthesis filter.
[0038] In one embodiment, the CELP coding memories are updated on a subframe basis using only the time-domain excitation contribution, while the full excitation is used to update these memories at frame boundaries.
[0039] In another possible implementation, the CELP coding memory is updated on a subframe basis and also at frame boundaries using only the time-domain excitation contribution. This results in an embedded structure in which the frequency-domain coded signal constitutes the upper quantized layer independently of the core CELP layer. In this particular case, a fixed codebook is always used to update the contents of the adaptive codebook. However, the frequency-domain coding mode can be applied to the entire frame. This embedded approach works for bit rates of approximately 12 kbps and above.
[0040] 1) Acoustic signal type classification 1 is a schematic block diagram illustrating simultaneously an overview of a joint time-domain / frequency-domain CELP coding method 150 and a corresponding joint time-domain / frequency-domain CELP coding device 100, e.g., an ACELP method and device. Of course, other types of CELP coding methods and devices may be implemented using the same concepts.
[0041] FIG. 2 is a schematic block diagram illustrating a more detailed structure of the joint time-domain / frequency-domain CELP coding method 150 and device 100 of FIG.
[0042] The joint time-domain / frequency-domain CELP coding device 100 includes a pre-processor 102 (FIG. 1) for performing an operation 152 of analyzing parameters of an input acoustic signal 101 (FIGS. 1 and 2). Referring to FIG. 2, the pre-processor 102 includes an LP analyzer 201 for performing an operation 251 of LP analysis of the input acoustic signal 101, a spectrum analyzer 202 for performing an operation 252 of spectral analysis, an open-loop pitch analyzer 203 for performing an operation 253 of open-loop pitch analysis, and a signal classifier 204 for performing an operation 254 of classification of the input acoustic signal. The analyzers 201 and 202 and the associated operations 251 and 252 perform the LP analysis and spectral analysis typically performed in CELP coding, as described, for example, in ITU-T Recommendation G.718, [5], sections 6.4 and 6.1.4, and therefore will not be further described in this disclosure.
[0043] The pre-processor 102 performs a first level of analysis to classify the input acoustic signal 101 into speech and non-speech (general audio (music or reverberant speech)), for example in a manner similar to that described in [6], the entire contents of which are incorporated herein by reference, or using any other reliable speech / non-speech discrimination method.
[0044] After this first level of analysis, the preprocessor 102 performs a second level of analysis of the input signal parameters, allowing the use of time-domain CELP coding (without frequency-domain coding) for some acoustic signals that have strong non-speech characteristics but are better encoded with the time-domain approach. When significant fluctuations in energy occur, this second level of analysis allows the joint time-domain / frequency-domain CELP coding device 100 to switch to a memoryless time-domain coding mode, commonly referred to as a transition mode in [7], the entire contents of which are incorporated herein by reference.
[0045] In this second level of analysis, the signal classifier 204 receives a smoothed version C of the open-loop pitch correlation from the open-loop pitch analyzer 203.st fluctuations, the current total frame energy E tot (the total energy of the input acoustic signal in the current frame), and the difference E between the current total frame energy and the previous total frame energy diff First, the signal classifier 204 calculates and uses the following relationship, for example:
[0046]
number
[0047] Calculate the smoothed open-loop pitch correlation variation using where: - C st teeth C st =0.9 C ol +0.1 C st is the smoothed open-loop pitch correlation defined as - C ol is the open-loop pitch correlation calculated by analyzer 203 using methods known to those skilled in the art of CELP coding, for example as described in ITU-T Recommendation G.718, [5], section 6.6, -
[0048]
number
[0049] is the smoothed open-loop pitch correlation C st is the average over the last 10 frames i of - σ c is the variation of the smoothed open-loop pitch correlation.
[0050] When the signal classifier 204 classifies a frame as non-voice during the first level of analysis, a subsequent verification is performed by the signal classifier 204 to determine whether it is truly safe to use the mixed time-domain / frequency-domain coding mode during the second level of analysis. However, there are times when it is better to encode the current frame exclusively in the time-domain coding mode using one of the time-domain approaches estimated by the pre-processing function of the time-domain coding mode. In particular, there are times when it is better to use the memoryless time-domain coding mode to at least reduce possible pre-echoes that may be introduced in the mixed time-domain / frequency-domain coding mode.
[0051] As a non-limiting implementation of a first verification of whether a mixed time-domain / frequency-domain coding mode should be used, the signal classifier 204 calculates the difference between the current total frame energy and the total energy of the previous frame. The current total frame energy E tot and the difference between the total energy of the previous frame, E diff is higher than, for example, 6 dB, this corresponds to a so-called "temporal attack" in the input acoustic signal 101. In such a situation, the voiced / non-voiced decision and the selected coding mode are overwritten, and the memoryless time-domain coding mode is forced. More specifically, the joint time-domain / frequency-domain CELP coding device 100 includes a time / time-frequency coding selector 103 (FIG. 1) for performing an operation 153 of selecting between time-domain-only coding and mixed time-domain / frequency-domain coding. For that purpose, the time / time-frequency coding selector 103 includes a voice / general audio selector 205 (FIG. 2) for performing an operation 255 of selecting either voice or general audio for classification of the input acoustic signal 101, a time attack detector 208 (FIG. 2) for performing an operation 258 of detecting a temporal attack in the input acoustic signal 101, and a selector 206 (FIG. 2) for performing an operation 256 of selecting the memoryless time-domain coding mode. In other words, Depending on the determination of the speech signal by the selector 205, a closed-loop CELP encoder 207 (FIG. 2) is used to perform an operation 257 of CELP coding the speech signal. - In response to both the determination of a non-speech signal (general audio) by the selector 205 and the detection of a time attack in the input acoustic signal 101 by the detector 208, the selector 206 forces the closed-loop CELP encoder 207 (Fig. 2) to encode the input acoustic signal using a memoryless time-domain coding mode. The closed-loop CELP encoder 207 forms part of the time-domain only encoder 104 of Figure 1. Closed-loop CELP encoders are well known to those skilled in the art and will not be described further herein.
[0052] As a non-limiting implementation of the second verification of whether a mixed time domain / frequency domain coding mode should be used, the current frame total energy E tot and the difference between the total energy of the previous frame, E diff is 6 dB or less, except - Smoothed open-loop pitch correlation C st is greater than 0.96, or - Smoothed open-loop pitch correlation C st is higher than 0.85 and the current total frame energy E tot and the difference between the total energy of the previous frame, E diff is less than 0.3 dB, or - smoothed open-loop pitch correlation variance σ C is less than 0.1 and the current total frame energy E tot and the difference between the total energy of the previous frame and the last frame E diff is less than 0.6 dB, or - Current layer frame energy E tot is less than 20 dB, When this is at least the second consecutive frame (cnt≧2) for which the first level analysis decision is changed, the speech / general audio selector 205 decides that the current frame is to be coded using the time domain only coding mode using the closed-loop CELP encoder 207 (FIG. 2).
[0053] Otherwise, the time / time-frequency coding selector 103 selects the mixed time-domain / frequency-domain coding mode as disclosed in the following description.
[0054] The second test can be summarized using the following pseudocode, for example, when the non-speech input acoustic signal is music: if (generic audio) if (E diff >6dB) Coding mode = Time domain memoryless cnt=1 else if (C st >0.96|(C st >0.85&E diff <0.3dB)|(σ c <0.1&E diff <0.6dB)|E tot <20dB) cnt++ if(cnt>=2) Coding mode = Time domain else Coding mode = mixed time / frequency domain cnt=0 where E tot teeth
[0055]
number
[0056] is the current total frame energy expressed as x(i) represents the sample of the input acoustic signal in the current frame, N is the number of samples of the input acoustic signal per frame, and E diffis the current total frame energy E tot and the total energy of the previous frame.
[0057] FIG. 7 is a schematic block diagram illustrating simultaneously an alternative implementation of a joint time-domain / frequency-domain CELP coding method 750 and a corresponding joint time-domain / frequency-domain CELP coding device 700, in which the pre-processor 702 also performs a first level of analysis to classify the input acoustic signal 101.
[0058] In particular, the joint time-domain / frequency-domain CELP coding method 750 includes an operation 752 of pre-processing the input acoustic signal 101 as described in [4] to obtain parameters necessary for classifying this input acoustic signal. To perform operation 752, the joint time-domain / frequency-domain CELP coding device 700 includes a pre-processor 702.
[0059] The joint time-domain / frequency-domain CELP coding method 750 includes an operation 751 of classifying the input acoustic signal 101 into speech, music, and unclear signal type categories using parameters from the pre-processor 702 in a manner similar to that described in [4] or using any other reliable speech / music and unclear signal type discrimination method. The unclear signal type category indicates that the nature of the input acoustic signal 101 is unclear, and in particular, that the input acoustic signal 101 cannot be classified as either speech or music. To perform operation 751, the joint time-domain / frequency-domain CELP coding device 700 includes an acoustic signal classifier 701.
[0060] If the audio signal classifier 701 classifies the input audio signal 101 into the music category, the frequency domain encoder 703 performs an operation 753 of encoding the input audio signal 101 using frequency domain coding, for example as described in [2]. The frequency domain encoded music signal can then be synthesized in a music synthesis operation 754, performed by the synthesizer 704, to recover the music signal.
[0061] In the same way, if the acoustic signal classifier 701 classifies the input acoustic signal 101 into the speech category, the time-domain encoder 705 performs an operation 755 of encoding the input acoustic signal 101 using time-domain coding, for example as described in [2]. The time-domain coded speech signal is then synthesized in a synthesis filtering operation 756, performed by a synthesizer 706 including a synthesis filter, to recover the speech signal.
[0062] Thus, the joint time-domain / frequency-domain coding device 700 and method 750 maximize the performance of time-domain coding alone and frequency-domain coding alone by limiting their use to input audio signals with clear speech characteristics and clear music characteristics, respectively, which improves the overall quality of all types of input audio signals at low to medium bit rates.
[0063] The coding submodes are designed as part of the joint time-domain and frequency-domain coding model to efficiently code input acoustic signals that are neither speech nor music (indistinct signal type category). Two bits are used to signal three coding submodes, identified by corresponding submode flags. The fourth submode enables backward interoperability to legacy joint time-domain and frequency-domain coding models (EVS).
[0064] 8, operation 751 of classifying input acoustic signal 101 includes operation 850 of selecting one of the coding sub-modes depending on the bit rate available for encoding input acoustic signal 101 and the characteristics of this input acoustic signal classified into an unclear signal type category. To perform operation 850, acoustic signal classifier 701 incorporates sub-mode selector 800.
[0065] The encoding submode is set by the submode flag F tfsm In the non-limiting implementation of FIG. 8, the sub-mode selector 800 selects the coding sub-mode as follows: The submode selector 800 selects the above-mentioned backward coding submode (see 803) if (a) the available bit rate for coding the input audio signal 101 is not higher than 9.2 kbps and (b) the input audio signal 101 is not classified as either speech or music. Then, the submode flag F tfsm is set to "0" (see 802). The selection of backward coding mode triggers the use of the legacy joint time and frequency domain coding model (EVS) of FIGS. The submode selector 800 selects the first coding submode (see 806) if (a) the input audio signal 101 is not classified as either speech or music by the classifier 701 and the available bit rate is high enough to allow adaptive and fixed codebook and gain coding, typically higher than 9.2 kbps (see 803), (b) the probability that the input audio signal 101 is music (weighted speech / music decision with music tendency, wdlp(n)) is less than or equal to zero (see 804), and (c) no possible time attack is detected in the current frame of the input audio signal (the transition counter is less than or equal to zero, as described in ITU-T Recommendation G.718 [5], sections 6.8.1.4 and 6.8.4.2). Then, the submode flag F tfsm is set to "1" (see 801). Although the input audio signal 101 is not classified as either speech or music by the classifier 701, the selector 800 detects "speech"-like characteristics in the input audio signal 101 and selects the first coding submode (submode flag F) since CELP is not optimal for coding such audio signals. tfsm =1). The submode selector 800 selects the second coding submode (see 806) if (a) the input audio signal 101 is not classified as either speech or music by the classifier 701 and the available bit rate is high enough to allow adaptive and fixed codebook and gain coding, typically 9.2 kbps (see 803), (b) the probability that the input audio signal 101 is music (weighted speech / music decision with music tendency, wdlp(n)) is less than or equal to zero (see 804), and (c) a possible time attack is detected in the current frame of the input audio signal (the transition counter is greater than zero, as described in ITU-T Recommendation G.718 [5], sections 6.8.1.4 and 6.8.4.2). Then, the submode flag F tfsm is set to "2" (see 807). As explained below, the second coding submode (submode flag F tfsm =2) allocates more bits to the lower part of the spectrum. The submode selector 800 selects the third coding submode (see 804) if (a) the input audio signal 101 is not classified as either speech or music by the classifier 701, the available bitrate is high enough to allow at least the encoding of the adaptive codebook and gain, and still has a significant amount of bits for frequency encoding, which typically means a bitrate higher than 9.2 kbps, and (b) the probability that the input audio signal 101 is music (weighted speech / music decision with a tendency to music, wdlp(n)) is greater than "0". Then, the submode flag F tfsm is set to "3" (see 808). Although the input audio signal 101 is not classified as either speech or music by the classifier 701, the selector 800 detects "music"-like characteristics in the input audio signal 101 and selects the third coding submode (submode flag F tfsm =3). Such audio signal segments are still considered non-musical, but the submode flag F tfsm is set to "3" (selection of the third coding submode) indicating that the sample contains high frequency or tonal content. The probability that the input audio signal 101 is speech, music, or something in between is explained in [4]. When the decision of speech or music classification is ambiguous, if the probability wdlp(n) is greater than 0, the signal is considered to have some musical characteristics. The table below shows the thresholds at which the probability is high enough to be considered music or speech:
[0066] [Table 1]
[0067] The selected coding submode, e.g., submode flag F tfsm are put into the bitstream and transmitted to the distant decoder. The path selected inside the decoder depends on the signaling bits included in the bitstream. After the decoder detects the presence of a frame coded using mixed time-domain / frequency-domain coding, it sets the submode flag F tfsm is decoded from the bitstream. The detected submode flag F tfsm If the submode flag F is '0', the EVS backward interoperable legacy joint time-domain and frequency-domain coding model is used to decode the rest of the bitstream. tfsm If is different from '0', then submode decoding follows. The decoder replicates the steps followed by the encoder as described later in Section 6.2, in particular the bit distribution between the time domain and the frequency domain, and the bit allocation in the different frequency bands.
[0068] 2) Determining the subframe length In a typical CELP system, input acoustic signal samples are processed in frames of 10 to 30 ms, which are then divided into subframes for adaptive and fixed codebook analysis. For example, a 20 ms frame (256 samples at an internal sampling rate of 12.8 kHz) may be used and divided into four 5 ms subframes. Variable subframe length is a feature used to integrate the time and frequency domains into a single coding mode. The subframe length can vary from the typical 1 / 4 of the frame length to half the frame length or the full frame length. Of course, other numbers of subframes (subframe lengths) may also be implemented.
[0069] The parameter analysis operation 152 of the joint time domain / frequency domain CELP coding method 150 includes an operation 259 of determining high-level spectral dynamics of the input acoustic signal 101 and an operation 260 of calculating the number of subframes per frame, as illustrated in Figure 2. To perform operations 259 and 260, the pre-processor 102 of the joint time domain / frequency domain CELP coding device 100 includes a high-level spectral dynamics analyzer 209 and a number of subframes calculator 210, respectively.
[0070] The decision regarding the subframe length (number of subframes), i.e., the time support, depends on the available bit rate and the input audio signal analysis, in particular the high spectral dynamics of the input audio signal 101 from analyzer 209 and the smoothed open-loop pitch correlation C from analyzer 203. stThe high spectral dynamics analyzer 209 responds to the information from the spectrum analyzer 202 to determine the high spectral dynamics of the input acoustic signal 101. The high spectral dynamics are calculated as the input spectrum without the noise floor, which provides a representation of the input spectral dynamics, as described, for example, in ITU-T Recommendation G.718, [5], Section 6.7.2.2. When the average spectral dynamics of the input acoustic signal 101 in the frequency band from 4.4 kHz to 6.4 kHz as determined by the analyzer 209 is, for example, less than 9.6 dB, and the last frame is considered to have high spectral dynamics, the input acoustic signal 101 is no longer considered to have high spectral dynamics. In that case, more bits can be allocated to frequencies below 4 kHz, for example, by adding more subframes in time-domain coding mode or forcing more pulses into the lower frequency portion in frequency-domain coding mode.
[0071] On the other hand, if the increase in the average spectral dynamics of the input acoustic signal 101 relative to the average spectral dynamics of the last frame that was not deemed to have high spectral dynamics as determined by the analyzer 209 is greater than, for example, 4.5 dB, then the input acoustic signal 101 is considered to have high spectral dynamics content, for example, higher than 4 kHz. In that case, depending on the available bit rate, some additional bits may be used to encode the higher frequencies of the input acoustic signal 101, allowing for one or more frequency pulse encodings.
[0072] The subframe length as determined by calculator 210 (FIG. 2) also depends on the bit budget available for coding the input acoustic signal 101. At very low bit rates, e.g., below 9 kbps, only one subframe is available for time-domain coding; otherwise, the number of available bits is insufficient for frequency-domain coding. At moderate bit rates, e.g., between 9 kbps and 16 kbps, one subframe is used if the high frequencies contain high spectral dynamic content, and two subframes are used otherwise. For moderate to high bit rates, e.g., above about 16 kbps, the smoothed open-loop pitch correlation C defined above is used. st The four subframe case is also available if is higher than, for example, 0.8.
[0073] In the case of one or two subframes, time-domain coding is limited to only the adaptive codebook contribution (with coded pitch lag and pitch gain), i.e., the fixed codebook is not used in that case. However, in the case of four subframes, adaptive and fixed codebook contributions are allowed if the available bit budget is sufficient. The four subframe case is allowed at bit rates starting from approximately 16 kbps. Due to the bit budget limitations, at low bit rates, the time-domain excitation contribution consists of only the adaptive codebook contribution at lower bit rates. At higher bit rates, e.g., starting from 24 kbps, a fixed codebook contribution can be added. In all cases, the time-domain coding efficiency is evaluated later to determine up to which frequency (the cutoff frequency mentioned above) such time-domain coding is beneficial.
[0074] 7 and 8, the input acoustic signal 101 is classified into the unclear signal type category by the classifier 701, and the submode flag F tfsm When is greater than zero "0", the first, second, or third coding submode defined above is used.
[0075] The acoustic signal classifier 701 determines the submode flag F tfsm is set to "1" or "2" (selection of the first or second coding submode), the number of subframes is determined to be 4, which means that the content of the input audio signal 101 is close to speech (a "speech"-like characteristic or a possible time attack is detected in the input audio signal 101) and the available bit rate is less than 15 kbps. In particular: - First or second coding submode (submode flag F tfsm is set to "1" or "2"), the acoustic signal classifier 701 determines the number of subframes to be four, and then a coding mode using two subframes is selected, unless the bit rate available for encoding the input acoustic signal 101 is less than 15 kbps. In both cases, a corresponding number of fixed codebooks, i.e., fixed codebooks of two or four in number, is used. - The third encoding mode (submode flag F tfsm is set to 3, meaning that the content of the input acoustic signal 101 is close to music (i.e., characteristics similar to "music" are detected in the input acoustic signal 101), the acoustic signal classifier 701 determines that the number of subframes is 4 but the fixed codebook contribution is not used in order to keep more bits available for the frequency-domain excitation contribution, unless the bit rate available for encoding the input acoustic signal 101 is 22.6 kbps or higher.
[0076] 3) Closed-loop pitch analysis In the joint time-domain / frequency-domain CELP coding device 100 and method 150 (FIG. 1), the mixed time-domain / frequency-domain coding method 170 and corresponding mixed time-domain / frequency-domain encoder 120 are used when the selector 205 selects general audio as the classification of the input acoustic signal 101 and no temporal attack is detected in the detector 208. Alternatively, in the joint time-domain / frequency-domain CELP coding device 700 and method 750 (FIG. 7), the acoustic signal classifier 701 classifies the input acoustic signal 101 into the "ambiguous signal type" category and one of the first, second and third coding sub-modes defined above is selected (sub-mode flag F tfsm is set to "1", "2", or "3").
[0077] When a mixed time-domain / frequency-domain coding mode is used, closed-loop pitch analysis is performed, followed, if necessary, by a fixed algebraic codebook search. To that end, the mixed time-domain / frequency-domain coding method 170 / 770 includes an operation 155 for calculating a time-domain excitation contribution. To perform operation 155, the mixed time-domain / frequency-domain encoder 120 / 720 includes a time-domain excitation contribution calculator 105. The calculator 105 itself includes an analyzer 211 (FIG. 2) that is responsive to the open-loop pitch analysis performed in the open-loop pitch analyzer 203 (or preprocessor 702) and the subframe length (or number of subframes in a frame) determined by the calculator 210 or the acoustic signal classifier 701 to perform an operation 261 of closed-loop pitch analysis. Closed-loop pitch analysis is well known to those skilled in the art, and implementation examples are described, for example, in ITU-T G.718 Recommendation [5], Section 6.8.4.1.4.1. The closed-loop pitch analysis results in the calculation of pitch parameters, also known as adaptive codebook parameters, which consist mainly of a pitch lag (adaptive codebook index T) and a pitch gain (adaptive codebook gain b). The adaptive codebook contribution is typically the past excitation at delay T or an interpolated version of it. The adaptive codebook index T is encoded and transmitted to the far decoder. The pitch gain b is also quantized and transmitted to the far decoder.
[0078] When the closed-loop pitch analysis is completed in operation 261 and a fixed codebook contribution is used, the time-domain excitation contribution calculator 105 includes a fixed algebraic codebook 212 that is searched during a fixed codebook search operation 262 to find the best fixed codebook parameters, typically including a fixed codebook index and a fixed codebook gain. The fixed codebook index and gain form the fixed codebook contribution. The fixed codebook index is encoded and transmitted to the distant decoder. The fixed codebook gain is also quantized and transmitted to the distant decoder. Fixed algebraic codebooks and their search are believed to be well known to those skilled in the art of CELP coding and therefore will not be further described in this disclosure.
[0079] The adaptive codebook index and gain, and, if used, the fixed codebook index and gain, form the time-domain CELP excitation contribution.
[0080] 4) Frequency conversion In frequency domain coding of the mixed time domain / frequency domain coding mode, the two signals are represented in a transform domain, e.g., the frequency domain. In one embodiment, the time-frequency transformation may be achieved using a 256-point Type II (or Type IV) DCT (Discrete Cosine Transform), which provides a resolution of 25 Hz at an internal sampling rate of 12.8 kHz, although any other suitable transform may also be used. If other transforms are used, the frequency resolution (defined above), the number of frequency bands, and the number of frequency bins per band (defined further below) may need to be modified accordingly.
[0081] As indicated in the preceding description, in the joint time-domain / frequency-domain CELP coding device 100 and method 150 (FIGS. 1 and 2), the mixed time-domain / frequency-domain coding mode is used when the selector 205 selects general audio as the classification of the input acoustic signal 101 and no temporal attack is detected by the detector 208. Alternatively, in the joint time-domain / frequency-domain CELP coding device 700 and method 750 (FIG. 7), the mixed time-domain / frequency-domain coding mode is used when the acoustic signal classifier 701 classifies the input acoustic signal 101 into the "unclear signal type" category. The mixed time-domain / frequency-domain encoder 120 / 720 generates an input LP residual r resulting from the LP analysis operation 251 of the input acoustic signal 101 performed by the analyzer 201 (and pre-processor 702). es 2, the frequency domain excitation contribution calculator 107 (FIGS. 1 and 7) performs operation 157 of calculating the frequency domain excitation contribution according to the input LP residual r es(n), the input LP residual f res and the time-domain CELP excitation contribution f exc The frequency transformation of
[0082]
number
[0083] and
[0084]
number
[0085] It can be calculated using
[0086] where r es (n) is the input LP residual, and e td (n) is the time domain excitation contribution and N is the frame length. In one possible implementation, the frame length is 256 samples for a corresponding internal sampling rate of 12.8 kHz. The time domain excitation contribution is given by the relationship e td (n)=bv(n)+gc(n) is given by
[0087] where v(n) is the adaptive codebook contribution, b is the adaptive codebook gain, c(n) is the fixed codebook contribution, and g is the fixed codebook gain. Note that the time-domain excitation contribution may consist solely of the adaptive codebook contribution as explained in the preceding discussion.
[0088] 5) Cutoff frequency of the time domain contribution For acoustic signal samples classified as general audio (FIG. 1) or the "indistinct signal type" category (FIG. 7), the time-domain excitation contribution does not necessarily contribute significantly to the coding improvement compared to frequency-domain coding. Often, this improves the coding of the lower part of the spectrum, but minimizes the coding improvement of the higher part of the spectrum. The mixed time-domain / frequency-domain encoder 120 / 720 includes a cutoff frequency finder and filter 108 (FIGS. 1 and 7) for performing operation 158 to determine the cutoff frequency at which the coding improvement provided by the time-domain excitation contribution becomes too low to be useful. The cutoff frequency finder and filter 108 is composed of a cutoff frequency calculator 215 and a filter 216, as illustrated in FIG. 2.
[0089] The operation 265 of estimating the cutoff frequency of the time domain excitation contribution is performed by using the f res and f exc This is first accomplished by calculator 215 (FIG. 2) using computer 303 (FIGS. 3 and 4) which performs a normalized cross-correlation operation 353 for each frequency band between the frequency transform of the input LP residual 301 from calculator 107 and the frequency transform of the time-domain excitation contribution 302 from calculator 106, designated as f teeth,
[0090]
number
[0091] It is defined in Hz as follows:
[0092] For this illustrative example, the number of frequency bins j per band, B b , cumulative frequency bins per band C Bb , and the normalized cross-correlation C per frequency band i c (i) is defined as follows, for example, for a 20 ms frame with an internal sampling rate of 12.8 kHz:
[0093]
number
[0094]
number
[0095] however
[0096]
number
[0097] and
[0098]
number
[0099] is.
[0100] where B b is band B b is the number of frequency bins j per Bb is the cumulative frequency bin per band, and C c (i) is the normalized cross-correlation per frequency band i,
[0101]
number
[0102] is the excitation energy for the band, and similarly
[0103]
number
[0104] is the residual energy per band.
[0105] The cutoff frequency calculator 215 includes a frequency band cross-correlation smoother 304 (FIGS. 3 and 4) that performs several operations 354 to smooth the cross-correlation vectors between different frequency bands. More specifically, the frequency band cross-correlation smoother 304 smooths the new cross-correlation vectors
[0106]
number
[0107] For example, the following relation
[0108]
number
[0109] Calculate using In one exemplary embodiment, α=0.95, δ=(1-α), Nb=13, β=δ / 2 is.
[0110] The cutoff frequency calculator 215 calculates the first N b bands (e.g., N b = 13 represents 5575 Hz)
[0111]
number
[0112] 3 and 4. The calculator 305 (FIGS. 3 and 4) performs an operation 355 of calculating the average of
[0113] The cutoff frequency calculator 215 also includes a cutoff frequency module 306 (FIG. 3) that includes a cross-correlation limiter 406, a cross-correlation normalizer 407, and a frequency band finder 408 where the cross-correlation is lowest, as illustrated in FIG. 4. More specifically, the limiter 406 determines the cross-correlation vector
[0114]
number
[0115] and the normalizer 407 performs an operation 456 of limiting the mean of the cross-correlation vector
[0116]
number
[0117] The finder 408 performs an operation 457 of normalizing the limited mean of L between 0 and 1. f and half the internal sampling rate of the input acoustic signal 101 (F s / 2)
[0118]
number
[0119] However, the cross-correlation vector is
[0120]
number
[0121] is the last frequency L of frequency band i that minimizes the difference between f performing an operation 458 of obtaining a first estimate of the cutoff frequency by determining:
[0122]
number
[0123] where:
[0124]
number
[0125] is.
[0126] In the above relation,
[0127]
number
[0128] represents a first estimate of the cutoff frequency.
[0129] At lower bit rates, the normalized average
[0130]
number
[0131] is never really high (as is the case with the joint time-domain / frequency-domain coding device 100 and method 150 of FIG. 1), or if the submode flag F tfsm is greater than "0", i.e., when the input acoustic signal falls into the "indistinct signal type" category (as in the case of the joint time-domain / frequency-domain coding device 700 and method 750 of FIG. 7), or when
[0132]
number
[0133] To increase the value of and give more weight to the time domain excitation contribution, a normalizer 407 is used to obtain the average normalized by a fixed scaling factor.
[0134]
number
[0135] As a non-limiting example, for bit rates below 8 kbps, the cutoff frequency
[0136]
number
[0137] The first estimate of is multiplied by 2.
[0138] The accuracy of the cutoff frequency can be improved by adding the following factor to the calculation: To that end, the cutoff frequency module 306 calculates, in a corresponding operation 460, from the minimum or minimum pitch lag value of the time domain excitation contributions of the subframes of the frame, e.g., in accordance with the following relation:
[0139]
number
[0140] an extrapolator 410 (FIG. 4) for the 8th harmonic, which is calculated using F s = 12800Hz is the internal sampling rate or frequency, and N sub is the number of subframes in a frame, and T(i) is the adaptive codebook index or pitch lag for subframe i.
[0141] The cutoff frequency module 306 is the 8th harmonic
[0142]
number
[0143] More specifically, the frequency band finder 409 (FIG. 4) for the subframe i <N sub For example, the finder 409 may find the following inequality
[0144]
number
[0145] 4. The next step is to perform an operation 459 of searching for the highest frequency band in which the The band index is
[0146]
number
[0147] This indicates the band in which the 8th harmonic is likely to be located.
[0148] The cutoff frequency module 306 finally determines the final cutoff frequency f tc More specifically, the selector 411 selects the initial estimate of the cutoff frequency f from the finder 408. tc1 and the last frequency in the frequency band where the 8th harmonic from finder 409 is located.
[0149]
number
[0150] The operation 461 for maintaining the higher frequency between
[0151]
number
[0152] Run it using:
[0153] When the coding submode is used, for the joint time domain / frequency domain coding device 700 and method 750 of FIG. 7, the cutoff frequency f tc For example, the following relation
[0154]
number
[0155] It is further thresholded using
[0156] As illustrated in Figs. 3 and 4, the cut-off frequency calculator 215 further comprises a determiner 307 (FIG. 3) for performing an operation 357 of determining the number of frequency bins of the frequency band to be zeroed; the determiner 307 itself comprises an analyzer 415 (FIG. 4) for carrying out an operation 465 of analysis of the parameters and a selector 416 (FIG. 4) for carrying out an operation 466 of selecting the frequency bins to be zeroed, Filter 216 (FIG. 2) operates in the frequency domain and includes a zeroer 308 (FIG. 3) to perform filtering operation 266. A corresponding operation 358 zeros the frequency bins determined to be zeroed in determiner 307. Zeroer 308 either (a) zeros all frequency bins (zeroer 417 and corresponding zeroing operation 467 in FIG. 4), or (b) zeros the cutoff frequency f , which is complemented by a smooth transition region (filter 418 and corresponding filtering operation 468 in FIG. 4). tc The transition region is defined as the cutoff frequency f tc The cutoff frequency f tc It allows for a smooth spectral transition between the unchanged spectrum below and the zeroed bins at higher frequencies.
[0157] As a non-limiting illustrative example, the cutoff frequency f tc The analyzer 415 considers the cost of the time-domain excitation contribution to be too high when f is less than or equal to 775 Hz. The selector 416 selects all frequency bins of the frequency representation of the time-domain excitation contribution to be zeroed, and the zeroer 417 forces all frequency bins to zero, resulting in a cutoff frequency f tcAlso, the analyzer 415 forces the selector 416 to select high frequency bins above the cutoff frequency to be zeroed by the filter (zeroer) 418. Then, all bits allocated to the time domain excitation contribution are reallocated to the frequency domain coding mode. Otherwise, the analyzer 415 forces the selector 416 to select high frequency bins above the cutoff frequency to be zeroed by the filter (zeroer) 418.
[0158] Finally, the cutoff frequency calculator 215 calculates the cutoff frequency f tc , a quantized version of this cutoff frequency f tcQ 3 and 4 for performing an operation 359 of quantizing the cutoff frequency parameter to f tcQ ={0,1175,1575,1975,2375,2775,3175,3575} It is defined in Hz as follows:
[0159] A number of mechanisms are in place to prevent the quantized version from switching between 0 and 1175 during inappropriate signal segments. tc , may be used by selector 411 to stabilize the selection of long-term average pitch gain G from closed-loop pitch analyzer 211 (FIG. 2), as a non-limiting example. lt 412, open-loop pitch correlation C from the open-loop pitch analyzer 203 ol 413, and smoothed open-loop pitch correlation C st 414. To prevent switching to only frequency domain coding, the analyzer 415 responds to, for example, the condition f tc >2375Hz or f tc >1175Hz and C ol >0.7 and G lt ≧0.6 or f tc ≧1175Hz and C st >0.8 and Glt ≧0.4 or f tcQ (t -1)!=0 and C ol >0.5 and C st >0.5 and G lt ≧0.6 is satisfied, that is, f tcQ Such frequency domain coding only is not allowed when σ cannot be set to 0.
[0160] where C ol is the open-loop pitch correlation,413 and C st is C st =0.9 C ol +0.1 C st corresponds to a smoothed version of the open-loop pitch correlation 414, which is defined as: lt (item 412 in FIG. 4) corresponds to the long-term average of the pitch gain obtained by the closed-loop pitch analyzer 211 within the time-domain excitation contribution. The long-term average value of pitch gain 412 is
[0161]
number
[0162] It is defined as
[0163]
number
[0164] is the average pitch gain over the current frame. To further reduce the switching rate between frequency-domain only and mixed time-domain / frequency-domain coding, a hangover can be added.
[0165] 6) Frequency domain coding 6.1) Creating the difference vector Cutoff frequency f of the time-domain excitation contribution tcAfter determining , frequency domain encoding is performed. To perform such frequency domain encoding, the mixed time domain / frequency domain encoding method 170 / 770 includes a subtraction operation 159, a frequency quantization operation 160, and an addition operation 161. The mixed time domain / frequency domain encoder 120 / 720 includes a subtractor or calculator 109, a frequency quantizer 110, and an adder 111 to perform operations 159, 160, and 161, respectively.
[0166] 5 is a schematic block diagram illustrating an overview of the frequency quantizer 110 and the corresponding frequency quantization operation 160. Also, FIG. 6 is a schematic block diagram illustrating a more detailed structure of the frequency quantizer 110 and the corresponding frequency quantization operation 160.
[0167] The subtractor or calculator 109 (FIGS. 1, 2, 5 and 6) subtracts the cutoff frequency f of the time domain excitation contribution from zero. tc The frequency transform f of the input LP residual from the DCT213 (Figure 2) up to res 502 (FIGS. 5 and 6) (or other frequency representation) and the frequency transform f of the time-domain excitation contribution from DCT 214 (FIG. 2). exc 501 (FIGS. 5 and 6) (or other frequency representation) d The downscale coefficients 603 (FIG. 6) form the first part of the frequency transform f res before each spectral portion of 502 is subtracted from it. trans = 2 kHz (80 frequency bins in this example implementation) exc 501 (see multiplier 604 and corresponding multiplication operation 654). The result of the subtraction is the cutoff frequency f tc From f tc +f trans The difference vector f represents the frequency range up to d The frequency transform of the input LP residual, f res 502 is the difference vector f d is used for the remaining third part of the
[0168] The difference vector f resulting from applying the downscaling factor 603 d The downscaled part of can be performed by any type of fade-out function and can be shortened to only a few frequency bins, but with a cut-off frequency f tc may be omitted when it is judged that the available bit budget is sufficient to prevent energy oscillation artifacts when f is varying. For example, for 25 Hz resolution, one frequency bin f in a 256-point DCT with an internal sampling rate of 12.8 kHz bin = 25Hz, the difference vector is f d (k)=f res (k)-f exc (k), where 0≦k≦f tc / f bin
[0169]
number
[0170] where f tc / f bin <k≦(f tc +f trans ) / f bin otherwise, f d (k)=f res (k) can be constructed as where f res , f exc , and f tc is already defined in the above description.
[0171] 6.2) Frequency Domain Bit Allocation for Coding Submodes 6.2.1) Allocating some of the available bits to lower frequencies As illustrated in FIG. 7, in the joint time domain / frequency domain CELP coding method 750, the mixed time domain / frequency domain encoder 720 includes a band selector and bit allocator 707, and the mixed time domain / frequency domain coding method 770 includes a corresponding operation of band selection and bit allocation detection 757.
[0172] FIG. 9 illustrates the available bit budget for an alternative implementation of the joint time-domain / frequency-domain CELP coding method 150 / 750 of FIGS. 7 and 8 when the input audio signal 101 is classified as neither speech nor music, using the difference vector f d 8 is a schematic block diagram illustrating simultaneously the band selection and bit allocation unit 707 of FIG. 7 and the corresponding band selection and bit allocation operation 757 for distributing to the frequency quantization of
[0173] In particular, Figure 9 illustrates an innovative way in which the band selector and bit allocator 707 allocates available bits to frequency quantization when the input audio signal 101 is not classified as either speech or music, but is classified as an "indistinct signal type" according to a previously selected coding sub-mode. In Figure 9, frequency quantization is performed band-by-band. For simplicity, the frequency bands have the same number of frequency bins, which in this illustrative example is 16 frequency bins, with an internal sampling rate of 12.8 kHz. Frequency band "0" represents the lower part of the spectrum, and frequency band "15" represents the higher part of the spectrum.
[0174] To make the best possible use of the available bits for frequency quantization, the band selection and bit allocation operation 757 selects the quantized cutoff frequency f from the cutoff frequency finder and filter 108. tcQ As a function of d The method includes a first operation 951 of pre-fixing a portion of the bit budget (see 900) available for quantizing the lower frequencies of the estimator 901. To perform operation 951, the estimator 901 may, for example, calculate the following relation:
[0175]
number
[0176] Use where P Blf is the difference vector f d is the percentage of available bits allocated to frequency quantization of the lower frequencies of the kHz band. In this example, the lower frequencies refer to the first five frequency bands, or the first 2 kHz. The term L f (f tcQ ) is the quantized cutoff frequency f tcQ Refers to the number of frequency bins up to
[0177] Next, the estimator 901 calculates the coding submode flag F tfsm The percentage of available bits P allocated to frequency quantization at lower frequencies based on Blf Adjust the encoding submode flag F tfsm is set to "2" (FIG. 8), i.e., if a possible temporal attack is detected in the current frame of the input acoustic signal 101, the proportion of bits P allocated to frequency quantization of low frequencies is Blf is increased by 10% of the available bits. If a "music"-like characteristic is detected in the content of the current frame, the submode coding flag F tfsm is set to "3", the percentage of bits allocated to frequency quantization at lower frequencies, P Blf is reduced by 10% of the available bits.
[0178] 6.2.2) Estimate the number of frequency bands to quantize difference vector f d Another parameter that affects the total number of bits per frequency band available for frequency quantization is the difference vector f d The estimated maximum number of frequency bands in Bmx In the currently described illustrative example, at an internal sampling rate of 12.8 kHz, the maximum total number of frequency bands N tt is 16.
[0179] When a coding submode is used, the band selection and bit allocation operation 757 selects the difference vector f d The maximum number of frequency bands N Bmx To perform operation 952, the estimator 902 estimates the coding submode flag F tfsm is set to "1" (the first coding submode is selected), the maximum number of frequency bands N Bmx Set to "10". Encoding submode flag F tfsm is set to "2" (the second coding submode is selected), the estimator 902 estimates the maximum number of frequency bands N Bmx Set the encoding submode flag F to "9". tfsm is set to "3" (the third coding submode is selected), the estimator 902 estimates the maximum number of frequency bands N Bmx is set to "13". The estimator 902 then calculates, for example, the following relationship:
[0180]
number
[0181] Using the difference vector f d The maximum number of frequency bands to quantize as a function of the bit budget available for frequency quantization of Bmx Readjust the where B F is the difference vector f d represents the number of bits available for frequency quantization (see 900), and B T is the total bit rate available for encoding the channels being processed (see 900), and F tfsm is the submode flag (see 900), and N tt is the maximum total number of frequency bands.
[0182] The estimator 902 calculates the difference vector f dThe difference vector f is quantized relative to the number of bits allocated to the quantization of the middle and higher frequency bands. d For the purpose of such a restriction, the last lower frequency band and the first frequency band thereafter are each assigned a similar number of bits m b , or the bits allocated to frequency quantization for lower frequencies P Blf For the last frequency band to be quantized, the minimum number m of 4.5 bits is used to quantize at least one frequency pulse. p is used. Available bitrate B T If m is 15 kbps or more, the minimum number of bits m must be set to allow for quantizing more pulses per frequency band. p becomes 9. However, the total available bitrate B T is less than 15 kbps, but the submode flag F tfsm is set to "3", that is, if the content has similarity to music, the number of bits of the last frequency band to be frequency quantized m p is 6.75 to allow for more accurate quantization. The estimator 902 then calculates the correct maximum number of frequency bands.
[0183]
number
[0184] For example, the following relation
[0185]
number
[0186] Calculate using where:
[0187]
number
[0188] corresponds to the maximum number of frequency bands to be quantized, and N Bmx is the estimated maximum number of frequency bands, the number "5" represents the minimum number of frequency bands, and B F is the difference vector f d represents the number of bits available for frequency quantization of P Blf is the fraction of bits allocated to quantizing the five lower frequency bands, and m p is the minimum number of bits allocated to frequency quantize a frequency band, and m b is the number of bits allocated to quantization of the first frequency band after the five lower frequency bands.
[0189] After calculating the maximum number of frequency bands, the estimator 902 calculates m p m b An additional verification may be performed such that it remains: This additional verification is an optional step, but at low bit rates this is d This helps allocate bits more efficiently among the frequency bands.
[0190] 6.2.3) Modify the number of bits allocated to lower frequencies The band selection and bit allocation operation 757 includes an operation 953 of calculating low frequency bits. A calculator 903 is provided to perform operation 953.
[0191]
number
[0192] If the number of frequency bands to be quantized is reduced as a result of the calculation of
[0193]
number
[0194] reallocating a portion of the bits previously allocated to the higher frequency bands such that they are no longer relevant to the quantization of the lower frequency bands using where B LF corresponds to the bits allocated to the five lower frequency bands, and B F is the difference vector f d corresponds to the number of bits available to frequency quantize the lower frequencies of Blf is the above-mentioned proportion of bits from the estimator 901 allocated to frequency quantization of, for example, the five lower frequency bands, and m p is the minimum number of bits allocated to quantize a frequency band, and m b is the number of bits allocated to quantize the first frequency band after the five (5) lower frequency bands.
[0195] 6.2.4) Double sorting of frequency bands The band selection and bit allocation operation 757 includes an operation 954 of characterizing frequency bands. To perform operation 954, the band selector and bit allocator 707 includes a frequency band characterizer 904 that performs a double sort of the frequency bands to determine the importance of each band after the lower bit rate frequency bands have been distributed between them and the remainder of those frequency bands. The first sort involves finding whether one or more bands have lower energy compared to adjacent frequency bands. When that occurs, the characterizer 904 determines whether one or more bands have a lower energy than a predetermined minimum number m, even if the available bit budget is high. p These bands are marked so that only bits of the low-energy frequency bands can be allocated to the frequency quantization of these bands. The second sorting involves, for example, performing a position sorting of the medium and high-energy frequency bands in descending order of energy. These first and second sortings (double sorting) are not performed for the lower frequency bands, but for the maximum number of frequency bands.
[0196]
number
[0197] The frequency band characterization operation 954 is performed by the formula
[0198]
number
[0199] It can be summarized as follows: where P pb (i) is the minimum number m p Only the bits in the ,are set to "1" for the frequency bands used,
[0200]
number
[0201] contains the positions of the mid- and higher-energy frequency bands in descending order of energy, and E(i) corresponds to the energy of each band. C Bb and B b is defined in Section 5 above. The difference vector f d is already defined in Section 6.1.
[0202] difference vector f d The energy E(i) of each frequency band of is calculated in calculator 708 and corresponding operation 758 of Figures 7 and 9. Calculator 708 and operation 758 also calculate the gain per frequency band as described with reference to calculator 615 and operation 665 of Figure 6. The difference vector f d 7 for the joint time-domain / frequency-domain coding device 700 and method 750, calculator 708 and operation 758 replace calculator 615 and operation 665 and quantizer 616 and operation 666.
[0203] 6.2.5) Distributing bits to selected bands The band selection and bit allocation operation 757 includes an operation of final allocation of bits per frequency band 955. To perform operation 955, the band selector and bit allocator 707 includes a final allocation of bits per frequency band 905.
[0204] After the frequency bands are characterized, divider 905 divides the difference vector f between the selected frequency bands. d The bit rate or number of bits available to frequency quantize B F Assign.
[0205] In a non-limiting example, for the first five lower frequency bands, the divider 905 divides the bits B allocated for frequency quantization of the lower frequencies into LF The first lowest frequency band is the bit B LF The fifth lower frequency band receives 23% of the bit B LF In this way, the difference vector f d The lower frequencies of the spectrum can be quantized with sufficient precision to recover a higher quality synthesis of the input acoustic signal 101.
[0206] The divider 905 divides the difference vector f d The remaining bits B allocated for frequency quantization of F , as a linear function across the other mid- and higher frequency bands, again taking into account the energy characterization of the previous frequency band (operation 954) so that more bits can be allocated to frequency bands with higher energy and fewer bits can be allocated to frequency bands that have lower energy compared to the energy of their adjacent frequency bands, thereby distributing the difference vector f d By more accurately quantizing the more significant parts of the spectrum of , the available bits are better utilized. As a non-limiting example,
[0207]
number
[0208] illustrates how bit allocation (operation 955) can be performed: where B p (i) represents the number of allocated bits per frequency band i, and B F is the difference vector f d represents the number of bits available to frequency quantize B LF corresponds to the bit rate or bits allocated to the five lower frequency bands, and m p is the minimum number of bits to quantize a frequency pulse in a frequency band, and P pb (i) is the minimum number of bits m p including where
[0209]
number
[0210] is the maximum number of frequency bands to be quantized.
[0211] If there are any unallocated bits after operation 955, divider 905 allocates them to lower frequency bands. As a non-limiting example, divider 905 allocates one remaining bit per frequency band, starting with the fifth band and returning to the first band, repeating this procedure as necessary to allocate all remaining bits.
[0212] Subsequently, the divider 905 may have to perform a floor, truncate, or round on the number of bits per frequency band depending on the algorithm used to perform the quantization of the frequency pulses and the potential fixed-point implementation.
[0213] 6.3) Search for frequency pulses The mixed time-domain / frequency-domain CELP coding method 170 / 770 uses the difference vector fd 1, 2, and 7) to frequency quantize the CELP signal. To perform operation 160, the mixed time-domain / frequency-domain CELP encoder 120 / 720 includes a frequency quantizer 110 (219 in FIG. 2).
[0214] difference vector f d can be quantized using several methods. In each case, the frequency pulses must be searched and quantized. In one possible implementation, the frequency quantizer 110 calculates the difference vector f across the spectrum. d The method for searching for pulses can be as simple as dividing the spectrum into frequency bands and allowing a certain number of pulses per frequency band. The number of pulses per frequency band depends on the available bit budget and the position of the frequency band in the spectrum. Typically, more pulses are assigned to lower frequencies.
[0215] 6.4) Quantized Difference Vector Depending on the available bit rate, the quantization of the frequency pulses may be performed by the frequency quantizer 110 using different techniques. In one implementation, for bit rates below 12 kbps, a simple search and quantization scheme may be used to encode the position and sign of the pulses. This scheme is described herein below as a non-limiting example.
[0216] For frequencies below 3175 Hz, a simple search and quantization scheme uses an approach based on factorial pulse coding (FPC), as described in, for example, [8], the entire contents of which are incorporated herein by reference.
[0217] 5 and 6, frequency quantizer 110 includes a selector 504 that performs an operation 554 of determining whether all of the spectrum is to be quantized using FPC. As illustrated in FIG. 5, if selector 504 determines that all of the spectrum is not to be quantized using FPC, an operation 556 of FPC encoding and pulse position and sign encoding is performed in encoder 506.
[0218] 6, the FPC encoding and pulse position and sign encoding operation 556 includes a frequency pulse search operation 659, an FPC encoding operation 660, a most energetic pulse finder operation 661, and a frequency pulse position and sign quantizer operation 662. To perform operations 659-662, the encoder 506 includes a frequency pulse searcher 609, an FPC encoder 610, a most energetic pulse finder 611, and a frequency pulse position and sign quantizer 612, respectively.
[0219] The searcher 609 searches for frequency pulses throughout the frequency band for frequencies below 3175 Hz. The FPC encoder 610 then processes the frequency pulses. The finder 611 determines the most energetic pulse for frequencies above 3175 Hz, and the quantizer 612 encodes the position and sign of the most energetic pulse found. If more than one pulse is allowed within the frequency band, the amplitude of the previously found pulse is divided by two, and the search is performed again across the frequency band. Each time a pulse is found, its position and sign are stored for the quantization and bit-packing stages. The following pseudocode
[0220]
number
[0221] illustrates, as a non-limiting example, this simple search and quantization scheme: where N BDis the number of frequency bands (N BD =16), N p is the number of pulses i to be coded in frequency band k, and B b is the number of frequency bins per frequency band, and C Bp is the cumulative frequency bin per band as already defined in the previous section 5, and p p represents a vector containing the found pulse positions, and p s represents a vector containing the signs of the found pulses, and p max represents the energy of the pulse found.
[0222] For bit rates above 12 kbps, the selector 504 determines that all spectrum should be quantized using FPC (FIGS. 5 and 6). As illustrated in FIG. 5, an operation 555 of FPC encoding is then performed in the FPC encoder 505. Referring to FIG. 6, the encoder 505 includes a searcher 607 for a frequency pulse, and operation 555 includes a corresponding operation 667 for searching for a frequency pulse. The search for the frequency pulse is performed throughout the frequency band. Operation 555 includes an operation 668 for encoding the found frequency pulse, and the encoder 505 includes an FPC processor 608 for performing operation 668.
[0223] The FPC processor 608 or pulse position and sign quantizer 612 then calculates the pulse sign p s The number of pulses with nb_pulses is calculated by the position p s quantized difference vector f by adding dQ For each frequency band, we obtain the quantized difference vector f dQ For example, the following pseudocode for j=0,..., j <nb_pulses f dQ (p p (j))+=p s (j) can be described using
[0224] 6.5) Noise Filling Although frequency bands may be quantized with greater or less precision, the quantization methods described in the previous section do not guarantee that all frequency bins within a frequency band are quantized. This is especially true at low bit rates, where the number of quantized pulses per frequency band is relatively small. To prevent the appearance of audible artifacts due to these unquantized frequency bins, the frequency quantizer 110 includes a noise filler 507 (FIG. 5) that performs a corresponding operation 557 of adding some noise to the unquantized frequency bins to fill these gaps. This noise addition may be performed across the entire spectrum for bit rates below 12 kbps, for example, but may be limited to the cutoff frequency f of the time-domain excitation contribution at higher bit rates. tc For simplicity, the noise intensity only varies with the available bit rate: at high bit rates the noise level is low, but at low bit rates the noise level is high.
[0225] The noise filler 507 includes an adder 613 (FIG. 6) which adds the quantized difference vector f after the intensity or energy level of such added noise is determined. dQ To that end, the frequency quantization operation 160 includes an operation 664 of estimating the intensity or energy level of the added noise, and the frequency quantizer 110 includes a corresponding estimator 614 of the noise energy level for performing operation 664. The operation 664 of estimating the intensity or energy level of the added noise is performed by the estimator 614 before an operation 665 of determining the gain per frequency band in a gain per band calculator 615 of the frequency quantizer 110.
[0226] In the illustrated embodiment, the noise level in the estimator 614 is directly related to the coding bit rate. For example, at 6.60 kbps, the estimator 614 estimates the noise level N Lis set to 0.4 times the amplitude of the frequency pulse encoded in a particular frequency band, and gradually reduced to a value of 0.2 times the amplitude of the frequency pulse encoded in the frequency band at 24 kbps. The adder 613 injects noise only into parts of the spectrum where a certain number of consecutive frequency bins have very low energy, for example, when the energy of the cumulative bins of half the frequency band is less than 0.5. For a particular frequency band i, the noise is, for example,
[0227]
number
[0228] is injected as Here, for band i, C Bb is the cumulative number of frequency bins per frequency band, and B b is the number of frequency bins in a particular band i, and N L is the level of added noise, and r and is a random number generator bounded between -1 and 1.
[0229] 6.6) Gain Quantization per Band 5 and 6, the frequency quantization operation 160 of the joint time-domain / frequency-domain coding device 100 and method 150 includes an operation 665 of determining a gain per frequency band, followed by an operation 666 of quantizing the gain per band. The frequency quantizer 110 includes a per-band gain calculator 615 and a per-band gain quantizer 616 to perform operations 665 and 666.
[0230] quantized difference vector f with noise fill if necessary dQ After is found, calculator 615 calculates the per-band gain for each frequency band. The per-band gain G for a particular band b (i) is, for example, the following relation:
[0231]
number
[0232] to obtain the unquantized difference vector f in the logarithmic domain using d Energy of and quantized difference vector f dQ is defined as the ratio between the energy of where C Bb and B b is defined herein in section 5 above.
[0233] The vector of the per-band gain quantizer 616 quantizes the per-band frequency gains. Before vector quantization, at low bit rates, the last gain (corresponding to the last frequency band) is quantized separately, and the remaining 15 per-band gains (e.g., when 16 frequency bands are used) are divided by the last quantized gain. The normalized 15 remaining gains are then vector quantized by the quantizer 616. At higher bit rates, the average per-band gain is first quantized and then subtracted from all per-band gains, e.g., for 16 frequency bands, before vector quantization of those per-band gains. The vector quantization used can be a standard minimization in the logarithmic domain of the distance between the vector containing the per-band gains and a particular codebook entry.
[0234] In frequency domain coding mode, the gain is calculated in calculator 615 for each frequency band and the unquantized vector f d The energy of the quantized vector f dQ The gain is vector quantized in quantizer 616 and multiplied by multiplier 509 (FIGS. 5 and 6) to produce a quantized vector f dQ is applied to each frequency band (operation 559).
[0235] Alternatively, it is possible to use the FPC coding scheme at rates lower than 12 kbps for the entire spectrum by selecting only a portion of the frequency bands to be quantized. Before performing the frequency band selection, the unquantized difference vector f dEnergy E in the frequency band d is quantized using quantizer 616. The energy can be calculated, for example, by the following relationship:
[0236]
number
[0237] is calculated using where C Bb and B b is defined herein in section 5 above.
[0238] Frequency band energy E d To perform the quantization of ', first the average energy over the first 12 frequency bands of the 16 bands used is quantized and subtracted from the energy of all 16 bands. Then all frequency bands are vector quantized in groups of three or four bands. The vector quantization used can be a standard minimization in the logarithmic domain of the distance between a vector containing the gain per band and a particular codebook entry. If sufficient bits are not available, it is possible to quantize only the first 12 frequency bands and extrapolate the last four frequency bands using the average of the previous three frequency bands or by any other method.
[0239] After the energy of the frequency bands of the unquantized difference vector is quantized, it is possible to sort the energy in descending order in a way that can be replicated at the decoder side. During sorting, all energy bands below 2 kHz are always kept, and only the most energetic bands are then passed to the FPC scheme to encode the amplitude and sign of the frequency pulses. This approach allows the FPC scheme to encode smaller vectors but cover a wider frequency range. In other words, fewer bits are needed to cover significant energy events across the entire spectrum.
[0240] In the particular case of an implementation of the joint time domain / frequency domain coding device 700 and method 750 of FIG. 7, frequency band selection and bit allocation are instead performed as determined by the energy per band and gain per band calculator 708 and calculation operation 758 and the band selector and bit allocator 707 and band selection and bit allocation operation 757 of FIGS. 7 and 9, as described above in this specification.
[0241] After the pulse quantization process, a noise filter similar to that described earlier is performed. Then, the gain adjustment factor G a is calculated for each frequency band, and the quantized difference vector f dQ Energy E dQ the unquantized difference vector f d quantized energy E d Then, this per-band gain adjustment factor is adjusted to match the quantized difference vector f dQ This applies to
[0242]
number
[0243] can be expressed as where
[0244]
number
[0245] and E d ' is the unquantized difference vector f as defined previously d is the quantized energy per band.
[0246] After the frequency domain coding stage is completed, the full time / frequency domain excitation is found. To that end, the mixed time / frequency domain CELP coding method 170 / 770 uses adder 111 (FIGS. 1, 2, 5, and 6) of the mixed time / frequency domain CELP encoder 120 / 720 to add the frequency quantized difference vector f from frequency quantizer 110. dQ , the filtered frequency transformed time domain excitation contribution f excF When the joint time-domain / frequency-domain coding device 100 / 700 changes its bit allocation from a time-domain-only coding mode to a mixed time-domain / frequency-domain coding mode, the excitation spectral energy per frequency band in the time-domain-only coding mode does not match the excitation spectral energy per frequency band in the mixed time-domain / frequency-domain coding mode. This energy mismatch can cause switching artifacts that are more audible at low bit rates. To reduce the sound quality degradation caused by this bit reallocation, a long-term gain can be calculated for each band and applied to the summed excitation to correct the energy of each frequency band for several frames after the reallocation. The mixed time-domain / frequency-domain CELP coding method 170 / 770 then calculates the frequency quantized difference vector f dQ and the frequency-transformed and filtered time-domain excitation contribution f excF 1, 5, and 6, which transforms the sum of .times. ...
[0247] The joint time-domain / frequency-domain encoding method 150 / 750 includes an operation 163 / 756 of generating a synthesis signal by filtering the total time-domain / frequency-domain excitation from the IDCT 220 through the LP synthesis filter 113 / 706 (FIGS. 1, 2, and 7) of the encoding device 100 / 700.
[0248] quantized difference vector f dQ The quantized positions and signs of the frequency pulses forming .times. ...
[0249] In one non-limiting embodiment, the CELP coding memories are updated on a subframe basis using only the time-domain excitation contribution, while the full time-domain / frequency-domain excitation is used to update these memories at frame boundaries. In another possible implementation, the CELP coding memories are updated on a subframe basis and also at frame boundaries using only the time-domain excitation contribution. This results in an embedded structure in which the frequency-domain quantized signal constitutes the upper quantized layer, independent of the core CELP layer. This provides advantages in certain applications. In this particular case, a fixed codebook is always used to maintain good perceptual quality, and for the same reason, the number of subframes is always four. However, the frequency-domain analysis can be applied to the entire frame. This embedded approach works for bit rates of approximately 12 kbps and above.
[0250] 7) Decoder Device and Method FIG. 11 is a schematic block diagram simultaneously illustrating a decoder device 1100 and a corresponding decoding method 1150 for decoding a bitstream 1101 from the joint time domain / frequency domain coding device 700 and the corresponding joint time domain / frequency domain coding method 750 described above.
[0251] The decoder device 1100 includes a receiver (not shown) for receiving a bitstream 1101 from the joint time-domain / frequency-domain coding device 700 .
[0252] If the audio signal coded by the joint time domain / frequency domain coding device 700 is classified as "music", this is indicated in the bitstream 1101 by a corresponding signaling bit and is detected by the decoder device 1100 (see 1102). The received bitstream 1101 is then decoded by a "music" decoder 1103, e.g. a frequency domain decoder.
[0253] If the acoustic signal coded by the joint time domain / frequency domain coding device 700 is classified as "speech", this is indicated in the bitstream 1101 by a corresponding signaling bit and is detected by the decoder device 1100 (see 1104). The received bitstream 1101 is then decoded by a "speech" decoder 1105, for example a time domain decoder using ACELP (Algebraic Code Excited Linear Prediction) or more generally CELP (Code Excited Linear Prediction).
[0254] If the audio signal coded by the joint time domain / frequency domain coding device 700 is not classified as either "music" or "speech" (see 1102 and 1104) and the bit rate available for coding the audio signal is 9.2 kbps or less (see 1106), this is reflected by the submode flag F set to "0". tfsm The received bitstream 1101 is then decoded using the backward coding mode, i.e., the legacy joint time domain and frequency domain coding model of Figures 1 and 2 (EVS), as shown at 1107.
[0255] Finally, if the audio signal coded by the joint time domain / frequency domain coding device 700 is not classified as either "music" or "speech" (see 1102 and 1104) and the bit rate available for coding the audio signal is higher than 9.2 kbps (see 1106), this indicates a submode flag F set to "1", "2", or "3". tfsm The received bitstream 1101 is then decoded using the audio signal decoder 1200 and corresponding audio signal decoding method 1250 of FIG.
[0256] 7.1) Audio signal decoder and decoding method FIG. 12 is a schematic block diagram illustrating simultaneously an audio signal decoder 1200 and a corresponding audio signal decoding method 1250 for decoding the bitstream from the joint time domain / frequency domain coding device 700 and the corresponding joint time domain / frequency domain coding method 750 described above in the case of an audio signal that falls into the unclear signal type category.
[0257] As mentioned in the preceding description, the adaptive codebook index T and the adaptive codebook gain b are quantized and transmitted, and therefore received in the bitstream by a receiver (not shown). Similarly, when used, the fixed codebook index and fixed codebook gain are also quantized and transmitted to the decoder, and therefore received in the bitstream 1101 by a receiver (not shown). The acoustic signal decoding method 1250 includes an operation 1256 of calculating a decoded time-domain excitation contribution using the adaptive codebook index and gain, and the fixed codebook index and gain, if used, as commonly done in the art of CELP coding. To perform operation 1256, the acoustic signal decoder 1200 includes a decoded time-domain excitation contribution calculator 126.
[0258] The acoustic signal decoding method 1250 also includes an operation 1257 of calculating a frequency transform of the decoded time domain excitation contribution using the same procedure as operation 156 using a DCT transform. To perform operation 1257, the acoustic signal decoder 1200 includes a calculator 1207 of the frequency transform of the decoded time domain excitation contribution.
[0259] As mentioned in the previous discussion, the quantized version of the cutoff frequency, f tcQ is transmitted to the decoder and thus received in the bitstream 1101 by a receiver (not shown). The audio signal decoding method 1250 converts the decoded cutoff frequency f recovered from the bitstream 1101 tcQand operation 1258 of filtering the frequency transform of the time-domain excitation contribution from calculator 1207 using the same or similar procedure as previously described filtering operation 266. To complete operation 1258, acoustic signal decoder 1200 filters the recovered cutoff frequency f tcQ 2. The filter 1208 includes a filter 1208 for frequency translation of the time domain excitation contribution using: Filter 1208 has the same, or at least a similar structure, as filter 216 of FIG.
[0260] The filtered frequency transform of the time domain excitation contribution from filter 1208 is fed to the positive input of adder 1209 which performs a corresponding summing operation 1259 .
[0261] The acoustic signal decoding method 1250 generates a difference vector f d 11. To perform operation 1260, the acoustic signal decoder 1200 includes a calculator 1210. In particular, the calculator 1210 dequantizes the quantized energy per frequency band and the quantized gain per frequency band received in the bitstream 1101 by the receiver (not shown) from the joint time-domain / frequency-domain coding device 700 using an inverse procedure to that described in this disclosure for quantization.
[0262] The acoustic signal decoding method 1250 includes a frequency quantized difference vector f dQ7. To perform operation 1261, the acoustic signal decoder 1200 includes a calculator 1211 that extracts the quantized positions and signs of the frequency pulses from the bitstream 1101 and replicates the selection of frequency bands to be used for quantization in the different frequency bands and the bit allocation in the different frequency bands as determined by operation 757 and the allocator 707 and employed by the joint time-domain / frequency-domain coding device 700 to encode the input acoustic signal. The calculator 1211 uses this replicated information to derive a frequency quantization difference vector f from the extracted frequency pulse quantization positions and signs. dQ In particular, for that purpose, the audio signal decoder 1200 recovers the frequency-quantized difference vector f dQ Depending on the number of bits (bit rate) available in the decoder 1200 for (see 1220), the total bit rate available for the channel being processed (see 1220), and the submode flag (see 1220), the procedure used in the joint time domain / frequency domain coding device 700 as illustrated in FIG. 9 is replicated.
[0263] In particular: The estimator 1201 and operation 1251 in FIG. 12 calculates the quantized cutoff frequency f tcQ The difference vector f as a function of d 9 corresponds to estimator 901 and operation 951 in FIG. 9, which pre-fixes a portion of the bit budget available for quantizing the lower frequencies of . The estimator 1202 and operation 1252 of FIG. 12 calculates the quantized difference vector f dQ The maximum number of frequency bands N Bmx , which corresponds to estimator 902 and operation 952 in FIG. Calculator 1203 and operation 1253 in FIG. 12 correspond to calculator 903 and operation 953 in FIG. 9, which calculate the lower frequency bits. Calculator 1204 and operation 1254 of FIG. 12 correspond to characterizer 904 and operation 954 of FIG. 9, which perform frequency band characterization. Distributor 1205 and operation 1255 of FIG. 12 correspond to distributor 905 and operation 955 of FIG. 9, which perform the final distribution of bits per frequency band.
[0264] The acoustic signal decoding method 1250 converts the recovered frequency quantized difference vector f from the calculator 1211 into dQ and the frequency-translated and filtered time-domain excitation contribution f from filter 1208 excF to form a mixed time domain / frequency domain excitation.
[0265] As can be seen, estimators 1201 and 1202, calculator 1203, characterizer 1204, distributor 1205, calculators 1206 and 1207, filter 1208, calculators 1210 and 1211, and summer 1212 form a mixed time-domain / frequency-domain excitation reconstructor using information conveyed in bitstream 1101, including a submode flag that identifies one of the coding submodes selected and used to code an acoustic signal classified into the unclear signal type category.
[0266] In the same manner, operations 1251 - 1261 form a method for reconstructing a mixed time-domain / frequency-domain excitation using the information conveyed in bitstream 1101 .
[0267] The acoustic signal decoder 1200 includes a transformer 1212 that performs an operation 1262 to transform the mixed time domain / frequency domain excitation back to the time domain using, for example, an IDCT (inverse DCT) 220 .
[0268] Finally, a synthesized acoustic signal is computed in the decoder 1200 by operation 1263 of filtering the total excitation from the transducer 1212 through an LP (Linear Prediction) synthesis filter 1213. Of course, the LP parameters needed by the decoder 1200 to reconstruct the synthesis filter 1213 are transmitted from the joint time-domain / frequency-domain coding device 700 and extracted from the bitstream 1101, as is well known in the art of CELP coding.
[0269] 8) Hardware implementation FIG. 10 is a simplified block diagram of an exemplary configuration of hardware components forming the joint time domain / frequency domain coding device 100 / 700 and method 150 / 750, decoder device 1100 and decoding method 1150 described above.
[0270] The joint time-domain / frequency-domain coding device 100 / 700 and decoder device 1100 may be implemented as part of a mobile terminal, as part of a portable media player, or in any similar device. The device 100 / 700 and decoder device 1100 (identified as 1000 in FIG. 10) includes an input 1002, an output 1003, a processor 1001, and a memory 1004.
[0271] The input 1002 is configured to receive the input audio signal 101 / bitstream 1101 of Figures 1 and 7 in digital or analog form. The output 1003 is configured to provide an output signal. The input 1002 and the output 1003 may be implemented in a common module, for example a serial input / output device.
[0272] The processor 1001 is operatively connected to an input 1002, an output 1003, and a memory 1004. The processor 1001 may be implemented as one or more processors for executing code instructions supporting the functionality of various components of the joint time-domain / frequency-domain coding device 100 / 700 for encoding an input audio signal as illustrated in FIGS. 1-9 or the decoder device 1100 of FIGS.
[0273] The memory 1004 may include non-transitory memory for storing code instructions executable by the processor 1001, in particular processor-readable memory containing / storing non-transitory instructions that, when executed, cause the processor to implement operations and components of the joint time-domain / frequency-domain coding device 100 / 700 and method 150 / 750 and decoder device 1100 and decoding method 1150 described in this disclosure. The memory 1004 may also include random access memory or buffers for storing intermediate processed data from various functions performed by the processor 1001.
[0274] Those skilled in the art will understand that the descriptions of the joint time-domain / frequency-domain coding device 100 / 700 and method 150 / 750, and decoder device 1100 and decoding method 1150 are illustrative only and are not intended to be limiting in any way. Other embodiments will be readily apparent to those skilled in the art having the benefit of this disclosure. Furthermore, the disclosed joint time-domain / frequency-domain coding device 100 / 700 and method 150 / 750, decoder device 1100 and decoding method 1150 can be customized to provide useful solutions to existing needs and problems in encoding and decoding audio.
[0275] For clarity, not all of the common features of implementations of the joint time domain / frequency domain coding device 100 / 700 and method 150 / 750 and decoder device 1100 and decoding method 1150 have been shown and described. Of course, it will be understood that in developing any such implementation of the joint time domain / frequency domain coding device 100 / 700 and method 150 / 750 and decoder device 1100 and decoding method 1150, numerous implementation-specific decisions may need to be made to achieve the developer's particular goals, such as compliance with application-related, system-related, network-related, and business-related constraints, and that these particular goals will vary from implementation to implementation and from developer to developer. Furthermore, it will be understood that the development effort may be complex and time-consuming, but is nevertheless a routine exercise in device design for those skilled in the art of audio processing having the benefit of this disclosure.
[0276] In accordance with this disclosure, the components / processors / modules, processing operations, and / or data structures described herein may be implemented using various types of operating systems, computing platforms, network devices, computer programs, and / or general-purpose machines. In addition, those skilled in the art will recognize that less general-purpose devices, such as hardwired devices, field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), or the like, may also be used. When a method including a series of operations and sub-operations is implemented by a processor, computer, or machine, and the operations and sub-operations can be stored as a series of non-transitory code instructions readable by the processor, computer, or machine, they can be stored on a tangible and / or non-transitory medium.
[0277] The joint time domain / frequency domain coding device 100 / 700 and method 150 / 750, and decoder device 1100 and decoding method 1150 described herein may use software, firmware, hardware, or any combination of software, firmware, and hardware suitable for the purposes described herein.
[0278] In the joint time domain / frequency domain coding device 100 / 700 and method 150 / 750 and decoder device 1100 and decoding method 1150 described in this specification, various operations and sub-operations are performed in various orders, and some of the operations and sub-operations may be optional.
[0279] Although the present disclosure has been described hereinabove using non-limiting exemplary embodiments thereof, these embodiments may be modified at will within the scope of the appended claims without departing from the spirit and nature of the present disclosure. 9)References This disclosure refers to the following references, the entire contents of which are incorporated herein by reference: JPEG0007764480000056.jpg189162 [Explanation of symbols]
[0280] 100 Integrated time-domain / frequency-domain CELP coding device 101 Input acoustic signal 102 Preprocessor 103 Time / Time-Frequency Encoding Selector 104 Time-Domain Only Encoder 105 Time Domain Excitation Contribution Calculator 106 Calculator 107 Calculator 108 Cutoff Frequency Finder and Filter 109 Subtractor or Calculator 110 Frequency quantizer 111 Adder 120, 720 mixed time domain / frequency domain encoder 126 Calculator 150 Integrated Time Domain / Frequency Domain CELP Coding Method 152 operation 155 operation 156 operation 158 operation 159 Subtraction Operation 160 Frequency Quantization Operation 161 Addition Operation 170, 770 mixed time-domain / frequency-domain encoding method 201 LP analyzer 202 Spectrum Analyzer 203 Open Loop Pitch Analyzer 204 Signal classifier 205 Voice / General Audio Selector 206 Selector 207 Closed-Loop CELP Encoder 208 Time Attack Detector 209 High Spectral Dynamics Analyzer 210 Subframe Calculator 211 Analyzer 212 Fixed Algebraic Codebook 213 DCT 214 DCT 215 Cutoff Frequency Calculator 216 filters 220 IDCT (inverse DCT) 251 operation 252 operation 253 operation 254 operation 255 operation 256 operation 257 operation 258 operation 259 operations 260 operations 261 operation 262 operation 265 operation 266 Filtering Actions 301 Input LP Residual 302 Time domain excitation contribution 303 Computer 304 Smoother 305 Calculator 306 Cutoff Frequency Module 307 Determiner 308 Zeroizer 309 Quantizer 353 operation 354 operation 355 operation 357 operation 358 operation 359 operation 406 Cross-correlation limiter 407 Normalizer 408 Finder 409 Finder 410 Extrapolator 411 Selector 412 Long-term average pitch gain G lt 413 Open Loop Pitch Correlation C ol 414 Smoothed Open-Loop Pitch Correlation C st 415 Analyzer 416 Selector 417 Zeroizer 418 filters 456 operation 457 operation 459 operation 460 operation 465 operation 466 operation 467 operation 468 Filtering Actions 501 Frequency Conversion f exc 502 Frequency Conversion fres 504 Selector 505 FPC encoder 506 encoder 507 Noise Filler 509 Multiplier 554 operation 555 operation 556 operation 557 operation 559 operation 603 Downscale Factor 604 Multiplier 607 Searcher 608 FPC Processor 609 Frequency Pulse Searcher 610 FPC encoder 611 Finder 612 Quantizer 613 Adder 614 Estimator 615 Calculator 616 Quantizer 654 Multiplication Operation 659 Frequency Pulse Search Operation 660 FPC encoding operation 661 operation 662 operation 663 operation 664 operation 665 operation 666 operation 667 operation 668 operation 700 Integrated Time-Domain / Frequency-Domain CELP Coding Device 701 Acoustic signal classifier 702 Preprocessor 703 Frequency Domain Encoder 704 Synthesizer 705 Time Domain Encoder 706 Synthesizer 707 Bit Allocator 708 Calculator 720 Mixed Time Domain / Frequency Domain Encoder 750 Integrated Time Domain / Frequency Domain CELP Coding Method 751 operation 752 operation 753 operation 754 Music Synthesis Actions 755 operation 756 Synthetic Filtering Operation 757 Bit Allocation Discovery 758 operation 770 Mixed Time Domain / Frequency Domain Coding Method 800 Submode Selector 850 operation 902 Estimator 903 Calculator 904 Characterizer 905 bits per frequency band final distributor 951 First Action 952 operation 953 operation 954 operation 955 operation 1001 processor 1002 input 1003 Output 1004 memory 1100 decoder device 1101 Bitstream 1103 "Music" Decoder 1105 "Audio" Decoder 1150 Decoding Method 1200 Acoustic Signal Decoder 1201 Estimator 1202 Estimator 1203 Calculator 1204 Calculator 1205 Distributor 1207 Calculator 1208 Filter 1209 Adder 1210 Calculator 1211 Calculator 1212 Adder, Converter 1213 LP (Linear Prediction) Synthesis Filter 1263 operation 1250 Acoustic signal decoding method 1251 operation 1252 operation 1253 operation 1254 operation 1255 operation 1256 operation 1257 operation 1258 operation 1259 Addition Operation 1260 operation 1261 operation 1262 operation
Claims
1. 1. A joint time domain / frequency domain coding device for coding an input audio signal, comprising: a classifier for classifying the input acoustic signal into one of a plurality of acoustic signal categories, the acoustic signal categories including an ambiguous signal type category indicating that the nature of the input acoustic signal is ambiguous; and a selector for selecting one of a plurality of coding sub-modes for coding the input acoustic signal if the input acoustic signal falls into the unclear signal type category; and a mixed time-domain / frequency-domain encoder for encoding the input audio signal using the selected coding sub-mode.
2. 2. The integrated time-domain / frequency-domain coding device of claim 1, wherein the audio signal categories include speech, music, and an unclear signal type indicating that the input audio signal is not classified as either speech or music.
3. 3. The integrated time-domain / frequency-domain coding device according to claim 1, wherein the selector selects the coding sub-mode depending on a bit rate for coding the input acoustic signal and characteristics of the input acoustic signal classified into the unclear signal type category.
4. 4. The joint time domain / frequency domain coding device according to claim 1, wherein the coding sub-modes are identified by respective sub-mode flags.
5. 5. The integrated time-domain / frequency-domain coding device of claim 3, wherein the selector selects a backward-coding sub-mode that uses a legacy integrated time-domain and frequency-domain coding model for coding the input audio signal if (a) an available bit rate for coding the input audio signal is less than or equal to a first given value, and (b) the input audio signal is not classified as either speech or music.
6. 6. The integrated time-domain / frequency-domain coding device of claim 3, wherein the selector selects a first coding sub-mode if "voice"-like characteristics are detected in the input acoustic signal.
7. 7. The integrated time-domain / frequency-domain coding device of claim 6, wherein the selector selects the first coding submode if (a) the input audio signal is not classified as either speech or music by the classifier, and the bit rate available for encoding the input audio signal is higher than a second given value, (b) a probability that the input audio signal is music is less than or equal to a third given value, and (c) no temporal attack is detected within a current frame of the input audio signal.
8. 8. The integrated time-domain / frequency-domain coding device according to claim 3, wherein the selector selects the second coding sub-mode if a temporal attack is detected in the input acoustic signal.
9. 9. The integrated time-domain / frequency-domain coding device of claim 8, wherein the selector selects the second coding submode if: (a) the input audio signal is not classified as either speech or music by the classifier, and the bit rate available for encoding the input audio signal is higher than a fourth given value; (b) a probability that the input audio signal is music is less than or equal to a fifth given value; and (c) a temporal attack is detected within a current frame of the input audio signal.
10. 10. The integrated time-domain / frequency-domain coding device of claim 3, wherein the selector selects a third coding sub-mode if a "music"-like characteristic is detected in the input audio signal.
11. 11. The integrated time-domain / frequency-domain coding device of claim 10, wherein the selector selects the third coding sub-mode if (a) the input audio signal is not classified by the classifier as either speech or music, an available bit rate for encoding the input audio signal is higher than a sixth given value, and (b) a probability that the input audio signal is music is greater than a seventh given value.
12. the selector selects a first coding sub-mode if "voice"-like characteristics are detected in the input audio signal; the selector selects a second coding sub-mode if a temporal attack is detected in the input acoustic signal; 6. The integrated time-domain / frequency-domain coding device of claim 1, wherein the selector selects a third coding sub-mode if a "music"-like characteristic is detected in the input audio signal.
13. 13. The integrated time-domain / frequency-domain coding device of claim 12, wherein the selector selects (a) in the third coding sub-mode a given number of sub-frames per frame for coding the input acoustic signal, and (b) in the first and second coding sub-modes a number of sub-frames less than the given number, the number depending on a bit rate available for coding the input acoustic signal.
14. 1. A joint time-domain / frequency-domain coding method for coding an input audio signal, comprising: classifying the input acoustic signal into one of a plurality of acoustic signal categories, the acoustic signal categories including an ambiguous signal type category indicating that the nature of the input acoustic signal is ambiguous; selecting one of a plurality of coding sub-modes for coding the input acoustic signal if the input acoustic signal falls into the unclear signal type category; and performing mixed time-domain / frequency-domain coding of the input audio signal using the selected coding sub-mode.
15. 15. The integrated time-domain / frequency-domain coding method of claim 14, wherein the audio signal categories include speech, music, and an unclear signal type indicating that the input audio signal is not classified as either speech or music.
16. 16. The integrated time-domain / frequency-domain coding method of claim 14 or 15, wherein selecting one of a plurality of coding sub-modes comprises selecting the coding sub-mode depending on a bit rate for encoding the input acoustic signal and characteristics of the input acoustic signal classified into the unclear signal type category.
17. 17. The joint time-domain / frequency-domain coding method according to any one of claims 14 to 16, comprising identifying said coding sub-modes by respective sub-mode flags.
18. 18. The integrated time-domain / frequency-domain coding method of claim 16 or 17, wherein selecting one of a plurality of coding sub-modes comprises selecting a backward coding sub-mode that uses a legacy integrated time-domain and frequency-domain coding model for coding the input audio signal if (a) an available bit rate for coding the input audio signal is less than or equal to a first given value, and (b) the input audio signal is not classified as either speech or music.
19. 19. The method of any one of claims 16 to 18, wherein selecting one of a plurality of coding sub-modes comprises selecting a first coding sub-mode if "voice"-like characteristics are detected in the input acoustic signal.
20. 20. The integrated time-domain / frequency-domain coding method of claim 19, wherein the first coding sub-mode is selected if (a) the input audio signal is classified as neither speech nor music, and the bit rate available for encoding the input audio signal is higher than a second given value, (b) a probability that the input audio signal is music is less than or equal to a third given value, and (c) no temporal attack is detected within a current frame of the input audio signal.
21. 21. The integrated time-domain / frequency-domain coding method of claim 16, wherein selecting one of a plurality of coding sub-modes comprises selecting a second coding sub-mode if a temporal attack is detected in the input acoustic signal.
22. 22. The integrated time-domain / frequency-domain coding method of claim 21, wherein the second coding sub-mode is selected if (a) the input audio signal is not classified as either speech or music, and the bit rate available for encoding the input audio signal is higher than a fourth given value, (b) a probability that the input audio signal is music is less than or equal to a fifth given value, and (c) a temporal attack is detected within a current frame of the input audio signal.
23. 23. The integrated time-domain / frequency-domain coding method of claim 16, wherein selecting one of a plurality of coding sub-modes comprises selecting a third coding sub-mode if a "music"-like characteristic is detected in the input audio signal.
24. 24. The integrated time-domain / frequency-domain coding method of claim 23, wherein the third coding sub-mode is selected if (a) the input audio signal is classified as neither speech nor music and the bit rate available for encoding the input audio signal is higher than a sixth given value, and (b) the probability that the input audio signal is music is greater than a seventh given value.
25. The method of claim 25, wherein a first coding submode is selected if "voice"-like characteristics are detected in the input audio signal; a second coding sub-mode is selected if a temporal attack is detected in the input audio signal; 19. The method of any one of claims 14 to 18, wherein a third coding sub-mode is selected if "music"-like characteristics are detected in the input audio signal.
26. 26. The integrated time-domain / frequency-domain coding method of claim 25, wherein selecting one of a plurality of coding sub-modes comprises: (a) selecting, in the third coding sub-mode, a given number of sub-frames per frame for coding the input acoustic signal; and (b) selecting, in the first and second coding sub-modes, a number of sub-frames less than the given number, the number depending on an available bit rate for coding the input acoustic signal.
27. 1. An audio signal decoder, comprising: a receiver for receiving a bitstream conveying information usable to reconstruct a mixed time-domain / frequency-domain excitation representing an input acoustic signal classified into an unclear signal type category indicating that the nature of the acoustic signal is unclear, the information including one of a plurality of coding sub-modes to be used for coding the input acoustic signal classified into the unclear signal type category; a reconstructor for reconstructing the mixed time-domain / frequency-domain excitation in response to the information conveyed in the bitstream, including the coding submode used to code the input acoustic signal; a transformer that transforms the mixed time domain / frequency domain excitation into the time domain; a synthesis filter that filters the mixed time-domain / frequency-domain excitation that has been transformed into the time domain to produce a synthesized version of the audio signal.
28. 28. The audio signal decoder of claim 27, wherein the coding sub-modes are identified in the bitstream by sub-mode flags.
29. 29. An audio signal decoder according to claim 27 or 28, wherein the encoding sub-modes include: (a) a first encoding sub-mode if the audio signal contains characteristics similar to "voice", (b) a second encoding sub-mode if the audio signal contains time attacks, and (c) a third encoding sub-mode if the audio signal contains characteristics similar to "music".
30. 30. An audio signal decoder according to any one of claims 27 to 29, wherein the reconstructor recovers a frequency representation of a time-domain excitation contribution from information conveyed in the bitstream, reconstructs a frequency-quantized difference vector between the frequency-domain excitation contribution and the frequency representation of the time-domain excitation contribution, and adds the frequency-quantized difference vector to the frequency representation of the time-domain excitation contribution to generate the mixed time-domain / frequency-domain excitation.
31. 1. A method for decoding an audio signal, comprising: receiving a bitstream conveying information usable to reconstruct a mixed time-domain / frequency-domain excitation representing an acoustic signal classified in an unclear signal type category indicating that the acoustic signal is unclear in nature, the information including one of a plurality of coding sub-modes to be used for coding the acoustic signal classified in the unclear signal type category; reconstructing the mixed time-domain / frequency-domain excitation in response to the information conveyed in the bitstream, including the coding sub-mode used to code an input acoustic signal; transforming the mixed time domain / frequency domain excitation into the time domain; and filtering the mixed time-domain / frequency-domain excitation that has been transformed into the time domain through a synthesis filter to generate a synthesized version of the audio signal.
32. 32. The method of claim 31, wherein the coding sub-mode is identified in the bitstream by a sub-mode flag.
33. 33. The audio signal decoding method of claim 31 or 32, wherein the encoding submodes include: (a) a first encoding submode if the audio signal contains characteristics similar to "voice," (b) a second encoding submode if the audio signal contains a time attack, and (c) a third encoding submode if the audio signal contains characteristics similar to "music."
34. 34. A method of decoding an acoustic signal according to any one of claims 31 to 33, wherein the step of reconstructing the mixed time domain / frequency domain excitation comprises the steps of recovering a frequency representation of a time domain excitation contribution from the information conveyed in the bitstream, reconstructing a frequency quantized difference vector between the frequency domain excitation contribution and the frequency representation of the time domain excitation contribution from the information conveyed in the bitstream, and adding the frequency quantized difference vector to the frequency representation of the time domain excitation contribution to generate the mixed time domain / frequency domain excitation.
Citation Information
Patent Citations
Smoothing of voice parameter based on existence of noise-like signal in voice signal
JP2011203737A
Audio signal processing method and apparatus
JP2011514558A
Coding generic audio signals at low bitrates and low delay
US9015038B2