Non-windowed DCT-based audio code processing using advanced quantization
By adopting back-to-back DCT transformation and modification quantization processes in audio compression, the resource waste problem caused by windowing in the prior art is solved, and more efficient audio compression and decoding is achieved.
Patent Information
- Application Number
- CN202280100722.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-20
- Publication Date
- 2025-05-09
AI Technical Summary
Existing audio compression techniques require overlapping windows when performing discrete cosine transforms (MDCTs), resulting in increased resource usage and complex compression processes.
Back-to-distance discrete cosine transform (DCT) is used to eliminate the need for windowed functions and modify the quantization and inverse quantization processes to achieve audio encoding and decoding without windowed.
Reduces the use of resources during audio compression, improves compression efficiency, and generates smaller compressed files, suitable for faster and more reliable streaming.
Smart Images

Figure CN119968676A_ABST
Abstract
Description
Technical Field
[0001] Embodiments relate to encoding and decoding audio. Background Art
[0002] It is common practice to transmit and / or store audio signals. For example, an audio signal can be streamed from a server to a user device so that the user can listen to a playback of the audio signal. The audio signal can be streamed alone or together with a video stream. The audio signal can also be stored in a storage medium (e.g., fixed and / or portable computer memory) for later consumption. Summary of the invention
[0003] Example implementations may achieve improved compression by using a back-to-back DCT to eliminate the need for a windowing function and modifying quantization associated with an audio encoder and / or inverse quantization at an audio decoder.
[0004] In general terms, an apparatus, a system, a non-transitory computer-readable medium having computer executable program code stored thereon that can be executed on a computer system, and / or a method may utilize a method to perform a process, the method comprising: receiving a time domain audio signal; generating a block of the time domain audio signal as a portion of the time domain audio signal; transforming the block of the time domain audio signal using a first non-windowed transform function to generate a first frequency domain audio signal; transforming the first frequency domain audio signal using a second non-windowed transform function to generate a second frequency domain audio signal; and compressing the second frequency domain audio signal to generate a compressed frequency domain audio signal.
[0005] In another general aspect, an apparatus, a system, a non-transitory computer-readable medium having computer-executable program code stored thereon that can be executed on a computer system, and / or a method may utilize a method to perform a process comprising: receiving a formatted data packet comprising a compressed frequency domain audio signal; generating a decompressed frequency domain audio signal by decompressing the compressed frequency domain audio signal; transforming the decompressed frequency domain audio signal using a first non-windowed transform function to generate a first time domain audio signal; transforming the first time domain audio signal using a second non-windowed transform function to generate a second time domain audio signal; and generating a reconstructed time domain audio signal based on the second time domain audio signal. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Example embodiments will be more fully understood from the detailed description given herein below and the accompanying drawings, in which like elements are represented by like reference numerals, which are given by way of illustration only and thus do not limit example embodiments, and in which:
[0007] Figure 1 A block diagram of an audio encoder and decoder system according to an example implementation is shown.
[0008] Figure 2 A block diagram of an audio encoder system according to an example implementation is shown.
[0009] Figure 3 Another block diagram of an audio decoder system according to an example implementation is shown.
[0010] Figure 4 A transform module associated with an audio encoder system according to an example implementation is shown.
[0011] Figure 5 A block diagram of a quantization module associated with an audio encoder system is shown according to an example implementation.
[0012] Figure 6 A block diagram of an inverse quantization module associated with an audio encoder system is shown according to an example implementation.
[0013] Figure 7 A method of compressing audio according to an example implementation is shown.
[0014] Figure 8 A method of decompressing audio according to an example implementation is shown.
[0015] Fig.9A A block diagram of an audio encoder according to an example implementation is shown.
[0016] Fig. 9B Another block diagram of an audio decoder according to an example implementation is shown.
[0017] It should be noted that these drawings are intended to illustrate the general characteristics of the methods and / or structures utilized in certain example embodiments and to supplement the written description provided below. However, these drawings are not drawn to scale and do not accurately reflect the precise structure or performance characteristics of any given embodiment and should not be interpreted as defining or limiting the range of values or properties covered by the example embodiments. For example, the positioning of modules and / or structural elements may be reduced or enlarged for clarity. The use of similar or identical reference numerals in the various drawings is intended to indicate the presence of similar or identical elements or features. DETAILED DESCRIPTION
[0018] Existing audio compression techniques use a modified discrete cosine transform (MDCT) to convert analog audio signals into digital signals. Typically, MDCT is performed on an audio signal in such a way that adjacent transform ranges overlap (e.g., overlapping windows) along the time axis, for example, by 50% in order to suppress distortion generated at the boundary portion between adjacent transform ranges. In other words, existing audio compression techniques use overlapping windows to perform audio code processing (code), in which consecutive block transforms codify the same signal twice. Overlapping windows are used to avoid discontinuities at, for example, block boundaries. Codifying the same signal twice may be a resource-demanding process (e.g., processor and memory) within an audio compression pipeline, which may be undesirable in many applications.
[0019] Example implementations described herein can reduce undesirable resource usage in audio compression by, for example, using back-to-back discrete cosine transform (DCT) transforms for audio code processing without windowing (e.g., without using the aforementioned overlapping windows). In other words, a time domain signal can be transformed into a frequency domain signal using a first non-windowed transform function (e.g., DCT). The frequency domain signal can then be transformed again using a second non-windowed transform function (e.g., DCT), except that the frequency domain signal remains as a frequency domain signal.
[0020] By reducing resource usage, using the techniques described herein, audio compression can be performed in less time, and / or resources can be used for other code processing processes. In some implementations of the techniques described herein, additional audio channels can be compressed within a specified time period or time range (e.g., requirement). Moreover, in some implementations of the techniques described herein, smaller compressed files can be generated, which may result in using less bandwidth to communicate, using less memory to store compressed files, etc. In some implementations of the techniques described herein, a user (e.g., using a playback device) can receive smaller compressed files during, for example, a streaming (e.g., receiving formatted data packets) operation. Therefore, streaming can be faster and / or more reliable. Alternatively, a user (e.g., using a playback device) can receive compressed files of approximately typical size during, for example, a streaming operation. Therefore, a user can receive additional audio channels, which can result in higher quality audio when played back, thereby improving the user experience.
[0021] Figure 1 A block diagram of an audio encoder and decoder system according to an example implementation is shown. Figure 1As shown, the system includes an audio encoder 105 and an audio decoder 110. The audio encoder 105 can be configured to generate a compressed audio 10 signal based on an input audio 5 signal. The audio 5 signal can be an analog audio signal, a time domain audio signal, etc. Therefore, the input audio 5 can be referred to as a time domain audio signal. The time domain audio 5 signal can be a live recording, a stored file of a recording, associated with a video, etc. The compressed audio 10 signal can be a digital audio signal, a frequency domain audio signal, etc. Therefore, the compressed audio 10 signal can be referred to as a compressed frequency domain audio signal. The audio decoder 110 can be configured to generate a reconstructed audio 15 signal based on the compressed frequency domain audio 10 signal. The reconstructed audio 15 signal can be an analog audio signal, a time domain audio signal, etc. Therefore, the reconstructed audio 15 signal can be referred to as a reconstructed time domain audio signal. The audio encoder 105 and the audio decoder 110 can be independent of each other. In other words, the audio encoder 105 may generate a compressed frequency-domain Audio 10 signal that may be decompressed using a decompression technique different from the decompression technique used by the audio decoder 110. Furthermore, the audio decoder 110 may be used to decompress a compressed frequency-domain audio signal compressed using a compression technique different from the compression technique used by the audio encoder 105.
[0022] In an example implementation, the audio encoder 105 and / or the audio decoder 110 may be configured to use properties associated with the DCT to reduce the maximum discontinuity at the block boundary to a fixed value. The maximum discontinuity at the block boundary may be based on the length of the block. For example, a property associated with the DCT may be that the signal at the first end of the block may be transformed based on the sum of the DCT coefficients. The sum of the DCT coefficients may be the sum of the values of the DCT coefficients in adjacent buckets of the frequency domain audio signal quantized by calculation (e.g., all, most, a portion). For example, another property associated with the DCT may be that the signal at the second end of the block may be transformed based on the interlaced sum of the DCT coefficients. The interlaced sum of the DCT coefficients may be the sum of the values of the DCT coefficients in every other bucket of the frequency domain audio signal quantized by calculation. In addition, the properties of the DCT apply to the derivatives of the signal within the block. For example, odd-order derivatives may be zero at the block boundary. In addition, even-order derivatives (e.g., second-order derivatives, fourth-order derivatives, etc.) of the signal within the block may be controlled because even-order derivatives are cosine functions.
[0023] Therefore, the encoder 5 can be configured to receive a time domain audio signal (e.g., audio 5), and the encoder 5 can be configured to generate a time domain audio signal that is divided into blocks. The time domain audio that is divided into blocks can be a part of the time domain audio signal (e.g., audio 5). As mentioned above, the encoder 5 can be configured to use the properties associated with DCT. The properties associated with DCT can be used to transform the time domain audio that is divided into blocks from the time domain to the frequency domain without using windowing. Therefore, the encoder 5 can be further configured to transform the time domain audio signal that is divided into blocks using a first non-windowed transform function (e.g., DCT) to generate a first frequency domain audio signal, and transform the first frequency domain audio signal using a second non-windowed transform function (e.g., DCT) to generate a second frequency domain audio signal. The audio encoder 5 can then compress (e.g., quantize and entropy encode) the second frequency domain audio signal to generate a compressed frequency domain audio signal (e.g., compressed audio 10). The compressed frequency-domain audio signal (eg, compressed audio 10) may be stored in a computer memory, streamed (eg, transmitting the formatted data packets to a remote device) for playback on a playback device, etc.
[0024] The audio decoder 110 may also be configured to use properties associated with DCT. The properties associated with DCT may be used to transform a compressed frequency domain audio signal (e.g., compressed audio 10) from the frequency domain to the time domain without using windowing. Therefore, the audio decoder 110 may be configured to receive a formatted data packet including a compressed frequency domain audio signal (e.g., compressed audio 10), and to generate a decompressed frequency domain audio signal by decompressing (e.g., inverse entropy coding and inverse quantization) the compressed frequency domain audio signal. The audio decoder 110 may further be configured to transform the decompressed frequency domain audio signal using a first non-windowed transform function (e.g., IDCT) to generate a first time domain audio signal, and to transform the first time domain audio signal using a second non-windowed transform function (e.g., IDCT) to generate a second time domain audio signal. The audio decoder 110 may then generate a reconstructed time domain audio signal (e.g., reconstructed audio 15) based on the second time domain audio signal. The user may then listen to (eg, by playing back on a playback device) the reconstructed time-domain audio signal.
[0025] Figure 2 A block diagram of an audio encoder system according to an example implementation is shown. Figure 2As shown, the audio encoder 105 includes an analysis and blocking module 205 block, a transform module 210 block, a quantization module 215 block, a code processing module 220 block, a perceptual model module 225 block, and a formatting module 230 block. In an example implementation, the audio encoder 105 can be configured to reduce discontinuities at, for example, block boundaries without using a DCT using a window.
[0026] The blocking module 205 may be configured to sample the input time domain audio 5 signal in sequential time frames, each frame may include a portion of the input time domain audio 5 signal, which is sometimes referred to as a blocked time domain audio signal based on the input time domain audio 5 signal. The analysis and blocking module 205 may be configured to pre-process the input time domain audio 5 signal. Pre-processing may include normalization, frequency weighting, frequency scaling, block sorting, dynamic range scaling, etc.
[0027] The transform module 210 may be configured to generate a frequency domain audio signal by transforming a block of an input time domain audio 5 signal (as a divided time domain audio signal) from analog (or time domain) to digital (or frequency domain). The transform module 210 may include an analog-to-digital converter (ADC). The ADC may use a Fourier transform (e.g., DCT, DFT, FFT). The ADC may be defined by an audio codec. The ADC typically has an associated bandwidth (or configurable bandwidth). The bandwidth may be the number of times per second that the input time domain audio 5 signal (e.g., as an analog source) is sampled and transformed by the transform module 210 to generate a discrete digital (or frequency domain) value. In an example implementation, the transform module 210 uses a DCT. In an example implementation, the transform module 210 uses two or more DCTs. The first DCT may be configured to generate a scalar value, which is an analog or frequency domain amplitude (without phase information) of the time domain audio 5 signal. The second DCT may be configured to generate a derivative of the scalar value. The derivatives of the scalar values (also frequency domain audio signals), called coefficients, may be quantized to generate a quantized frequency domain audio signal and code processed to generate a compressed frequency domain audio signal.
[0028] The quantization module 215 can be configured to generate a quantized frequency domain audio signal by quantizing the transformed audio. Quantization can be a process of mapping an input value from a large value set (e.g., a continuous set) to a value in a smaller value set with a limited number of elements (associated with the value). The elements of the smaller set are sometimes referred to as buckets, where each bucket represents a value range. For example, the quantization module 215 can be configured to quantize the coefficients, calculate the errors associated with the quantized coefficients, sort the quantized coefficients according to the amount of error, and reposition the quantized coefficients based on the errors. The quantized coefficients can be repositioned (e.g., moved to different buckets) until the error is within a threshold range. The repositioning of the quantized coefficients can reduce discontinuities at, for example, block boundaries by reducing the errors associated with the quantized coefficients.
[0029] The code processing module 220 may be configured to generate a compressed frequency domain audio signal (e.g., compressed audio 10) by entropy encoding the quantized frequency domain audio signal. Entropy encoding the quantized frequency domain audio signal may include creating and assigning a unique prefix-free code to each unique quantized coefficient or quantization level corresponding to the quantized frequency domain audio signal. Entropy encoding may include compressing data by replacing each quantized coefficient or quantization level with a corresponding variable length prefix-free output codeword to generate a compressed audio signal.
[0030] The perceptual model module 225 can be configured to analyze the input time domain audio 5 signal and determine the relevant perceptual signal aspects, most notably determining the masking capability (e.g., masking threshold) of the signal as a function of frequency and time. The result is transmitted to the quantization module 215 to control the code processing distortion, thereby rendering the distortion as essentially inaudible. The perceptual model module 225 can quantize the transformed audio using psychoacoustic criteria such as masking thresholds to maximize the audio quality perceived by human listeners. For example, the perceptual model module 225 can generate a spectral distortion profile with a perceptually shaped spectral distortion profile that provides improved subjective audio quality at the expense of a noise-based quality metric.
[0031] The formatting module 230 may be configured to generate a formatted data packet or file including a compressed frequency domain audio signal (e.g., compressed audio 10). The formatted data packet or file may be formatted based on a codec. For example, the formatted file may be formatted into an audio file format such as opus, mp3, ambisonic, advanced audio coding (AAC), etc.
[0032] Figure 3 A block diagram of an audio decoder system according to an example implementation is shown. Figure 3As shown, the audio decoder 110 includes a decoding module 305 block, an inverse quantization module 310 block, a transform module 315 block, and a synthesis module 320 block. In an example implementation, the audio decoder 110 can be configured to decode a compressed frequency domain audio signal (e.g., compressed audio 10) without windowing.
[0033] The decoding module 305 may be configured to perform the reverse operation of the code processing module 220. The decoding module 305 may be configured to generate a decompressed frequency domain audio signal. In other words, the decoding module 305 may be configured to perform inverse entropy coding on the compressed audio 10.
[0034] The inverse quantization module 310 may be configured to perform the reverse operation of the quantization module 215. The inverse quantization module 310 may be configured to generate an inversely quantized frequency domain audio signal based on the decompressed frequency domain audio signal. In an example implementation, the inverse quantization module 310 may be configured to calculate the interleaved sum of the previous block (e.g., the number of quantized transform coefficients in the interleaved buckets) and reposition the sum of the current block toward the interleaved sum of the previous block (within the quantization boundary). The repositioning may be configured to generate bandwidth-limited or frequency-response-manipulated continuity at block boundaries.
[0035] The transform module 315 may be configured to perform the reverse operation of the transform module 210. For example, the transform module 315 may be configured to generate a time domain audio signal by inversely transforming the inverse quantized frequency domain audio signal from digital (or frequency domain) to analog (or time domain). The transform module 210 may include a digital-to-analog converter (DAC). The DAC may use an inverse Fourier transform (e.g., IDCT, IDFT, IFFT). The DAC may be defined by an audio codec. In an example implementation, the transform module 315 uses an IDCT. In an example implementation, the transform module 315 uses two or more IDCTs. The first IDCT may be configured to generate an integral scalar value (e.g., in the time domain) associated with the inverse quantized frequency domain audio signal. The second IDCT may be configured to generate a scalar value (e.g., in the time domain), which is an analog or time domain amplitude associated with the reconstructed time domain audio signal (e.g., the reconstructed audio 15) (e.g., as its block or frame).
[0036] The synthesis module 320 is configured to generate a reconstructed time-domain audio signal (e.g., the reconstructed audio 15). For example, the synthesis module 320 is configured to combine adjacent time-domain audio blocks into a continuous output audio signal as the reconstructed time-domain audio signal (e.g., the reconstructed audio 15).
[0037] Figure 4 A transform module associated with an audio encoder system according to an example implementation is shown. Figure 4 As shown, the transform module 210 includes a DCT 405 (e.g., a first non-windowed transform function or DCT) block and a DCT 410 (e.g., a second non-windowed transform function or DCT) block. The DCT 405 and the DCT 410 together can transform the time-domain audio 5 signal (a portion of the time-domain audio 5 signal or a block of the time-domain audio 5 signal, etc.) without windowing the time-domain audio 5 signal. The DCT 410 can be configured to generate coefficients (e.g., in the frequency domain) that will be quantized into a quantized frequency-domain audio signal.
[0038] DCT 405 can be configured to generate a scalar value (e.g., digital or frequency domain), which is an analog or time domain amplitude (without phase information) of a time domain audio 5 signal. DCT 405 can be configured to generate a derivative of the scalar value (e.g., as a quantized frequency domain audio signal). The derivative of the scalar value, referred to as a coefficient, can be quantized and coded. DCT 405 can be expressed as shown in Equation 1. in: N is the length of the signal (e.g., block size); is the frequency being evaluated; C(0, ... N-1) are the transform coefficients; and .
[0039] Equation (1) can be simplified as shown in equation (2). in: n is the index of the current value in signal; and x n is the value at this index.
[0040] DCT 410 may be the derivative of equation (2) and expressed as shown in equation (3). where m is the index, and all terms in the summation except the term involving index m tend to zero as a constant.
[0041] Figure 5 A block diagram of a quantization module associated with an audio encoder system according to an example implementation is shown. Figure 5 As shown, the quantization module 215 may include a quantization module 505 block, an error calculation module 510 block, a quantization sorting module 515 block, and a position exchange module 520 block.
[0042] In an example implementation, the quantization module 215 may be configured to select a transform coefficient from a first bucket as a first mapped position, identify a second bucket adjacent to the first bucket as a second mapped position (e.g., to a bucket having a close value range to the first bucket), and map the transform coefficient to the second bucket. The selection of the transform coefficient may be based on an error associated with the quantized transform coefficient value corresponding to the transform coefficient. Mapping the transform coefficient to the second bucket may include repeatedly selecting the transform coefficient and / or mapping (or remapping) the transform coefficient to the second bucket until the sum of the errors is less than a threshold. In other words, repeatedly (e.g., over and over) operating the error calculation module 510, the quantization sorting module 515, and the position exchange module 520 until the sum of the errors is less than a threshold. Mapping the transform coefficient to the second bucket may include identifying a subset of multiple quantized transform coefficient values, identifying the first bucket as being within the subset of multiple quantized transform coefficient values, and the second bucket may be within the subset of multiple quantized transform coefficient values.
[0043] The quantization module 505 may be configured to generate a quantized frequency domain audio signal by quantizing the transform coefficients generated by the DCT 410. The quantization module 505 may be configured to reduce the number of bits used to represent the transform coefficients generated by the DCT 410. Quantization may be a process of mapping an input value from a large value set (e.g., a continuous set) to a value in a smaller value set having a limited number of elements (associated with the value). The elements of the smaller set are sometimes referred to as buckets, where each bucket represents an energy band (e.g., a range of values). Therefore, the quantization module 505 may be configured to map the transform coefficients generated by the DCT 410 to buckets representing a range of transform coefficient scalar values. The quantization module 505 may be configured to reduce the error of the sum and interlaced sum of each bucket by introducing a small amount of error into each quantization decision that causes the transform coefficient to be mapped to the bucket. If the introduced error varies based on the order of the coefficients, each energy band may have better continuity at the block boundary. In an example implementation, the quantization error may be modified by adjusting the quantization of nearby frequencies, resulting in maximizing the continuity at the block boundary.
[0044] The error calculation module 510 is configured to calculate the quantization error associated with the quantized transform coefficients. Quantizing a series of numbers produces a series of quantization errors. Allocating more bits to each frequency can result in the introduction of fewer errors (noise), but more space is required to store the results. Conversely, fewer bits allocated to each frequency result in more noise, but less space is required to store the results. In an example implementation, the quantization error can be the difference between the quantized transform coefficient value (e.g., the value range assigned to the bucket) and the transform coefficient value. The error calculation module 510 can be configured to calculate the error sum for each bucket. The error calculation module 510 can be configured to calculate the error sum of all buckets and the error sum of interleaved buckets.
[0045] The quantization sorting module 515 can be configured to sort the quantized transform coefficient values based on the amount of error introduced by each quantized transform coefficient value. The position exchange module 520 can be configured to map the transform coefficients to adjacent buckets (e.g., to buckets with a close value range). For example, if the transform coefficient has the most associated error, the transform coefficient can be mapped (or assigned) to an adjacent bucket. The process then returns to the error calculation module 510. If the error sum of all buckets and the error sum of the interleaved buckets are below a threshold, above a threshold, or within a threshold range, the process can end (or quantization is complete). The threshold and / or threshold range can be preconfigured.
[0046] Figure 6 A block diagram of an inverse quantization module associated with an audio encoder system according to an example implementation is shown. Figure 6 As shown, the inverse quantization module 310 may include a summing module 605 block, a remapping module 610 block, and an inverse quantization module 615 block.
[0047] The summing module 605 may be configured to calculate an interlaced sum of a first block and an interlaced sum of a second block. The first block may be a previous block, and the second block may be a current block. The interlaced sum may be stored in a memory associated with the summing module to form a queue. For example, the interlaced sum of the second block may be stored in a memory for use when inverse quantizing a third (e.g., next) block. The first (e.g., previous) block, the second (e.g., current) block, and the third (e.g., next) block may be sequential (e.g., temporal) blocks. In an example implementation, the interlaced sum may be the sum of the number of quantized transform coefficients in the interlaced buckets.
[0048] The remapping module 610 may be configured to remap the quantized transform coefficients from the first bucket to the second bucket. For example, the remapping module 610 may be configured to compare the interleaved sum of the first block with the interleaved sum of the second block. If the sum is within the threshold, the process may continue to the inverse quantization module 615. Otherwise, the quantized transform coefficients of the second block may be remapped from the first bucket to the second bucket, and the process may return to the summing module, in which the interleaved sum of the second block may be recalculated. In other words, the summing module 605 and the remapping module are repeatedly (e.g., over and over) operated until the sum increment is less than the threshold. The remapping may be configured to generate bandwidth-limited or manipulated continuity of the frequency response at the block boundary. For example, the first bucket and the second bucket may be within the range of the bucket (e.g., within the frequency range). In this implementation, the interleaved sum of the first block and the interleaved sum of the second block may be limited to the range of the bucket. The remapping module 610 may be configured to operate within the quantization boundary (e.g., to remap the quantized transform coefficients). The quantization boundaries may be a minimum range of values assigned to a bucket and a maximum range of values. In other words, the mapping associated with the quantization may be restricted to a minimum range of values and a maximum range of values.
[0049] In an example implementation, the elements (e.g., buckets) of every second (e.g., every other) block can be reversed. Then when the corresponding low ends (of the blocks) meet, the sums of the two blocks can match, and when the corresponding long ends meet, the interleaved sums can match.
[0050] Referring to Table 1, the initial elements of each block (eight (8) elements are shown, but there may be more, such as 1024) are in the order of C0, C1, C2, C3, C4, C5, C6, C7, where C0, ..., C7 represent the elements of the corresponding bucket (e.g., bucket). The reversed element order shows the reversed order of every second block (e.g., the 2nd block and the 4th block). The low end of the block may be when C0 is the last element and the first element of the consecutive blocks (e.g., the second block and the third block). The long end of the block may be when C7 is the last element and the first element of the consecutive blocks (e.g., the first block and the second block).
[0051] In an example implementation, the interleaved sums of the first block and the second block may match, the sums of the second block and the third block may match, the interleaved sums of the third block and the fourth block may match, the sums of the fourth block and the fifth block may match, and so on. As discussed above, matching blocks (via remapping module 610) may include repeatedly remapping values of a second block in consecutive blocks whose sum (or interleaved sum) is within a threshold of a sum (or interleaved sum) of a first block in the consecutive blocks.
[0052] The inverse quantization module 615 may be configured to generate an inversely quantized frequency domain audio signal by performing the reverse function of the quantization module 505. For example, the inverse quantization module 615 may be configured to map the quantized transform coefficient values to transform coefficient values. For example, each bucket may have an associated transform coefficient value, so that the quantized transform coefficient associated with the bucket may be mapped to the transform coefficient value.
[0053] Figure 7 A method of compressing audio according to an example implementation is shown. Figure 7 As shown, at step S705, a time domain audio signal is received. At step S710, a time domain audio signal that is a portion of the time domain audio signal is generated. The time domain audio signal can be processed as a block-based audio signal. The block-based audio signal can include samples (e.g., blocks) of the time domain audio signal. The samples can be sequential time frames, and each frame can include a portion of the time domain audio signal.
[0054] In step S715, the divided time domain audio signal is transformed using a first non-windowed transform function to generate a first frequency domain audio signal. The transformation can be based on an analog-to-digital converter (ADC). The ADC can use a Fourier transform (e.g., DCT, DFT, FFT). The ADC can be defined by an audio codec. For example, the ADC can be a direct ADC, a successive approximation ADC, a sigma-delta ADC, a pipeline ADC, a jump-comparison ADC, a Wilkinson ADC, an integral ADC, etc., just to name a few. In an example implementation, the first non-windowed transform function can be a DCT (e.g., a first DCT, a first non-windowed DCT, etc.).
[0055] At step S720, the first frequency domain audio signal is transformed using a second non-windowed transform function to generate a second frequency domain audio signal. In an example implementation, the second non-windowed transform function may be a DCT (e.g., a second DCT, a second non-windowed DCT, etc.). The transform may find a derivative of the first frequency domain audio signal. The first frequency domain audio signal may include a plurality of scalar values. The DCT may be configured to generate a derivative of the scalar value.
[0056] In step S725, the second frequency domain audio signal is compressed to generate a compressed frequency domain audio signal. In an example implementation, before compression, the second frequency domain audio signal may be quantized to generate the quantized frequency domain audio signal. Compressing the second frequency domain audio signal may include entropy encoding the second frequency domain audio signal and / or the quantized frequency domain audio signal. Entropy encoding the frequency domain (e.g., digital) audio signal and / or the quantized digital audio signal may include creating a unique prefix-free code and assigning it to each unique frequency domain level, quantized coefficient, or quantization level corresponding to the quantized digital audio signal. Entropy encoding may include compressing data by replacing each frequency domain level, quantized coefficient, or quantization level with a corresponding variable length prefix-free output codeword to generate a compressed audio signal.
[0057] Figure 8 A method of decompressing audio according to an example implementation is shown. Figure 8 As shown, in step S805, a formatted data packet including a compressed frequency domain audio signal is received. The formatted file can be formatted based on a codec. For example, the formatted file can be formatted into an audio file format such as opus, mp3, ambisonic, AAC, ASF, etc. The formatted file can include the compressed frequency domain audio signal and the power coefficient. In step S810, the decompressed frequency domain audio signal is decompressed. For example, the frequency domain audio signal can be inverse entropy decoded, and the inverse entropy decoded frequency domain audio signal can be inverse quantized.
[0058] In step S815, the compressed frequency domain audio signal is transformed using a first non-windowed transform function to generate a first time domain audio signal. The first non-windowed transform function may be an IDCT. An IDCT (e.g., a first IDCT, a first non-windowed IDCT, etc.) may be configured to generate a scalar value based on an integral of the frequency domain audio signal. Thus, the first time domain audio signal may include a plurality of scalar values.
[0059] In step S820, the first time-domain audio signal is transformed using a second non-windowed transform function to generate a second time-domain audio signal. The transform may be a digital-to-analog converter (DAC). The DAC may use an inverse Fourier transform (e.g., IDCT, IDFT, IFFT). The DAC may be defined by an audio codec. In an example implementation, the second non-windowed transform function is an IDCT. The IDCT (e.g., a first IDCT, a first non-windowed IDCT, etc.) may be configured to generate a scalar value, which is an analog or time-domain amplitude. The second time-domain audio signal is a block-based audio signal. The block-based audio signal may include samples (e.g., blocks) of the time-domain audio signal. The samples may be sequential time frames, each of which may include a portion of the time-domain audio signal. Therefore, the second time-domain audio signal may be a block-based audio signal among a plurality of block-based audio signals (e.g., analog or time-domain audio signals).
[0060] In step S825, a reconstructed time domain audio signal is generated based on the second time domain audio signal. For example, each block-based audio signal in a plurality of block-based audio signals may be concatenated together on a sequential basis. The sequentially concatenated time domain audio signals may be the reconstructed time domain audio signal.
[0061] Fig.9A A block diagram of an audio encoder according to an example implementation is shown. Fig.9A In the example of the present invention, the audio encoder system 900 can be at least one computing device and should be understood to represent almost any computing device configured to perform the methods described herein. Thus, the audio encoder system 900 can be understood to include various standard components or different or future versions thereof that can be used to implement the techniques described herein.
[0062] Fig.9A An audio encoder system 900 is shown according to at least one example embodiment. Fig.9A As shown, the audio encoder system 900 includes at least one processor 905, at least one memory 910, a controller 920, and an audio encoder 105. The at least one processor 905, at least one memory 910, the controller 920, and the audio encoder 105 are communicatively coupled via a bus 915.
[0063] Therefore, as can be understood, at least one processor 905 can be used to execute instructions stored on at least one memory 910, so as to thereby implement various features and functions described herein, or additional or alternative features and functions. Of course, at least one processor 905 and at least one memory 910 can be used for various other purposes. In particular, at least one memory 910 can be understood to represent examples of various types of memory and related hardware and software that can be used to implement any of the modules described herein.
[0064] At least one processor 905 can be configured to execute computer instructions associated with the controller 920 and / or the audio encoder 105. The at least one processor 905 can be a shared resource. For example, the audio encoder system 900 can be an element of a larger system (e.g., a streaming server). Therefore, the at least one processor 905 can be configured to execute computer instructions associated with other elements within the larger system (e.g., a streaming server that streams audio).
[0065] At least one memory 910 may be configured to store data and / or information associated with the audio encoder system 900. For example, the at least one memory 910 may be configured to store an audio codec. The controller 920 may be configured to generate various control signals and transmit the control signals to various blocks in the audio encoder system 900. The controller 920 may be configured to generate control signals according to the techniques described above.
[0066] Fig. 9B A block diagram of an audio decoder according to an example implementation is shown. Fig. 9B In the example of , the audio decoder system 950 can be at least one computing device and should be understood to represent substantially any computing device configured to perform the methods described herein. Thus, the audio decoder system 950 can be understood to include various standard components or different or future versions thereof that can be used to implement the techniques described herein. Fig. 9B As shown, the audio decoder system 950 includes at least one processor 955, at least one memory 960, a controller 970, and an audio decoder 110. The at least one processor 955, at least one memory 960, the controller 970, and the audio decoder 110 are communicatively coupled via a bus 965.
[0067] At least one processor 955 can be used to execute the instruction stored on at least one memory 960, so as to realize various features and functions described herein, or additional or alternative features and functions. Of course, at least one processor 955 and at least one memory 960 can be used for various other purposes. In particular, at least one memory 960 can be understood as representing various types of memory and related hardware and software examples that can be used to realize any module in the modules described herein. According to an example embodiment, audio encoder system 900 and audio decoder system 950 can be included in the same larger system. In addition, at least one processor 905 and at least one processor 955 can be the same at least one processor, and at least one memory 910 and at least one memory 960 can be the same at least one memory. In addition, controller 920 and controller 970 can be the same controller.
[0068] At least one processor 955 may be configured to execute computer instructions associated with the controller 970 and / or the audio decoder 110. The at least one processor 955 may be a shared resource. For example, the audio decoder system 950 may be an element of a larger system (e.g., a mobile device). Thus, the at least one processor 955 may be configured to execute computer instructions associated with other elements within the larger system (e.g., web browsing or wireless communication).
[0069] At least one memory 960 may be configured to store data and / or information associated with the audio decoder system 950. The controller 970 may be configured to generate various control signals and transmit the control signals to various blocks in the audio decoder system 950. The controller 970 may be configured to generate the control signals according to the techniques described above.
[0070] Implementations may include one or more of the following examples and / or combinations thereof.
[0071] Example 1. A method comprising: receiving a time domain audio signal; generating a blocked time domain audio signal as a portion of the time domain audio signal; transforming the blocked time domain audio signal using a first non-windowed transform function to generate a first frequency domain audio signal; transforming the first frequency domain audio signal using a second non-windowed transform function to generate a second frequency domain audio signal; and compressing the second frequency domain audio signal to generate a compressed frequency domain audio signal.
[0072] Example 2. The method of Example 1, wherein the first non-windowed transform function may be a discrete cosine transform (DCT) transform.
[0073] Example 3. The method described in Example 1 or Example 2 may further include generating a quantized frequency domain audio signal by quantizing the second frequency domain audio signal, wherein the compression of the second frequency domain audio signal may include compressing the quantized frequency domain audio signal, the second frequency domain audio signal may include multiple transform coefficient values, quantizing the second frequency domain audio signal may include mapping each of the multiple transform coefficient values to one of the multiple quantized transform coefficient values, and mapping each of the multiple transform coefficient values to one of the quantized transform coefficient values may include introducing an error into each quantized transform coefficient value.
[0074] Example 4. A method as described in Example 3, wherein the quantization of the second frequency-domain audio signal may include selecting a transform coefficient from a first mapped position, identifying a second mapped position adjacent to the first mapped position, and mapping the transform coefficient to the second mapped position.
[0075] Example 5. The method of Example 4, wherein the selection of the transform coefficient may be based on an error associated with the quantized transform coefficient value corresponding to the transform coefficient.
[0076] Example 6. The method of Example 4, wherein mapping the transform coefficient to the second mapped position may include repeatedly selecting the transform coefficient and mapping the transform coefficient to the second mapped position until an error is less than a threshold.
[0077] Example 7. A method as described in Example 4, wherein mapping the first transform coefficient to the second mapped position may include identifying a subset of the multiple quantized transform coefficient values, identifying the first mapped position as being within the subset of the multiple quantized transform coefficient values, and the second mapped position being within the subset of the multiple quantized transform coefficient values.
[0078] Example 8. The method as described in any one of Examples 1 to 7 may further include one of the following: storing the compressed frequency domain audio signal in a computer memory or streaming the compressed frequency domain audio signal.
[0079] Example 9. A method comprising: receiving a formatted data packet including a compressed frequency domain audio signal; generating a decompressed frequency domain audio signal by decompressing the compressed frequency domain audio signal; transforming the decompressed frequency domain audio signal using a first non-windowed transform function to generate a first time domain audio signal; transforming the first time domain audio signal using a second non-windowed transform function to generate a second time domain audio signal; and generating a reconstructed time domain audio signal based on the second time domain audio signal.
[0080] Example 10. The method of Example 9, wherein the first non-windowed transform function may be a discrete cosine transform (DCT) transform.
[0081] Example 11. The method as described in Example 9 or Example 10 may further include generating an inverse quantized frequency domain audio signal by inverse quantizing the decompressed frequency domain audio signal, wherein the quantization of the decompressed frequency domain audio signal may include calculating the interleaved sum of a first block of the decompressed frequency domain audio signal, calculating the sum of a second block of the decompressed frequency domain audio signal, and repeatedly remapping the value of the second block of the decompressed frequency domain audio signal until the sum of the second block of the decompressed frequency domain audio signal is within a threshold of the interleaved sum of the first block of the decompressed frequency domain audio signal.
[0082] Example 12. A method as described in Example 11, wherein, before the calculation of the interleaved sum of the first block of the decompressed frequency domain audio signal, the method may further include identifying a frequency range associated with the decompressed frequency domain audio signal, the calculation of the interleaved sum of the first block of the decompressed frequency domain audio signal being calculated within the frequency range, and the calculation of the interleaved sum of the second block of the decompressed frequency domain audio signal being calculated within the frequency range.
[0083] Example 13. The method as described in Example 9 or Example 10 may further include generating an inverse quantized frequency domain audio signal by inverse quantizing the decompressed frequency domain audio signal, wherein the quantization of the decompressed frequency domain audio signal includes calculating the interleaved sum of a first block of the decompressed frequency domain audio signal, reversing the order of elements of a second block of the decompressed frequency domain audio signal, calculating the interleaved sum of the second block of the decompressed frequency domain audio signal, and repeatedly remapping the values of the second block of the decompressed frequency domain audio signal until the sum of the second block of the decompressed frequency domain audio signal is within a threshold of the interleaved sum of the first block of the decompressed frequency domain audio signal.
[0084] Example 14. The method as described in Example 9 or Example 10 may further include generating an inverse quantized frequency domain audio signal by inverse quantizing the decompressed frequency domain audio signal, wherein the quantization of the decompressed frequency domain audio signal includes reversing the order of elements of a first block of the decompressed frequency domain audio signal, calculating the sum of the first block of the decompressed frequency domain audio signal, calculating the sum of the second block of the decompressed frequency domain audio signal, and repeatedly remapping the values of the second block of the decompressed frequency domain audio signal until the sum of the second block of the decompressed frequency domain audio signal is within a threshold of the interleaved sum of the first block of the decompressed frequency domain audio signal.
[0085] Example 15 The method as described in Examples 9 to 12 may further include playing back the reconstructed time-domain audio signal.
[0086] Example 16. A method may include any combination of one or more of Examples 1 to 15.
[0087] Example 17. A non-transitory computer-readable storage medium comprising instructions stored thereon, which when executed by at least one processor are configured to cause a computing system to perform the method of any one of Examples 1 to 15.
[0088] Example 18. An apparatus comprising means for performing the method of any one of Examples 1 to 15.
[0089] Example 19. A device comprising at least one processor and at least one memory comprising computer program code, wherein the at least one memory and the computer program code are configured to cause the device to perform at least the method as described in any one of Examples 1 to 15 through the at least one processor.
[0090] An example implementation may include a non-transitory computer-readable storage medium including instructions stored thereon, which when executed by at least one processor are configured to cause a computing system to perform any of the methods described above. An example implementation may include a device including means for performing any of the methods described above. An example implementation may include a device including at least one processor and at least one memory including computer program code, the at least one memory and the computer program code being configured to cause the device, through at least one processor, to perform at least any of the methods described above.
[0091] Various implementations of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system that includes at least one programmable processor that can be either special purpose or general purpose and can be coupled to receive data and instructions from and transmit data and instructions to a storage system, at least one input device, and at least one output device.
[0092] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in high-level procedural and / or object-oriented programming languages and / or in assembly / machine languages. As used herein, the terms "machine-readable medium," "computer-readable medium," refer to any computer program product, apparatus, and / or device (e.g., a disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0093] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having a display device (LED (light emitting diode), or OLED (organic LED), or LCD (liquid crystal display) monitor / screen) for displaying information to the user and a keyboard and pointing device (e.g., a mouse or trackball) that the user can use to provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including sound, voice, or tactile input.
[0094] The systems and techniques described herein may be implemented in a computing system that includes a back-end component (e.g., as a data server), or includes a middleware component (e.g., an application server), or includes a front-end component (e.g., a client computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or any combination of such back-end, middleware, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), and the Internet.
[0095] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises through computer programs running on the respective computers and having a client-server relationship to each other.
[0096] A number of embodiments have been described, however, it will be appreciated that various modifications can be made without departing from the spirit and scope of the specification.
[0097] In addition, the logic flows depicted in the figures do not require the particular order or sequential order shown to achieve the desired results. In addition, other steps can be provided, or steps can be deleted from the described processes, and other components can be added to or removed from the described systems. Therefore, other embodiments are within the scope of the appended claims.
[0098] Although certain features of the described implementations have been shown as described herein, those skilled in the art will now expect many modifications, substitutions, variations and equivalents. Therefore, it should be understood that the appended claims are intended to cover all these modifications and variations that fall within the scope of the implementations. It should be understood that they are presented only in an exemplary and non-restrictive manner, and various changes in form and detail may be made. Except for mutually exclusive combinations, any part of the devices and / or methods described herein may be combined in any combination. The implementations described herein may include various combinations and / or sub-combinations of the functions, components and / or features of the different implementations described.
[0099] Although the example embodiments may include various modifications and alternative forms, embodiments thereof are shown by way of example in the drawings and will be described in detail herein. However, it should be understood that it is not intended to limit the example embodiments to the particular forms disclosed, but rather, the example embodiments will cover all modifications, equivalents, and alternatives falling within the scope of the claims. The same numbers refer to the same elements in the description of the drawings.
[0100] Some of the above example embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe the operations as sequential processes, many operations may be performed in parallel, concurrently, or simultaneously. In addition, the order of the operations may be rearranged. Processes may terminate when their operations are completed, but may also have additional steps not included in the figure. A process may correspond to a method, function, procedure, subroutine, subprogram, etc.
[0101] The methods discussed above (some of which are illustrated by flow charts) may be implemented by hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented by software, firmware, middleware, or microcode, the program code or code segments for performing the necessary tasks may be stored in a machine or computer readable medium, such as a storage medium. One or more processors may perform the necessary tasks.
[0102] The specific structural and functional details disclosed herein are merely representative for describing example embodiments. Example embodiments may, however, be embodied in many alternate forms and should not be construed as limited to only the embodiments set forth herein.
[0103] It should be understood that although the terms first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another element. For example, without departing from the scope of the exemplary embodiments, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element. As used herein, the terms and / or include any and all combinations of one or more of the associated listed items.
[0104] It should be understood that when an element is referred to as being connected or coupled to another element, the element may be directly connected or coupled to the other element, or there may be intervening elements. Conversely, when an element is referred to as being directly connected or coupled to another element, there are no intervening elements. Other words used to describe the relationship between elements should be interpreted in a similar manner (e.g., between versus directly between, adjacent versus directly adjacent, etc.).
[0105] The terms used herein are only used for the purpose of describing specific embodiments and are not intended to limit the exemplary embodiments. As used herein, unless the context clearly indicates otherwise, the singular forms one, a kind and the are also intended to include plural forms. It should be further understood that the terms include, contain, contain and / or have when used herein to specify the presence of stated features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups.
[0106] It should also be noted that in some alternative implementations, the functions / actions noted may not occur in the order noted in the figures. For example, two figures shown in succession may actually be executed simultaneously or may sometimes be executed in the reverse order, depending on the functionality / behaviors involved.
[0107] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the example embodiments belong. It should be further understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art, and should not be interpreted in an idealized or overly formal sense unless explicitly defined as such herein.
[0108] Parts of the above exemplary embodiments and corresponding detailed descriptions are presented in terms of software or algorithms and symbolic representations of the operations of the data bits in the computer memory. These descriptions and representations are descriptions and representations that effectively convey the essence of their work to other persons of ordinary skill in the art. As a term used herein and as a commonly used term, an algorithm is considered to be a self-consistent sequence of steps leading to a desired result. A step is a step that requires physical manipulation of a physical quantity. Typically, although not necessarily, these quantities are in the form of optical, electrical or magnetic signals that can be stored, transmitted, combined, compared and otherwise manipulated. Mainly for general reasons, it has sometimes been proven to be convenient to refer to these signals as bits, values, elements, symbols, characters, items, numbers, etc.
[0109] In the above illustrative embodiments, references to symbolic representations of actions and operations that can be implemented as program modules or functional processes (e.g., in the form of flowcharts) include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types, and can be described and / or implemented using existing hardware at existing structural elements. Such existing hardware may include one or more central processing units (CPUs), digital signal processors (DSPs), application specific integrated circuits, field programmable gate arrays (FPGAs), computers, etc.
[0110] It should be borne in mind, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless otherwise specifically noted or as is apparent from the discussion, terms such as processing or computing or calculating or determining display refer to the actions and processes of a computer system or similar electronic computing device that manipulates data represented as physical electronic quantities within the computer system's registers and memories and transforms that data into other data similarly represented as physical quantities within the computer system's memories or registers or other such information storage, transmission or display devices.
[0111] It is also noted that the software-implemented aspects of the example embodiments are typically encoded on some form of non-transitory program storage medium or implemented over some type of transmission medium. The program storage medium may be magnetic (e.g., a floppy disk or hard disk) or optical (e.g., a compact disk read-only memory or CD ROM), and may be read-only or random access. Similarly, the transmission medium may be twisted pair, coaxial cable, optical fiber, or some other suitable transmission medium known in the art. The example embodiments are not limited by these aspects of any given implementation.
[0112] Finally, it should also be noted that although the appended claims set forth specific combinations of features described herein, the scope of the present disclosure is not limited to the specific combinations claimed below, but extends to cover any combination of features or embodiments disclosed herein, regardless of whether the specific combination is specifically listed in the appended claims.
Claims
1. A method comprising: receiving a time domain audio signal; generating a blocked time-domain audio signal as a portion of the time-domain audio signal; transforming the blocked time-domain audio signal using a first non-windowed transform function to generate a first frequency-domain audio signal; transforming the first frequency-domain audio signal using a second non-windowed transform function to generate a second frequency-domain audio signal; as well as The second frequency domain audio signal is compressed to generate a compressed frequency domain audio signal.
2. The method of claim 1, wherein: The first non-windowed transform function is a discrete cosine transform (DCT) transform.
3. The method of claim 1 or claim 2, further comprising: A quantized frequency domain audio signal is generated by quantizing the second frequency domain audio signal, wherein: said compressing of said second frequency domain audio signal comprises compressing said quantized frequency domain audio signal, The second frequency domain audio signal comprises a plurality of transform coefficient values, quantizing the second frequency domain audio signal comprises mapping each transform coefficient value of the plurality of transform coefficient values to a quantized transform coefficient value of a plurality of quantized transform coefficient values, and Mapping each transform coefficient value of the plurality of transform coefficient values to one of the quantized transform coefficient values comprises introducing an error to each quantized transform coefficient value.
4. The method of claim 3, wherein: The quantization of the second frequency domain audio signal comprises: selecting a transform coefficient from a first mapped position, identifying a second mapped position adjacent to the first mapped position, and The transform coefficients are mapped to the second mapped positions.
5. The method of claim 4, wherein: The selection of the transform coefficient is based on an error associated with the quantized transform coefficient value corresponding to the transform coefficient.
6. The method of claim 4, wherein: Mapping the transform coefficients to the second mapped positions includes repeatedly selecting the transform coefficients and mapping the transform coefficients to the second mapped positions until an error is less than a threshold.
7. The method of claim 4, wherein: Mapping the first transform coefficient to the second mapped position comprises: identifying a subset of the plurality of quantized transform coefficient values, identifying the first mapped position as being within the subset of the plurality of quantized transform coefficient values, and The second mapped position is within the subset of the plurality of quantized transform coefficient values.
8. The method of any one of claims 1 to 7, further comprising one of the following: storing the compressed frequency domain audio signal in a computer memory, or The compressed frequency domain audio signal is streamed.
9. A method comprising: receiving a formatted data packet comprising a compressed frequency-domain audio signal; generating a decompressed frequency domain audio signal by decompressing the compressed frequency domain audio signal; transforming the decompressed frequency-domain audio signal using a first non-windowed transform function to generate a first time-domain audio signal; transforming the first time-domain audio signal using a second non-windowed transform function to generate a second time-domain audio signal; as well as A reconstructed time domain audio signal is generated based on the second time domain audio signal.
10. The method of claim 9, wherein: The first non-windowed transform function is a discrete cosine transform (DCT) transform.
11. The method of claim 9 or claim 10, further comprising: Generate an inverse quantized frequency domain audio signal by inverse quantizing the decompressed frequency domain audio signal, wherein the quantization of the decompressed frequency domain audio signal comprises: computing an interleaved sum of a first block of the decompressed frequency domain audio signal, calculating the sum of a second block of the decompressed frequency domain audio signal, and The values of the second block of the decompressed frequency-domain audio signal are iteratively remapped until the sum of the second block of the decompressed frequency-domain audio signal is within a threshold of the interleaved sum of the first block of the decompressed frequency-domain audio signal.
12. The method of claim 11, wherein: Prior to said calculation of said interleaved sum of said first block of said decompressed frequency domain audio signal, said method further comprises: identifying a frequency range associated with the decompressed frequency-domain audio signal, said calculation of said interleaved sum of said first block of said decompressed frequency-domain audio signal is calculated over said frequency range, and The calculation of the interleaved sum of the second block of the decompressed frequency domain audio signal is calculated over the frequency range.
13. The method of any one of claims 9 to 12, further comprising: Generate an inverse quantized frequency domain audio signal by inverse quantizing the decompressed frequency domain audio signal, wherein the quantization of the decompressed frequency domain audio signal comprises: computing an interleaved sum of a first block of the decompressed frequency domain audio signal, reversing the order of elements of the second block of the decompressed frequency-domain audio signal, calculating an interleaved sum of the second block of the decompressed frequency domain audio signal, and The values of the second block of the decompressed frequency-domain audio signal are iteratively remapped until the sum of the second block of the decompressed frequency-domain audio signal is within a threshold of the interleaved sum of the first block of the decompressed frequency-domain audio signal.
14. The method of any one of claims 9 to 13, further comprising: Generate an inverse quantized frequency domain audio signal by inverse quantizing the decompressed frequency domain audio signal, wherein the quantization of the decompressed frequency domain audio signal comprises: reversing the order of elements of the first block of the decompressed frequency-domain audio signal, calculating a sum of said first block of said decompressed frequency domain audio signals, calculating the sum of a second block of the decompressed frequency domain audio signal, and The values of the second block of the decompressed frequency-domain audio signal are iteratively remapped until the sum of the second block of the decompressed frequency-domain audio signal is within a threshold of the sum of the first block of the decompressed frequency-domain audio signal.
15. The method of any one of claims 9 to 14, further comprising playing back the reconstructed time-domain audio signal.
16. A method comprising: generating a blocked time-domain audio signal as a portion of the time-domain audio signal; transforming the blocked time-domain audio signal using a first non-windowed transform function to generate a first frequency-domain audio signal; transforming the first frequency-domain audio signal using a second non-windowed transform function to generate a second frequency-domain audio signal; compressing the second frequency-domain audio signal to generate a compressed frequency-domain audio signal; generating a decompressed frequency domain audio signal by decompressing the compressed frequency domain audio signal; transforming the decompressed frequency-domain audio signal using a third non-windowed transform function to generate a third time-domain audio signal; transforming the third time domain audio signal using a fourth non-windowed transform function to generate a fourth time domain audio signal; and A reconstructed time domain audio signal is generated based on the fourth time domain audio signal.
17. The method of claim 16, further comprising: A quantized frequency domain audio signal is generated by quantizing the second frequency domain audio signal, wherein: said compressing of said second frequency domain audio signal comprises compressing said quantized frequency domain audio signal, The second frequency domain audio signal comprises a plurality of transform coefficient values, quantizing the second frequency domain audio signal comprises mapping each transform coefficient value of the plurality of transform coefficient values to a quantized transform coefficient value of a plurality of quantized transform coefficient values, and Mapping each transform coefficient value of the plurality of transform coefficient values to one of the quantized transform coefficient values comprises introducing an error to each quantized transform coefficient value.
18. The method of claim 17, wherein: The quantization of the second frequency domain audio signal comprises: selecting a transform coefficient from a first mapped position, identifying a second mapped position adjacent to the first mapped position, and The transform coefficients are mapped to the second mapped positions.
19. The method of claim 18, wherein: The selection of the transform coefficient is based on an error associated with the quantized transform coefficient value corresponding to the transform coefficient.
20. The method of any one of claims 16 to 19, further comprising: Generate an inverse quantized frequency domain audio signal by inverse quantizing the decompressed frequency domain audio signal, wherein the quantization of the decompressed frequency domain audio signal comprises: computing an interleaved sum of a first block of the decompressed frequency domain audio signal, calculating the sum of a second block of the decompressed frequency domain audio signal, and The values of the second block of the decompressed frequency-domain audio signal are iteratively remapped until the sum of the second block of the decompressed frequency-domain audio signal is within a threshold of the interleaved sum of the first block of the decompressed frequency-domain audio signal.
21. The method of claim 20, wherein: Prior to said calculation of said interleaved sum of said first block of said decompressed frequency-domain audio signal, identifying a frequency range associated with the decompressed frequency-domain audio signal, said calculation of said interleaved sum of said first block of said decompressed frequency-domain audio signal is calculated over said frequency range, and The calculation of the interleaved sum of the second block of the decompressed frequency-domain audio signal is calculated over the frequency range.
22. The method according to claim 20, further comprising: Generate an inverse quantized frequency domain audio signal by inverse quantizing the decompressed frequency domain audio signal, wherein the quantization of the decompressed frequency domain audio signal comprises: computing an interleaved sum of a first block of the decompressed frequency domain audio signal, reversing the order of elements of the second block of the decompressed frequency-domain audio signal, calculating an interleaved sum of the second block of the decompressed frequency domain audio signal, and The values of the second block of the decompressed frequency-domain audio signal are iteratively remapped until the sum of the second block of the decompressed frequency-domain audio signal is within a threshold of the interleaved sum of the first block of the decompressed frequency-domain audio signal.
23. The method of claim 20, further comprising: Generate an inverse quantized frequency domain audio signal by inverse quantizing the decompressed frequency domain audio signal, wherein the quantization of the decompressed frequency domain audio signal comprises: reversing the order of elements of the first block of the decompressed frequency-domain audio signal, calculating a sum of said first block of said decompressed frequency domain audio signals, calculating the sum of a second block of the decompressed frequency domain audio signal, and The values of the second block of the decompressed frequency-domain audio signal are iteratively remapped until the sum of the second block of the decompressed frequency-domain audio signal is within a threshold of the sum of the first block of the decompressed frequency-domain audio signal.
24. The method of any one of claims 16 to 23, further comprising playing back the reconstructed time domain audio signal.