Frame segmentation and grouping for audio coding

By combining auxiliary information bit rate and spectral coefficient quantization distortion modification, audio blocks are grouped and processed, solving the problem of increased signal distortion during block combination in existing technologies and achieving more efficient audio coding results.

CN120937075APending Publication Date: 2025-11-11DOLBY LABORATORIES LICENSING CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202480021059.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-01
Filing Date
2024-03-18
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing audio coding techniques fail to effectively consider the distortion changes caused by spectral coefficient quantization during block assembly, resulting in increased overall signal distortion. Furthermore, existing methods are computationally intensive, making it difficult to optimize the signal processing efficiency of each frame.

Method used

By combining the distortion changes caused by auxiliary information bit rate and spectral coefficient quantization, an audio block grouping and processing method based on quality metric and distortion estimation is adopted. Greedy merging and fast optimization methods are used to optimize block combination, thereby reducing quantization distortion and improving coding efficiency.

Benefits of technology

This reduces quantization distortion while decreasing auxiliary information rate, thereby improving signal processing efficiency and overall quality of audio coding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120937075A_ABST
    Figure CN120937075A_ABST
Patent Text Reader

Abstract

Devices, systems, and methods for encoding an audio block having content into a frame are described. An example method includes receiving an input signal including a block of audio information. The audio information block includes a set of block groups for respective frames. Some methods include obtaining a first quality metric for each respective block group, and obtaining a second quality metric for each respective block group. The first quality metric indicates a cost associated with merging two or more blocks of audio information to form a respective block group. The second quality metric indicates an estimated distortion associated with merging two or more blocks of audio information to form a respective block group. The method includes merging the at least two block groups based on the first quality metric and the second quality metric to generate an encoded signal.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to U.S. Provisional Application No. 63 / 560,563, filed March 1, 2024, and U.S. Provisional Application No. 63 / 491,839, filed March 23, 2024, all of which are incorporated herein by reference in their entirety. Technical Field

[0003] This application generally relates to audio and speech encoding and decoding, and more specifically to transform and subband encoding and decoding. Background Technology

[0004] Unless otherwise indicated herein, the material described in this section is not prior art to the claims of this application and is not considered prior art by virtue of its inclusion in this section.

[0005] Many audio processing systems operate by dividing a stream of audio information into frames and further dividing the frames into blocks of sequential data representing portions of audio information within specific time intervals. Some type of signal processing is applied to each block in the stream. An example of an audio processing system that applies a perceptual coding process to each block is a system conforming to the Advanced Audio Codec (AAC) standard, described in ISO / IEC 13818-7, “MPEG-2 advanced audio coding, AAC,” International Standard, 1997; ISO / IEC JTCI / SC29, “Information technology—very low bitrate audio—visual coding,” and ISO / IEC IS-14496 (Part 3, Audio), 2019; and the so-called AC-3 system conforming to the codec standard described in the Advanced Television Systems Committee (ATSC) document A / 52A, published on August 20, 2001, entitled “Revision A to Digital Audio Compression (AC-3) Standard.”

[0006] Many perceptual audio and speech codecs operate in the frequency domain, thereby transforming segments of time-domain samples into sets of spectral coefficients (e.g., blocks). Each block is represented in the encoded bitstream by a combination of the set of spectral coefficients and some control parameters (e.g., auxiliary information). Sets of one or more blocks, along with associated auxiliary information, are combined into a single frame. In most encoders, the number of bits allocated to the spectral coefficients and auxiliary information within each frame depends on the characteristics of the input audio signal. The auxiliary information typically represents high-level information about the signal, while the spectral coefficients convey fine-grained information. For signals where the high-level information does not change much across a frame (e.g., short-term stationary signals), a low auxiliary information rate can be used. This allows for more bits to be allocated to the spectral coefficients, resulting in a corresponding reduction in overall quantization noise compared to time-invariant allocation schemes. For signals where the high-level information changes significantly across a frame (e.g., short-term non-stationary signals), a relatively high auxiliary information rate is required to preserve temporal characteristics.

[0007] The auxiliary information rate is typically controlled by dividing each frame into one or more separate block groups. For example, in MPEG-4 AAC, a frame may consist of a single Modified Discrete Cosine Transform (MDCT) block of length 1024 or eight blocks of length 128. In the latter case, these eight blocks may be transmitted as eight separate groups of size 1, as a single large group of size 8, or any group arrangement in between. Each group includes associated auxiliary information, and thus the auxiliary information rate increases with the number of groups. This disclosure recognizes various limitations in AAC. For example, the primary auxiliary information rate, or cost, in AAC is allocated to a scaling factor value. Because the scaling factor is shared across all blocks in a group, adding a new group will increase the auxiliary information rate by the amount of additional bits required to represent the scaling factor and associated information. For codecs operating at a constant bit rate, increasing the auxiliary information reduces the number of bits available for the spectral coefficients. Therefore, for audio encoders, it is advantageous to carefully select the grouping of blocks to optimize the signal processing efficiency of each frame.

[0008] The only known fully optimal solution is based on exhaustive search; however, this method is computationally very intensive for most codec applications. Greedy merging does not require as much computation as exhaustive search and often achieves near-optimal results. This disclosure recognizes various limitations of previously implemented search methods. For example, a limitation of previous optimization methods is that distortion is defined only based on auxiliary information (e.g., scaling factors or spectral envelopes), as described in, for example, U.S. Patent No. 7,840,410, “Audio Coding Based on Block Grouping,” which is incorporated herein in its entirety. This type of method attempts to preserve the original spectral shape while reducing the amount of auxiliary information required. For example, the cost metric is expressed as a measure of the error between the logarithmic energy of the spectral coefficients from two separate block groups and the logarithmic energy of the spectral coefficients from the candidate merged group. However, these previous grouping methods do not account for changes in distortion due to spectral coefficient quantization. For a fixed number of spectral coefficient bits in a frame, merging two groups can at most preserve the same spectral coefficient quantization error as when the two groups are transmitted separately. However, in many merging scenarios, the quantization error will increase. By observing only the spectral envelope energy, existing methods only partially explain the impact of grouping decisions on overall signal distortion at the decoder output.

[0009] This article is being published precisely because of these and other considerations. Summary of the Invention

[0010] Techniques for processing audio signals are described. The various embodiments described herein provide a system and method for grouping and processing audio blocks based on both the auxiliary information bit rate and changes in distortion due to coefficient quantization.

[0011] Apparatus, systems, and methods for encoding audio blocks having content into frames are described. Example methods include receiving an input signal comprising audio information blocks. The audio information blocks comprise a set of block groups for a corresponding frame. Some methods include obtaining a first quality metric for each corresponding block group and a second quality metric for each corresponding block group. The first quality metric indicates the cost associated with merging two or more audio information blocks to form the corresponding block group. The second quality metric indicates an estimated distortion associated with merging two or more audio information blocks to form the corresponding block group. Methods include merging at least two block groups based on the first and second quality metrics to generate an encoded signal.

[0012] According to an example embodiment, a method for encoding audio blocks in a frame is provided, each frame comprising a set of block groups, and each block comprising content. The method includes receiving an input signal comprising audio information blocks. The audio information blocks comprise a set of block groups for a given frame. The method includes obtaining a first quality metric for each given block group and obtaining a second quality metric for each given block group. The first quality metric indicates the cost associated with merging two or more audio information blocks to form a given block group. The second quality metric indicates an estimated distortion associated with merging two or more audio information blocks to form a given block group. The method includes merging at least two block groups from the set of block groups based on the first and second quality metrics to generate an encoded signal representing the content of the input signal and associated control parameters for each block group in the set, and outputting the encoded signal.

[0013] According to another example embodiment, an apparatus for processing audio information blocks arranged in frames is provided. The apparatus includes an electronic processor configured to receive an input signal comprising audio information blocks. The audio information blocks comprise a set of block groups for a corresponding frame. The electronic processor is configured to obtain a first quality metric for each corresponding block group and a second quality metric for each corresponding block group. The first quality metric indicates the cost associated with merging two or more audio information blocks to form a corresponding block group. The second quality metric indicates an estimated distortion associated with merging two or more audio information blocks to form a corresponding block group. The electronic processor is configured to merge at least two block groups from the set of block groups based on the first and second quality metrics to generate an encoded signal representing the content of the input signal and associated control parameters for each block group in the set, and output the encoded signal.

[0014] According to another example embodiment, a non-transitory computer-readable storage medium is provided, the medium recording a program executable by a device to perform instructions for processing audio information blocks arranged in frames. The method includes receiving an input signal comprising audio information blocks. The audio information blocks comprise a set of block groups for a corresponding frame. The method includes obtaining a first quality metric for each corresponding block group and obtaining a second quality metric for each corresponding block group. The first quality metric indicates the cost associated with merging two or more audio information blocks to form a corresponding block group. The second quality metric indicates an estimated distortion associated with merging two or more audio information blocks to form a corresponding block group. The method includes merging at least two block groups from the set of block groups based on the first and second quality metrics to generate an encoded signal representing the content of the input signal and associated control parameters for each block group in the set, and outputting the encoded signal.

[0015] According to another example embodiment, a method for encoding audio blocks in a frame is provided. The method includes receiving an input signal comprising a set of block groups in the frame, wherein each block comprises content. In a first loop, the method includes sequentially selecting each block group from the set of block groups as a selected first block group for potential merging. In a second loop, the method includes sequentially selecting each block group different from the first block group from the set of block groups as a selected second block group for potential merging, obtaining a first quality metric associated with merging the selected first block group and the selected second block group, obtaining a second quality metric associated with merging the selected first block group and the selected second block group, and comparing the first and second quality metrics of the second block group with previous iterations of the second loop to selectively identify the second block group as a merging candidate block. The method includes, after all iterations of the second loop have been completed, selectively merging the selected first block group and the identified merging candidate blocks, assembling the frame into an encoded signal, and outputting the encoded signal.

[0016] According to another example embodiment, an apparatus is provided for encoding audio information blocks arranged in frames, wherein each frame includes a set of block groups, and wherein each block includes content. The apparatus includes an electronic processor configured to perform an operation of a method comprising receiving an input signal including the set of block groups in the frame, wherein each block includes content. In a first loop, the method includes sequentially selecting each block group from the set of block groups as a selected first block group for potential merging. In a second loop, the method includes sequentially selecting each block group different from the first block group from the set of block groups as a selected second block group for potential merging, obtaining a first quality metric associated with merging the selected first block group and the selected second block group, obtaining a second quality metric associated with merging the selected first block group and the selected second block group, and comparing the first and second quality metrics of the second block group with previous iterations of the second loop to selectively identify the second block group as a merging candidate block. The method includes, after all iterations of the second loop have been completed, selectively merging the selected first block group with the identified merging candidate blocks, assembling the frame into an encoded signal, and outputting the encoded signal.

[0017] According to another example embodiment, a non-transitory computer-readable storage medium is provided, the medium recording a program executable by a device to perform a method comprising receiving an input signal comprising a set of block groups in a frame, wherein each block comprises content. In a first loop, the method comprises sequentially selecting each block group from the set of block groups as a selected first block group for potential merging. In a second loop, the method comprises sequentially selecting each block group different from the first block group from the set of block groups as a selected second block group for potential merging, obtaining a first quality metric associated with merging the selected first block group and the selected second block group, obtaining a second quality metric associated with merging the selected first block group and the selected second block group, and comparing the first and second quality metrics of the second block group with previous iterations of the second loop to selectively identify the second block group as a merging candidate block. The method comprises, after all iterations of the second loop have been completed, selectively merging the selected first block group and the identified merging candidate blocks, assembling the frame into an encoded signal, and outputting the encoded signal.

[0018] In some examples, the method for encoding audio blocks in a frame includes receiving an input signal comprising audio information blocks and merging the set of audio information blocks based on (i) the cost associated with merging two or more audio information blocks and (ii) the estimated distortion resulting from merging two or more audio information blocks.

[0019] In some additional examples, the cost associated with merging two or more audio information blocks and / or the estimated distortion resulting from merging two or more audio information blocks is determined based on the weighted dB cost of those two or more blocks. The weighted dB cost is implemented to identify whether merging two or more blocks is more advantageous (e.g., more cost-efficient) than not merging them. The weighted dB cost can be based on the average power level of each scaling factor band and the power level of each audio information block.

[0020] In some other examples, the cost associated with merging two or more audio information blocks and / or the estimated distortion resulting from merging two or more audio information blocks are determined based on the bit cost of the transmitted audio information block calculated using perceptual entropy.

[0021] In some other examples, merging two or more audio information blocks includes merging adjacent blocks. An audio information block may include time-domain samples of audio information, frequency-domain coefficients of audio information, or a combination thereof.

[0022] In this way, various aspects of this disclosure provide processing of audio blocks and achieve improvements at least in the technical fields of audio encoding, audio decoding, etc.

[0023] The embodiments described herein can generally be described as technology, wherein the term “technology” can refer to one or more systems, devices, methods, computer-readable instructions, modules, components, hardware logic and / or operations as suggested in the context to which this is applied.

[0024] Other features and technical benefits beyond those explicitly described above will become apparent from the following detailed description and by reviewing the associated drawings. This summary is provided to illustrate a range of techniques in a simplified form and is not intended to identify key or essential features of the claimed subject matter, as defined by the appended claims. Attached Figure Description

[0025] These and other more detailed and specific features of the various embodiments are disclosed more fully in the following description with reference to the accompanying drawings, wherein:

[0026] Figure 1 A block diagram of an example audio codec system in which various aspects of the present invention can be incorporated is illustrated.

[0027] Figure 2A A block diagram illustrating an example electronic device architecture suitable for implementing various aspects of the present invention is shown.

[0028] Figure 2B The illustrations depict various aspects that can be used to implement this disclosure. Figure 2A A schematic block diagram of an example CPU implemented in the device architecture.

[0029] Figure 3 The diagram illustrates flowcharts for various example methods for merging blocks based on quality metrics.

[0030] Figure 4 The diagram illustrates flowcharts for some example methods of merging blocks based on 6dB / bit estimation rules.

[0031] Figure 5 The flowchart illustrates an example method for perceptual entropy-based merging blocks.

[0032] Figure 6 A block diagram illustrating an example of a greedy merging process applied to four blocks according to various aspects of the present invention is shown.

[0033] Figure 7 A flowchart illustrating an example method performed by a decoding device according to various aspects of the present invention is shown. Detailed Implementation

[0034] To provide an understanding of one or more aspects of this disclosure, various details such as audio device configuration, timing, and operation are set forth in the following description. It will be readily recognized by those skilled in the art that these specific details are merely examples and are not intended to limit the scope of this application.

[0035] As used herein, the term “comprising” and its variations shall be understood as open-ended terms meaning “including but not limited to”. Unless the context clearly indicates otherwise, the term “or” shall be understood as “and / or”. The term “based on” shall be understood as “at least partially based on”. The terms “one example implementation” and “example implementation” shall be understood as “at least one example implementation”. The term “another implementation” shall be understood as “at least one other implementation”. The terms “determined,” “determines,” or “determining” shall be understood as obtaining, receiving, computing, calculating, estimating, predicting, or deriving. Furthermore, in the following description and claims, unless otherwise defined, all technical and scientific terms used herein shall have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0036] Figure 1 A block diagram of an example audio codec system 100 in which various aspects of the present invention can be incorporated is illustrated. The example audio codec system 100 includes an encoder 110 and a decoder 120. The input of the encoder 100 corresponds to a first signal path 105, and the output of the encoder 100 corresponds to a second signal path 115. The input of the decoder 120 corresponds to the second signal path 115, and the output of the decoder 120 corresponds to a third signal path 125.

[0037] Encoder 110 is configured to receive one or more streams of audio information representing one or more channels of an audio signal from a first signal path 105. Encoder 110 is also configured to process the audio information streams to generate an encoded signal, which can be output to a second signal path 115. At the second signal path 115, the encoded signal can be stored (e.g., captured, buffered, and / or recorded) or transmitted (e.g., via a wired or wireless communication medium). Decoder 120 is configured to receive the encoded signal from the second signal path 115. Decoder 120 is also configured to process the encoded signal and generate a decoded signal, which can be output to a third signal path 125. The decoded signal generated by decoder 120 corresponds to a copy of the audio information previously received by encoder 110 from the first signal path 105. At the third signal path, the decoded signal can be stored (e.g., captured and / or recorded), transmitted (e.g., via a wireless or wired electronic communication medium), or output to a listening device (e.g., an audio processing device, such as a receiver, speaker, soundbar, etc.).

[0038] In the various examples described herein, the terms "replica" and "replica signal" are not intended to indicate that the stream of audio information is "identical." Instead, the term "replica" can indicate that the stream of audio information is approximately identical to the original audio information. For example, when encoder 110 generates an encoded signal using a lossless encoding technique, decoder 120 can, in principle, recover a lossless version from the stream that is approximately identical to the original audio information. However, in the example where encoder 110 uses a lossy encoding technique (such as perceptual encoding / decoding), the content of the recovered replica signal is generally not exactly the same as the content of the original stream, but perceptually indistinguishable from the original content. Therefore, as used herein, the terms "replica" and "replica signal" are intended to cover both lossless and lossy encoding techniques.

[0039] Figure 2A A block diagram of an example electronic device architecture 200 (e.g., device 200) suitable for implementing various aspects of this disclosure is shown. Architecture 200 includes, but is not limited to, references to... Figure 3-5The server and client devices, systems, and methods described herein. As shown in the figure, architecture 200 includes a central processing unit (CPU) 201, which is capable of performing various processes based on a program stored, for example, in read-only memory (ROM) 202 or a program loaded from, for example, storage unit 208 into random access memory (RAM) 203. The CPU 201 may be, for example, an electronic processor 201. Data required by the CPU 201 to perform various processes is also stored as needed in RAM 203. The CPU 201, ROM 202, and RAM 203 are connected to each other via bus 204. The CPU 201 can execute instructions stored in ROM 202, RAM 203, or both to perform the methods described herein related to frame segmentation and data encoding and decoding. Input / output (I / O) interface 205 is also connected to bus 204.

[0040] The following components are connected to I / O interface 205: input unit 206, which may include a keyboard, mouse, etc.; output unit 207, which may include a display such as a liquid crystal display (LCD) and one or more speakers; storage unit 208, which includes a hard disk or other suitable storage device; and communication unit 209, which includes a network interface card, such as a network card (e.g., wired or wireless).

[0041] In some implementations, the input unit 206 includes one or more microphones located at different positions (depending on the host device), enabling the capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).

[0042] In some implementations, output unit 207 includes a system with a different number of speakers. Output unit 207 (depending on the capabilities of the host device) can render audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats).

[0043] In some embodiments, communication unit 209 is configured to communicate with other devices (e.g., via a network). Drive 210 is also connected to I / O interface 205 as needed. Removable media 211 (such as a disk, optical disk, magneto-optical disk, flash drive, or other suitable removable media) is mounted on drive 210 to install computer programs read therefrom into storage unit 208 as needed. Those skilled in the art will understand that while apparatus 200 is described as including the components described above, in practice, it is possible to add, remove, and / or replace some of these components, and all such modifications or changes fall within the scope of this disclosure.

[0044] According to exemplary embodiments of this disclosure, the above processes can be implemented as computer software programs or on computer-readable storage media. For example, embodiments of this disclosure include computer program products comprising computer programs tangibly implemented on machine-readable media, the computer programs including program code for performing methods. In such embodiments, the computer program can be downloaded and installed from a network via communication unit 209, and / or installed from removable media 211, such as... Figure 2A As shown in the image.

[0045] Figure 2B The illustrations depict various aspects that can be used to implement this disclosure. Figure 2A A schematic block diagram of an example CPU 201 implemented in device architecture 200. CPU 201 includes an electronic processor 220 and a memory 221. The electronic processor 220 is electrically and / or communicatively connected to the memory 221 to enable bidirectional communication. The memory 221 stores encoding software 222 and / or decoding software 223. In some examples, the memory 221 may be located internally to the electronic processor 220, such as internal cache memory or other internally located ROM, RAM, or flash memory. In other examples, the memory 221 may be located, for example, in ROM 202, RAM 203, flash memory, or removable media 211, or in another non-transitory computer-readable medium contemplated by device architecture 200. In some cases, the electronic processor 220 may implement the encoding software 222 stored in the memory 221 to perform, in particular, Figure 3 Method 300 Figure 4 Method 400 and / or Figure 5 Method 500. Furthermore, the electronic processor 220 can implement the decoding software 223 stored in the memory 221 to perform, in particular... Figure 7 Method 700.

[0046] Generally, the various example embodiments of this disclosure can be implemented in hardware or dedicated circuitry (e.g., a control circuitry system), software, logic, or any combination thereof. For example, the units discussed above can be implemented by a control circuitry system (e.g., CPU 201 and...). Figure 2A(Other component combinations) are executed, and therefore, the control circuitry system can perform the actions described in this disclosure. Some aspects may be implemented in hardware, while others may be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device (e.g., the control circuitry system). Although various aspects of the exemplary embodiments of this disclosure are illustrated and described as block diagrams, flowcharts, or using other graphical representations, it will be appreciated that, as non-limiting examples, the blocks, apparatuses, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.

[0047] Furthermore, the various blocks shown in the flowchart can be considered as method steps, and / or operations resulting from the operation of computer program code, and / or multiple coupled logic circuit elements configured to perform one or more associated functions. For example, embodiments of this disclosure include a computer program product comprising a computer program tangibly implemented on a machine-readable medium, the computer program containing program code configured to perform the methods described above.

[0048] In the context of this disclosure, a machine-readable medium can be any tangible medium that can contain or store a program that can be used by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be non-transitory and can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable optical disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0049] Computer program code used to perform the methods of this disclosure may be written in any combination of one or more programming languages. This computer program code may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus having a system of control circuitry, such that when executed by the processor of the computer or other programmable data processing apparatus, it causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a computer, partially on a computer, as a standalone software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server, or distributed across one or more remote computers and / or servers.

[0050] The embodiments described herein perform frame segmentation by jointly considering the grouping effect based on both auxiliary information cost and distortion caused by spectral coefficient quantization. Indices have been developed for these two cost forms and are typically expressed in bits. For a given pair of adjacent groups as merging candidates, the number of auxiliary information bits saved by merging is compared to the number of bits lost in spectral coefficient quantization. If the number of saved bits is greater than the number of lost bits, then the corresponding adjacent groups can be merged, thereby net reducing quantization distortion.

[0051] The embodiments described herein can be implemented using any cost-based search method to find grouped solutions. These cost-based search methods include exhaustive search, greedy merging, and fast optimization. As an example, greedy merging begins with each block in a frame represented by its own auxiliary information, and then iteratively combines the blocks into groups that minimize a suitable cost metric. In each iteration step, the cost of merging each pair of adjacent groups is compared to the cost of keeping them separate. The pairs of groups with the most favorable cost are merged into one group. This iterative process continues until no adjacent groups can be merged to produce a better solution.

[0052] In some implementations, the auxiliary information primarily consists of a scaling factor, also known as spectral envelope data for AC-3 and split-render codecs. The scaling factor can be derived from the perceptual masking threshold in the context of MPEG AAC. However, other forms of auxiliary information, such as Huffman codebook assignments and segmentation information, can also be used in perceptual audio codecs, as described by M. Bosi et al. in "ISO / IEC MPEG-2 Advanced Audio Coding" (J. Audio Eng. Soc., Vol. 45, No. 10, October 1997). The cost of the auxiliary information can be calculated by the encoder 110 using conventional bit-counting methods. Alternatively, the cost of the auxiliary information can be estimated using the scaling factor, the Huffman codebook, and the long-term average of the segmented data.

[0053] Figure 3 The diagram illustrates flowcharts of various example methods 300 for merging blocks based on quality metrics. Example method 300 can be executed by a processor (such as, for example, CPU 201), which can be configured to execute method 300 via machine-executable instructions. Example method 300 can be divided into multiple blocks or partitions, such as blocks 302, 304, 306, and 308. Figure 3The various process blocks shown provide examples of the various methods disclosed herein, and it should be understood that some blocks may be removed, added, combined, or modified without departing from the spirit of this disclosure. For some examples, processing of various blocks (which may be described as processes, methods, steps, blocks / modules, operations, or functions) may begin at block 302.

[0054] In block 302, "Receive Audio Information Block," CPU 201 receives audio information blocks. The audio information blocks are arranged in frames. In some embodiments, each block of audio information includes content representing a corresponding time interval of the audio information. Furthermore, a block may include time-domain samples of the audio information, frequency-domain coefficients of the audio information, or a combination thereof. Processing can continue from block 302 to block 304.

[0055] In block 304, “a first quality metric is obtained for each block group,” CPU 201 obtains a first quality metric for each block group. For example, CPU 201 can obtain a first quality metric for each possible block group that can be formed by merging various blocks of audio information, as a metric for the cost of auxiliary information. The metric for the cost of auxiliary information can be obtained or determined, for example, through one or more estimations, calculations, or other methods, as will be described in further detail below. Processing can continue from block 304 to block 306.

[0056] In block 306, “A second quality metric is obtained for each block group,” CPU 201 obtains a second quality metric for each block group. For example, CPU 201 may obtain a second quality metric for each possible block group as a measure of distortion caused by quantization of spectral coefficients. The distortion metric may be obtained or determined, for example, via one or more estimations, calculations, or other methods, as will be described in further detail below. Processing may continue from block 306 to block 308.

[0057] In block 308, “A set of audio information blocks merged based on a first quality metric and a second quality metric,” CPU 201 merges a set of audio information blocks based on a first quality metric and a second quality metric. For example, CPU 201 can iteratively process the first and second quality metrics for each block group that can be formed, and selectively merge audio information blocks, which reduces cost while limiting distortion. In some cases, the block group set includes at least two block groups. In other cases, the block group set may include at least one block group, at least three block groups, at least four block groups, etc.

[0058] In some implementations, a rule of 6dB per bit is used to correlate the distortion metric with the change in the number of bits allocated to the spectral coefficients. For example, if in a spectrum containing b kIf, within a specific frequency band of a given spectral coefficient, the scaling factor increases by 3 dB due to merging (relative to encoding / decoding the block individually), then the lost bit depth is approximately 3 dB. k / 6. The resulting distortion in this frequency band will increase. Alternatively, if the scaling factor is reduced by 4dB due to combining, then the number of bits obtained by encoding and decoding this block alone is 4b. k / 6. In this case, the allocation gain can be viewed as over-encoding and decoding of the spectral coefficients.

[0059] Figure 4 The diagram illustrates a flowchart of some example methods 400 for merging blocks based on a 6dB / bit estimation rule. Example methods 400 can be executed by a processor (such as, for example, CPU 201), which can be configured to execute method 400 via machine-executable instructions. Example methods 400 can be divided into multiple blocks or partitions, such as blocks 402, 404, 406, 408, 410, 412, and 414. Figure 4 The various process blocks shown provide examples of the various methods disclosed herein, and it should be understood that some blocks may be removed, added, combined, or modified without departing from the spirit of this disclosure. In some examples, processing of individual blocks (which may be described as processes, methods, steps, blocks / modules, operations, or functions) may begin at block 402.

[0060] In block 402, “Receiving a Pair of Audio Block Groups,” CPU 201 receives a pair of block groups. For example, multiple audio information blocks are received. The audio information blocks form block groups, and each block group contains at least one audio block. CPU 201 considers a first block group and a second block group for merging. In some embodiments, the first block group and the second block group are adjacent block groups. The first block group includes N1 blocks, while the merged group (e.g., a group formed by merging or combining the first block group and the second block group) includes N blocks. Each audio information block includes K scaling factor bands per block. Processing can continue from block 402 to block 404.

[0061] In block 404, “Calculate the total weighted dB cost of the first block group”, CPU 201 calculates the total weighted dB cost of the first block group relative to the first block group without a merged group. For example, Equation 1 provides the calculation of the first total weighted cost C1:

[0062]

[0063] in:

[0064] S nk =10*log 10 (s nk )

[0065]

[0066] And among them:

[0067] w nk =The subjective importance weight of each scaling factor, between 0 and 1;

[0068] b k = The number of spectral coefficients in the k-th frequency band; and

[0069] s nk = Power domain scaling factor for each block in the nth and kth frequency bands.

[0070] Processing can continue from block 404 to block 406.

[0071] In block 406, “Calculate the total weighted dB cost of the second block group”, CPU 201 calculates the total weighted dB cost of the second block group relative to the block group that was not merged. For example, Equation 2 provides the calculation of the second total weighted cost C2:

[0072]

[0073] in:

[0074]

[0075] Processing can continue from block 406 to block 408.

[0076] In block 408, “Calculate the total weighted dB cost of the merged group,” CPU 201 calculates the total weighted dB cost relative to the merged group that was not created. For example, Equation 3 provides the calculation of the total weighted cost C after merging:

[0077]

[0078] in:

[0079]

[0080] Processing can continue from block 408 to block 410.

[0081] In block 410, “Calculating the Merging Ratio Based on Total Weighted Cost,” CPU 201 calculates the merging ratio based on the total weighted dB cost (e.g., based on the calculated first total weighted dB cost, the calculated second total weighted dB cost, and the calculated post-merging total weighted dB cost). For example, CPU 201 calculates the merging ratio R, which represents the estimated number of spectral bits lost during merging divided by the number of bits of auxiliary information saved, as shown in Equation 4:

[0082]

[0083] in:

[0084] b SI = Average number of bits per frequency band used to convey auxiliary information.

[0085] In the consolidation ratio R, the components An estimate of the number of bits "lost" during the merger is provided, including spectral coefficients for under-allocation and over-allocation due to the merger. Component b SI *K provides the average number of bits saved during the candidate merge. Processing can continue from block 410 to block 412.

[0086] In block 412, under the heading "Have all block groups been considered?", CPU 201 determines whether all block groups included in the multiple audio information blocks have been considered. Specifically, CPU 201 determines whether the merge ratio has been calculated for all block groups that can be merged. If not all block groups have been considered, processing proceeds from block 412 to block 414 and CPU 201 selects the next pair of audio block groups. Processing can then return from block 414 to block 404, and CPU 201 continues calculating the merge ratio for the new block group.

[0087] Once all block groups included in the multiple audio information blocks have been considered, processing proceeds from block 412 to block 416. In block 416, regarding the question "Are there any merges that meet the threshold?", CPU 201 determines whether any merges of block groups meet the threshold. For example, CPU 201 compares the calculated merge ratio of each pair of audio block groups to the threshold. For instance, when the calculated merge ratio is lower than the threshold to meet the threshold, a pair of audio block groups associated with the calculated merge ratio having the lowest value is identified. In another example, when the calculated merge ratio is higher than the threshold to meet the threshold, a pair of audio block groups associated with the calculated merge ratio having the highest value is identified.

[0088] In some examples, the threshold can be 1. When the merge ratio R is less than 1, merging the first block group and the second block group into a merge group is considered more advantageous than encoding the first block group and the second block group separately. When the merge ratio R is greater than or equal to 1, merging the first block group and the second block group into a merge group is considered disadvantageous, and the first block group and the second block group are retained as separate groups.

[0089] The numerator can be non-negative because the combined set of scaling factors produces a higher quantizer precision loss than two separate sets. For example, weights w can be calculated based on the sensation-level masking effect. nkThis results in the spectral components that are subjectively louder than other components contributing more to the total cost than less important components. Further details regarding the sensory-level masking effect are described in U.S. Patent Publication No. 2022 / 0415334, “A Psychoacoustic Model for Audio Processing,” which is incorporated herein by reference in its entirety.

[0090] When at least one merge satisfies the threshold, processing can continue from block 416 to block 418. In block 418, "Identifying the Block Groups Most Satisfying the Threshold," CPU 201 identifies which block groups best meet the threshold. For example, if the CPU determines whether the calculated merge ratio R is less than 1, CPU 201 identifies which pair of audio block groups has a calculated merge ratio R that is "minimum" 1. In another example, if the CPU determines whether the calculated merge ratio R is greater than 1, CPU 201 identifies which pair of audio block groups has a calculated merge ratio R that is "maximum" 1. Processing can continue from block 418 to block 420. In block 420, "Merging the Identified First and Second Block Groups," CPU 201 merges the identified first and second block groups identified at block 418.

[0091] In another case, CPU 201 calculates the merge difference S, which provides the difference between the bits "lost" due to merging and the bits saved, as provided in Equation 5:

[0092] In this case, when S is less than 0, CPU 201 merges the first block group and the second block group.

[0093] A variant of estimating spectral coefficient bit loss using a 6dB rule per bit is based on a backward adaptive bit allocation scheme, such as the one described by G. Davidson et al. in “Parametric Bit Allocation in a Perceptual Audio Coder” presented at the 97th AES conference in San Francisco in November 1994. In this case, the scaling factor represents the spectral envelope (e.g., the strip RMS or peak signal energy). Furthermore, if the bit allocation is derived based on perceptual level masking, then two offsetting factors affect the bit loss estimation. As mentioned above, the number of estimated spectral coefficient bits decreases as the spectral envelope increases. Alternatively, using a masking threshold derived from the increased envelope using a perceptual level rule will increase the number of estimated spectral coefficient bits. The effects of these two changes can be combined to estimate the net change in the estimated spectral bits.

[0094] In the example of a backward adaptive codec based on sensory level masking, the first total weighted cost C1, the second total weighted cost C2, and the absolute value operator of C can be adjusted to account for the offsetting effect of sensory level masking. Furthermore, the average auxiliary information difference estimate b can be... SI *K is replaced with the actual bit count difference to improve accuracy. For example, b SI The *K term can be replaced with a value corresponding to the difference between the auxiliary information bit counts of the individual two groups and the combined group, as shown in Equation 6:

[0095]

[0096] Among them, B1, B2 and B M These represent the actual auxiliary information bit counts for the first group, the second group, and the merged group, respectively.

[0097] Referring again to block 416, in some cases, none of the calculated merge ratios meet the threshold. In this case, processing can proceed from block 416 to block 422. In block 422, "Terminate Merging Operation," CPU 201 terminates the merging operation, and all audio block groups are not merged. Various example methods 400 can continue execution until none of the calculated merge ratios meet the threshold. For example, once merging occurs, the process continues to consider each possible merge group and merge the adjacent groups that produce the greatest improvement until no two adjacent groups can be merged to produce an improvement in codec accuracy. Since multiple candidates that meet the threshold can exist, considering each group before performing the merging provides the group that produces the largest amount of improvement.

[0098] Another example used to estimate the bit loss of spectral coefficients during combining is based on perceptual entropy. In some encoders (such as MPEG-4 AAC and Dolby AC-4), the scaling factor is derived from an estimate of a noise masking threshold. The noise masking threshold can be used to estimate the number of bits required to convey the spectral coefficients from each block in a perceptually transparent manner (e.g., perceptual entropy). The estimated number of bits lost after the combining operation is then derived based on the per-band difference accumulated between the perceptual entropy and the actual number of bits used.

[0099] Figure 5 A flowchart illustrating an example method 500 for merging blocks based on perceptual entropy is shown. Example method 500 can be executed by a processor (such as, for example, CPU 201), which can be configured to execute method 500 via machine-executable instructions. Example method 500 can be divided into individual blocks or partitions, such as blocks 502, 504, 506, 508, 510, 512, and 514. Figure 5The various process blocks shown provide examples of the various methods disclosed herein, and it should be understood that some blocks may be removed, added, combined, or modified without departing from the spirit of this disclosure. For some examples, processing of individual blocks (which may be described as processes, methods, steps, blocks / modules, operations, or functions) may begin at block 502.

[0100] In block 502, “Receiving a Pair of Audio Block Groups,” CPU 201 receives a pair of audio block groups. For example, multiple audio information blocks are received. The audio information blocks form block groups, each containing at least one audio block. CPU 201 considers a first block group and a second block group for merging. In some implementations, the first block group and the second block group are adjacent block groups. The first block group comprises N1 blocks, while the merged group (e.g., a group formed by merging or combining the first and second block groups) comprises N blocks. Each block of audio information includes K scaling factor bands per block. Processing can continue from block 502 to block 504.

[0101] In block 504, “Calculate the cost of sending individual blocks”, CPU 201 calculates the cost of sending the first block group and the second block group individually. For example, CPU 201 calculates the cost Cs (in bits) of sending individual blocks relative to baseline conditions, as provided in Equation 7:

[0102]

[0103] in:

[0104] PE nk = Baseline sensing entropy of the nth block and the kth frequency band (each block is encoded and decoded using its own scaling factor);

[0105] G1 nk =Estimation of the bit count of the spectral coefficients of the nth block and the kth frequency band in the first group; and

[0106] G2 nk =Estimation of the bit count of the spectral coefficients of the nth block and the kth frequency band in Group 2. Processing can continue from block 504 to block 506.

[0107] In block 506, “Calculate the cost of sending the merged group”, CPU 201 calculates the cost of sending the first block group and the second block group as a merged group. For example, CPU 201 calculates the cost C of sending the merged group relative to the baseline conditions. m (in bits), as provided in Equation 8:

[0108]

[0109] in:

[0110] M nk= The bit depth estimate of the spectral coefficients of the nth block and the kth frequency band in the merged group.

[0111] Processing can continue from block 506 to block 508.

[0112] In block 508, “Calculate the cost difference between merging and not merging”, CPU 201 calculates the cost difference between pairs of merged and unmerged audio blocks. For example, CPU 201 calculates the cost C. m The difference between the cost Cs and the cost Cs is provided by Equation 9:

[0113] C = C m -C s [Equation 9] Processing can continue from block 508 to block 510.

[0114] In block 510, under the question "Have all block groups been considered?", CPU 201 determines whether all block groups included in multiple audio information blocks have been considered. Specifically, CPU 201 determines whether the cost difference has been calculated for all merging block groups. If not all block groups have been considered, processing can proceed from block 510 to block 512. In block 512, under the question "Select the next pair of audio block groups", CPU 201 selects the next pair of audio block groups. Processing can then return to block 504, and CPU 201 continues calculating the cost difference for the new block group.

[0115] Once all block groups included in multiple audio information blocks have been considered, processing can proceed from block 510 to block 514. In block 514, under the heading "Are there any merges that meet the threshold?", CPU 201 determines whether any merges meet the threshold. For example, CPU 201 compares the calculated cost difference for each pair of audio block groups to the threshold. For instance, when the calculated cost difference is below the threshold and meets the threshold, a pair of audio block groups associated with the lowest cost difference is identified. In another example, when the calculated cost difference is above the threshold and meets the threshold, a pair of audio block groups associated with the highest cost difference is identified.

[0116] In some implementations, the threshold is 0. When the cost difference C is less than 0, merging is considered advantageous, and CPU 201 identifies the first block group and the second block group as candidates for generating the merged group. When the cost difference C is greater than or equal to 0, CPU 201 does not merge the first block group and the second block group.

[0117] Processing can continue from block 514 to block 516. In block 516, "Identifying the pairs of audio block groups that best meet the threshold," CPU 201 identifies which block groups best meet the threshold. For example, when the threshold is 0, CPU 201 identifies which pair of audio block groups is "less than" 0. Processing can continue from block 516 to block 518. In block 518, "Merging the identified first and second block groups," CPU 201 merges the first and second block groups identified at block 516.

[0118] Referring again to block 514, in some cases, none of the calculated merge ratios meet the threshold. In this case, processing can proceed from block 514 to block 520. In block 520, "Terminate Merging Operation," CPU 201 terminates the merging operation, and all audio block groups are not merged. Method 500 can continue execution until none of the calculated merge ratios meet the threshold. For example, once merging occurs, the process continues to consider each possible merge group and merge the adjacent groups that produce the greatest improvement until no two adjacent groups can be merged to produce an improvement in codec accuracy. Since there may be multiple candidates that meet the threshold, considering each group before performing the merging provides the group that produces the largest amount of improvement.

[0119] In some cases, the MDCT transform is used to calculate the spectral coefficients. In this case, Equation 10 can be used to calculate the sensing entropy PE for a block n and a frequency band k. nk Sum of Digits Estimation G1 nk G2 nk M nk

[0120] PE nk =L nk *log2(1.5+0.35*E k / T k [Equation 10] Where:

[0121] L nk =f nk / (E k 0.25 )

[0122]

[0123] And among them:

[0124] L nk This refers to the number of rows related to block n and frequency band k;

[0125] E k It is the average MDCT coefficient energy for all blocks and frequency band k in the group;

[0126] Tk It is the masking threshold for frequency band k, which is the average of all blocks across the group;

[0127] f nk It is a shape factor for block n and frequency band k;

[0128] b is an interval index of an MDCT spectral coefficient; and

[0129] X nb It is an MDCT coefficient from the set of all coefficients from block n and frequency band k.

[0130] PE nk Only when E k >T k PE is calculated only when the time comes; otherwise, it is not calculated. nk =0.

[0131] Figure 5 Only a few example methods for calculating perceptual entropy are provided. Other example methods for calculating perceptual entropy, in conjunction with this disclosure, are welcome and can be implemented in this paper.

[0132] Figure 6 An example of a greedy merging process applied to four blocks according to various aspects of the present invention is illustrated. The greedy merging process can be, for example, method 300, method 400, or method 500. Figure 6 In the example, four blocks are initially arranged into four groups a, b, c, and d, with one block in each group. These four groups a, b, c, and d are the original block groups from the input audio data. Once the CPU 201 begins determining whether block groups should be merged, these groups are candidate block groups. The method then searches for (e.g., identifies or determines) two adjacent groups that should be merged, as determined using method 400 or method 500. In the first iteration, the method finds that groups b and c should be merged because the cost J is less than a threshold T; therefore, groups b and c are merged into a new group to obtain three groups a, bc, and d, where bc is the merged group. In the second iteration, groups a, bc, and d are candidate groups. The method finds (e.g., identifies or determines) that adjacent groups a and bc should be merged because the cost J is less than a threshold T. Groups a and bc are merged into a new group to give a total of two groups abc and d. In the third iteration, the method finds (e.g., identifies or determines) that the cost J of the last remaining pair of groups is greater than the threshold T, and the method terminates, leaving the last two groups abc and d.

[0133] Figure 7A flowchart illustrating an example method 700 for decoding an encoded signal is shown. Example method 700 can be executed by a processor (such as, for example, CPU 201), which can be configured to execute method 700 via machine-executable instructions. Example method 700 can be divided into multiple blocks or partitions, such as blocks 702 and 704. Figure 7 The various process blocks shown provide examples of the various methods disclosed herein, and it should be understood that some blocks may be removed, added, combined, or modified without departing from the spirit of this disclosure. For some examples, processing of individual blocks (which may be described as processes, methods, steps, blocks / modules, operations, or functions) may begin at block 702.

[0134] At block 702, "Receive encoded signal including merged audio information block", CPU 201 receives encoded signal including merged audio information block. The audio information block is merged based on a first quality metric and a second quality metric, as previously described. Processing can continue from block 702 to block 704.

[0135] In block 704, "Decoding the Encoded Signal," CPU 201 decodes the encoded signal. In some cases, the decoded signal specifies the grouping of block groups calculated during encoding in the grouping information (provided as metadata). Accordingly, the decoder reconstructs the audio information by ungrouping the merged groups based on the grouping information. In some cases, to reconstruct the audio information, a scaling factor is decoded from the encoded signal and applied to the spectral coefficients.

[0136] Those skilled in the art will recognize that the present invention is by no means limited to the embodiments described above. Rather, many modifications and variations are possible and are considered to be within the scope of the appended claims. Various aspects and embodiments of this disclosure can also be understood through the following enumerated exemplary embodiments (EEE), which are not claims and may represent systems, methods, and apparatuses arranged entirely according to various aspects of this disclosure.

[0137] EEE1. A method for encoding audio blocks in a frame, wherein each frame includes a set of block groups, and wherein each block includes content, the method comprising: receiving an input signal including audio information blocks, wherein the audio information blocks include a set of block groups for a corresponding frame; obtaining a first quality metric for each corresponding block group, wherein the first quality metric indicates a cost associated with merging two or more audio information blocks to form the corresponding block group; obtaining a second quality metric for each corresponding block group, wherein the second quality metric indicates an estimated distortion associated with merging two or more audio information blocks to form the corresponding block group; merging at least two block groups from the set of block groups based on the first quality metric and the second quality metric to generate an encoded signal, the encoded signal representing the content of the input signal and associated control parameters for each block group in the set; and outputting the encoded signal.

[0138] EEE2. According to the method of EEE1, the steps of obtaining a first quality metric for each corresponding block group and obtaining a second quality metric for each corresponding block group include: in a first loop, sequentially selecting each block group from the block group set as a selected first block group for potential merging; and in a second loop, sequentially selecting each block group different from the selected first block group from the block group set as a selected second block group for potential merging.

[0139] EEE3. According to the method of EEE2, the steps of obtaining a first quality metric for each corresponding block group and obtaining a second quality metric for each corresponding block group include: comparing the first and second metrics of the second block group with previous iterations of the second loop to selectively identify the second block group as a candidate block for merging.

[0140] EEE4. According to the method of EEE3, the step of merging at least two block groups in the block group set to generate an encoded signal and outputting the encoded signal includes: selectively merging the selected first block group with the identified merge candidate block after all iterations of the second loop have been completed; and outputting the encoded signal after all iterations of the second loop have been completed.

[0141] EEE5. A method for encoding audio blocks in a frame, the method comprising: receiving an input signal including a set of block groups in the frame, wherein each block includes content; in a first loop, sequentially selecting each block group from the set of block groups as a selected first block group for potential merging; in a second loop, sequentially selecting each block group different from the first block group from the set of block groups as a selected second block group for potential merging; obtaining a first quality metric associated with merging the selected first block group and the selected second block group; obtaining a second quality metric associated with merging the selected first block group and the selected second block group; comparing the first quality metric and the second quality metric of the second block group with previous iterations of the second loop to selectively identify the second block group as a merging candidate block; after all iterations of the second loop have been completed, selectively merging the selected first block group with the identified merging candidate blocks; and after all iterations of the second loop have been completed, assembling the frame into an encoded signal and outputting the encoded signal.

[0142] EEE6. The method according to any one of EEE1 to EEE5, wherein obtaining the first quality metric and obtaining the second quality metric comprises: calculating a first total weighted dB cost relative to the first block group and the second block group not being merged; calculating a second total weighted dB cost relative to the second block group not being merged; calculating a total weighted cost relative to the merged group not being merged; and determining whether to merge the first block group and the second block group to form a merged group based on the first total weighted dB cost, the second total weighted dB cost, and the total weighted cost.

[0143] EEE7. According to the method of EEE6, determining whether to merge the first block group and the second block group includes: calculating a merging ratio value based on the first total weighted dB cost, the second total weighted dB cost and the total weighted cost; and comparing the merging ratio value with a threshold.

[0144] EEE8. According to the method of EEE7, the merging ratio value is based on the bit count difference between the auxiliary information bit counts of the first block group, the second block group, and the merged group.

[0145] EEE9. According to the method of EEE6, the first total weighted dB cost is calculated based on the average power level of each scaling factor band and the scaling factor of each power domain of each scaling factor band.

[0146] EEE10. According to any one of EEE6 to EEE9, wherein the first block group and the second block group are adjacent block groups.

[0147] EEE11. The method according to any one of EEE1 to EEE5, wherein obtaining the first quality metric and obtaining the second quality metric comprises: calculating the first cost of separately transmitting the first block group and the second block group relative to a baseline perception entropy using perception entropy; calculating the second cost of transmitting the first block group and the second block group as a merged group relative to a baseline perception entropy using perception entropy; and calculating the cost difference between the first cost and the second cost.

[0148] EEE12. According to the method of EEE11, it also includes: determining whether to merge the first block group and the second block group by comparing the cost difference with a threshold.

[0149] EEE13. The method according to any one of EEE1 to EEE12, wherein the block comprises temporal samples of audio.

[0150] EEE14. The method according to any one of EEE1 to EEE12, wherein the block includes the frequency domain coefficients of the audio.

[0151] EEE15. The method according to any one of EEE5 to EEE14 further includes: terminating the potential merging of the selected first block group and the selected second block group after all iterations have been completed in the first loop and after all iterations have been completed in the second loop.

[0152] EEE16. An apparatus for encoding audio information blocks arranged in frames, wherein each frame includes a set of block groups and wherein each block includes content, the apparatus comprising: an electronic processor configured to perform operations including a method according to any one of EEE5 to EEE15.

[0153] EEE17. A non-transitory computer-readable storage medium containing a program of instructions, the program of instructions being executable by a device to perform a method according to any one of EEE5 to EEE15.

[0154] EEE18. An apparatus for encoding audio information blocks arranged in frames, wherein each frame includes a set of block groups and each block includes content, the apparatus comprising: an electronic processor configured to: receive an input signal including audio information blocks, wherein the audio information blocks include a set of block groups for a corresponding frame; obtain a first quality metric for each corresponding block group, wherein the first quality metric indicates the cost associated with merging two or more audio information blocks to form the corresponding block group; obtain a second quality metric for each corresponding block group, wherein the second quality metric indicates an estimated distortion associated with merging two or more audio information blocks to form the corresponding block group; merge at least two block groups from the set of block groups based on the first quality metric and the second quality metric to generate an encoded signal, the encoded signal representing the content of the input signal and associated control parameters for each block group in the set; and output the encoded signal.

[0155] EEE19. The apparatus according to EEE18, wherein, in order to obtain a first quality metric and a second quality metric, the electronic processor is configured to: calculate a first total weighted dB cost relative to the first block group and the second block group that are not merged; calculate a second total weighted dB cost relative to the second block group that are not merged; calculate a total weighted cost relative to the merged group that is not merged; and determine whether to merge the first block group and the second block group to form a merged group based on the first total weighted dB cost, the second total weighted dB cost, and the total weighted cost.

[0156] EEE20. According to the apparatus of EEE19, in order to determine whether to merge the first block group and the second block group, the electronic processor is configured to: calculate a merging ratio value based on a first total weighted dB cost, a second total weighted dB cost and a total weighted cost, and compare the merging ratio value with a threshold.

[0157] EEE21. The apparatus according to EEE20, wherein the merging ratio value is based on the bit count difference between the auxiliary information bit counts of the first block group, the second block group, and the merging group.

[0158] EEE22. The apparatus according to EEE19, wherein the first total weighted dB cost is calculated based on the energy of each scaling factor band and the scaling factor of each power domain of each scaling factor band.

[0159] EEE23. An apparatus according to any one of EEE18 to EEE22, wherein the first block group and the second block group are adjacent block groups.

[0160] EEE24. According to the apparatus of EEE18, in order to obtain a first quality metric and a second quality metric, the electronic processor is configured to: calculate a first cost of separately transmitting a first block group and a second block group relative to a baseline perception entropy using perception entropy; calculate a second cost of transmitting the first block group and the second block group as a merged group relative to a baseline perception entropy using perception entropy; and calculate the cost difference between the first cost and the second cost.

[0161] EEE25. According to the device of EEE24, the electronic processor is configured to determine whether to merge the first block group and the second block group by comparing the cost difference with a threshold.

[0162] EEE26. An apparatus according to any one of EEE18 to EEE25, wherein the block comprises temporal samples of audio.

[0163] EEE27. A non-transitory computer-readable storage medium containing a program of instructions executable by a device to perform a method of processing audio information blocks arranged in frames, the method comprising: receiving an input signal including audio information blocks, wherein the audio information blocks include a set of block groups for a corresponding frame; obtaining a first quality metric for each corresponding block group, wherein the first quality metric indicates a cost associated with merging two or more audio information blocks to form the corresponding block group; obtaining a second quality metric for each corresponding block group, wherein the second quality metric indicates an estimated distortion associated with merging two or more audio information blocks to form the corresponding block group; merging at least two block groups from the set of block groups based on the first quality metric and the second quality metric to generate an encoded signal representing the content of the input signal and associated control parameters including each block group in the set; and outputting the encoded signal.

[0164] EEE28. A method for encoding audio blocks in a frame, the method comprising: receiving an input signal comprising a set of two or more audio block groups in the frame, wherein each block comprises content; selecting one or more pairs of audio block groups from the frame as merging candidates; for each merging candidate, obtaining a first quality metric associated with merging the audio block groups; for each merging candidate, obtaining a second quality metric associated with merging the audio block groups; based on the first and second quality metrics, selecting from the merging candidates a pair of audio block groups that produces the greatest improvement in encoding / decoding accuracy; merging the selected pair of audio block groups; and assembling the frame into an encoded signal and outputting the encoded signal.

[0165] Regarding the processes, systems, methods, heuristics, etc., described herein, it should be understood that although the steps of such processes, etc., are described as occurring in a certain ordered sequence, such processes may be performed in an order other than that described herein. It should also be understood that some steps may be performed simultaneously, other steps may be added, or some steps described herein may be replaced, revised, or omitted. In other words, the description of processes herein is provided to illustrate certain embodiments and should in no way be construed as limiting the claims.

[0166] Accordingly, it should be understood that the above description is intended to be illustrative and not limiting. Many other embodiments and applications besides the examples provided will become apparent upon reading the above description. The scope should not be determined by reference to the above description, but rather by reference to the appended claims and the full scope of their equivalents. The techniques discussed herein are contemplated and intended to be developed in the future, and the disclosed systems and methods will be incorporated into such future embodiments. In conclusion, it should be understood that modifications and variations are possible with this application.

[0167] Unless expressly indicated herein, all terms used in the claims shall be given the broadest reasonable construction and their ordinary meaning as understood by one of skill in the art described herein. In particular, unless the claims expressly limit them to the contrary, the use of singular articles such as “a,” “the,” and “the” shall be understood to refer to one or more of the indicated elements.

[0168] The abstract is provided to allow readers a quick understanding of the nature of the technical disclosure. When submitting the abstract, it should be understood that it will not be used to interpret or limit the scope or meaning of the claims. Furthermore, as can be seen in the foregoing detailed description, various features are grouped together in various embodiments for the purpose of simplifying the disclosure. This method of disclosure should not be construed as reflecting an intention to combine more features than are expressly set forth in each claim. Rather, as reflected in the following claims, the inventive subject matter lies in fewer than all features of a single disclosed embodiment. Therefore, the following claims are hereby incorporated into the detailed description, each claim being a separate claim.

Claims

1. A method for encoding audio blocks in a frame, wherein each frame includes a set of block groups, and wherein each block includes content, the method comprising: Receive (302) an input signal including an audio information block, wherein the audio information block includes a set of block groups for a corresponding frame; For each corresponding block group, a first quality metric is obtained (304), wherein the first quality metric indicates the cost associated with merging two or more audio information blocks to form the corresponding block group; For each corresponding block group, a second quality metric is obtained (306), wherein the second quality metric indicates the estimated distortion associated with merging two or more audio information blocks to form the corresponding block group; At least two block groups in the (308) block group set are merged based on a first quality metric and a second quality metric to generate an encoded signal, the encoded signal representing the content of the input signal and the associated control parameters included in each block group in the merge; as well as Output the encoded signal described in (310).

2. The method of claim 1, wherein the steps of obtaining a first quality metric for each corresponding block group and obtaining a second quality metric for each corresponding block group include: In the first loop, each block group is selected sequentially from the block group set as the first selected block group for potential merging; as well as In the second loop, each block group that is different from the first block group is selected from the block group set as the second block group to be selected for potential merging.

3. The method of claim 2, wherein the steps of obtaining a first quality metric for each corresponding block group and obtaining a second quality metric for each corresponding block group include: The first and second metrics of the second block group are compared with previous iterations of the second loop to selectively identify the second block group as a candidate block for merging.

4. The method of claim 3, wherein the step of merging at least two block groups in the block group set to generate an encoded signal and outputting the encoded signal comprises: After all iterations have been completed in the second loop, the selected first block group is selectively merged with the identified merge candidate blocks; as well as After all iterations have been completed in the second loop, the encoded signal is output.

5. A method for encoding audio blocks in a frame, the method comprising: Receive (302) an input signal, the input signal comprising a set of block groups in a frame, wherein each block comprises content; In the first loop, each block group is selected sequentially from the block group set as the first selected block group for potential merging; In the second loop, each block group that is different from the first block group is selected from the block group set as the second block group to be used for potential merging; Obtain (304) a first quality metric associated with the combination of the selected first block group and the selected second block group; Obtain (306) a second quality metric associated with the combination of the selected first block group and the selected second block group; The first and second quality metrics of the second block group are compared with the previous iterations of the second loop to selectively identify the second block group as a candidate block for merging; After all iterations have been completed in the second loop, the selected first block group is selectively merged with the identified merge candidate blocks; as well as After all iterations have been completed in the second loop, the frames are assembled into an encoded signal and the encoded signal is output.

6. The method of any one of claims 1 to 5, wherein obtaining the first quality metric and obtaining the second quality metric comprises: Calculate (404) the first total weighted dB cost of the first block group relative to the first block group and the second block group not being merged; Calculate (406) the second total weighted dB cost of the second block group relative to the first and second block groups without merging; Calculate (408) the total weighted cost of the merged group relative to the first and second block groups that are not merged; as well as The decision to merge the first block group and the second block group to form a merged group is based on the first total weighted dB cost, the second total weighted dB cost, and the total weighted cost.

7. The method of claim 6, wherein determining whether to merge the first block group and the second block group comprises: The consolidation ratio value (410) is calculated based on the first total weighted dB cost, the second total weighted dB cost, and the total weighted cost; as well as The merging ratio value is compared with the threshold (412).

8. The method of claim 7, wherein the merging ratio value is based on the bit count difference between the auxiliary information bit counts of the first block group, the second block group, and the merged group.

9. The method of claim 6, wherein the calculation of the first total weighted dB cost is based on the average power level of each scaling factor band and the scaling factor of each power domain of each scaling factor band.

10. The method of any one of claims 6 to 9, wherein the first block group and the second block group are adjacent block groups.

11. The method of any one of claims 1 to 5, wherein obtaining (304) the first quality metric and obtaining (306) the second quality metric comprises: The first cost of sending the first block group and the second block group separately relative to the baseline perception entropy is calculated using perception entropy (504). The second bit cost of sending the first block group and the second block group as a merged group is calculated (506) relative to the baseline perception entropy using the perception entropy; and Calculate the cost difference between the first cost and the second cost (508).

12. The method of claim 11, further comprising: Whether to merge the first block group and the second block group is determined by comparing the cost difference with a threshold.

13. The method of any one of claims 1 to 12, wherein the block comprises temporal samples of audio.

14. The method of any one of claims 1 to 12, wherein the block comprises frequency domain coefficients of the audio.

15. The method of any one of claims 5 to 14, further comprising: After all iterations have been completed in the first loop and after all iterations have been completed in the second loop, terminate the potential merge of the selected first block group and the selected second block group.

16. An apparatus for encoding audio information blocks arranged in frames, wherein each frame includes a set of block groups, and wherein each block includes content, the apparatus comprising: An electronic processor configured to perform operations including the method as described in any one of claims 5 to 15.

17. A non-transitory computer-readable storage medium containing a program of instructions, the program of instructions being executable by a device to perform the method as claimed in any one of claims 5 to 15.

18. An apparatus (200) for encoding audio information blocks arranged in frames, wherein each frame includes a set of block groups and each block includes content, the apparatus comprising: An electronic processor (201) is configured to: Receive an input signal including audio information blocks, wherein the audio information blocks include a set of block groups for a corresponding frame; A first quality metric is obtained for each corresponding block group, wherein the first quality metric indicates the cost associated with merging two or more audio information blocks to form the corresponding block group; A second quality metric is obtained for each corresponding block group, wherein the second quality metric indicates the estimated distortion associated with merging two or more audio information blocks to form the corresponding block group; At least two block groups in the block group set are merged based on a first quality metric and a second quality metric to generate an encoded signal, the encoded signal representing the content of the input signal and the associated control parameters of each block group in the set; as well as Output the encoded signal.

19. The apparatus (200) as claimed in claim 18, wherein, In order to obtain the first mass metric and the second mass metric, the electronic processor (201) is configured to: Calculate the first total weighted dB cost of the first block group relative to the first block group and the second block group not being merged. Calculate the second total weighted dB cost of the second block group relative to the second block group without merging the first and second block groups. Calculate the total weighted cost of the merged group relative to the first and second block groups that are not merged, and The decision to merge the first block group and the second block group to form a merged group is based on the first total weighted dB cost, the second total weighted dB cost, and the total weighted cost.

20. The apparatus (200) as claimed in claim 19, wherein, In order to determine whether to merge the first block group and the second block group, the electronic processor (201) is configured to: The consolidation ratio is calculated based on the first total weighted dB cost, the second total weighted dB cost, and the total weighted cost. The merging ratio value is compared with the threshold.

21. The apparatus (200) of claim 20, wherein the merging ratio value is based on the bit count difference between the auxiliary information bit counts of the first block group, the second block group, and the merged group.

22. The apparatus (200) of claim 19, wherein the calculation of the first total weighted dB cost is based on the energy of each scaling factor band and the scaling factor of each power domain of each scaling factor band.

23. The apparatus (200) according to any one of claims 18 to 22, wherein the first block group and the second block group are adjacent block groups.

24. The apparatus (200) as claimed in claim 18, wherein, In order to obtain a first mass metric and a second mass metric, the electronic processor (201) is configured to: The first cost of sending the first and second block groups separately relative to the baseline perception entropy is calculated using perception entropy. The second bit cost of sending the first and second block groups as merged groups is calculated using the perception entropy relative to the baseline perception entropy, and Calculate the cost difference between the first cost and the second cost.

25. The apparatus (200) of claim 24, wherein the electronic processor (201) is configured to: The decision to merge the first and second block groups is made by comparing the cost differences with a threshold.

26. The apparatus (200) of any one of claims 18 to 25, wherein the block comprises temporal samples of audio.

27. A non-transitory computer-readable storage medium containing a program of instructions, the program of instructions being executable by a device to perform a method (300) of processing audio information blocks arranged in frames, the method (300) comprising: Receive (302) an input signal including an audio information block, wherein the audio information block includes a set of block groups for a corresponding frame; For each corresponding block group, a first quality metric is obtained (304), wherein the first quality metric indicates the cost associated with merging two or more audio information blocks to form the corresponding block group; For each corresponding block group, a second quality metric is obtained (306), wherein the second quality metric indicates the estimated distortion associated with merging two or more audio information blocks to form the corresponding block group; At least two block groups in a set of block groups are merged (308) based on a first quality metric and a second quality metric to generate an encoded signal, the encoded signal representing the content of the input signal and the associated control parameters of each block group in the set; and Output the encoded signal.

28. A method for encoding an audio block in a frame, the method comprising: Receive an input signal, the input signal comprising a set of two or more audio block groups in a frame, wherein each block comprises content; Select one or more pairs of audio block groups from the frame as merging candidates; For each merge candidate, obtain the first quality metric associated with the merge of the audio block group; For each merge candidate, obtain a second quality metric associated with the merge of the audio block group; Based on the first and second quality metrics, select a pair of audio blocks from the merging candidates that produce the greatest improvement in codec accuracy; Merge the selected pair of audio block groups; as well as The frames are assembled into an encoded signal and the encoded signal is output.

Citation Information

Patent Citations

  • A psychoacoustic model for audio processing

    US20220415334A1

  • Audio coding based on block grouping

    US7840410B2