Spectral classifier for audio coding mode selection

A novel metric for audio coding technologies addresses the challenge of distinguishing between harmonic and noise-like music signals by using spectral peakiness and local energy concentration measures, improving encoding mode selection in complex codecs.

JP2025157307APending Publication Date: 2025-10-15TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025114186
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-10-15

Smart Images

  • Figure 2025157307000001_ABST
    Figure 2025157307000001_ABST
Patent Text Reader

Abstract

To provide a communication method in an encoder for determining which of two coding modes or groups of coding modes to use, along with related devices and nodes.SOLUTION: A method comprises: deriving a frequency spectrum of an input audio signal (1001); obtaining a magnitude of a critical frequency region in the frequency spectrum (1003); obtaining a frame peakedness measure (1005); obtaining a noise band detection measure (1007); determining which of two encoding modes or groups of encoding modes to use based on at least the peakedness measure and the noise band detection measure (1009); and encoding the input audio signal based on the determined encoding mode to use (1011).SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates generally to communications, and more particularly to communication methods and associated devices and nodes supporting wireless communications. [Background technology]

[0002] Modern audio codecs consist of multiple compression schemes optimized for signals with different characteristics. Speech-like signals are typically processed with codecs operating in the time domain, while music signals are processed with codecs operating in the transform domain. Coding schemes intended to handle both speech and music signals require a mechanism to recognize the input signal (a speech / music classifier) ​​and switch to the appropriate codec mode. A schematic diagram of a multi-mode audio codec that uses mode decision logic based on the input signal is shown in Figure 1.

[0003] Similarly, between classes of music signals, we can distinguish between noise-like and harmonic music signals and build classifiers and optimal coding schemes for each of these groups. In particular, identifying signals with sparse and peaky structures is of great interest, as transform domain codecs are well suited to handling these types of signals. Crest C or spectral flatness f determined according to TIFF2025157307000002.tif17170 There are several known signal measures that aim to identify peak signal structures such as TIFF2025157307000003.tif18170.

[0004] A high spectral flatness or crest may indicate that a coding mode suitable for such a spectrum may be selected. Summary of the Invention

[0005] Currently, a particular challenge exists: In the field of audio coding, various speech and music classifiers are used. However, these speech and music classifiers may not be able to distinguish between different classes in the space of music signals. Many speech and music classifiers do not provide sufficient resolution to distinguish between classes, which is required in complex multi-mode codecs.

[0006] The problem of harmonic and noise-like music segment discrimination is solved by a novel metric computed directly on the frequency-domain coefficients. The metric is based on a spectral peakiness measure and a measure of local concentration of energy that indicates the noisy components of the spectrum.

[0007] Various embodiments of the inventive concept address these challenges by performing a frequency-domain analysis of critical bands of the spectrum. The analysis includes at least a peakiness measure, and various embodiments provide additional measures that provide an indication of noisy bands in the spectrum. Based on these measures, a decision is made whether to use at least one coding mode that targets signals with strong peaks while avoiding signals with noisy bands.

[0008] According to some embodiments of the inventive concept, there is provided a method in an encoder for determining which of two coding modes or a group of coding modes to use. The method includes deriving a frequency spectrum of an input audio signal. The method further includes obtaining a magnitude of a critical frequency range of the frequency spectrum. The method further includes obtaining a peakiness measure. The method further includes obtaining a noise band detection measure. The method further includes determining which of the two coding modes or the group of coding modes to use based on at least the peakiness measure and the noise band detection measure. The method further includes encoding the input audio signal based on the coding mode determined to be used.

[0009] Similar encoders, computer programs, and computer program products are provided.

[0010] According to another embodiment of the inventive concept, there is provided a method in an encoder for determining whether an input audio signal has high peakiness and low energy density. The method includes deriving a frequency spectrum of the input audio signal. The method further includes obtaining a magnitude of a critical frequency range of the frequency spectrum. The method further includes obtaining a peakiness measure. The method further includes obtaining a noise band detection measure. The method further includes determining a harmonic condition based on at least the peakiness measure and the noise band detection measure. The method includes outputting an indication of whether the harmonic condition is true or false.

[0011] Similar encoders, computer programs, and computer program products are provided.

[0012] The accompanying drawings, which are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this specification, illustrate certain non-limiting embodiments of the inventive concepts. [Brief explanation of the drawings]

[0013] [Figure 1] FIG. 1 is a block diagram illustrating a multi-mode audio codec that uses a mode decision log based on an audio input signal. [Figure 2] FIG. 1 is a diagram of acceptable and unacceptable spectra in accordance with some embodiments of the inventive concepts. [Figure 3] FIG. 1 is a classification diagram illustrating desired and undesired signals in accordance with some embodiments of the inventive concepts. [Figure 4] 1 is a flowchart illustrating the operation of an encoder in accordance with some embodiments of the inventive concept. [Figure 5]1 is a block diagram illustrating a multi-mode audio codec that uses a mode decision log based on an audio input signal, in accordance with some embodiments of the inventive concept. [Figure 6A] 1 is an illustration of a decision tree according to some embodiments of the inventive concepts. [Figure 6B] 1 is an illustration of a decision tree according to some embodiments of the inventive concepts. [Figure 7] 1 is a block diagram illustrating an example operating environment in accordance with some embodiments of the inventive concepts. [Figure 8] FIG. 1 is a block diagram illustrating a block diagram of a virtualization environment in accordance with some embodiments of the inventive concepts. [Figure 9] 1 is a block diagram illustrating an encoder in accordance with some embodiments of the inventive concept. [Figure 10] 1 is a flowchart illustrating the operation of an encoder in accordance with some embodiments of the inventive concept. [Figure 11] 1 is a flowchart illustrating the operation of an encoder in accordance with some embodiments of the inventive concept. [Figure 12] 1 is a flowchart illustrating the operation of an encoder in accordance with some embodiments of the inventive concept. Modes for carrying out the invention

[0014] Some of the embodiments contemplated herein will now be described more fully with reference to the accompanying drawings. The embodiments are provided as examples to convey the scope of the subject matter to those skilled in the art, and illustrate examples of embodiments of the inventive concepts. However, the inventive concepts may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be comprehensive and complete, and will fully convey the scope of the inventive concepts to those skilled in the art. It should also be noted that these embodiments are not mutually exclusive. Elements from one embodiment may be implicitly assumed to be present / used in another embodiment.

[0015] Before describing the embodiments in further detail, FIG. 7 illustrates an example operating environment for an encoder 500 that may be used to encode a bitstream as described herein. The encoder 500 receives audio from a network 702 and / or from a storage 704 and / or from an audio recorder 706, encodes the audio into a bitstream as described below, and transmits the encoded audio to a decoder 708 via a network 710. In some embodiments in which the encoder 500 is a distributed encoder, a transmitting entity 5001 may transmit the encoded audio to the decoder 708 via the network 710, as indicated by the dashed line. The storage device 704 may be part of a storage repository for multi-channel audio signals, such as a store or streaming audio service storage repository, a separate storage component, a component of a mobile device, etc. The decoder 708 may be part of a device 712 having a media player 714. The device 712 may be a mobile device, a set-top device, a desktop computer, etc.

[0016] FIG. 8 is a block diagram illustrating a virtualization environment 800 in which functionality implemented by some embodiments may be virtualized. In this context, virtualizing means creating a virtual version of an apparatus or device, such as the encoder 500, which may include virtualizing the hardware platform, storage devices, and networking resources. As used herein, virtualization may apply to any device or component thereof described herein and relates to implementations in which at least a portion of functionality is implemented as one or more virtual components. Some or all of the functionality described herein may be implemented as virtual components executed by one or more virtual machines (VMs) implemented in one or more virtualization environments 800 hosted by one or more hardware nodes, such as a network node, a UE, a core network node, or a hardware computing device acting as a host. Furthermore, in embodiments in which the virtualized node does not require wireless connectivity (e.g., to a core network node or host), the node may be fully virtualized.

[0017] An application 802 (alternatively referred to as a software instance, a virtual appliance, a network function, a virtual node, a virtual network function, etc.) may be run in the virtualized environment 800 to implement some of the features, functionality, and / or advantages of some of the embodiments disclosed herein.

[0018] The hardware 804 includes processing circuitry, memory that stores software and / or instructions executable by the hardware processing circuitry, and / or other hardware devices described herein, such as network interfaces, input / output interfaces, etc. Software can be executed by the processing circuitry to instantiate one or more virtualization layers 806 (also referred to as hypervisors or virtual machine monitors (VMMs)), provide VMs 808A and 808B (one or more of which may be generally referred to as VMs 808), and / or perform any of the functions, features, and / or benefits described in connection with some embodiments described herein. The virtualization layer 806 may present a virtual operating platform that appears to be networking hardware to the VMs 808.

[0019] A VM 808 may include virtual processing, virtual memory, virtual networking or interfaces, and virtual storage and may be run by a corresponding virtualization layer 806. Different embodiments of the virtual appliance 802 instance may be implemented in one or more of the VMs 808, and the implementation may be done in different ways. Hardware virtualization is referred to in some contexts as network functions virtualization (NFV). NFV may be used to consolidate many network equipment types onto industry-standard high-volume server hardware, physical switches, and physical storage that may be located in data centers and customer premises equipment.

[0020] In the context of NFV, a VM 808 may be a software implementation of a physical machine that runs a program as if the program were running on a physical, non-virtualized machine. Each VM 808, and the portion of the hardware 804 on which it runs, even the hardware dedicated to that VM and / or shared by that VM with other VMs, forms a separate virtual network element. Still in the context of NFV, a virtual network function runs within one or more VMs 808 on the hardware 804 and is responsible for handling the specific network functions corresponding to the application 802.

[0021] The hardware 804 may be implemented in a standalone network node having general or specific components. The hardware 804 may implement some functions via virtualization. Alternatively, the hardware 804 may be part of a larger cluster of hardware (e.g., in a data center or CPE) where many hardware nodes cooperate and are managed via a management and orchestration 810 that oversees, among other things, the lifecycle management of the application 802. In some embodiments, the hardware 804 is coupled to one or more radio units, each including one or more transmitters and one or more receivers, which may be coupled to one or more antennas. The radio units may communicate directly with other hardware nodes via one or more appropriate network interfaces or may be used in combination with virtual components to provide radio capabilities to virtual nodes, such as radio access nodes or base stations. In some embodiments, the control system 812 may be used to provide some signaling that may alternatively be used for communication between the hardware nodes and the radio units.

[0022] 9 is a block diagram illustrating elements of an encoder 500 configured to encode audio frames in accordance with some embodiments of the inventive concepts. As shown, the encoder 500 may include a network interface circuit 905 (also referred to as a network interface) configured to provide communication with other devices / entities / functions, etc. The encoder 500 may also include a processor circuit 901 (also referred to as a processor) coupled to the network interface circuit 905 and a memory circuit 903 (also referred to as a memory) coupled to the processor circuit. The memory circuit 903 may include computer-readable program code that, when executed by the processor circuit 901, causes the processor circuit to perform operations in accordance with embodiments disclosed herein.

[0023] According to other embodiments, the processor circuitry 901 may be defined to include memory such that a separate memory circuit is not required. As discussed herein, operations of the encoder 500 may be performed by the processor 901 and / or the network interface 905. For example, the processor 901 may control the network interface 905 to send communications to the decoder 708 and / or receive communications from one or more other network nodes / entities / servers, such as other encoder nodes, depository servers, etc. Additionally, modules may be stored in the memory 903, and these modules may provide instructions to the processor 901 to perform their respective operations when their instructions are executed by the processor 901.

[0024] As mentioned above, various speech and music classifiers are used in the field of audio coding. However, these classifiers may not be able to distinguish between different classes in the music signal space. Many classifiers do not provide sufficient resolution to distinguish between classes, which is required in complex multi-mode codecs. In particular, spectral flatness and crest values ​​do not capture the energy spread or sparsity across the spectrum. Figure 2 shows two exemplary spectra, A and B. Spectrum A is a sparse spectrum suitable for a certain coding mode, while spectrum B is not suitable for this coding mode. However, the spectral flatness and crest measures cannot distinguish between these spectra because they both yield the same value. Figure 3 further shows the spectra of signals that are not suitable for coding by a certain coding mode.

[0025] Figure 4 shows an abstraction for creating a classifier to determine the class of a signal that controls subsequent mode decisions. These implementations address finding better classifiers for distinguishing between harmonic and noise-like music signals.

[0026] In some embodiments, the inventive concept is part of an audio encoding and decoding system. The audio encoder is a multi-mode audio encoder, and the method improves the selection of an appropriate encoding mode for a signal. To clarify that this is the encoding mode selected by the encoder, this will be referred to hereinafter as the encoding mode, but it will be understood by those skilled in the art that these terms may be used interchangeably. An input signal x(m,n), n=0,1,2,...L-1, is segmented into audio frames of length L, where m represents the frame index and n represents the sample index within the frame. The input signal is transformed into a frequency-domain representation such as the modified discrete cosine transform (MDCT) or the discrete Fourier transform (DFT). Other frequency-domain representations, such as filter banks, are also possible, but they should provide a suitably high frequency resolution for the target analysis range. In this embodiment, at least one of the audio coding modes operates in the MDCT domain. Therefore, it is beneficial to reuse the same transform for frequency-domain analysis. The MDCT is expressed by the following relationship: TIFF2025157307000004.tif13170, where X(m,k) represents the MDCT spectrum of frame m at frequency index k, and w a (n) is the analysis window. The frequency index k is sometimes called the frequency bin. Usually, audio frames are extracted with time overlap. The analysis window is chosen to give a good trade-off between, for example, algorithm delay, frequency resolution, and quantization noise shaping. If the frequency domain representation is based on the DFT, the spectrum can be written as It is specified according to TIFF2025157307000005.tif13170.

[0027] Note that in this case the frame length L may be different to provide a frame length suitable for DFT analysis.

[0028] Signal classification aims to select a coding mode that can best represent the input audio file. In particular, classification aims to identify signals with high peakiness and low energy density. The analysis can focus on critical frequency regions where the choice of coding method has a significant impact. Here, the frequency index k=k start …k end In some of these embodiments, the critical range is the upper half of the frequency spectrum that is coded with the bandwidth extension technique. start =320 and k end = 639, the operating sampling rate is 32 kHz, and the frame length is L = 640. The bandwidth extensions of the different coding modes have differences in their spectral signatures that are important for mode selection. More specifically, the aim is to identify signals that have a peaked structure in the high frequency range but do not have noisy components characterized by broad bands of high energy coefficients in the spectrum. Examples of desired and undesired signals can be found in Figure 3. This is done by comparing the features crest(m) and the novel feature crest mod Implemented by parsing (m).

[0029] 4 illustrates the operations performed by an encoder in some embodiments of the inventive concept. Referring to FIG. 4, in step 410, the magnitude or absolute spectrum A of the critical region is calculated. i (m) is obtained by the encoder 500. The magnitude of the critical region or absolute spectrum A i (m) is TIFF2025157307000006.tif30170, where M=k end -k start +1 is the number of bins or frequency indices in the critical band. In step 420, the crest value for frame m is calculated as is derived by the encoder 500 according to TIFF2025157307000007.tif15170, where crest(m) gives the measure of peakiness of frame m. The complementary peakiness measure t(m) is also given by TIFF2025157307000008.tif29170, where A thr is the appropriate value A thr = 0.1 or a relative threshold that can be in the range [0.01, 0.4]. In step 430, the noise band detection measure is TIFF2025157307000009.tif18170, where movmean(A i (m),W) is the absolute spectrum A using a window size of W i (m) is the moving average of the window size. A suitable value for the window size may be W=21, or any odd number in the range [7,31].

[0030] In one embodiment of the inventive concept, movemean(A i (m),W) are TIFF2025157307000010.tif15170a=max(0,i-(W-1) / 2) b=min(M-1,i+(W-1) / 2) It is prescribed in accordance with Here, the absolute spectrum A i The average at the edge of (m) is A i It is formed using only values ​​in the range (m).

[0031] Alternatively, the specification may be written in a recursive form that requires fewer computational operations. TIFF2025157307000011.tif63170

[0032] In another embodiment, movmean(A i The definition of (m), W) can be assumed to have an absolute spectrum of 0 outside the range i = 0...M-1, which means that TIFF2025157307000012.tif15170a=max(0,i-(W-1) / 2) b=min(M-1,i+(W-1) / 2) Simplify the numerator of the equation according to

[0033] movmean(A i Note that the definition of movmean(A(m),W) assumes that the window length W is odd and extends an equal number of samples in both the positive and negative directions from the current frequency bin i. It is possible to use an even window length W with appropriate adaptations to the above formula. For example, i (m), W) when using only backward shifted windows. TIFF2025157307000013.tif15170a=max(0,iW / 2) b=min(M-1,i+W / 2-1) Or if you want to calculate the average between the rear and front alignment of the window, TIFF2025157307000014.tif17170a=max(0,iW / 2) b=min(M-1,i+W / 2-1) c=max(0,iW / 2) d=min(M-1,i+W / 2-1) In general, the moving average operation is performed using a moving average filter of the form It can be implemented using TIFF2025157307000015.tif16170, where w j are the filter coefficients.

[0034] crest mod (m) gives a measure of the local concentration of energy and indicates the noise band in the spectrum. To stabilize the decision, crest(m) and crest mod (m) may be low-pass filtered by the encoder 500. For example, crest LP (m)=(1-α)·crest(m)+α·crest LP (m-1) crest mod,LP (m)=(1-β)·crest mod (m)+β·crest mod,LP (m-1) where α and β are filter coefficients. A suitable value for α may be α=0.97 or in the range [0.5, 1), and equivalently, a suitable value for β may be β=0.97 or in the range [0.5, 1).

[0035] Coding modes targeting peaked spectra without noisy components are based on the following conditions: TIFF2025157307000016.tif10170 is satisfied, and in the formula, crest thr , crest mod,thr and t thr are the decision thresholds. Appropriate values ​​for these thresholds are thr =7, crest mod,thr =2.128 and t thr = 220. More generally, suitable values ​​are in the range crest thr ∈[3,12], crest mod,thr ∈[1,4] and t thr ∈[150,300]. These conditions can be decomposed into crest mod,LP (m)>crest mod,thr ensures that the coding mode is disabled for noisy components, while crest LP (m)>crest thr and t(m)>t thr limits the impact of this decision to signals with peaked spectra.

[0036] Alternatively, the condition on t(m) can be omitted and the test is The file will be TIFF2025157307000017.tif12170.

[0037] In another embodiment of the inventive concept, the determination is made by LP (m) is high while crest LP,mod If (m) is low, TIFF2025157307000018.tif12170, where the threshold crest thr2 and crest mod,thr2 crest thr and crest mod,thr It may be the same as:

[0038] In step 440, a coding mode, including at least the decision Harmonic_decision(m), is selected by the encoder 500. Finally, in step 450, the encoder 500 performs encoding using the selected coding mode.

[0039] Figure 5 is a schematic diagram of a multi-mode audio codec that uses mode decision logic based on an input signal. Referring to Figure 5, an absolute value calculator 510 receives the input audio and converts it to a frequency-domain representation, such as a modified discrete cosine transform (MDCT). If the signal has already been converted to the frequency domain used in the multi-mode encoder, the frequency-domain representation may be reused in this step. The absolute value calculator 510 then determines the absolute value (e.g., magnitude) of the MDCT. The absolute magnitude is used by a peakiness measure 520 to determine a peakiness measure of the MDCT. An additional peakiness measure 530 may also be derived.

[0040] The noise detection measure 540 receives the absolute value of the MDCT and determines the noisiness of the input audio signal. The mode enable decision 550 receives the peakiness measure and the noise detection measure and determines whether to enable mode selection. For example, if two coding modes exist, the mode enable decision 550 determines which of the two coding modes can be used.

[0041] The mode selector 560 determines which encoding mode to use and indicates which mode to use to the multi-mode encoder 580. The multi-mode encoder 580 encodes the input audio signal to generate coded audio 590. The determined mode decision 570, combined with the coded audio 590, is transmitted or stored for use by a multi-mode decoder.

[0042] FIG. 6A shows an example of which coding mode or group of coding modes is used. In response to a harmonic condition being true (e.g., harmonic_decision(m) being true), coding mode C is determined to be used. In response to a harmonic condition being false (e.g., harmonic_decision(m) being false), coding mode D or coding mode E is determined to be used. FIG. 6B shows an example in which both branches have groups of coding modes. In FIG. 6B, in response to a harmonic condition being true (e.g., harmonic_decision(m) being true), coding mode C or coding mode F is determined to be used. In response to a harmonic condition being false (e.g., harmonic_decision(m) being false), coding mode D or coding mode E is determined to be used.

[0043] The operation of encoder 500 (implemented using the block diagram structure of FIG. 9) will now be described with reference to the flowchart of FIG. 10, in accordance with some embodiments of the inventive concepts. For example, modules may be stored in memory 903 of FIG. 9, and these modules may provide instructions such that, when the instructions of the modules are executed by respective communications device processing circuitry 901, the processing circuitry 901 performs the respective operations of the flowchart.

[0044] 10, in block 1001, the processing circuit 901 derives a frequency spectrum of an input audio signal. In some embodiments of the inventive concept, the processing circuit 901 derives the frequency spectrum by segmenting the input audio signal x(m,n), n=0, 1, 2, ...L-1, into audio frames of length L, where m represents a frame index and n represents a sample index within the frame, and the input audio signal is Convert to frequency domain representation according to TIFF2025157307000019.tif13170.

[0045] In block 1003, the processing circuit 901 obtains the magnitude of the critical frequency region of the frequency spectrum. The critical frequency region is the region of the frequency index k=k start …k end where the critical frequency range is the upper half of X(m,k). In some embodiments, the critical frequency range is k start =320 and k end = 639, the operating sampling rate is 32 kHz, and the frame length is L = 640.

[0046] In some embodiments of the inventive concept, the processing circuitry 901 includes: Obtain the magnitude of the critical frequency region according to TIFF2025157307000020.tif30170, where M = k end -k start +1 is the number of bins in the critical band associated with the critical frequency region.

[0047] In block 1005, processing circuitry 901 obtains a peakiness measure. In some embodiments of the inventive concept, processing circuitry 901: Obtain the peakiness measure according to TIFF2025157307000021.tif18170, where crest(m) gives the peakiness measure of frame m.

[0048] In another embodiment of the inventive concept, the processing circuitry 901 includes: Obtain the peakness measure of the frame according to TIFF2025157307000022.tif29170, where A thr is the relative threshold.

[0049] In some embodiments, A thr = 0.1. thr is in the range [0.01,0.4].

[0050] In block 1007, processing circuitry 901 obtains a noise band detection measure. In some embodiments of the inventive concept, processing circuitry 901: Obtain the noise band detection measure according to TIFF2025157307000023.tif18170, where crest mod (m) is the noise band detection measure, and movmean(A i (m),W) is the absolute spectrum A using a window size of W i (m) is the moving average.

[0051] In some embodiments, the processing circuitry 901 includes: TIFF2025157307000024.tif15170a=max(0,i-(W-1) / 2) b=min(M-1,i+(W-1) / 2) According to movmean(A i (m), W) are determined.

[0052] In block 1009, the processing circuit determines which of two coding modes or groups of coding modes to use based on at least the peakiness measure and the noise band detection measure. For example, a sparse spectrum may be suitable for a first coding mode or set of coding modes but not for a second coding mode or set of coding modes.

[0053] In some embodiments of the inventive concept, the processing circuitry 901 determines which of the two coding modes or groups of coding modes to use based on at least a peakiness measure, a noise band detection measure, and by determining which of the two coding modes to use based on if Harmonic_decision(m) is true, where Harmonic_decision(m) is TIFF2025157307000025.tif10170, where crest thr , crest mod,thr and t thr is the decision threshold, and crest LP (m) is the low-pass filtered crest(m), and crest mod,LP (m) is the low-pass filtered crest mod (m).

[0054] The processing circuit 901 crest LP (m)=(1-α)·crest(m)+α·crest LP (m-1) crest mod,LP (m)=(1-β)·crest mod (m)+β·crest mod,LP (m-1) The low-pass filtered crest(m) and the low-pass filtered crest mod Harmonic_decision(m) can be determined, where α and β are filter coefficients. In some embodiments, α is in the range [0.5,1) and β is in the range [0.5,1). In other embodiments, Harmonic_decision(m) can be determined by Determined according to TIFF2025157307000026.tif12170.

[0055] In another embodiment, the processing circuit 901 determines which of the two coding modes to use based on at least the peakiness measure, the noise band detection measure, and by determining which of the two coding modes to use based on if Harmonic_enabled(m) is true, where Harmonic_enabled(m) is TIFF2025157307000027.tif12170, where crest thr2 and crest mod,thr2 is the decision threshold.

[0056] Thus, processing circuitry 901 determines the coding mode based on at least the peakiness measure, the noise band detection measure, and Harmonic_decision(m).

[0057] 11, in some embodiments of the inventive concept, processing circuit 901 determines to use a first encoding mode of two encoding modes or a first encoding mode from a group of encoding modes in response to Harmonic_decision(m) being TRUE in block 1101. In block 1103, processing circuit 901 determines to use a second encoding mode of two encoding modes or a second encoding mode from the group of encoding modes in response to Harmonic_decision(m) being FALSE.

[0058] Returning to FIG. 10, in block 1011, the processing circuit 901 encodes the input audio signal based on the encoding mode determined to be used.

[0059] In other embodiments of the inventive concepts, the inventive concepts described herein can be used to determine whether an input audio signal has high peakiness and low energy density. Figure 12 illustrates one embodiment of determining whether an input audio signal has high peakiness and low energy density.

[0060] 12, the processing circuit 901 derives the frequency spectrum of the input audio signal in block 1201. Block 1201 is similar to block 1001 described above.

[0061] In block 1203, the processing circuit 901 obtains the magnitude of the critical frequency region of the frequency spectrum. Block 1203 is similar to block 1003 described above.

[0062] In block 1205, processing circuit 901 obtains a peakiness measure. Block 1205 is similar to block 1005 described above.

[0063] In block 1207, the processing circuit 901 obtains a noise band detection measure. Block 1207 is similar to block 1007 described above.

[0064] In block 1209, processing circuitry 901 determines harmonic conditions based on at least the peakiness measure and the noise band detection measure.

[0065] In block 1211, processing circuitry 901 outputs an indication of whether the harmonic condition is true or false.

[0066] In some embodiments, the processing circuit 901 determines whether the low-pass filtered crest(m) is greater than the crest threshold and whether the low-pass filtered crest(m) is greater than the crest threshold. mod (m) is crest mod In response to being greater than the threshold, it is determined that the harmonic condition is true, where crest(m) is a measure of peakiness of frame m, and crest mod (m) is a measure of the local concentration of energy.

[0067] The processing circuitry 901, in some embodiments of the inventive concept, crest(m) and crest according to TIFF2025157307000028.tif34170 mod(m) is determined, where A i (m) is the magnitude of the modified discrete cosine transform (MDCT) of the audio signal at frame m, M is the number of frequency indices in the critical region, and movmean(A i (m), W) is the A using window size W i (m) is the moving average.

[0068] The processing circuit 901 According to TIFF2025157307000029.tif30170 A i (m), where X(m,k) represents the MDCT spectrum of frame m at frequency index k, and M=k end -k start +1 and k end and k startは is the frequency index of the critical region of X(m,k).

[0069] movmean(A i Various embodiments for determining (m), W) are described above.

[0070] The processing circuit 901 Determine X(m,k) according to TIFF2025157307000030.tif15170, where L is the frame length of frame m.

[0071] While the computing devices (e.g., UEs, network nodes, hosts) described herein may include the illustrated combination of hardware components, other embodiments may include computing devices having different combinations of components. It should be understood that these computing devices comprise any suitable combination of hardware and / or software necessary to perform the tasks, features, functions, and methods disclosed herein. The determining, calculating, obtaining, or similar operations described herein may be performed by processing circuitry, which may, for example, transform obtained information into other information, compare the obtained or converted information with information stored in the network node, and / or perform one or more operations based on the obtained or converted information, and make a decision as a result of said processing. Furthermore, while components are shown as a single box disposed within a larger box or a single box nested within multiple boxes, in reality, the computing device may comprise multiple different physical components that make up a single illustrated component, and functionality may be divided among the separate components. For example, a communication interface may be configured to include any of the components described herein, and / or the functionality of a component may be divided between the processing circuitry and the communication interface. In another example, non-computationally intensive functions of any such components may be implemented in software or firmware, and computationally intensive functions may be implemented in hardware.

[0072] In particular embodiments, some or all of the functionality described herein may be provided by a processing circuit executing instructions stored in a memory, which in particular embodiments may be a computer program product in the form of a non-transitory computer-readable storage medium. In alternative embodiments, some or all of the functionality may be provided by a processing circuit without executing instructions stored on a separate or distinct device-readable storage medium, such as in a hardwired manner. In any of these particular embodiments, the processing circuit can be configured to perform the above-described functionality, regardless of whether it executes instructions stored on a non-transitory computer-readable storage medium. Benefits provided by such functionality are not limited to the processing circuit alone or other components of the computing device, but are enjoyed by the computing device as a whole and / or by end users and wireless networks in general.

[0073] Further definitions and embodiments are described below.

[0074] In the above description of various embodiments of the inventive concept, it should be understood that the terminology used herein is merely for the purpose of describing specific embodiments and is not intended to limit the inventive concept. Unless otherwise specified, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art to which the inventive concept belongs. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning in accordance with the meaning of those terms in the context of this specification and related art, and should not be interpreted in an idealized or overly formal sense unless expressly so defined herein.

[0075] When an element is referred to as being "connected," "coupled," "responsive," or variations thereof, to another element, the element may be directly connected, coupled, or responsive to the other element, or intervening elements may be present. In contrast, when an element is referred to as being "directly connected," "directly coupled," "directly responsive," or variations thereof, to another element, there are no intervening elements present. Like numbers refer to like elements throughout. Furthermore, as used herein, "coupled," "connected," "responsive," or variations thereof may include wirelessly coupled, wirelessly connected, or wirelessly responsive. As used herein, the singular forms "a," "an," and "the" are intended to include the plural unless the context clearly dictates otherwise. For the sake of brevity and / or clarity, well-known features or structures may not be described in detail. The term "and / or" (abbreviated " / ") includes any and all combinations of one or more of the associated listed items.

[0076] Although terms such as first, second, third, etc. may be used herein to describe various elements / operations, it will be understood that these elements / operations should not be limited by these terms. These terms are used only to distinguish one element / operation from another. Thus, a first element / operation in some embodiments may be referred to as a second element / operation in other embodiments without departing from the teachings of the inventive concept. The same reference numbers or characters denote the same or similar elements throughout this specification.

[0077] As used herein, the terms "comprise," "comprising," "comprises," "include," "including," "includes," "have," "has," "having," or variations thereof, are open-ended and refer to the inclusion of one or more stated features, integers, elements, steps, components, or functions, but do not exclude the presence or addition of one or more other features, integers, elements, steps, components, functions, or groups thereof. Furthermore, as used herein, the common abbreviation "eg," derived from the Latin phrase "exempli gratia," may be used to introduce or specifically name one or more general examples of the aforementioned items, without limiting such items. The common abbreviation "ie," derived from the Latin phrase "id est," may be used to specifically name a particular item from a more general statement.

[0078] Exemplary embodiments are described herein with reference to block diagrams and / or flowchart illustrations of computer-implemented methods, apparatus (systems and / or devices), and / or computer program products. It should be understood that blocks of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by computer program instructions executed by one or more computer circuits. These computer program instructions can be provided to processor circuits of general-purpose computer circuits, special-purpose computer circuits, and / or other programmable data processing circuits to create a machine whereby the instructions executed via the processor of the computer and / or other programmable data processing apparatus transform and control transistors, values ​​stored in memory locations, and other hardware components within such circuits to implement the functions / acts specified in one or more blocks of the block diagrams and / or flowcharts, and thereby create means (functions) and / or structures for implementing the functions / acts specified in the block diagrams and / or flowchart blocks.

[0079] The computer program instructions may also be stored on a tangible computer-readable medium that can direct a computer or other programmable data processing apparatus to function in a particular manner, whereby the instructions stored on the computer-readable medium produce an article of manufacture comprising instructions that implement the functions / acts specified in one or more blocks of the block diagrams and / or flowcharts. Thus, embodiments of the inventive concepts may be embodied in the form of hardware and / or software (including firmware, resident software, microcode, etc.) running on a processor, such as a digital signal processor, which may be collectively referred to as a "circuit," "module," or variations thereof.

[0080] It should also be noted that in some alternative implementations, the functions / acts noted in the blocks may occur in an order other than that noted in the flowcharts. For example, depending on the functionality / acts involved, two blocks shown in succession may in fact be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order. Furthermore, the functionality of a given block in the flowcharts and / or block diagrams may be separated into multiple blocks, and / or the functionality of two or more blocks in the flowcharts and / or block diagrams may be at least partially integrated. Finally, other blocks may be added / inserted between the illustrated blocks, and / or blocks / acts may be omitted without departing from the scope of the inventive concepts. Furthermore, while some of the figures include arrows on communication paths to indicate a primary direction of communication, it should be understood that communication may occur in the opposite direction to that of the illustrated arrows.

[0081] Many variations and modifications can be made to the embodiments without substantially departing from the principles of the inventive concept. All such variations and modifications are intended to be included herein within the scope of the inventive concept. Accordingly, the subject matter disclosed above should be considered illustrative and not limiting, and the example embodiments are intended to cover all such modifications, extensions, and other embodiments that fall within the spirit and scope of the inventive concept. Therefore, to the fullest extent permitted by law, the scope of the inventive concept should be determined by the broadest permissible interpretation of this disclosure, including example embodiments and their equivalents, and should not be limited or restricted by the above detailed description.

[0082] Embodiment 1. 1. A method in an encoder for determining which of two coding modes or groups of coding modes to use, comprising: Deriving (901) a frequency spectrum of an input audio signal; Obtaining (903) the magnitude of a critical frequency region of the frequency spectrum; Obtaining a peakiness measure for the frame (905); Obtaining a noise band detection measure (907); determining (909) which of two coding modes or groups of coding modes to use based on at least a peakiness measure and a noise band detection measure; encoding (911) the input audio signal based on the encoding mode determined to be used; A method comprising: 2. Encoding the input audio signal based on the encoding mode determined to be used; selecting, in response to determining to use the group of encoding modes, one encoding mode from the group of encoding modes to use for encoding the input audio signal; 2. The method of embodiment 1, comprising: 3. A method according to any one of embodiments 1 to 2, wherein deriving the frequency spectrum comprises deriving a frequency spectrum X(m,k), where X(m,k) represents the frequency spectrum of frame m at frequency index k. 4. Deriving the frequency spectrum Segmenting an input audio signal x(m,n), n=0, 1, 2, ... L-1, into audio frames of length L, where m represents a frame index and n represents a sample index within the frame; Transforming an input audio signal in a frequency domain representation according to TIFF2025157307000031.tif17170, where X(m,k) represents the frequency spectrum of the Modified Discrete Cosine Transform (MDCT) of frame m at frequency index k, and w a transforming the input audio signal, where (n) is an analysis window; Frequency index k=k start …k endand obtaining a magnitude spectrum of X(m,k) defined by: 4. The method of any one of embodiments 1 to 3, comprising: 5. The critical frequency range is k start =320 and k end 5. The method of any one of embodiments 3 to 4, wherein the input sampling rate is 32 kHz and the frame length is L=640, corresponding to L=639. 6. Obtaining the magnitude of the critical frequency region and obtaining the magnitude of the critical frequency region according to TIFF2025157307000032.tif35170, where M=k end -k start 6. A method according to any one of embodiments 3 to 5, wherein +1 is the number of frequency indices of the critical band associated with the critical frequency region. 7. Obtaining a peakiness scale 7. The method of embodiment 6, comprising obtaining a peakiness measure according to TIFF2025157307000033.tif18170, where crest(m) gives the peakiness measure for frame m. 8. Obtaining a peakiness scale TIFF2025157307000034.tif36170, wherein A thr 7. The method of embodiment 6, wherein: 9.A thr 9. The method of embodiment 8, wherein: 10.A thr 9. The method of embodiment 8, wherein: 11. Obtaining a noise band detection measure obtaining a noise band detection measure according to TIFF2025157307000035.tif18170, wherein crest mod (m) is the noise band detection measure, and movmean(A i (m),W) is the absolute spectrum A using a window size of W i11. The method of any preceding embodiment, wherein (m) is a moving average of 12.movmean(A i (m),W) TIFF2025157307000036.tif14170a=max(0,i-(W-1) / 2) b=min(M-1,i+(W-1) / 2) 12. The method of embodiment 11, wherein the method is determined according to 13. crest LP (m)=(1-α)·crest(m)+α·crest LP (m-1) crest mod,LP (m)=(1-β)·crest mod (m)+β·crest mod,LP (m-1) According to crest(m) and crest mod 13. The method of any one of embodiments 7 to 12, further comprising low-pass filtering (m), wherein α and β are filter coefficients. 14. The method of embodiment 13, wherein α is in the range [0.5, 1) and β is in the range [0.5, 1). 15. Determining which of the two coding modes to use based on at least the peakness measure and the noise band detection measure includes determining which of the two coding modes to use based on if Harmonic_decision(m) is true, and Harmonic_decision(m) is TIFF2025157307000037.tif16170, where crest thr , crest mod,thr and t thr 15. The embodiment of any of claims 1 to 14, wherein is the decision threshold. 16. Determining which of the two coding modes to use based on at least the peakness measure and the noise band detection measure includes determining which of the two coding modes to use based on if Harmonic_decision(m) is true, and Harmonic_decision(m) is TIFF2025157307000038.tif16170, where crest thr and crest mod,thr 15. The method of any preceding embodiment, wherein: 17. Determining the coding mode based on at least the peakiness measure and the noise band detection measure includes determining which of two coding modes to use if Harmonic_decision(m) is true, and TIFF2025157307000039.tif16170, where crest thr2 and crest mod,thr2 15. The method of any preceding embodiment, wherein: 18. A method according to any one of embodiments 15 to 17, wherein determining the coding mode based on at least the peakiness measure and the noise band detection measure comprises determining the coding mode based on at least the peakiness measure, the noise band detection measure, and Harmonic_decision(m). 19. Determining the coding mode based on at least a peakiness measure, a noise band detection measure, and Harmonic_disabled(m), determining (1101) to use a first encoding mode of the two encoding modes in response to Harmonic_decision(m) being TRUE; determining (1103) to use a second encoding mode of the two encoding modes in response to Harmonic_decision(m) being FALSE; 19. The method of embodiment 18, comprising: 20. A method in an encoder for determining whether an input audio signal has high peakiness and low energy density, comprising: Deriving (1201) a frequency spectrum of an input audio signal; Obtaining (1203) the magnitude of a critical frequency region of the frequency spectrum; Obtaining a peakiness measure (1205); Obtaining a noise band detection measure (1207); determining (1209) harmonic conditions based on at least a peakiness measure and a noise band detection measure; Outputting an indication of whether the harmonic condition is true or false (1211); A method comprising: 21. The low-pass filtered crest(m) is greater than the crest threshold, and the low-pass filtered crest mod (m) is crest mod and determining that a harmonic condition is true in response to the peak value being greater than the threshold value, wherein crest(m) is a measure of peakiness of frame m, and crest mod 21. The method of embodiment 20, wherein (m) is a measure of the local concentration of energy. twenty two. crest(m) and crest according to TIFF2025157307000040.tif38170 mod (m), wherein A i where (m) is the magnitude of the frequency spectrum of the audio signal at frame m, M is the number of frequency indices in the critical region, and movmean(A i (m), W) is A using window size W i 22. The method of embodiment 21, wherein (m) is a moving average of (m). twenty three. According to TIFF2025157307000041.tif35170 A i (m), where X(m,k) represents the frequency spectrum of frame m at frequency index k, and M=k end-k start +1 and k end and k start 23. The method of embodiment 22, wherein x is the frequency index of the critical region of X(m,k). twenty four. 24. The method of embodiment 23, further comprising determining X(m,k) according to TIFF2025157307000042.tif17170, where L is the frame length of frame m. 25. A processing circuit (901); a memory (905) coupled to the processing circuit, the memory (905) including instructions that, when executed by the processing circuit, cause the communication device to perform the operations of any one of embodiments 1 to 24; An encoder device (500) comprising: 26. An encoder device (500) adapted to perform according to any of embodiments 1 to 24. 27. A computer program comprising program code to be executed by a processing circuit (901) of an encoder device (500), thereby causing the encoder device (500) to perform the operations described in any of embodiments 1 to 24 by executing the program code. 28. A computer program product comprising a non-transitory storage medium containing program code to be executed by a processing circuit (901) of an encoder device (500), thereby causing the encoder device (500) to perform the operations described in any of embodiments 1 to 24 by executing the program code.

[0083] Explanations of various abbreviations / acronyms used in this disclosure are provided below. Abbreviation Explanation MDCT Modified Discrete Cosine Transform DFT Discrete Fourier Transform

Claims

1. 1. A method in an encoder for determining which of two coding modes or groups of coding modes to use, comprising: Deriving (901) a frequency spectrum of an input audio signal; Obtaining the magnitude of the frequency spectrum in the critical frequency region (903); Obtaining a peakiness measure (905); Obtaining a noise band detection measure (907); determining (909) which of the two coding modes or groups of coding modes to use based on at least the peakiness measure and the noise band detection measure; encoding (911) the input audio signal based on the encoding mode determined to be used; A method comprising:

2. encoding the input audio signal based on the encoding mode determined to be used; selecting, in response to determining to use a group of encoding modes, one encoding mode from the group of encoding modes to use for encoding the input audio signal. The method of claim 1 , comprising:

3. 3. The method of claim 1, wherein deriving the frequency spectrum comprises deriving a frequency spectrum X(m,k), where X(m,k) represents the frequency spectrum for frame m at frequency index k.

4. deriving the frequency spectrum segmenting the input audio signal x(m,n), n=0, 1, 2, ... L-1, into audio frames of length L, where m represents a frame index and n represents a sample index within a frame; where X(m,k) represents the frequency spectrum of a Modified Discrete Cosine Transform (MDCT) of frame m at frequency index k, and w a (n) transforming the input audio signal, where n is an analysis window; Frequency index k=k start …k end and obtaining the magnitude spectrum of X(m,k) defined by:

4. The method of claim 1, comprising:

5. The critical frequency range is k start = 320 and k end 5. The method of claim 3, wherein the input sampling rate is 32 kHz and the frame length is L=640, corresponding to L=639.

6. Obtaining the magnitude of the frequency spectrum of the critical frequency region obtaining the magnitude of the frequency spectrum of the critical frequency region according to end -k start 6. The method of claim 3, wherein +1 is the number of frequency indices of the critical band associated with the critical frequency region.

7. obtaining the peakiness measure, 7. The method of claim 6, comprising obtaining the peakiness measure according to: where crest(m) gives the peakiness measure for frame m.

8. obtaining the peakiness measure, obtaining the peakiness measure according to thr The method of claim 6 , wherein: is a relative threshold.

9. A thr 9. The method of claim 8, wherein: = 0.

1.

10. A thr The method of claim 8 , wherein is in the range [0.01, 0.4].

11. obtaining the noise band detection measure, obtaining the noise band detection measure according to mod (m) is the noise band detection measure, and movmean(A i (m), W) is the absolute spectrum A using a window size of W i 11. The method of claim 1, wherein the (m) is a moving average of (m).

12. movmean (A i (m), W) is a=max(0,i-(W-1) / 2) b=min(M-1, i+(W-1) / 2) The method of claim 11 , wherein the value is determined according to

13. crest LP (m)=(1-α)・crest(m)+α・crest LP (m-1) crest mod,LP (m)=(1-β)・crest mod (m)+β・crest mod,LP (m-1) According to crest(m) and crest mod 13. The method of claim 7, further comprising low-pass filtering (m), where α and β are filter coefficients.

14. 14. The method of claim 13, wherein α is in the range [0.5, 1) and β is in the range [0.5, 1).

15. determining which of two coding modes or groups of coding modes to use based on at least the peakness measure and the noise band detection measure comprises determining the one of the two coding modes or groups of coding modes if Harmonic_decision(m) is true, wherein Harmonic_decision(m) is where crest is determined according to the formula thr , crest mod,thr and t thr 15. The method of claim 1, wherein is a decision threshold.

16. determining which of two coding modes or groups of coding modes to use based on at least the peakness measure and the noise band detection measure comprises determining the one of the two coding modes or groups of coding modes if Harmonic_decision(m) is true, wherein Harmonic_decision(m) is where crest is determined according to the formula thr and crest mod,thr 15. The method of claim 1, wherein is a decision threshold.

17. determining the coding mode based on at least the peakiness measure and the noise band detection measure includes enabling the coding mode determination if Harmonic_decision(m) is true, and where crest is determined according to the formula thr2 and crest mod,thr2 15. The method of claim 1, wherein is a decision threshold.

18. 18. The method of claim 15, wherein determining the coding mode based on at least the peakiness measure and the noise band detection measure comprises determining the coding mode based on at least the Harmonic_decision(m).

19. determining the coding mode based on the Harmonic_decision(m); In response to the Harmonic_decision(m) being TRUE, deciding to use a first encoding mode of the two encoding modes (1101); determining (1103) to use a second encoding mode of the two encoding modes in response to the Harmonic_disabled(m) being FALSE; 20. The method of claim 18, comprising:

20. 1. A method in an encoder for determining whether an input audio signal has high peakiness and low energy density, comprising: Deriving (1201) a frequency spectrum of an input audio signal; Obtaining (1203) the magnitude of a critical frequency region of the frequency spectrum; Obtaining a peakiness measure for the frame (1205); Obtaining a noise band detection measure (1207); determining 1209 harmonic conditions based on at least the peakiness measure and the noise band detection measure; transmitting (1211) an indication of whether the harmonic condition is true or false; A method comprising:

21. If the low-pass filtered crest(m) is greater than the crest threshold and the low-pass filtered crest mod (m) is crest mod and determining that the harmonic condition is true in response to the peakiness measure being greater than a threshold value, wherein crest(m) is the peakiness measure for frame m, and crest mod 21. The method of claim 20, further comprising: (m) being a measure of the local concentration of energy.

22. According to crest(m) and crest mod (m), wherein A i (m) is the magnitude of the frequency spectrum of the audio signal at frame m, M is the number of frequency indices in the critical region, and movmean(A i (m), W) is the A using window size W i 22. The method of claim 21, wherein the (m) is a moving average of According to claim 23, i (m), where X(m,k) represents the frequency spectrum of frame m at frequency index k, and M=k end -k start +1, and k end and k start 23. The method of claim 22, wherein X(m,k) is the frequency index of the critical region of X(m,k).

24. The method of claim 23, further comprising determining X(m, k) according to: where L is the frame length of frame m.

25. A processing circuit (901); a memory (905) coupled to the processing circuitry, the memory (905) containing instructions that, when executed by the processing circuitry, cause the communications device to perform the operations of any one of claims 1 to 24; An encoder device (500) comprising:

26. An encoder device (500) adapted to perform the method according to at least one of claims 1 to 24.

27. 25. A computer program comprising program code to be executed by a processing circuit (901) of an encoder device (500), whereby executing said program code causes said encoder device (500) to perform the operations of any one of claims 1 to 24.

28. 25. A computer program product comprising a non-transitory storage medium containing program code to be executed by a processing circuit (901) of an encoder device (500), whereby executing the program code causes the encoder device (500) to perform the operations of any one of claims 1 to 24.