Spectrum classifier for audio coding mode selection
A new metric for audio encoding in the frequency domain addresses the challenge of distinguishing harmonic and noise-like music segments by analyzing peakiness and noise bands, enhancing encoding mode selection for audio signals.
Patent Information
- Application Number
- JP2023580679
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-06-29
- Publication Date
- 2025-07-17
- Estimated Expiration
- 2041-06-29
AI Technical Summary
Existing audio music classifiers struggle to distinguish different classes within the space of music signals, lacking sufficient resolution for complex multi-mode codecs, particularly in distinguishing harmonic and noise-like music segments.
A new metric based on peakiness and noise band detection measures is calculated directly on the frequency domain coefficients to determine the appropriate encoding mode, analyzing the critical frequency region of the spectrum to identify signals with high peakiness and low energy concentration.
Enhances the selection of encoding modes by accurately distinguishing between harmonic and noise-like music signals, improving the encoding process for audio signals.
Smart Images

Figure 0007710054000051 
Figure 0007710054000052 
Figure 0007710054000053
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to communications, and more particularly, to a communication method for supporting wireless communications and related devices and nodes.
Background Art
[0002] Modern audio codecs consist of multiple compression methods optimized for signals with different characteristics. Usually, voice-like signals are processed by a codec operating in the time domain, and music signals are processed by a codec operating in the transform domain. An encoding method aimed at handling both voice signals and music signals requires a mechanism to recognize the input signal (voice / music classifier) and switch the appropriate codec mode. A schematic diagram of a multi-mode audio codec using mode determination logic based on the input signal is shown in FIG. 1.
[0003] Similarly, among classes of music signals, noise-like music signals and high-pitched music signals can be distinguished, and a classifier and an optimal encoding method can be constructed for each of these groups. In particular, the identification of signals having a sparse and peaky structure is very interesting because transform domain codecs are suitable for handling these types of signals. Crest C or spectral flatness f determined according to TIFF0007710054000001.tif17170 There are several known signal metrics aimed at identifying peak signal structures such as TIFF0007710054000002.tif18170.
[0004] A high spectral flatness or crest may indicate that an encoding mode suitable for such a spectrum can be selected.
Summary of the Invention
[0005] Currently, there are specific problems. In the field of audio encoding, various audio music classifiers are used. However, these audio music classifiers may not be able to distinguish different classes within the space of music signals. Many audio music classifiers do not provide sufficient resolution to distinguish the classes required by complex multi-mode codecs.
[0006] The problem of distinguishing harmonic and noise-like music segments is solved by a new metric calculated directly on the frequency domain coefficients. This metric is based on the peakiness measure of the spectrum and the measure of the local concentration of energy indicating the noisy components of the spectrum.
[0007] Various embodiments of the concept of the present invention to address these problems involve analysis in the frequency domain at the critical bands of the spectrum. The analysis includes at least the peakiness measure, and various embodiments provide additional measures that give an indication of the noisy bands within the spectrum. Based on these measures, a determination is formed as to whether to use at least one encoding mode that targets signals with strong peaks while avoiding signals with noisy bands.
[0008] According to some embodiments of the concept of the present invention, a method in an encoder is provided for determining which of two encoding modes or groups of encoding modes to use. The method includes deriving the frequency spectrum of the input audio signal. The method further includes obtaining the size of the critical frequency range of the frequency spectrum. The method further includes obtaining a peakiness measure. The method further includes obtaining a noise band detection measure. The method further includes determining which of two encoding modes or groups of encoding modes to use based on at least the peakiness measure and the noise band detection measure. The method further includes encoding the input audio signal based on the encoding mode determined to be used.
[0009] Similar encoders, computer programs, and computer program products are provided.
[0010] According to other embodiments of the concepts of the present invention, a method in an encoder for determining whether an input audio signal has a high peakiness and a low energy concentration is provided. The method includes deriving a frequency spectrum of the input audio signal. The method further includes obtaining a magnitude of a critical frequency range of the frequency spectrum. The method further includes obtaining a peakiness measure. The method further includes obtaining a noise band detection measure. The method further includes determining a harmonic condition based on at least the peakiness measure and the noise band detection measure. The method includes outputting an indication of whether the harmonic condition is true or false.
[0011] Similar encoders, computer programs, and computer program products are provided.
[0012] Included to provide a further understanding of the disclosure, the accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate certain non-limiting specific embodiments of the concepts of the present invention.
Brief Description of the Drawings
[0013]
FIG. 1
FIG. 2
FIG. 3
FIG. 4
FIG. 5
FIG. 6A
FIG. 6B
FIG. 7
FIG. 8
FIG. 9
FIG. 10
FIG. 11
FIG. 12
[0014] Next, some of the embodiments contemplated herein will be more fully described with reference to the accompanying drawings. The embodiments are provided as examples to convey the scope of the subject matter to those skilled in the art, and examples of embodiments of the concepts of the present invention are shown. However, the concepts of the present invention can be embodied in many different forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the concepts of the present invention to those skilled in the art. It should also be noted that these embodiments are not mutually exclusive. Components from one embodiment can be implicitly assumed to be present / used in another embodiment.
[0015] Before describing the embodiments in more detail, FIG. 7 shows an example of an operating environment of an encoder 500 that can be used to encode a bitstream as described herein. The encoder 500 receives audio from the network 702 and / or the storage 704 and / or the audio recorder 706, encodes the audio into a bitstream as described below, and transmits the encoded audio to the decoder 708 via the network 710. In some embodiments where the encoder 500 is a distributed encoder, the transmitting entity 5001 can transmit the encoded audio to the decoder 708 via the network 710, as indicated by the dashed line. The storage device 704 may be part of a storage repository for multi-channel audio signals such as a store or streaming audio service's storage repository, a separate storage component, a component of a mobile device, etc. The decoder 708 may be part of a device 712 having a media player 714. The device 712 may be a mobile device, a set-top device, a desktop computer, or the like.
[0016] FIG. 8 is a block diagram showing a virtualized environment 800 in which functions implemented according to some embodiments may be virtualized. In the present context, virtualizing may include virtualizing a hardware platform, memory devices, and networking resources, and means creating a virtual version of an apparatus or device, such as encoder 500. As used herein, virtualization can be applied to any device or its components described herein, and relates to an implementation in which at least some of the functions are implemented as one or more virtual components. Some or all of the functions described herein may be implemented as virtual components executed by one or more virtual machines (VMs) implemented in one or more virtualized environments 800 hosted by one or more of the hardware nodes such as network nodes, UEs, core network nodes, or hardware computing devices operating as hosts. Further, in embodiments where the virtual node does not require wireless connectivity (e.g., a core network node or a host), the node may be fully virtualized.
[0017] Application 802 (alternatively, may be referred to as a software instance, virtual appliance, network function, virtual node, virtual network function, etc.) may be operated in the virtualized environment 800 to implement some of the features, functions, and / or advantages of some of the embodiments disclosed herein.
[0018] Hardware 804 includes a processing circuit, a memory storing software and / or instructions executable by the hardware processing circuit, and / or other hardware devices described herein such as a network interface, an input / output interface. The software is executed by the processing circuit to instantiate one or more virtualization layers 806 (also referred to as a hypervisor or a virtual machine monitor (VMM)), provide VMs 808A and 808B (one or more of which may generally be referred to as VM 808), and / or perform any of the functions, features, and / or benefits described in connection with some of the embodiments described herein. The virtualization layer 806 may present a virtual operating platform that appears to the VMs 808 as networking hardware.
[0019] The VMs 808 include virtual processing, virtual memory, virtual networking or interfaces, and virtual storage and may be operated by the corresponding virtualization layer 806. Different embodiments of instances of the virtual appliance 802 may be implemented in one or more of the VMs 808, and the implementation may be done in different ways. The virtualization of hardware is, in some contexts, referred to as network function virtualization (NFV). NFV may be used to consolidate many network device types onto industry-standard high-volume server hardware, physical switches, and physical storage that may be located within data centers and customer premise equipment.
[0020] In the context of NFV, VM808 can be a software implementation of a physical machine that runs a program as if the program were running on a non-physically virtualized machine. Each of the VM808s, and that portion of the hardware 804 that executes the VM, forms a distinct virtual network element, whether it is hardware dedicated to that VM and / or hardware shared with other VMs by that VM. Still in the context of NFV, the virtual network functions are run within one or more VMs808 on the hardware 804 and serve to handle specific network functions corresponding to the application 802.
[0021] The hardware 804 may be implemented in a stand-alone network node having general or specific components. The hardware 804 can implement some functions via virtualization. Alternatively, the hardware 804 may be part of a larger class of hardware (such as in a data center or CPE, etc.) that is managed via management and orchestration 810 where multiple hardware nodes cooperate, and in particular, oversee the lifecycle management of the application 802. In some embodiments, the hardware 804 is coupled to one or more radio units each including one or more transmitters and one or more receivers that can be coupled to one or more antennas. The radio units may communicate directly with other hardware nodes via one or more appropriate network interfaces and may be used in combination with virtual components to provide wireless capabilities to virtual nodes, such as a wireless access node or a base station. In some embodiments, the use of the control system 812 can alternatively provide some signaling that can be used for communication between the hardware node and the radio unit.
[0022] FIG. 9 is a block diagram showing elements of an encoder 500 configured to encode an audio frame according to some embodiments of the concepts of the present invention. As shown, encoder 500 may include a network interface circuit 905 (also referred to as a network interface) configured to provide communication with other devices / entities / functions, etc. Encoder 500 may also include a processor circuit 901 (also referred to as a processor) coupled to network interface circuit 905, and a memory circuit 903 (also referred to as a memory) coupled to the processor circuit. Memory circuit 903 may include computer-readable program code that, when executed by processor circuit 901, causes the processor circuit to perform operations according to the embodiments disclosed herein.
[0023] According to other embodiments, processor circuit 901 may be defined to include memory such that a separate memory circuit is not required. As discussed herein, the operations of encoder 500 may be performed by processor 901 and / or network interface 905. For example, processor 901 may control network interface 905 to send communications to decoder 708 and / or receive communications from one or more other network nodes / entities / servers such as other encoder nodes, repository servers, etc. via network interface 905. Further, modules may be stored in memory 903, and these modules may provide instructions that cause processor 901 to perform respective operations when the instructions of the modules are executed by processor 901.
[0024] As described above, in the field of audio encoding, various audio music classifiers are used. However, these classifiers may not be able to distinguish different classes within the space of music signals. Many classifiers do not provide sufficient resolution to distinguish the classes required by complex multi-mode codecs. In particular, spectral flatness and crest value do not capture the spread or sparsity of energy across the spectrum. FIG. 2 shows two exemplary spectra A and B. Spectrum A is a sparse spectrum suitable for a certain encoding mode, while spectrum B is not suitable for this encoding mode. However, since both spectral flatness and crest measure result in the same value, these spectra cannot be distinguished. FIG. 3 further shows the spectra of signals that are each undesirably encoded by a certain encoding mode.
[0025] FIG. 4 shows an abstraction for creating a classifier to determine the class of a signal that controls subsequent mode determination. These implementations address finding better classifiers for distinguishing harmonic and noise-like music signals.
[0026] In some embodiments, the concepts of the present invention are part of an audio encoding and decoding system. The audio encoder is a multi-mode audio encoder, and the method improves the selection of an appropriate encoding mode for a signal. To clarify that this is the encoding mode selected by the encoder, hereinafter this will be referred to as the encoding mode, although those skilled in the art will understand that these terms can be used interchangeably. The input signal x(m,n), where n = 0, 1, 2, … L-1, is segmented into audio frames of length L, m represents the frame index, and n represents the sample index within the frame. The input signal is converted into a frequency domain representation such as a modified discrete cosine transform (MDCT) or a discrete Fourier transform (DFT). Other frequency domain representations such as filter banks are possible, but they should provide a moderately high frequency resolution for the target analysis range. In this embodiment, at least one of the audio encoding modes operates in the MDCT domain. Therefore, it is beneficial to reuse the same transform for frequency domain analysis. The MDCT is defined by the following relationship TIFF0007710054000003.tif13170, where X(m,k) represents the MDCT spectrum of frame m at frequency index k, and w a (n) is the analysis window. The frequency index k may sometimes be referred to as a frequency bin. Usually, audio frames are extracted with time overlap. The analysis window is selected to provide a good trade-off, for example, between algorithm delay, frequency resolution, and quantization noise shaping. When the frequency domain representation is based on the DFT, the spectrum is defined as follows TIFF0007710054000004.tif13170.
[0027] Note that in this case, the frame length L may be different to provide a frame length suitable for DFT analysis.
[0028] Signal classification aims to select an encoding mode that can represent the input audio file in the best way. In particular, the classification aims to identify signals with high peakiness and low energy concentration. The analysis can focus on the critical frequency region where the choice of encoding method has a significant impact. Here, the frequency index \(k = k start …k end and pay attention to the range of the spectrum \(X(m,k)\) defined by. In some of these embodiments, the critical range is the upper half of the frequency spectrum encoded by the bandwidth expansion technique. This corresponds to \(k start = 320 and \(k end = 639, the operating sampling rate is 32 kHz, and the frame length is \(L = 640\). The bandwidth expansion of different encoding modes has differences in spectral signatures that are important for mode selection. More specifically, it aims to identify signals that have a peaked structure in the high-frequency range but do not have a noisy component characterized by a wideband of high-energy coefficients within the spectrum. Examples of desired and undesired signals can be found in Figure 3. This is implemented by analyzing the feature \(crest(m)\) and the new feature \(crest mod (m).
[0029] Figure 4 shows the operations performed by the encoder in some embodiments of the concept of the present invention. Referring to Figure 4, in step 410, the size of the critical region or the absolute spectrum \(A i (m) is obtained by the encoder 500. The size of the critical region or the absolute spectrum \(A i (m) can be obtained by the encoder 500 according to TIFF0007710054000005.tif30170, where \(M = k end -k start + 1 is the number of bins or frequency indices in the critical band. In step 420, the crest value of frame \(m\) is Derived by the encoder 500 according to TIFF0007710054000006.tif15170, where crest(m) gives a measure of the peakness of frame m. A complementary peakness measure t(m) can also be Obtained by the encoder 500 according to TIFF0007710054000007.tif29170, where A thr Is a relative threshold value where an appropriate value can be A thr = 0.1 or in the range [0.01, 0.4]. In step 430, the detection measure of the noise band is Calculable by the encoder 500 according to TIFF0007710054000008.tif18170, where movmean(A i (m), W) is the moving average of the absolute spectrum A i (m) using a window size of W. A value suitable for the window size may be W = 21, or any odd number within the range [7, 31].
[0030] In one embodiment of the concept of the present invention, movemean(A i (m), W) is Defined according to TIFF0007710054000009.tif15170a = max(0, i - (W - 1) / 2) b = min(M - 1, i + (W - 1) / 2) Here, the average at the edge of the absolute spectrum A (m) is formed using only the values within the range of A i (m). i (m).
[0031] Alternatively, the definition may be described in a recursive form that requires fewer arithmetic operations. TIFF0007710054000010.tif63170
[0032] In another embodiment, the definition of movmean(A i (m), W) can assume that the absolute spectrum is 0 outside the range of i = 0…M - 1, which TIFF0007710054000011.tif15170a = max(0, i - (W - 1) / 2) b = min(M - 1, i + (W - 1) / 2) Simplify the numerator of the formula according to
[0033] movmean(A i (m), W). It should be noted that the definition of movmean(A i (m), W) assumes that the window length W is odd and extends the same number of samples in the positive and negative directions from the current frequency bin i. It is possible to use an even window length W by appropriately adapting the above formula. For example, when using only the shifted window for even W, movmean(A TIFF0007710054000012.tif15170a = max(0, i - W / 2) b = min(M - 1, i + W / 2 - 1) As in, or when calculating the average of the backward and forward alignments of the window, TIFF0007710054000013.tif17170a = max(0, i - W / 2) b = min(M - 1, i + W / 2 - 1) c = max(0, i - W / 2) d = min(M - 1, i + W / 2 - 1) It can be written as follows. Generally, the moving average operation can be implemented using a moving average filter of the following form TIFF0007710054000014.tif16170, where w j are the filter coefficients.
[0034] crest mod (m) gives a measure of the local concentration of energy and indicates the noise band in the spectrum. To stabilize the decision, crest(m) and crest mod (m) may be low-pass filtered by the encoder 500. For example, crest LP (m) = (1 - α)·crest(m) + α·crest LP (m - 1) crest mod,LP (m) = (1 - β)·crest mod (m) + β·crest mod,LP (m - 1) where α and β are filter coefficients. A suitable value of α may be α = 0.97 or in the range [0.5, 1), and equivalently, a suitable value of β may be β = 0.97 or in the range [0.5, 1).
[0035] The encoding mode targeting a spectrum with peaks and no noisy components is subject to the following condition: TIFF0007710054000015.tif10170 is satisfied, in which case it is invalidated, where crest thr 、crest mod,thr and t thr are decision thresholds. Suitable values for these thresholds may be crest thr = 7, crest mod,thr = 2.128 and t thr = 220. More generally, suitable values can be found in the ranges crest thr ∈ [3, 12], crest mod,thr ∈ [1, 4] and t thr ∈ [150, 300]. Decomposing these conditions, crest mod,LP (m) > crest mod,thr ensures that the encoding mode is invalidated for noisy components, while crest LP (m) > crest thr and t(m) > t thr restricts the impact of this decision to signals with a peak spectrum.
[0036] Alternatively, the condition regarding t(m) can be omitted, and the decision is TIFF0007710054000016.tif12170.
[0037] In another embodiment of the concept of the present invention, the decision is that when crest LP (m) is high while crest LP,mod (m) is low, TIFF0007710054000017.tif may be formed to enable a harmonic mode according to 12170, where the threshold crest thr2 and crest mod,thr2 are, crest thr and crest mod,thr may be the same as.
[0038] In step 440, the encoding mode is selected by the encoder 500, including at least determining Harmonic_decision(m). Finally, in step 450, the encoder 500 performs encoding using the selected encoding mode.
[0039] FIG. 5 is a schematic diagram of a multi-mode audio codec using mode determination logic based on an input signal. Referring to FIG. 5, the absolute value calculator 510 receives the input audio and converts the input audio into a frequency domain representation such as a modified discrete cosine transform (MDCT). If the signal has already been converted into the frequency domain used by the multi-mode encoder, the frequency domain representation can be reused in this step. Then, the absolute value calculator 510 determines the absolute value (e.g., magnitude) of the MDCT. The absolute magnitude is used by the peakness scale 520 to determine the peakness scale of the MDCT. An additional peakness scale 530 can also be derived.
[0040] The noise detection scale 540 receives the absolute value of the MDCT and determines the noise level of the input audio signal. The mode enable determination 550 receives the peakness scale and the noise detection scale and determines whether to enable the selection of a mode. For example, if there are two encoding modes, the mode enable determination 550 determines which of the two encoding modes can be used.
[0041] The mode selector 560 determines the encoding mode to be used and indicates to the multi-mode encoder 580 which mode to use. The multi-mode encoder 580 encodes the input audio signal and generates the encoded audio 590. The determined mode decision 570 is combined with the encoded audio 590 and transmitted or stored for the multi-mode decoder.
[0042] FIG. 6A shows an example of which encoding mode or group of encoding modes is used. In response to the harmonic condition being true (e.g., harmonic_decision(m) being true), it is determined that encoding mode C is to be used. In response to the harmonic condition being false (e.g., harmonic_decision(m) being false), it is determined that encoding mode D or encoding mode E is to be used. FIG. 6B shows an example where both branches have a group of encoding modes. In FIG. 6B, in response to the harmonic condition being true (e.g., harmonic_decision(m) being true), it is determined that encoding mode C or encoding mode F is to be used. In response to the harmonic condition being false (e.g., harmonic_decision(m) being false), it is determined that encoding mode D or encoding mode E is to be used.
[0043] Next, with reference to the flowchart of FIG. 10, according to some embodiments of the concepts of the present invention, the operation of the encoder 500 (implemented using the block diagram structure of FIG. 9) will be described. For example, the modules may be stored in the memory 903 of FIG. 9, and these modules may provide instructions such that when the instructions of the modules are executed by the respective communication device processing circuits 901, the processing circuit 901 executes the respective operations of the flowchart.
[0044] Referring to FIG. 10, in block 1001, the processing circuit 901 derives the frequency spectrum of the input audio signal. In some embodiments of the concept of the present invention, the processing circuit 901 derives the frequency spectrum by segmenting the input audio signal x(m,n), where n = 0, 1, 2, … L-1 into audio frames of length L, where m represents the frame index and n represents the sample index within the frame, and the input audio signal is converted to a frequency-domain representation according to TIFF0007710054000018.tif13170.
[0045] In block 1003, the processing circuit 901 obtains the size of the critical frequency region of the frequency spectrum. The critical frequency region is defined by the frequency index k = k start …k end and the critical frequency range is the upper half of X(m,k). In some embodiments, the critical frequency range corresponds to k start = 320 and k end = 639, the operating sampling rate is 32 kHz, and the frame length is L = 640.
[0046] In some embodiments of the concept of the present invention, the processing circuit 901 obtains the size of the critical frequency region according to TIFF0007710054000019.tif30170, where M = k end -k start +1 is the number of bins in the critical band related to the critical frequency region.
[0047] In block 1005, the processing circuit 901 obtains a peakness measure. In some embodiments of the concept of the present invention, the processing circuit 901 obtains the peakness measure according to TIFF0007710054000020.tif18170, where crest(m) gives the peakness measure for frame m.
[0048] In other embodiments of the concept of the present invention, the processing circuit 901 Obtain the peak sharpness measure of the frame according to TIFF0007710054000021.tif29170, where A thr is a relative threshold value.
[0049] In some embodiments, A thr = 0.1. In other embodiments, A thr is within the range of [0.01, 0.4].
[0050] In block 1007, the processing circuit 901 obtains a noise band detection measure. In some embodiments of the concept of the present invention, the processing circuit 901 obtains the noise band detection measure according to TIFF0007710054000022.tif18170, where crest mod (m) is the noise band detection measure, and movmean(A i (m), W) is the moving average of the absolute spectrum A i (m) using a window size of W.
[0051] In some embodiments, the processing circuit 901 determines movmean(A (m), W) according to TIFF0007710054000023.tif15170a = max(0, i - (W - 1) / 2) b = min(M - 1, i + (W - 1) / 2) i .
[0052] In block 1009, the processing circuit determines which of two encoding modes or a group of encoding modes to use based on at least the peak sharpness measure and the noise band detection measure. For example, a sparse spectrum may be suitable for the first encoding mode or set of encoding modes but not for the second encoding mode or set of encoding modes.
[0053] In some embodiments of the concepts of the present invention, the processing circuit 901 determines which of two coding modes or a group of coding modes to use by determining which of the two coding modes to use based on at least the peakness measure and the noise band detection measure, when Harmonic_decision(m) is true, where Harmonic_decision(m) is determined according to TIFF0007710054000024.tif10170, where crest thr , crest mod,thr and t thr are decision thresholds, crest LP (m) is the low-pass filtered crest(m), and crest mod,LP (m) is the low-pass filtered crest mod (m).
[0054] The processing circuit 901 crest LP (m) = (1 - α)·crest(m) + α·crest LP (m - 1) crest mod,LP (m) = (1 - β)·crest mod (m) + β·crest mod,LP (m - 1) can determine the low-pass filtered crest(m) and the low-pass filtered crest mod (m) according to, where α and β are filter coefficients. In some embodiments, α is in the range of [0.5, 1), and β is in the range of [0.5, 1). In other embodiments, Harmonic_decision(m) is determined according to TIFF0007710054000025.tif12170.
[0055] In other embodiments, the processing circuit 901 determines which of two encoding modes to use based on at least a peakiness measure, a noise band detection measure, and based on whether Harmonic_enabled(m) is true, where Harmonic_enabled(m) is determined according to TIFF0007710054000026.tif12170, where crest thr2 and crest mod,thr2 are decision thresholds.
[0056] Accordingly, the processing circuit 901 determines the encoding mode based on at least a peakiness measure, a noise band detection measure, and Harmonic_decision(m).
[0057] Referring to FIG. 11, in some embodiments of the concepts of the present invention, in block 1101, the processing circuit 901 determines to use the first encoding mode of the two encoding modes or the first encoding mode from a group of encoding modes in response to Harmonic_decision(m) being TRUE. In block 1103, the processing circuit 901 determines to use the second encoding mode of the two encoding modes or the second encoding mode from a group of encoding modes in response to Harmonic_decision(m) being FALSE.
[0058] Returning to FIG. 10, in block 1011, the processing circuit 901 encodes the input audio signal based on the encoding mode determined to be used.
[0059] In other embodiments of the concepts of the present invention, using the concepts of the present invention described herein, it is possible to determine whether an input audio signal has a high peakiness and a low energy concentration. FIG. 12 shows one embodiment of determining whether an input audio signal has a high peakiness and a low energy concentration.
[0060] Referring to FIG. 12, in block 1201, the processing circuit 901 derives the frequency spectrum of the input audio signal. Block 1201 is similar to block 1001 described above.
[0061] In block 1203, the processing circuit 901 obtains the magnitude of the critical frequency region of the frequency spectrum. Block 1203 is similar to block 1003 described above.
[0062] In block 1205, the processing circuit 901 obtains the peakiness measure. Block 1205 is similar to block 1005 described above.
[0063] In block 1207, the processing circuit 901 obtains the noise band detection measure. Block 1207 is similar to block 1007 described above.
[0064] In block 1209, the processing circuit 901 determines the harmonic condition based on at least the peakiness measure and the noise band detection measure.
[0065] In block 1211, the processing circuit 901 outputs an indication of whether the harmonic condition is true or false.
[0066] In some embodiments, the processing circuit 901 determines that the harmonic condition is true in response to the low-pass filtered crest(m) being greater than the crest threshold and the low-pass filtered crest mod (m) being greater than the crest mod threshold, where crest(m) is a measure of the peakiness of frame m and crest mod (m) is a measure of the local concentration of energy.
[0067] In some embodiments of the concept of the present invention, the processing circuit 901 According to TIFF0007710054000027.tif34170, crest(m) and crest mod(m) is determined, where A i (m) is the magnitude of the modified discrete cosine transform (MDCT) of the audio signal in frame m, M is the number of frequency indices in the critical region, movmean(A i (m),W) is the moving average of A i (m) using window size W.
[0068] Processing circuit 901 determines A i (m) according to TIFF0007710054000028.tif30170, where X(m,k) represents the MDCT spectrum of frame m at frequency index k, and M = k end - k start + 1, and k end and k startは are the frequency indices of the critical region of X(m,k).
[0069] movmean(A i (m),W) are described above in various embodiments.
[0070] Processing circuit 901 determines X(m,k) according to TIFF0007710054000029.tif15170, where L is the frame length of frame m.
[0071] The computing devices (e.g., UEs, network nodes, hosts) described in this specification can include the illustrated combinations of hardware components, but other embodiments can include computing devices with different combinations of components. It should be understood that these computing devices comprise any suitable combination of hardware and / or software necessary to perform the tasks, features, functions, and methods disclosed herein. The determinations, calculations, acquisitions, or similar operations described herein may be performed by a processing circuit, which, for example, converts acquired information into other information, compares the acquired or converted information with information stored in a network node, and / or performs one or more operations based on the acquired or converted information, and, as a result of said processing, may make a determination. Further, the components are shown as a single box placed within a larger box, or a single box nested within multiple boxes, but in reality, the computing device may comprise a plurality of different physical components that make up a single illustrated component, and the functions may be divided among separate components. For example, a communication interface may be configured to include any of the components described herein, and / or the functions of the components may be divided between the processing circuit and the communication interface. In another example, the non-computation-intensive functions of any of such components may be implemented in software or firmware, and the computation-intensive functions may be implemented in hardware.
[0072] In certain embodiments, some or all of the functions described herein may be provided by a processing circuit that executes instructions stored in a memory, and in certain embodiments, may be a computer program product in the form of a non-transitory computer-readable storage medium. In alternative embodiments, some or all of the functions may be provided by a processing circuit without executing instructions stored in a separate or discrete device-readable storage medium, such as in a hardwired fashion. In any of these particular embodiments, whether or not executing instructions stored in a non-transitory computer-readable storage medium, the processing circuit can be configured to perform the above functions. The advantages provided by such functions are not limited to the processing circuit alone or other components of a computing device, but are enjoyed by the entire computing device and / or by the end user and the wireless network generally.
[0073] Further provisions and embodiments are described below.
[0074] In the foregoing description of various embodiments of the concepts of the present invention, it should be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the concepts of the present invention. Unless otherwise specified, all terms (including technical and scientific terms) used herein shall have the same meaning as commonly understood by one of ordinary skill in the art to which the concepts of the present invention pertain. Terms such as those defined in commonly used dictionaries shall be interpreted as having a meaning that conforms to their meaning in the context of the present specification and the relevant art, and shall not be interpreted in an idealized or overly formal sense unless expressly so defined herein.
[0075] When an element is referred to as being "connected to", "coupled to", "responsive to", or variations thereof, another element, that element can be directly connected to, coupled to, or responsive to the other element, or intervening elements may be present. In contrast, when an element is referred to as being "directly connected to", "directly coupled to", "directly responsive to", or variations thereof, another element, no intervening elements are present. Like numbers refer to like elements throughout. Further, as used herein, "coupled to", "connected to", "responsive to", or variations thereof may include wirelessly coupled to, wirelessly connected to, or wirelessly responsive to. As used herein, the singular forms "a", "an", and "the" include the plural forms as well, unless the context clearly indicates otherwise. For brevity and / or clarity, well-known functions or structures may not be described in detail. The term "and / or" (abbreviated " / ") includes any and all combinations of one or more of the associated listed items.
[0076] To describe various elements / acts, the terms first, second, third, etc. may be used herein, but it will be understood that these elements / acts should not be limited by these terms. These terms are only used to distinguish one element / act from another. Thus, a first element / act in some embodiments may, without departing from the teachings of the concepts of the present invention, be referred to as a second element / act in other embodiments. The same reference numbers or the same reference signs represent the same or similar elements throughout this specification.
[0077] As used herein, the terms "comprise," "comprising," "comprises," "include," "including," "includes," "have," "has," "having," or variations thereof are open-ended and include one or more recited features, integers, elements, steps, components, or functions, but do not exclude the presence or addition of one or more other features, integers, elements, steps, components, functions, or groups thereof. Further, as used herein, the common abbreviation "e.g.", which is derived from the Latin phrase "exempli gratia," may be used to introduce or specifically list one or more general examples of the foregoing items and is not limiting of such items. The common abbreviation "i.e.", which is derived from the Latin phrase "id est," may be used to specifically list particular items from a more general recitation.
[0078] Exemplary embodiments are described herein with reference to block diagrams and / or flowchart diagrams of a computer-implemented method, apparatus (system and / or device) and / or computer program product. It should be understood that the blocks of the block diagrams and / or flowchart diagrams, as well as combinations of blocks in the block diagrams and / or flowchart diagrams, can be implemented by computer program instructions executed by one or more computer circuits. These computer program instructions can be provided to a processor circuit of a general-purpose computer circuit, a special-purpose computer circuit, and / or other programmable data processing circuits to create a machine, whereby the instructions executed via the processor of the computer and / or other programmable data processing device implement the functions / acts specified in one or more blocks of the block diagram and / or flowchart, and thereby transform and control transistors, values stored in memory locations, and other hardware components within such circuits to create means (functions) and / or structures for implementing the functions / acts specified in the blocks of the block diagram and / or flowchart.
[0079] These computer program instructions can also be stored in a tangible computer-readable medium that can direct a computer or other programmable data processing device to function in a particular manner, whereby the instructions stored in the computer-readable medium create a manufacture including instructions for implementing the functions / acts specified in one or more blocks of the block diagram and / or flowchart. Accordingly, embodiments of the concepts of the present invention may be embodied in hardware and / or in software (including firmware, resident software, microcode, etc.) running on a processor such as a digital signal processor, which may sometimes be generically referred to as a "circuit", "module" or variations thereof.
[0080] Also, note that in some alternative implementations, the functions / acts recited in a block may occur in any order other than that recited in a flowchart. For example, depending on the functionality / acts involved, two blocks shown in succession may in fact be executed substantially concurrently, or the blocks may sometimes be executed in reverse order. Further, the functionality of a given block of a flowchart and / or block diagram may be separated into multiple blocks, and / or the functionality of two or more blocks of a flowchart and / or block diagram may be at least partially integrated. Finally, other blocks may be added / inserted between the blocks illustrated, and / or blocks / acts may be omitted without departing from the scope of the inventive concept. Further, although some of the figures include arrows on communication paths to indicate a primary direction of communication, it should be understood that communication may occur in a direction opposite to that shown by the illustrated arrows.
[0081] Many variations and modifications can be made to the embodiments without substantially departing from the principles of the inventive concept. All such variations and modifications are intended to be included herein within the scope of the inventive concept. Accordingly, the subject matter disclosed above should be regarded as illustrative and not restrictive, and the examples of embodiments should be considered to cover all such modifications, extensions, and other embodiments that fall within the spirit and scope of the inventive concept. Accordingly, to the maximum extent permitted by law, the scope of the inventive concept should be determined by the broadest permissible interpretation of this disclosure, including examples of embodiments and their equivalents, and should not be limited or restricted by the foregoing detailed description.
[0082] Embodiments 1. A method in an encoder for determining which of two encoding modes or groups of encoding modes to use, the method comprising: deriving a frequency spectrum of an input audio signal (901); Obtaining the size of the critical frequency region of the frequency spectrum (903) and Obtaining the peakiness measure of the frame (905), and Obtaining the noise band detection measure (907), and Determining which of two encoding modes or a group of encoding modes to use based on at least the peakiness measure and the noise band detection measure (909), and Encoding the input audio signal based on the encoding mode determined to be used (911) and A method including. 2. Encoding the input audio signal based on the encoding mode determined to be used is Selecting one encoding mode from the group of encoding modes used to encode the input audio signal in response to determining that a group of encoding modes is to be used The method according to embodiment 1, including. 3. Deriving the frequency spectrum includes deriving the frequency spectrum X(m,k), where X(m,k) represents the frequency spectrum of frame m at frequency index k, the method according to any one of embodiments 1 to 2. 4. Deriving the frequency spectrum is Segmenting the input audio signal x(m,n), n = 0, 1, 2, … L-1 into audio frames of length L, where m represents the frame index and n represents the sample index within the frame, segmenting, and Converting the input audio signal into a frequency domain representation according to TIFF0007710054000030.tif17170, where X(m,k) represents the frequency spectrum of the modified discrete cosine transform (MDCT) of frame m at frequency index k, and w a (n) is the analysis window, converting the input audio signal, and Frequency index k = k start …k endTo obtain the magnitude spectrum of X(m,k) defined by, the magnitude spectrum is obtained with the critical frequency range being the upper half of X(m,k), and The method according to any one of Embodiments 1 to 3, including 5. The critical frequency range is k start = 320 and k end = 639, the input sampling rate is 32 kHz, and the frame length is L = 640. The method according to any one of Embodiments 3 to 4 6. Obtaining the magnitude of the critical frequency region is Including obtaining the magnitude of the critical frequency region according to TIFF0007710054000031.tif35170, where M = k end - k start + 1 is the number of frequency indices of the critical band related to the critical frequency region. The method according to any one of Embodiments 3 to 5 7. Obtaining the peakness measure is Including obtaining the peakness measure according to TIFF0007710054000032.tif18170, where crest(m) gives the peakness measure of frame m. The method according to Embodiment 6 8. Obtaining the peakness measure is Including obtaining the peakness measure according to TIFF0007710054000033.tif36170, where A thr Is the relative threshold. The method according to Embodiment 6 9. A thr = 0.1. The method according to Embodiment 8 10. A thr Is within the range [0.01, 0.4]. The method according to Embodiment 8 11. Obtaining the noise band detection measure is Including obtaining the noise band detection measure according to TIFF0007710054000034.tif18170, where crest mod (m) is the noise band detection measure, and movmean(A i (m), W) is the absolute spectrum A using the window size of W iThe method according to any one of Embodiments 1 to 10, which is the moving average of (m). 12.movmean(A i (m),W) is TIFF0007710054000035.tif14170a = max(0, i - (W - 1) / 2) b = min(M - 1, i + (W - 1) / 2) The method according to Embodiment 11, which is determined according to 13. crest LP (m) = (1 - α)·crest(m) + α·crest LP (m - 1) crest mod,LP (m) = (1 - β)·crest mod (m) + β·crest mod,LP (m - 1) According to, further including low - pass filtering crest(m) and crest mod (m), where α and β are filter coefficients. The method according to any one of Embodiments 7 to 12. 14. The method according to Embodiment 13, where α is in the range of [0.5, 1) and β is in the range of [0.5, 1). 15. Determining which of two encoding modes to use based on at least a peakness measure and a noise band detection measure, including determining which of two encoding modes to use based on the case where Harmonic_decision(m) is true, and Harmonic_decision(m) is TIFF0007710054000036.tif16170 determined according to, where crest thr , crest mod,thr and t thr are decision thresholds. The embodiment according to any one of Claims 1 to 14. 16. Determining which of two encoding modes to use based on at least a peakness measure and a noise band detection measure includes determining which of two encoding modes to use based on the case where Harmonic_decision(m) is true, and Harmonic_decision(m) is determined according to TIFF0007710054000037.tif16170, where crest thr and crest mod,thr are decision thresholds, the method according to any one of Embodiments 1 to 14. 17. Determining an encoding mode based on at least a peakness measure and a noise band detection measure includes determining which of two encoding modes to use when Harmonic_decision(m) is true, and Harmonic_decision(m) is determined according to TIFF0007710054000038.tif16170, where crest thr2 and crest mod,thr2 are decision thresholds, the method according to any one of Embodiments 1 to 14. 18. Determining an encoding mode based on at least a peakness measure and a noise band detection measure includes determining an encoding mode based on at least a peakness measure, a noise band detection measure, and Harmonic_decision(m), the method according to any one of Embodiments 15 to 17. 19. Determining an encoding mode based on at least a peakness measure, a noise band detection measure, and Harmonic_disabled(m) includes determining to use a first encoding mode of two encoding modes in response to Harmonic_decision(m) being TRUE (1101), and determining to use a second encoding mode of two encoding modes in response to Harmonic_decision(m) being FALSE (1103) including, the method according to Embodiment 18. A method in an encoder for determining whether an input audio signal has a high peakiness and a low energy concentration, comprising: deriving a frequency spectrum of the input audio signal (1201); obtaining a size of a critical frequency region of the frequency spectrum (1203); obtaining a peakiness measure (1205); obtaining a noise band detection measure (1207); determining a harmonic condition based on at least the peakiness measure and the noise band detection measure (1209); outputting an indication of whether the harmonic condition is true or false (1211); and a method including the above steps. 21. The method according to embodiment 20, further comprising determining that the harmonic condition is true in response to the low-pass filtered crest(m) being greater than a crest threshold and the low-pass filtered crest mod (m) being greater than a crest mod threshold, where crest(m) is a measure of the peakiness of frame m and crest mod (m) is a measure of the local concentration of energy. 22. The method according to embodiment 21, further comprising determining crest(m) and crest mod (m) according to TIFF0007710054000039.tif38170, where A i (m) is the magnitude of the frequency spectrum of the audio signal in frame m, M is the number of frequency indices in the critical region, and movmean(A i (m), W) is the moving average of A i (m) using a window size W. 23. The method according to embodiment 21, further comprising determining A i (m) according to TIFF0007710054000040.tif35170, where X(m,k) represents the frequency spectrum of frame m at frequency index k, and M = k end-k start is +1, and k end and k start are the frequency indices of the critical region of X(m,k), the method according to Embodiment 22. 24. Further including determining X(m,k) according to TIFF0007710054000041.tif17170, where L is the frame length of frame m, the method according to Embodiment 23. 25. A processing circuit (901), A memory (905) coupled to the processing circuit, which includes instructions that, when executed by the processing circuit, cause the communication device to perform the operations according to any one of Embodiments 1 to 24, the memory (905) An encoder device (500) comprising. 26. An encoder device (500) adapted to execute according to any one of Embodiments 1 to 24. 27. A computer program comprising program code to be executed by the processing circuit (901) of the encoder device (500), whereby, by executing the program code, the encoder device (500) is caused to perform the operations according to any one of Embodiments 1 to 24, the computer program. 28. A computer program product comprising a non - transitory storage medium including program code to be executed by the processing circuit (901) of the encoder device (500), whereby, by executing the program code, the encoder device (500) is caused to perform the operations according to any one of Embodiments 1 to 24, the computer program product.
[0083] Explanations for various abbreviations / acronyms used in this disclosure are provided below. Abbreviation Explanation MDCT Modified Discrete Cosine Transform DFT Discrete Fourier Transform
Claims
1. A method in an encoder for determining which of two encoding modes or groups of encoding modes to use, comprising: deriving a frequency spectrum of an input audio signal (901); obtaining a magnitude of the frequency spectrum of a critical frequency region (903), wherein the critical frequency region is the upper half of the frequency spectrum of the input audio signal; obtaining a peakiness measure of the frequency spectrum of the critical frequency region (905); obtaining a noise band detection measure (907); determining which of the two encoding modes or groups of encoding modes to use based at least on the peakiness measure and the noise band detection measure (909); encoding the input audio signal based on the encoding mode determined to be used (911); and obtaining the noise band detection measure comprises obtaining the noise band detection measure according to, where crest_mod(m) is the noise band detection measure and movmean(Ai(m), W) is the moving average of the absolute spectrum Ai(m) using a window size W.
2. Encoding the input audio signal based on the encoding mode determined to be used comprises selecting one encoding mode from the group of encoding modes to be used for encoding the input audio signal in response to determining that a group of encoding modes is to be used The method according to claim 1.
3. Deriving the frequency spectrum comprises deriving a frequency spectrum X(m, k), where X(m, k) represents the frequency spectrum of frame m at frequency index k. The method according to claim 1 or 2.
4. Deriving the frequency spectrum comprises segmenting the input audio signal x(m, n), n = 0, 1, 2,... L-1 into audio frames of length L, where m represents a frame index and n represents a sample index within the frame. Converting the input audio signal in the frequency domain representation according to, where X(m, k) represents the frequency spectrum of the modified discrete cosine transform (MDCT) of frame m at frequency index k, and w a (n) is an analysis window, converting the input audio signal, Frequency index k = k start …k end obtaining the magnitude spectrum of X(m, k) defined by...k, wherein the critical frequency region is the upper half of X(m, k), and obtaining the magnitude spectrum The method according to any one of claims 1 to 3.
5. wherein the critical frequency region corresponds to k start = 320 and k end = 639, the input sampling rate is 32 kHz, and the length of the frame is L = 640, the method according to claim 4.
6. Obtaining the magnitude of the frequency spectrum of the critical frequency region comprises obtaining the magnitude of the frequency spectrum in the critical frequency region according to, where M = k end -k start +1 is the number of frequency indices of the critical band related to the critical frequency region, the method according to any one of claims 3 to 5.
7. Obtaining the peakiness measure comprises obtaining the peakness measure according to, wherein crest(m) gives the peakness measure of frame m, the method according to claim 6.
8. obtaining the peakness measure is obtaining the sharpness measure according to the formula, where t(m) is the sharpness measure and A thr is a relative threshold value, the method according to claim 6.
9. A thr The method according to claim 8, wherein A = 0.
1.
10. A thr The method according to claim 8, wherein A is within the range [0.01, 0.4].
11. movmean(A i (m, W) is determined according to, where a = max(0, i - (W - 1) / 2) b = min(M - 1, i + (W - 1) / 2) the method according to claim 1.
12. crest LP (m) = (1 - α) · crest(m) + α · crest LP (m - 1) crest mod,LP (m) = (1 - β) · crest mod (m) + β · crest mod,LP (m - 1) According to which crest(m) and crest mod The method according to any one of claims 7 to 11, further comprising low-pass filtering (m), where α and β are filter coefficients.
13. α is in the range of [0.5, 1) and β is in the range of [0.5, 1), the method according to claim 12.
14. Determining which of two coding modes or a group of coding modes to use based at least on the peakness measure and the noise band detection measure includes determining one of the two coding modes or the group of coding modes when Harmonic_decision(m) is true, where Harmonic_decision(m) is determined according to, where crest thr , crest mod,thr and t thr are determination thresholds, the method according to any one of claims 1 to 13.
15. Determining which of two coding modes or a group of coding modes to use based at least on the peakness measure and the noise band detection measure includes determining one of the two coding modes or the group of coding modes when Harmonic_decision(m) is true, where Harmonic_decision(m) is determined according to, wherein crest thr and crest mod,thr is a determination threshold value, the method according to any one of claims 1 to 13.
16. Determining the coding mode based at least on the peakness measure and the noise band detection measure includes enabling the determination of the coding mode when Harmonic_decision(m) is true, where Harmonic_decision(m) is determined according to, wherein crest thr2 and crest mod,thr2 is a determination threshold value, the method according to any one of claims 1 to 13.
17. Determining the coding mode based at least on the peakness measure and the noise band detection measure includes determining the coding mode based at least on the Harmonic_decision(m), the method according to any one of claims 14 to 16.
18. determining the coding mode based on the Harmonic_decision(m) is in response to Harmonic_decision(m) being TRUE, determining to use the first coding mode of the two coding modes (1101) In response to the Harmonic_decision(m) being FALSE, determining (1103) to use a second coding mode of the two coding modes The method according to claim 17, comprising: **Claim 19** A processing circuit (901); A memory (905) coupled to the processing circuit, the memory (905) including instructions that, when executed by the processing circuit, cause a communication device to execute the method according to any one of claims 1 to 18 An encoder device (500) comprising:
Citation Information
Patent Citations
Classification of Fast and Slow Signal
US20100063806A1
High-band signal generation
US20160372126A1
Audio coding method and related apparatus
US20170125031A1