Signal generation system and signal generation method

The signal generation system addresses computational inefficiencies and artifacts in HFR by using block-based processing, enhancing audio quality and reducing computational load.

JP7844735B2Active Publication Date: 2026-04-13DOLBY INTERNATIONAL AB
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-11-04
Publication Date
2026-04-13

AI Technical Summary

Technical Problem

Existing high-frequency reconstruction (HFR) methods in audio coding suffer from computational burden, unpleasant artifacts, and ghost pitches, particularly in signals with prominent periodic structures, leading to suboptimal audio quality.

Method used

A signal generation system that processes input signals using an analysis filter bank, subband processing unit, and composite filter bank, employing block extraction, nonlinear frame processing, and overlap addition to generate time-stretched or frequency-transposed signals, reducing computational effort while maintaining high quality.

Benefits of technology

The method achieves superior audio playback with reduced computational burden and minimal artifacts, preserving the natural harmonics of the original signal, resulting in a psychologically pleasing acoustic output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007844735000034
    Figure 0007844735000034
  • Figure 0007844735000035
    Figure 0007844735000035
  • Figure 0007844735000036
    Figure 0007844735000036
Patent Text Reader

Abstract

The present invention provides an efficient scheme for improved cross-product high frequency reconstruction (HFR), where a new frequency component of Q Ω + r Ω 0 is generated based on the existing components Ω and Ω + Ω 0.SOLUTION: The present invention performs a block-wise harmonic transposition, wherein time blocks of complex subband samples are processed with a common phase modification. The superposition of the modified samples has the advantageous effect of limiting unwanted intermodulation products, thereby allowing a coarser frequency resolution and / or a lower degree of oversampling to be used. The present invention in one embodiment uses a window function suitable for use with an improved block-based cross product HFR. A hardware embodiment of the present invention includes an analysis filter bank (101), a subband processor (102) configured by control data (104), and a synthesis filter bank (103).SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an audio source coding system using a harmonic transposition method for high-frequency reconstruction (HFR) in a digital effects processor, the digital effects processor being, for example, an exciter that introduces the resulting harmonic distortion into the luminance of the signal being processed, or a time stretcher or time extender that extends the signal duration along with the preserved spectral content. [Background technology]

[0002] Patent Document 1 describes the concept of transposition as a method for constructing high-frequency bands from low-frequency bands of an audio signal. Using this concept in audio coding can significantly reduce bitrate. In HFR-based audio coding systems, a narrow-bandwidth signal is fed to the core waveform encoder, and higher frequencies are reconstructed using additional side information (information describing the target spectral waveform on the decoder side) and transposition at a very low bitrate. At low bitrates, the bandwidth of the core coded signal is narrow, and reconstructing a perceptually pleasing high band is becoming increasingly important. The harmonic transposition disclosed in Patent Document 1 works very well for complex musical signals in situations with low crossover frequencies. The principle of harmonic transposition is that a sine wave of frequency ω is transposed to a sine wave of frequency Q. φ This involves associating or mapping ω to a sine wave, and Q φ>1 is an integer that determines the order of the transposition. In contrast, HFR based on single sideband modulation (SSB) associates a sine wave of frequency ω with a sine wave of frequency ω + Δω, where Δω is a constant frequency deviation or frequency shift. In the case of low-bandwidth core signals, unpleasant ringing artifacts can occur due to SSB transposition.

[0003] To achieve the best possible audio quality, modern high-quality HFR methods utilize complex modulation frequency banks, along with a large degree of oversampling and very fine frequency resolution to obtain the required audio quality. Fine resolution is necessary to avoid unwanted intermodulation distortion arising from the nonlinearity associated with the synthesis of sine waves. For sufficiently narrow subbands, high-quality methods are intended to have at most one sine wave in each subband. A large degree of temporal oversampling is necessary to avoid aliasing distortion, and some degree of frequency oversampling is also required to avoid pre-echoes in transient signals. The obvious drawback in this case is the extremely heavy computational burden.

[0004] Another common drawback associated with harmonic transposition becomes apparent in the case of signals with a prominent periodic structure. Such signals are superpositions of harmonics of the frequency Ω, 2Ω, 3Ω, ..., where Ω is the fundamental frequency and the order is Q. φ In the case of harmonic transposition, the output sine wave group is Q φ Ω, 2Q φ Ω, 3Q φ Ω, ... has a frequency of Q φ If >1, they become part of the desired complete harmonic set. From the perspective of the resulting audio quality, the fundamental frequency Q of the transpose. φA "ghost" pitch corresponding to Ω is commonly perceived. Harmonic transpositions often introduce a "metallic" sounding character into the encoded and decoded audio signal.

[0005] Patent Document 2, incorporated as a reference in this application, describes an improved cross-products method to address the ghost pitch problem that occurs in the case of high-quality transposition. By transmitting all or part of the fundamental frequency values ​​of the dominant harmonic portion of the signal being transposed with high fidelity, the correction of the nonlinear subbands is supplemented along with the nonlinear coupling of at least two different analysis subbands. As a result, the missing portion in the transposed output is reconstructed, but this incurs considerable computational costs. [Overview of the project] [Problems that the invention aims to solve]

[0006] In view of the above-mentioned shortcomings of existing available HFR methods, the object of the present invention is to provide an improved and more effective cross-product HFR method. In particular, the object of the present invention is to provide a method that enables superior audio playback with less computational burden compared to existing methods.

[0007] The present invention mitigates or eliminates at least one of the above-mentioned problems by the invention described in the claims. [Means for solving the problem]

[0008] The signal processing system according to the disclosed invention is: A signal generation system that generates a time-stretched signal and / or a frequency-transposed signal from an input signal, An analysis filter bank is derived from the input signal, each of which Y (Y≧1) analysis subband signals has multiple complex analysis samples with phase and amplitude, and Y analysis subband signals are derived from the input signal. A subband processing unit that generates a composite subband signal from the Y analyzed subband signals using a subband transposition factor Q and a subband stretching factor S, A composite filter bank that generates the aforementioned time-stretched signal and / or frequency-transposed signal from the composite subband signal. The subband processing unit has a block extraction unit, a nonlinear frame processing unit, and an overlap addition unit, The aforementioned block extraction unit is i) Generate Y frames from L input samples, each of which is extracted from multiple complex analysis samples of the analyzed subband signal, and the length of each frame is L (L > 1). ii) Before generating subsequent frames of L input samples, a series of frames of input samples are generated by applying a block hop size of h samples to multiple complex analysis samples. The nonlinear frame processing unit determines the phase and amplitude of each processed sample (processed sample) of the frame, and generates a frame of the processed sample based on the Y corresponding frames of the input sample generated by the block extraction unit, and for at least one processed sample, i) The phase of the processed sample is based on the phase of each corresponding input sample in each of the Y frames of the input sample, ii) The amplitude of the processed sample is based on the phase of each corresponding input sample in each of the Y frames of the input sample. The overlap summing unit generates the composite subband signal by adding samples from a series of frames of the processing sample while overlapping them. The signal generation system in question is a signal generation system that operates at least when Y=2. [Brief explanation of the drawing]

[0009] [Figure 1]A diagram illustrating the principle of harmonic transposition based on subband blocks. [Figure 2] A diagram illustrating the nonlinear subband blocking process for a single subband input. [Figure 3] A diagram illustrating the nonlinear subband blocking process for two subband inputs. [Figure 4] A diagram illustrating the operation of harmonic transpositions based on improved mutual product subbandblocks. [Figure 5] This figure illustrates an application example of performing transposition based on subbandblocks using several orders of transposition in an improved HFR audio coder. [Figure 6] This figure illustrates an application example of performing transposition based on multiple subbandblocks using a 64-band QMF analysis filter bank. [Figure 7] A diagram illustrating the results of using the transposition method based on the disclosed subbandblock. [Figure 8] A diagram illustrating the results of using the transposition method based on the disclosed subbandblock. [Figure 9] Figure 2 shows a detailed view of the nonlinear processing unit (including the pre-normalization unit and the multiplication unit). [Modes for carrying out the invention]

[0010] <Summary of the Invention> The signal generation system according to the first embodiment of the disclosed invention is: A signal generation system that generates a time-stretched signal and / or a frequency-transposed signal from an input signal, An analysis filter bank is derived from the input signal, each of which Y (Y≧1) analysis subband signals has multiple complex analysis samples with phase and amplitude, and Y analysis subband signals are derived from the input signal. A subband processing unit that generates a composite subband signal from the Y analyzed subband signals using a subband transposition factor Q and a subband stretching factor S, A composite filter bank that generates the aforementioned time-stretched signal and / or frequency-transposed signal from the composite subband signal. The subband processing unit has a block extraction unit, a nonlinear frame processing unit, and an overlap addition unit, The aforementioned block extraction unit is i) Generate Y frames from L input samples, each of which is extracted from multiple complex analysis samples of the analyzed subband signal, and the length of each frame is L (L > 1). ii) Before generating subsequent frames of L input samples, a series of frames of input samples are generated by applying a block hop size of h samples to multiple complex analysis samples. The nonlinear frame processing unit determines the phase and amplitude of each processed sample (processed sample) of the frame, and generates a frame of the processed sample based on the Y corresponding frames of the input sample generated by the block extraction unit, and for at least one processed sample, i) The phase of the processed sample is based on the phase of each corresponding input sample in each of the Y frames of the input sample, ii) The amplitude of the processed sample is based on the phase of each corresponding input sample in each of the Y frames of the input sample. The overlap summing unit generates the composite subband signal by adding samples from a series of frames of the processing sample while overlapping them. The signal generation system in question is a signal generation system that operates at least when Y=2.

[0011] The signal generation method according to the second embodiment of the disclosed invention is: A signal generation method for generating a time-stretched signal and / or a frequency-transposed signal from an input signal, A step of deriving Y (Y≧2) analysis subband signals from the input signal, wherein each of the analysis subband signals has multiple complex analysis samples having phase and amplitude, A step of forming Y frames of L input samples, wherein each frame is extracted from the plurality of complex analysis samples of the analysis subband signal, and the length of the frame is L. The steps include generating a sequence of frames for input samples by applying the block hop size of h samples to the plurality of analysis samples before deriving subsequent frames for L input samples, A step in which a frame of a processed sample is generated by determining the phase and amplitude for each processed sample of the Y corresponding frames of the input sample, and for at least one processed frame, i) the phase of the processed sample is based on the phase of each corresponding input sample in each of the Y frames of the input sample, and ii) the amplitude of the processed sample is based on the amplitude of each corresponding input sample in each of the Y frames of the input sample. The steps include determining the composite subband signal by adding the samples in the sequence of processing sample frames while overlapping them, and A step of generating the time-stretched signal and / or frequency-transposed signal from the composite subband signal. This is a signal generation method that has [a certain characteristic].

[0012] In this case, Y is any integer greater than 1. The signal generation system according to the first embodiment performs the above method at least when Y=2.

[0013] A third embodiment of the disclosed invention is a computer program or software having software instructions for causing a programmable computer to execute the signal generation method according to the second embodiment.

[0014] A fourth embodiment of the disclosed invention is a storage medium (or data carrier) for storing software instructions that cause a programmable computer to execute the signal generation method according to the second embodiment.

[0015] This invention is based on the recognition that the general concept of improved cross-product HFR yields superior results when data is arranged and processed in blocks of complex subband samples. In particular, it makes it possible to apply a frame-by-frame phase offset to the samples. It also makes it possible to adjust the amplitude or magnitude, which also provides similar benefits. Embodiments of the improved cross-product HFR according to the present invention perform subband-block harmonic transposition and can significantly reduce intermodulation. Therefore, it is possible to use filter banks (e.g., QMF filter banks) with coarser frequency resolution and / or less oversampling while maintaining excellent output quality. In the case of subband-block processing, time blocks of complex subband samples are processed with a common phase correction value, and the effect of reducing intermodulation products is obtained by superimposing multiple corrected samples that form the output subband samples. If the input subband signal consists of multiple sine waves, intermodulation products would occur if the present invention were not used. Block-based transposition requires significantly less computational effort than high-resolution transposers, while achieving nearly the same high quality for many signals.

[0016] For the sake of this explanation, it should be noted that in the embodiments, Y ≥ 2, and the nonlinear processing unit uses Y "corresponding" frames from the input samples as input, meaning that the frames are synchronized or nearly synchronized. For example, the samples within each frame relate to time intervals that have a lot of temporal overlap (or overlap or convergence) between frames. The term "corresponding" is used to indicate that they are synchronized or approximately so. Furthermore, the term "frame" may be used interchangeably with "block". Thus, the "block hop size" may be equal to or shorter than the frame length (adjusted for downsampling if downsampling is performed), meaning that the input samples may belong to more than one frame. By determining the phase and amplitude based on the phase and amplitude of all Y corresponding frames of the input sample, the system does not need to generate all processed samples in a frame; without departing from the present invention, the system may generate the phase and / or amplitude of some processed samples based on a smaller number of corresponding input samples or on only one sample.

[0017] In one embodiment, the analysis filter bank is a quadrature mirror filter (QMF) bank or a pseudo-QMF bank with any number of taps and points. It may be, for example, a 64-point QMF bank. The analysis filter bank may be selected from a class such as a windowed discrete Fourier transform or a wavelet transform. Advantageously, the synthesis filter bank coincides with the analysis filter bank by an inverse QMF bank, an inverse pseudo-QMF bank, etc. Such filters are known to have a relatively coarse frequency resolution and / or a relatively low oversampling rate. Different from the prior art, the present invention can be realized using such relatively simple components without suffering from the influence of output degradation, and such embodiments of the present invention exhibit better economy than the prior art.

[0018] In one embodiment, one or more of the following hold for the analysis filter bank: ● The analysis time progression width is Δt A ; ● The analysis frequency interval is Δf A ; ● The analysis filter bank has N>1 analysis sub-bands, and the analysis sub-bands are specified by an analysis sub-band index n = 0,..., N - 1; ● The analysis sub-bands are associated with the frequency band of the input signal.

[0019] In one embodiment, one or more of the following hold for the synthesis filter bank: ● The synthesis time progression width is Δt s ; ● The synthesis frequency interval is Δf s ; ● The synthesis filter bank has M>1 synthesis sub-bands, and the synthesis sub-bands are specified by a synthesis sub-band index m = 0,..., M - 1; ● The composite subband is associated with a time-stretched signal and / or a frequency-transposed signal.

[0020] In one embodiment, a nonlinear processing unit is applied to two input frames (Y=2) to generate one frame that constitutes the frame to be processed, and a subband processing unit includes a cross-processing control unit that generates cross-processing control data. By clarifying the qualitative and / or quantitative nature of the subband processing, the present invention can be greatly enhanced in terms of flexibility and applicability. The control data specifies subbands that differ on the frequency axis by the fundamental frequency of the input signal (for example, specified by an index). In other words, the index specifying the subbands may differ by an integer that approximates the ratio obtained by dividing such fundamental frequency by the analysis frequency interval. Since the new spectral components generated by harmonic transposition become comparable to naturally occurring harmonics, this results in a psychologically pleasing acoustic output.

[0021] In a further improvement to the above embodiment, the (input) analysis and (output) composite subband indices are selected to satisfy equation (16) described below. The parameter σ in this equation is applicable to both the odd-numbered and even-numbered filter banks. When the subband indices are specified as an approximate solution (e.g., least-squares error) to equation (16), the new spectral components obtained by harmonic transposition will be comparable to the series of natural harmonics. Thus, HFR yields a faithful reconstruction of the original signal with the removed high-frequency components.

[0022] A further improvement to the above embodiment provides a method for selecting the parameter r, which appears in equation (16) and represents the order of the cross-product transposition. Given an output subband index m, each value of the transposition order r determines two analytical subband indices n1 and n2. This further improved embodiment evaluates the magnitude or amplitude of the two subbands for a number of r choices and selects the value that maximizes the smaller of the two analytical subband amplitudes. This method of selecting indices eliminates the need to restore numerous amplitudes by amplifying weak components (leading to poor output quality) in the input signal. In this regard, the magnitude or amplitude of the subbands may be calculated in a known way, for example, by the square root of the squares of the input samples forming a frame (block) or part of a frame. The magnitude or amplitude of the subbands may be calculated as the amplitude of the central sample or a sample near the center within the frame. Such calculations easily yield appropriate amplitude measurements.

[0023] In a further improvement to the above embodiment, the analysis subband receives contributions from the harmonic transposition instance according to both direct processing and cross-product-based processing. In this regard, a criterion is applied to determine whether a particular possibility of reconstructing the missing portion by cross-product-based processing is used. For example, this improvement may be configured to refrain from using one or more cross-subband processing units if any of the following conditions (a)-(c) are met.

[0024] Condition (a) is the amplitude M of the direct source term analysis subband that gives rise to the composite subband. S The minimum amplitude value M in the optimal pair of cross-source terms that yield the combined subband. C The ratio to is greater than a predetermined value q, The condition for (b) is that the composite subband receives a large contribution from the direct processing unit. Condition (c) is that the fundamental frequency Ω0 is the interval Δf of the analysis filter bank. A It is smaller.

[0025] In one embodiment, the present invention may perform downsampling or decimation of the input signal. In fact, one or more frames of the input sample may be determined by downsampling the complex analysis sample within a subband, as performed by the block extraction unit.

[0026] According to a further improvement to the above embodiment, the applied downsampling factors satisfy equation (15), which will be described later. It is not permissible for all downsampling factors to be zero, and if they are all zero, it corresponds to a trivial or meaningless case. Equation (15) not only defines the relationship between the downsampling factors D1 and D2, the subband stretching factor S, and the subband transposition factor Q, but also defines the relationship between them and the phase coefficients T1 and T2 that appear in equation (13), which determines the phase of the sample being processed (processed sample). This ensures that the phase of the sample being processed matches the other components of the input signal to which the sample being processed is added.

[0027] In one embodiment, a window function is applied to the sample frames to be processed (windowed) before the frames to be processed are superimposed and added together (overlap added). The windowing unit applies a finite-length window function to the sample frames to be processed. A suitable window function is defined in the claims at the time of filing.

[0028] The inventors recognized that the type of mutual product method described in Patent Document 2 is not entirely suitable from the outset for processing methods based on subband blocks. While such a method may be perfectly applicable to any subband sample in a given block, directly extending it to other samples within the block leads to aliasing artifacts. Therefore, in one embodiment, a window function is applied that includes window samples that fit into a substantially constant sequence (when weighted by complex weights and shifted by the hop size). The hop size may be the product of the block hop size h and the subband extension factor S. Using such a window function can significantly reduce aliasing artifacts. Alternatively or additionally, such a window function also reduces artifacts related to other quantities, such as the phase rotation of the sample being processed.

[0029] Preferably, a set of complex weights or complex weighting coefficients applied to evaluate the state of a window sample differ by a constant phase rotation angle. More preferably, this constant phase rotation angle is proportional to the fundamental frequency of the input signal. The phase rotation angle may be proportional to (the order of the applied cross-product transposition) and / or (the difference in the downsampling factor) and / or (the analysis time progression). The phase rotation angle may be given, at least in an approximate sense, by equation (21).

[0030] According to one embodiment, the present invention enables improved cross-product enhanced harmonic transposition by changing the synthesis window processing according to the fundamental frequency parameters.

[0031] In one embodiment, a series of frames of the sample to be processed are added with a certain degree of overlap or overlap. To achieve appropriate overlap, the frames belonging to the processed frame are appropriately shifted by a hop size, which is the block size h upscaled or stretched by a subband stretching factor S. If the overlap of the series of frames belonging to the input sample is Lh, then the overlap of the consecutive frames belonging to the processed sample becomes S(Lh).

[0032] A system according to one embodiment of the present invention may generate processing samples based on input samples with Y=2, or it may be based on only Y=1 samples. That is, the system can restore or regenerate missing portions not only by a cross-product method (e.g., equation (13), etc.) but also by a direct subband method (e.g., equations (5) and (11), etc.). Preferably, a control unit controls the operation of the system, and this control includes specifying which method should be used to restore a particular missing portion.

[0033] A further improvement to the above embodiment generates a processing sample based on more than three samples (i.e., Y≧3). For example, the sample to be processed may be obtained by multiple harmonic transpositions based on mutual products contributing to the processing sample, multiple direct subband processing, or a combination of mutual product transpositions and direct transpositions. This method of applying the transposition method results in a powerful and flexible HFR. That is, the embodiment can operate to perform the method according to the second embodiment for Y=3, 4, 5, etc.

[0034] In one embodiment, the sample to be processed is determined as a complex number with amplitude, and its amplitude is the average of the amplitude values ​​of each corresponding input sample. The mean may be a (weighted) arithmetic mean, a (weighted) geometric mean, or a (weighted) harmonic mean of two or more samples. When Y=2, the mean is based on two complex input samples. Preferably, the amplitude of the processed sample is a weighted geometric mean. More preferably, the geometric mean is weighted by parameters ρ and 1-ρ as shown in equation (13). In this case, the weighting parameter ρ of the geometric mean is a real number inversely proportional to the subband transposition factor Q. The parameter ρ may also be inversely proportional to the stretching factor S.

[0035] In one embodiment, the system determines the processing sample as a complex number having a phase, the phase being a linear combination of the phases of each corresponding input sample in the frame of the input samples. In particular, the linear combination may be the phase associated with two input samples (Y=2). The linear combination of the two phases may use non-zero integer coefficients, the sum of which is equal to the expansion factor S multiplied by the subband transposition factor Q. Alternatively, the phase obtained by such a linear combination may be further adjusted by a certain phase correction parameter. The phase of the processing sample may be given by equation (13).

[0036] In one embodiment, the block extraction unit (or the corresponding step in the method according to the present invention) may interpolate two or more analytical samples in the analytical subband signal to obtain one input sample that will be included in a frame (block). Such interpolation enables downmixing of the input signal by non-integer factors. The interpolated analytical samples may be continuous or not.

[0037] In one embodiment, the subband processing configuration may be controlled by control data provided by an external means controlling the processing. The control data relates to the acoustic characteristics of the input signal at that point in time. For example, the system itself may have means for determining the acoustic characteristics of the signal at that point in time (e.g., the (dominant) fundamental frequency in the signal). The fundamental frequency information serves as a criterion or guidance when selecting the analysis subbands from which to acquire processing samples. Preferably, the spacing of the analysis subbands is proportional to such fundamental frequencies of the input signal. Alternatively, the control data may be provided from outside the system and preferably contained in an encoded format suitable for communication over a digital communication network as a bitstream. In addition to the control data, such an encoded format may also contain information about the low-frequency components of the signal (e.g., the frequency components in portion 701 of Figure 7). However, from the viewpoint of economically using bandwidth, it is preferable that the encoded format does not contain complete information about the high-frequency components (portion 702 of Figure 7) (in this invention, high-frequency components are reconstructed from low-frequency components). In particular, the present invention provides a decoding system equipped with a control data receiving unit suitable for receiving such control data, wherein the control data may be included in a received bitstream encoded with an input signal, or it may be received as a separate signal or bitstream.

[0038] One embodiment provides a technique for efficiently performing calculations performed by the method according to the present invention. For this purpose, the hardware implementation includes a prenormalizer or pre-normalizer that readjusts (rescales) the amplitude of the corresponding input stream in a portion of the Y frame on which the frame of the sample to be processed is based. After such readjustment, the processed sample can be calculated as a (weighted) complex product of the readjusted input sample, or possibly the unreadjusted input sample. Input samples that appear as readjusted factors in the product do not usually need to appear as unreadjusted factors. With a possible exception relating to the phase correction parameter θ, it is possible to compute equation (13) as a product of (possibly rescaled) complex input samples. This is advantageous in terms of computational burden compared to treating the amplitude and phase of the processed sample separately.

[0039] In one embodiment, a system set to Y=2 has two block extraction units that perform the formation of one frame of input samples in parallel.

[0040] In another embodiment for Y≧3, the system has a plurality of subband processing units, each of which determines an intermediate composite subband signal using various subband transposition factors and / or various subband stretching factors and / or transition methods based on mutual product or different from direct. The plurality of subband processing units may be arranged in parallel and operate in parallel. In this embodiment, the system further has a combining unit located downstream of the subband processing units and upstream of the combining filter bank. The combining unit combines the relevant intermediate composite subband signals (e.g., by combining them together) to produce a composite subband signal. As described above, the intermediate composite subband to be combined may be obtained by both direct and mutual product-based harmonic transposition. The system according to one embodiment may further have a core decoder that decodes the bitstream into an input signal. This forms an HFR processing unit configured to apply spectral band information, in particular by performing spectral shaping. The operation of the HFR processing unit may be controlled by the information encoded in the bitstream.

[0041] One embodiment provides high frequency frequency (HFR) for multidimensional signals in a system that reproduces audio signals in a stereo format forming Z channels such as left, right, center, and surround. In one embodiment that processes an input signal with multiple channels, the stretching factor S and transposition factor Q for each band may differ between channels, but each channel is based on the same number of input samples. For this purpose, the embodiment has an analysis filter bank that generates Y analog subband signals from each channel, a subband processing unit that generates Z subband signals, and Z time-stretched and frequency-transposed signals (forming the output signal).

[0042] In modifications of the above embodiment, the output signal may have output channels based on a different number of analysis subband signals. For example, it is desirable to allocate more computational resources to the HFR of acoustically prominent channels, and for example, it is desirable that multiple channels reproduced from an audio source in front of the listener be surround or near surround channels.

[0043] It should be noted that the present invention relates to all combinations of the above features, even if they are described in different claims within the patent claims.

[0044] <Overview of Drawings> Hereinafter, embodiments of the present invention, which do not limit the scope or spirit of the invention, will be described with reference to the attached drawings.

[0045] Figure 1 illustrates the principle of harmonic transposition based on subband blocks.

[0046] Figure 2 shows the nonlinear subband blocking process for a single subband input.

[0047] Figure 3 shows the nonlinear subband blocking process for two subband inputs.

[0048] Figure 4 shows the behavior of harmonic transposition based on the improved cross-product enhanced subband block.

[0049] Figure 5 shows an example of an application in an improved HFR audio coder where transposition is performed based on subbandblocks using several orders of transposition.

[0050] Figure 6 shows an example application of performing transposition based on multiple subband blocks using a 64-band QMF analysis filter bank.

[0051] Figure 7 illustrates the results of using the disclosed subbandblock-based transposition method.

[0052] Figure 8 illustrates the results of using the disclosed subbandblock-based transposition method.

[0053] Figure 9 shows in detail the nonlinear processing unit (including the pre-normalization unit and multiplication unit) shown in Figure 2.

[0054] <Description of Preferred Embodiments> The embodiments described below merely illustrate the principles of the present invention relating to harmonic transposition based on improved mutual product subbandblocks. It will be understood that variations and modifications of the apparatus, methods and specific details described herein will be apparent to those skilled in the art. Accordingly, the present invention is intended to be defined solely by the appended claims and not by the specific details shown in the description of the specification and drawings.

[0055] Figure 1 illustrates the operating principle of subband-block based transposition, time stretching, or a combination of transposition and time stretching. The input time-domain signal is fed to the analysis filter bank 101, which provides multiple complex-valued subband signals (complex subband signals). These are fed to the subband processing unit 102, and the operation of the subband processing unit 102 is controlled by control data 104. Each output subband may be obtained by processing one or two input subbands, or as a superposition of several subbands processed in this way. The multiple complex-valued output subbands (complex subband signals) are fed to the synthesis filter bank 103, which outputs a modified time-domain signal. Selective control data 104 indicates the method and parameters of subband processing performed on the signal to be transposed. In the case of improved cross-product transposition, the data includes information about the dominant fundamental frequency.

[0056] Figure 2 illustrates the operation of nonlinear subband blocking when there is one subband input. Using the target values ​​for physical time stretching and transposition, and the physical parameters of the analysis filter bank 101 and the synthesis filter bank 103, not only the parameters for subband time stretching and transposition, but also the source subband index is derived for each target subband index. The purpose of subband blocking is to generate a target subband signal by performing transposition, time stretching, or a combination of transposition and time stretching corresponding to the complex-valued source subband signal.

[0057] The block extraction unit 201 samples a finite number of frames from the input complex signal. The frames are defined by the input pointer position and the subband transposition factor. These frames undergo nonlinear processing by the processing unit 202, followed by window processing by the window processing unit 213, which performs window processing of a finite, possible variable length. The resulting samples are pre-added to the output samples in the overlap addition unit 204, and the output frame position is defined by the output pointer position. The input pointer is incremented by a fixed value, and the output pointer is incremented by the amount obtained by multiplying this fixed value by the subband stretching factor. Through the repetition of this series of processes, an output signal is generated with a duration equal to the subband stretching factor multiplied by the input subband signal duration, along with the complex frequency transposed by the subband transposition position, and this duration is within the length of the composite window. The control signal 104 affects (controls) each of the three processing units 201, 202, and 203.

[0058] Figure 3 illustrates the operation of nonlinear subband blocking when there are two subband inputs. Using the target values ​​for physical time stretching and transposition, and the physical parameters of the analysis filter bank 101 and the synthesis filter bank 103, not only the parameters for subband time stretching and transposition, but also the source subband index is derived for each target subband index. If the nonlinear subband blocking is to generate the missing portion by cross product addition, then not only the settings of the processing units 301-1, 301-2, 302, and 303, but also the values ​​of the two source band indices depend on the output 403 of the cross processing control unit 404. The purpose of subband blocking is to generate a target subband signal by performing corresponding transposition, time stretching, or a combination of transposition and time stretching on the two complex source subband signals. The first block extraction unit 301-1 samples a finite time frame from the first complex source band, and the second block extraction unit 301-2 samples a finite frame from the second complex source band. The frames are defined by a common input pointer position and a subband transposition factor. These two frames proceed to the nonlinear processing unit 302, where they are then windowed by the window processing unit 303 using a finite-length window. The overlap summing unit 204 is identical or similar to that shown in Figure 2. Through iteration of this series of processes, an output signal is generated with a duration equal to the length of the longer of the two subband signals (but within the length of the composite window) multiplied by a subband stretching factor. If the two input subband signals have the same frequency, the output signal will have a complex frequency transposed by the subband transition factor. If the two subband signals have different frequencies, the window processing unit 303 can be used to generate an output signal with a target frequency suitable for generating the missing portion in the transposed signal.

[0059] Figure 4 is a diagram illustrating the principles of transposition, time stretching, or a combination of transposition and time stretching based on improved cross-product subbandblocking. The direct sub-band processing unit 401 may have already been described with reference to Figure 2 (processing unit 202) or Figure 3. The cross sub-band processing unit 402 performs nonlinear subbandblocking on the two subband inputs shown in Figure 3, and the output target subband is added with that from the direct subband processing unit 401 in the summing unit. The cross-processing control data 403 is different for each input pointer position, and • Selected list of target subband indices • Each selected target subband index is a pair of source subband indices, and • Finite-length composition window It includes at least information indicating this.

[0060] The cross-processing control unit 404 provides cross-processing control data 403 based on a portion of the control data 104 that shows multiple complex subband signals output from the analysis filter bank 101 and the fundamental frequency. The control data 104 also includes other signal-dependent setting parameters that affect the cross-product process.

[0061] The principles of time stretching and transposition based on improved mutual product subbandblocks will be explained below, along with appropriate mathematical methods, with reference to Figure 1-4.

[0062] The two main setting parameters for the harmonic transposer and / or time stretching as a whole are: ·S φ : Desired physical time stretching factor, and Q φ : Desired physical transposition factor That is the case.

[0063] Filter banks 101 and 103 may be of any complex exponential modulated type, such as QMF, windowed DFT, or wavelet transform. The analytical filter bank 101 and the composite filter bank 103 are stacked in even-numbered or odd-numbered bands during modulation, and are defined from a wider range of prototype filters and / or windows. All of these quadratic selections affect subsequent design details such as phase correction and subband mapping management, but the main system design parameters for subband processing are generally the following four filter bank parameters (all measured in physical units): Δt s / Δt A and Δf s / Δf A It is derived from the following two quotients. In the above quotient, ·Δt A This is the subband sample time step or temporal stride, progression length, step size, or step length of the analysis filter bank 101 (for example, measured in seconds), ·Δf A This is the subband frequency interval of the analysis filter bank 101 (e.g., measured in Hertz [1 / s]), ·Δt s This is the subband sample time step or temporal stride, progression length, step size, or step length of the synthetic filter bank 103 (e.g., measured in seconds), ·Δf s is the subband frequency interval of the composite filter bank 103 (measured, for example, in Hertz [1 / s]).

[0064] Based on the configuration of the subband processing unit 102, the following parameters should be calculated: • S: Subband stretching factor. The subband stretching factor is applied to the subband processing unit 102 as the ratio of input to time sample, S φ This is intended to perform an overall physical time stretching of time-domain signals.

[0065] • Q: Subband transposition factor. The subband transposition factor is applied to the subband processing unit 102, and factor Q φ This is intended to perform an overall physical frequency transposition of a time-domain signal.

[0066] • Correspondence between source and target subband indices. n represents the index of the analyzed subband that enters the subband processing unit 102, and m represents the index of the corresponding synthesized subband in the output of the subband processing unit 102.

[0067] To determine the subband expansion factor S, the input signal to the analysis subband during physical period D is used as input to the subband processing unit 102, where D / Δt is the number of analysis subband samples. A Confirm that these correspond to D / Δt. A Each sample is processed by a subband processing unit 102 that applies a subband stretching factor S, S·D / Δt A It is extended to the output of the composite filter bank 103, these S·D / Δt A The number of samples, Δt s ·S·D / Δt A The output signal has a physical period of length S. φ Since it must match a specific value called D, that is, the duration of the time-domain output signal is determined by the physical time extension factor S. φ Since it should be extended for time-domain input signals, the following design rules are obtained.

[0068]

number

[0069]

number

[0070]

number

[0071] Below, the subband processing shown in Figure 2 for one source subband is described as a function of the subband processing parameters S and Q. x(k) is the input signal to the block extraction unit 201, and h is the stride, progression length, step length, or step size of the input block. That is, x(k) is the complex analysis subband signal of the analysis subband with index n. The block extracted by the block extraction unit 201 is generally considered to be defined by L=R1+R2 samples, so there is no loss.

[0072]

number

[0073] An interesting special case of equation (4) is when R1=0 and R2=1, where the extracted block consists of one sample, i.e., the block length L is L=1.

[0074] In the case of the polar coordinate representation of a complex number z = |z|exp(j∠z), |z| represents the amplitude of a complex number, and ∠z represents the phase or phase angle of a complex number. The nonlinear processing unit 202 that generates the output frame yl from the input frame xl is advantageously defined by the phase correction factor T = SQ given by the following formula.

[0075]

number

[0076] Equation (5) above shows that the phase of an output frame sample is determined by shifting the phase of the corresponding input sample by a certain offset value. This constant offset value depends on a correction factor T, which itself depends on the subband stretching factor and / or subband transposition factor. Furthermore, the constant offset value depends on the phase of a sample in a particular input frame among the input frames. This particular input frame sample is fixed when determining the phase of all output frame samples in a given block. In the case of equation (5), the phase of the central sample of the input frame is used as the phase of the sample in the particular input frame.

[0077] The second line of equation (5) shows that the amplitude of a sample in the output frame depends on the amplitude of the corresponding sample in the input frame. Furthermore, the amplitude of a sample in the output frame may depend on the amplitude of a specific input frame sample. That specific input frame sample may be used when determining the amplitude of all output frame samples. In the case of equation (5), the center sample of the input frame is used as the specific input frame sample. In one embodiment, the amplitude of a sample in the output frame may correspond to the geometric mean of the amplitudes of the corresponding sample in the input frame and the specific input frame sample.

[0078] In the windowing processing unit 203, a window w of length L is applied to the output frame, and the following output frame with windowing applied is obtained.

[0079]

number

[0080]

number

[0081] When a complex sine wave is used as the input to the subband processing unit 102, the analyzed subband signal corresponds to the complex sine wave.

[0082]

number

[0083]

number

[0084]

number

[0085]

number

[0086] The following explanation of the subband processing unit will be extended to apply to the case shown in Figure 3, where there are two subband inputs. (1) (k) is the input subband signal for the first block extraction unit 301-1, and x (2) Let (k) be the input subband signal for the second block extraction unit 301-2. Since each extraction unit can use a different downsampling factor, the extracted blocks are as follows:

[0087]

number

[0088]

number

[0089] The definitions of the non-negative real parameters D1, D2, ρ, the non-negative integer parameters T1, T2, and the composite window w depend on the desired operating mode. When the same subband is given to both inputs, x (1) (k) = x (2) (k) When D1=Q, D2=0, T1=1, and T2=T-1, it should be noted that the processing related to equations (12) and (13) reduces to equations (4) and (5) for the case of one input.

[0090] In one embodiment, the frequency interval Δf of the composite filter bank 103 s and the frequency interval Δf of the analysis filter bank 101 AIf the ratio differs from the desired physical transposition factor Q, it is useful to determine a sample of a composite subband with index m from two analytical subbands with indices n and n+1, respectively. For a given index m, the corresponding index n is given by an integer value obtained by truncating the analytical index value n given by equation (3). For example, one analytical subband signal, such as the analytical subband signal corresponding to index n, is given to the first block extraction unit 301-1, and the other analytical subband signal, such as the analytical subband signal corresponding to index n+1, is given to the second block extraction unit 301-2. Based on these two analytical subband signals, the composite subband signal corresponding to index m is determined according to the above process. The way in which the two block extraction units 301-1 and 301-2 specify adjacent analytical subband signals may be based on the remainder obtained when truncating the index value of equation (3), that is, on the difference between the extraction index value given by equation (3) and the truncated integer value n obtained from equation (3). If the remaining value is greater than 0.5, the analysis subband signal corresponding to index n is assigned to the second block extraction unit 301-2; otherwise, the analysis subband signal may be assigned to the first block extraction unit 301-1. In this operating mode, the parameters are designed so that the input subband signals share the same complex frequencies.

[0091]

number

[0092]

number

[0093] The following describes the method for cross-processing control 404. Given an output subband index m, the parameters r = 1, ..., Q φ For -1 and the fundamental frequency Ω0, approximate source subband indices n1 and n2 can be approximated by approximately solving the following equations.

[0094]

number

[0095] Under these definitions, the following equation holds:

[0096] p = Ω0 / Δf A :Fundamental frequency measured in units of frequency interval of the analysis filter bank, F=Δf s / Δf A : The quotient of the frequency interval of the composite filter bank relative to the frequency interval of the analysis filter bank. · n f =[(m+σ)F-rp] / Q φ -σ: Real-valued target for source indices with low integer values.

[0097] A concrete example of a favorable approximate solution to equation (16) is to change n1 to n f Let n2 be the integer closest to n f It can be obtained by selecting the integer closest to +p.

[0098] If the fundamental frequency is smaller than the analysis filter bank interval, i.e., p < 1, it is advantageous to cancel out or offset the sum of the cross products.

[0099] As taught in Patent Document 2, cross products should not be added to output subbands from transpositions without cross products, where a significantly large contribution has already been obtained. Furthermore, in at most one case, r=1,...,Q φ -1 should contribute to the cross-product output. Here, these rules may be made by performing the following three steps for each of the target output subband indices m: 1. Amplitude |x of the candidate source subband calculated at the central time slot k=hk (1) | and |x (2) For all r=1,...,Q, the minimum value of | φ -1, with the maximum value M C Calculate the source subband x. (1) and x (2) These are given as indices n1 and n2 in equation (16).

[0100] 2. Index n ≈ (F / Q) φ Calculate the corresponding magnitude or amplitude Ms for the direct source term |x| obtained from the source subband along with m (see Equation 3).

[0101] 3. Only if Mc > qMs as described above, then M in point 1 (step 1) above. C A cross-term is selected from the remaining candidates. Here, q is a predetermined threshold.

[0102] Modifications of the above procedure are desirable to depend on specific system configuration parameters. One such modification is to change the fixed threshold in point 3 (step 3) to M C / M S This involves substituting with a relaxed rule that depends on the quotient of Q. Another variation is to maximize at point 1 (step 1) Q φ This involves extending the range beyond -1, for example, to a finite list of candidate values ​​for the fundamental frequency measured in an analytical frequency interval unit p. Another variation involves using a different quantity for the subband amplitude, for example, the amplitude of a fixed sample, the maximum amplitude, the average amplitude, l p Amplitudes based on norms may also be used.

[0103] The list of target subbands m selected to be added to the cross product, along with the values ​​n1 and n2, forms the main part of the cross-processing control data 403. The remaining discussion concerns the setting parameters or configuration parameters D1, D2, ρ, the non-negative integer parameters T1, T2 appearing in the phase rotation (13), and the synthesis window w used in the cross-subband processing unit 402. Using a sinusoidal model for the cross product situation, the following source band signals are obtained.

[0104]

number

[0105]

number

[0106]

number

[0107]

number

[0108]

number

[0109] It should be noted that the above algorithm for calculating cross-processing control data 403 based on input parameters such as the target output subband index m and the fundamental frequency Ω0 merely illustrates the nature of the present invention and does not limit its scope. Modifications of the present disclosure—for example, another subband-blocking method that provides a signal such as output (18) in response to an input signal (17)—are also within the scope of the present invention, as is common technical knowledge and everyday experience of those skilled in the art.

[0110] FIG. 5 shows a specific example of applying sub-band block-based transposition using some degree of transposition in an improved HFR audio codec. The transmitted bitstream is received by a core decoder 501, and the core decoder provides a low-bandwidth decoded core signal at a sampling frequency of fs. The low-bandwidth decoded core signal is resampled (resampled) to an output sampling frequency of 2fs by a complex modulation 32-band QMF analysis bank 502, and after the complex modulation 32-band QMF analysis bank 502, a 64-band QMF synthesis bank (inverse QMF, IQMF) 505 follows (via the HFR processing unit). The two filter banks 502 and 505 share the same physical parameters Δt s =Δt A and Δf s =Δf A and the HFR processing unit 504 passes the unmodified low sub-bands corresponding to the low-bandwidth core signal. By the spectral shaping and modification performed by the HFR processing unit 504, high-frequency components of the output signal are obtained by providing the high-frequency sub-bands of the 64QMF synthesis bank 505 together with the output bands from the multiple transposer processing unit 503. The multiple transposer processing unit 503 takes the decoded core signal as an input and outputs a plurality of sub-band signals, and the plurality of sub-band signals represent a 64QAM band analysis by superposition or summation of a plurality of transposed signal components. The purpose or policy is that when the HFR processing is bypassed or bypassed, each of the signal components corresponds to an integer physical transposition (Q φ =2, 3,... and S φ =1) without time stretching of the core signal. In an embodiment of the present invention, the transposer control signal 404 includes data indicating the fundamental frequency. This data may be transmitted by the bitstream from the corresponding audio encoder (the decoder performs pitch detection), or may be obtained from a combination of transmitted and detected information.

[0111] FIG. 6 is a diagram for explaining the operation of a multi-subband block-based transposition that applies a single 64-band QMF analysis filter bank. Three transpositions or orders Q φ = 2, 3, 4 are generated and given in the region of a 64-band QMF operating at an output sampling rate of 2fs.

[0112] The multiplexing unit, synthesis unit or merging unit 603 selects and synthesizes related subbands from one transposition factor branch of a plurality of QMF subbands given to the HFR processing unit. The specific purpose or policy is that a series of processes of the 64-band QMF analysis unit 601, subband processing unit 602-Q φ 、64-band QMF synthesis unit 505 results in a physical transposition of Q φ = 1 (i.e., without scaling) together with Q φ Identifying these three blocks together with 101, 102, 103 in FIG. 1, Δt s / Δt A = 1 / 2 and F = Δf s / Δf A = 2, so that Δt A = 64f s and Δf A = f s / 128. The design of the specific setting parameters for 602-Q φ will be described separately for each of Q φ = 2, 3, 4. In all cases, the analysis stride is selected to be h = 1, and the normalized fundamental frequency parameter p = Ω0 / Δf A = 128Ω0 / f s is assumed to be known.

[0113] First, Q φLet us consider the case where =2. In this case, 602-2 must perform subband stretching with S=2 and subband transposition with Q=1 (i.e., no stretching), and the correspondence between source n and target subband m is given by n=m for direct subband processing. In the process of mutual product addition, there is only one mutual product to consider (i.e., r=1) (see equation (15) and subsequent equations above), and equation (20) is simplified to T1=T2=1 and D1+D2=1. One example of a solution is to select D1=0 and D2=1. For the direct processing composite window, a rectangular window of length L=10 with R1=R2=5 may be used to satisfy condition (10). For the cross processing composite window, a short tap window of L=2 with R1=R2=1 is used to minimize the additional complexity of mutual product addition. Furthermore, the advantageous effect of using long blocks for subband processing is most pronounced in the case of complex audio signals, in which case unwanted intermodulation terms are suppressed, and the probability of such artifacts occurring is low for the dominant pitch. The tap window L=2 is the smallest possible to satisfy equation (10), since h=1 and S=2. However, the present invention can also satisfy equation (21). In that case, the parameters are defined as follows:

[0114]

number

[0115] Q φ When =3, the specification or procedure for 602-3 according to equations (1)-(3) is to perform subband extension of S=2 and subband transposition of Q=3 / 2, and the relationship between the target m subband and source n subband regarding the processing of the direct terms is given by n≈2m / 3. There are two types of mutual product terms r=1,2, and equation (20) is simplified as follows.

[0116]

number

[0117] • D1=0 and D2=3 / 2 (when r=1) • D1 = 3 / 2 and D2 = 0 (when r = 2) For a direct processing and blending window, a rectangular window of length L=8 may be used with R1=R2=4. For a cross-processing and blending window, a short window of L=2 taps may be used with R1=R2=1, satisfying the following equation.

[0118]

number

[0119] Q φ When =4, the specification or action of 602-4 according to equations (1)-(3) is to perform subband extension with S=2 and subband transposition with Q=2, and the relationship between the target m subband and source n subband regarding the processing of the direct terms is given by n≈2m. There are three types of mutual product terms r=1,2,3, and equation (20) is simplified as follows.

[0120]

number

[0121] • D1=0 and D2=2 (when r=1) • D1=0 and D2=1 (when r=2) • D1=2 and D2=0 (when r=3) For a direct processing composite window, a rectangular window with length L=6 may be used along with R1=R2=3. For a cross processing composite window, a short window with L=2 taps may be used along with R1=R2=1, satisfying the following equation.

[0122]

number

[0123] In each of the above examples where a value of r greater than 1 is applicable, there are options similar to the three-step procedure described earlier, for example, before equation (17).

[0124] Figure 7 shows the amplitude spectrum of a harmonic signal with a fundamental frequency Ω0 = 564.7 Hz. The low-frequency portion 701 of this signal is used as the input to multiple transposers. The purpose of the transposers is to generate a signal that is as close as possible to the high-frequency portion 702 of the input signal, so that transmission of the high-frequency portion 702 is not essential and the available bitrate can be used economically.

[0125] Figure 8 shows the amplitude spectrum of the output from a transposer that takes the low-frequency component 701 of the signal in Figure 7 as input. As explained in Figure 5, multiple transposers are constructed using a 64-band QMF filter bank with an input sampling frequency fs = 14400 Hz. However, for the sake of simplification, we have two transposition orders Q φ We will only consider values ​​2 and 3. The three different spectra 801-803 represent the final output obtained using cross-processing control data with different settings.

[0126] The upper spectrum 801 shows the output spectrum obtained when all cross-processing is canceled and only direct subband processing 401 is performed. This is when the cross-processing control data 404 receives p=0 (instruction without pitch). Q φ The transposition of =2 generates an output in the range of 4 to 8 kHz, Q φ The transposition of =3 generates an output in the range of 8 to 12 kHz. As illustrated, the generated portion is far removed, and the output deviates significantly from the (original) high-frequency portion 702. Audible 2x and 3x "ghost pitch" artifacts occur in the resulting audio output.

[0127] In the middle section of spectrum 802, cross-processing is performed and a pitch parameter p=5 is used (approximately equal to 128Ω0 / fs=5.0196), but although equation (10) is satisfied, a simple two-tap composite window w(0)=w(-1)=1 is used for cross-subband processing. This is due to a direct combination of subband block-based processing and improved cross-product harmonic transposition. As illustrated, additional output signal components not present in 801 do not match the desired harmonic sequence. This indicates that using the procedure described above results in insufficient audio quality to offset the effects of the direct subband processing design by cross-product processing.

[0128] The lower spectrum 803 shows an output spectrum similar to the middle spectrum 802, but the Q in Figure 5 φ The difference lies in the use of a cross-subband processing synthesis window given by formulas relating =2,3. That is, a two-tap window with w(0)=1 and w(-1)=exp(iα) satisfies formula (21), employing the features of the present invention that depend on the value of p. As shown in the figure, the synthesized output signal is well matched to the desired harmonic portion 702.

[0129] Figure 9 shows a portion of the nonlinear processing frame processing unit 202. The nonlinear processing frame processing unit 202 receives two input samples u1 and u2 and generates a processing sample (processing sample) w based on them. The amplitude of the processing sample is given by the geometric mean of the amplitudes of the input samples, and the phase of the processing sample is a linear combination of the phases of the input samples. That is, it can be expressed as follows:

[0130]

number

[0131] Further embodiments relating to the present invention will be obvious to those skilled in the art upon understanding the above description. Although this description and drawings illustrate embodiments and specific examples, the present invention is not limited to these specific examples. Numerous modifications and variations are possible without departing from the scope of the present invention as defined by the appended claims.

[0132] The systems and methods disclosed herein may be implemented as software, firmware, hardware, or a combination thereof. All or part of the elements may be implemented as software executed by a digital signal processor or microprocessor, or as hardware or as an application-specific integrated circuit. Such software may be stored on a computer-readable storage medium, the storage medium being a concept that includes computer-readable media (or non-temporary media), but the medium itself includes communication media (temporary media). As is known to those skilled in the art, computer storage media include volatile media, non-volatile media, removable media, non-removable media, etc., and are implemented by any method or technique for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media may be, but are not limited to, RAM, ROM, EEPROM, flash memory or other types of memory, CD-ROM, digital versatile disk (DVD) or other optical media, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible by a computer. Furthermore, those skilled in the art will understand that the communication medium may typically be implemented by computer-readable instructions, data structures, or program modules, or by other data in a modulated data signal such as a carrier wave or transmission means, and may include any information transport means. [Prior art documents] [Patent Documents]

[0133] [Patent Document 1] International Publication No. 98 / 57436 [Patent Document 2] International Publication No. 2010 / 081892 [Patent Document 3] International Publication No. 2004 / 097794 [Patent Document 4] International Publication No. 2007 / 085275

[0134] (Note 1) A signal generation system that generates a time-stretched signal and / or a frequency-transposed signal from an input signal, An analysis filter bank is derived from the input signal, each of which Y (Y≧1) analysis subband signals has multiple complex analysis samples with phase and amplitude, and Y analysis subband signals are derived from the input signal. A subband processing unit that generates a composite subband signal from the Y analyzed subband signals using a subband transposition factor Q and a subband stretching factor S, A composite filter bank that generates the aforementioned time-stretched signal and / or frequency-transposed signal from the composite subband signal. The subband processing unit has a block extraction unit, a nonlinear frame processing unit, and an overlap addition unit, The aforementioned block extraction unit is i) Generate Y frames from L input samples, each of which is extracted from multiple complex analysis samples of the analyzed subband signal, and the length of each frame is L (L > 1). ii) Before generating subsequent frames of L input samples, a series of frames of input samples are generated by applying a block hop size of h samples to multiple complex analysis samples. The nonlinear frame processing unit determines the phase and amplitude of each processed sample (processed sample) of the frame, and generates a frame of the processed sample based on the Y corresponding frames of the input sample generated by the block extraction unit, and for at least one processed sample, i) The phase of the processed sample is based on the phase of each corresponding input sample in each of the Y frames of the input sample, ii) The amplitude of the processed sample is based on the phase of each corresponding input sample in each of the Y frames of the input sample. The overlap summing unit generates the composite subband signal by adding samples from a series of frames of the processing sample while overlapping them. The signal generation system is a signal generation system that operates at least when Y=2. (Note 2) The aforementioned analysis filter bank is one of the following: an orthogonal mirror filter bank, a windowed discrete Fourier transform, or a wavelet transform. The signal generation system described in Appendix 1, wherein the composite filter bank is a corresponding inverse filter bank or transform. (Note 3) The signal generation system as described in Appendix 2, wherein the analysis filter bank is a 64-point orthogonal mirror filter bank, and the synthesis filter bank is an inverse 64-point orthogonal mirror filter bank. (Note 4) The aforementioned analysis filter bank is the analysis time progression Δt A Apply the above input signal, The analysis filter bank has an analysis frequency interval Δf A Use n=0,...,N-1 is the analysis subband index, and the analysis filter bank has N analysis subbands. A certain analysis subband belonging to the N analysis subbands is associated with the frequency band of the input signal. The aforementioned composite filter bank is composed of a composite time progression width Δt sApply the above-mentioned composite subband signal, The composite filter bank is a composite frequency interval Δf s Use m=0,...,M-1 is the composite subband index, and the composite filter bank has M composite subbands. A signal generation system according to any one of the appendices 1-3, wherein a composite subband belonging to the M composite subbands is associated with the frequency band of the time-stretched signal and / or the frequency-transposed signal. (Note 5) The subband processing unit is formed for Y=2 and further comprises a cross-processing control unit, the cross-processing control unit controls the fundamental frequency Ω0 of the input signal and the analysis frequency interval Δf A The signal generation system described in Appendix 4 generates cross-processing control data that defines subband indices n1 and n2 related to the analyzed subband signal, such that the subband indices differ by an integer p, which is an approximation of the ratio. (Note 6) The subband processing unit is formed for Y=2 and further comprises a cross-processing control unit, the cross-processing control unit generates cross-processing control data that defines subband indices n1 and n2 related to the analyzed subband signal and the analyzed subband index m, and the subband indices are related to the approximate solution of the following equation,

number

number

number

number

number

Claims

1. A system configured to generate a time-stretched and / or frequency-transposed signal from an input signal, comprising a control data receiving unit, an analysis filter bank, a subband processing unit, and a synthesis filter bank, The control data receiving unit is configured to receive control data; The analysis filter bank is configured to derive Y (Y≧2) analysis subband signals from the input signal, and each analysis subband signal has multiple complex analysis samples, each having phase and amplitude; The subband processing unit is configured to generate a composite subband signal from the Y analyzed subband signals using a subband transposition factor Q and a subband stretching factor S, wherein at least one of Q and S is greater than 1, and the subband processing unit is configured to determine the composite subband signal taking the control data into consideration, and the subband processing unit comprises a block extraction unit, a nonlinear frame processing unit, and an overlap addition unit; The block extraction unit is configured to form Y frames of L input samples, each frame being extracted from the plurality of complex analysis samples of the analysis subband signal, where L is a frame length greater than 1, and the block extraction unit is configured to apply a block hop size of h samples to the plurality of complex analysis samples before forming subsequent frames of the L input samples, thereby generating a sequence of frames of the input samples; The nonlinear frame processing unit is configured to generate a frame of a processing sample based on Y corresponding frames of the input sample formed by the block extraction unit, by determining the phase and amplitude of each processing sample of the frame, with respect to at least one processing sample: i) The phase of the processed sample is based on the phase of each corresponding input sample in each of the Y frames of the input sample; and ii) The amplitude of the processed sample is based on the amplitude of the corresponding input sample in each of the Y frames of the input sample; The overlap summing unit is configured to determine the composite subband signal by overlapping and adding samples from a sequence of frames of the processed samples; The composite filter bank is configured to generate the time-stretched and / or frequency-transposed signal from the composite subband signal; The system is configured such that the block extraction unit derives at least one frame of the input sample by downsampling the complex analysis sample of the analysis subband signal.

2. The aforementioned analysis filter bank has an analysis time progression width Δt A Apply the above input signal, The analysis filter bank has an analysis frequency interval Δf A Use n = 0 , . . , N-1 is the analysis subband index, and when N > 1, the analysis filter bank has N analysis subbands, One of the N analysis subbands is associated with the frequency band of the input signal. The aforementioned composite filter bank is composed of a composite time progression width Δt S Apply the above-mentioned composite subband signal, The composite filter bank has a composite frequency interval Δf S Use m = 0 , . . , When M-1 is the composite subband index and M > 1, the composite filter bank has M composite subbands, The system according to claim 1, wherein one of the M composite subbands is associated with the frequency band of the time-stretched and / or frequency-transposed signal.

3. The subband processing unit is configured for Y=2 and further comprises a cross-processing control unit, the cross-processing control unit processes the analyzed subband signal and the subband index n related to the analyzed subband index m. 1 ,n 2 It is configured to generate cross-processing control data that defines the following, and the subband indices n1 and n2 are associated with approximate integer solutions to the following equations: [Math 1] Ω 0 Ω is the fundamental frequency belonging to the dominant pitch component of the input signal, and Ω is the frequency of the input signal for the analysis filter bank. σ = 0 or 1 / 2, Q = (Δt S / Δt A )Q φ where Q φ It is a transposition factor, r is 1 ≤ r ≤ Q φ The system according to claim 2, wherein the integer is -1.

4. The value of r that maximizes the minimum amplitude of the subbands of the two samples formed by extracting the analysis sample from the analysis subband signal is determined by the subband index n. 1 ,n 2 The system according to claim 3, wherein the cross-processing control unit is configured to generate cross-processing control data, based on the above.

5. The system according to claim 4, wherein the amplitude of the subband in each frame of L input samples is the amplitude of the central or near-central sample.

6. The system further comprises a plurality of subband processing units and a combining unit provided downstream of the plurality of subband processing units and upstream of the combining filter bank, Each of the plurality of subband processing units is configured to determine an intermediate composite subband signal using different values ​​of the subband transposition factor Q and / or the subband stretching factor S. The system according to any one of claims 1 to 5, wherein the combining unit is configured to combine a corresponding intermediate combining subband signal in order to determine the combining subband signal.

7. The analysis filter bank is configured to form Y × Z analysis subband signals from the input signal. The subband processing unit generates Z composite subband signals from the Y × Z analyzed subband signals, and applies pairs of S and Q values ​​to each group of Y analyzed subband signals that form the basis of a given composite subband signal. The system according to any one of claims 1 to 6, wherein the composite filter bank is configured to generate Z time-stretched and / or frequency-transposed signals from the Z composite subband signals.

8. A method performed by one or more processing units to generate a time-stretched and / or frequency-transposed signal from an input signal, Step of receiving control data; A step of deriving Y (Y≧2) analysis subband signals from the input signal, wherein each analysis subband signal has a plurality of complex analysis samples, each having a phase and amplitude; A step of forming Y frames of L input samples, wherein each frame is extracted from the plurality of complex analysis samples of the analysis subband signal, and L is a frame length greater than 1; Before deriving subsequent frames of L input samples, the step of applying a block-hop size of h samples to the plurality of complex analysis samples to generate a sequence of frames of the input samples; A step of generating a frame of a processing sample based on Y corresponding frames of an input sample and taking into account the control data, by determining the phase and amplitude of each processing sample of the frame, wherein with respect to at least one processing sample: i) The phase of the processed sample is based on the phase of each corresponding input sample in each of the Y frames of the input sample; and ii) The amplitude of the processed sample is based on the amplitude of the corresponding input sample in each of the Y frames of the input sample; A step of determining a composite subband signal by overlapping and adding samples from a sequence of frames of the processed sample; and A step of generating the time-stretched and / or frequency-transposed signal from the composite subband signal, wherein forming the frame of the input sample includes downsampling the complex analysis sample of the analysis subband signal; A method that includes this.

9. A storage medium for storing computer-readable instructions for performing the method according to claim 8.

Citation Information

Patent Citations

  • Audio signal interpolation method and audio signal interpolation device

    JP2007316254A

  • Source coding enhancement using spectral-band replication

    WO1998057436A2

  • Advanced processing based on a complex-exponential-modulated filterbank and adaptive time signalling methods

    WO2004097794A2

  • Efficient filtering with a complex modulated filterbank

    WO2007085275A1

  • Cross product enhanced harmonic transposition

    WO2010081892A2