Quantization of spatial speech parameters
By quantizing and indexing spatial audio direction parameters using averaging and Golomb-Rice coding, the method addresses the high bitrate issue in spatial audio encoding, achieving efficient transmission and storage of spatial audio metadata.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- NOKIA TECHNOLOGIES OY
- Filing Date
- 2024-12-23
- Publication Date
- 2026-05-08
AI Technical Summary
Existing spatial audio encoding technologies face challenges in efficiently encoding spatial audio parameters, particularly spatial audio direction parameters, leading to high bitrates that are not optimized for efficient transmission and storage.
A method and apparatus for quantizing and indexing spatial speech direction parameters, including averaging and weighting spatial audio direction parameters across time subframes, and using Golomb-Rice coding to reduce the number of bits required for encoding.
The proposed method significantly reduces the bitrate required for transmitting and storing spatial audio metadata, optimizing the encoding process for efficient spatial audio reconstruction.
Smart Images

Figure 0007855672000008 
Figure 0007855672000009 
Figure 0007855672000010
Abstract
Description
[Technical Field]
[0001] This application relates to an apparatus and method for sound field-related parameter coding, but not limited to time-frequency-domain direction-related parameter coding for speech encoders. [Background technology]
[0002] Parametric spatial speech processing is a field of speech signal processing in which the spatial aspects of acoustics are described using a set of parameters. For example, in parametric spatial speech acquisition from a microphone array, a typical and effective selection is to estimate a set of parameters from the microphone array signal, such as the direction of the acoustics within a frequency band and the ratio between the directional and omnidirectional portions of the acquired acoustics within the frequency band. These parameters are known to well describe the perceptual spatial characteristics of the acoustics acquired at the position of the microphone array. Therefore, these parameters can be used in the synthesis of spatial acoustics for binaural headphones, loudspeakers, or other forms such as ambisonics.
[0003] Therefore, the direct-to-total energy ratio within the frequency band is a particularly effective parameter representation for spatial speech capture.
[0004] A parameter set consisting of frequency band and directional parameters within time subframes, and energy ratio parameters within the frequency band (indicating acoustic directivity), can also be used as spatial metadata for an audio codec (which may also include other parameters such as surround coherence, spread coherence, directionality, distance, etc.). For example, these parameters can be estimated from an audio signal captured by a microphone array, and a stereo or mono signal to be transmitted using spatial metadata can be generated from the microphone array signal. The stereo signal can be encoded using, for example, an AAC encoder, and the mono signal can be encoded using an EVS encoder. A decoder can decode the audio signal into a PCM signal, process the acoustics within the frequency band (using spatial metadata), and obtain a spatial output, such as a binaural output.
[0005] The solutions described above are particularly well-suited for encoding spatial acoustics captured from microphone arrays (e.g., in mobile phones, VR cameras, or standalone microphone arrays). However, it may be desirable for such encoders to also accept input formats other than signals captured by microphone arrays, such as loudspeaker signals, audio object signals, or ambisonic signals.
[0006] The analysis of first-order Ambisonics (FOA) inputs for spatial metadata extraction is fully documented in scientific literature related to Directional Audio Coding (DirAC) and Harmonic planewave expansion (Harpex). This is because there are microphone arrays that directly provide FOA signals (more precisely, variants thereof, B-format signals), and thus analyzing such inputs has been the focus of research in this technical field. Furthermore, the analysis of higher-order Ambisonics (HOA) inputs for multi-directional spatial metadata extraction is documented in scientific literature related to higher-order directional audio coding (HO-DirAC).
[0007] Further inputs for the encoder are also multi-channel loudspeaker inputs such as 5.1 or 7.1 channel surround inputs and audio objects.
[0008] However, with regard to the components of spatial metadata, there is a high interest in the compression and encoding of spatial audio parameters (such as spatial audio direction parameters) in order to minimize the total number of bits required to represent the spatial audio parameters. Summary of the Invention
[0009] According to the first embodiment, there is a device for spatial speech coding that includes means for quantizing and indexing spatial speech direction parameters to form a quantized spatial audio difference index, wherein the spatial speech direction parameters are associated with time subframes of frequency subbands of speech frames, and for determining a quantized spatial audio difference index by calculating the difference between the quantized spatial audio direction index and the quantized average spatial audio direction index.
[0010] A quantized average spatial audio direction index can be determined by the device having means for averaging at least two spatial audio direction parameters to provide an average spatial audio direction parameter, wherein at least two spatial audio direction parameters are associated with consecutive time subframes of a preceding frequency subband, and the preceding frequency subband is a frequency subband lower than the frequency subband; and for quantizing and indexing the average spatial audio direction.
[0011] The apparatus may further comprise means for determining an initial average spatial speech direction parameter for a frequency subband, the determination being performed by weighting the average spatial speech direction parameter with a first weight; weighting the average spatial speech direction parameter associated with at least two spatial speech direction parameters from a corresponding previous frequency subband from a previous speech frame with a second weight; averaging the first weighted average spatial speech direction parameter and the second weighted average spatial speech direction parameter to provide an initial average spatial speech direction parameter for the frequency subband.
[0012] The apparatus may further comprise means for quantizing and indexing further spatial speech direction parameters and for forming quantized further spatial speech direction indices, wherein the further spatial speech direction parameters are associated with the progressing time subframes of the frequency subband, and the quantized average spatial speech direction indices may be determined by the apparatus having means for averaging the spatial speech direction parameters and the previous spatial speech direction parameters for the frequency subband, wherein the previous spatial speech direction parameters are associated with the time subframes prior to the time subframes associated with the spatial speech direction parameters, and for quantizing and indexing the average of the spatial speech direction parameters and the previous spatial speech direction parameters.
[0013] The apparatus may further comprise means for averaging a spatial speech direction parameter and at least one further spatial speech direction parameter, wherein the at least one further spatial speech direction parameter is associated with at least one further time subframe of a frequency subband; determining the variance of the spatial speech direction parameter and at least one further spatial speech direction parameter; determining an index as a ratio of the variance to the mean of the spatial speech direction parameter and at least one further spatial speech direction parameter; and comparing the index to a threshold.
[0014] When the index is below a threshold, the apparatus may include means for: quantizing and indexing the average of the spatial speech direction parameter and at least one further spatial speech direction parameter to provide a quantized average spatial speech direction index; quantizing and indexing at least one further spatial speech direction parameter to provide a quantized at least one further spatial speech direction index; and determining a quantized at least one further spatial speech difference index by calculating the difference between the quantized at least one further spatial speech direction index and the quantized average spatial speech direction index.
[0015] The apparatus may further include means for encoding further quantized spatial speech direction indices, quantized spatial speech difference indices, and quantized average spatial speech direction indices using Golomb-Rice coding.
[0016] The spatial audio direction parameter can be a spherical coordinate azimuth value.
[0017] Means for averaging may include means for converting spatial speech direction parameters from the spherical domain to Cartesian domain parameters, averaging the parameters in the Cartesian domain, and converting the averaged Cartesian domain parameters back to the spherical domain.
[0018] According to a second aspect, there is a method for spatial speech coding, which includes quantizing and indexing a spatial speech direction parameter to form a quantized spatial speech direction index, wherein the spatial speech direction parameter is associated with a time subframe of a frequency subband of a speech frame, and determining a quantized spatial speech difference index by calculating the difference between the quantized spatial speech direction index and the quantized average spatial speech direction index.
[0019] A quantized average spatial speech direction index may include providing an average spatial speech direction parameter by averaging at least two spatial speech direction parameters, wherein at least two spatial speech direction parameters are associated with consecutive time subframes of a preceding frequency subband, and the preceding frequency subband is a frequency subband lower than the frequency subband; and quantizing and indexing the average spatial speech direction.
[0020] The method may further include determining an initial mean spatial speech direction parameter for a frequency subband, the determination of which is done by weighting the mean spatial speech direction parameter with a first weight, weighting the mean spatial speech direction parameter associated with at least two spatial speech direction parameters from the corresponding previous frequency subband from a previous speech frame with a second weight, averaging the first weighted mean spatial speech direction parameter and the second weighted mean spatial speech direction parameter to provide an initial mean spatial speech direction parameter for the frequency subband.
[0021] The method may further comprise means for quantizing and indexing further spatial speech direction parameters to form a quantized further spatial speech direction index, wherein the further spatial speech direction parameters are associated with the time subframes progressing in the frequency subband, and the quantized average spatial speech direction index can be determined by averaging the spatial speech direction parameters for the frequency subband and the previous spatial speech direction parameters, wherein the previous spatial speech direction parameters are associated with the time subframes preceding the time subframes associated with the spatial speech direction parameters, and by quantizing and indexing the average of the spatial speech direction parameters and the previous spatial speech direction parameters.
[0022] The method may further include averaging a spatial speech direction parameter and at least one further spatial speech direction parameter, wherein the at least one further spatial speech direction parameter is associated with at least one further time subframe of a frequency subband; determining the variance of the spatial speech direction parameter and at least one further spatial speech direction parameter; determining an index as a ratio of the variance to the mean of the spatial speech direction parameter and at least one further spatial speech direction parameter; and comparing the index to a threshold.
[0023] When the index is less than a threshold, the method may include: quantizing and indexing the mean of the spatial speech direction parameter and at least one further spatial speech direction parameter to provide a quantized mean spatial speech direction index; quantizing and indexing at least one further spatial speech direction parameter to provide a quantized at least one further spatial speech direction index; and determining a quantized at least one further spatial speech difference index by calculating the difference between the quantized at least one further spatial speech direction index and the quantized mean spatial speech direction index.
[0024] This method may further include encoding quantized spatial speech direction indices, quantized spatial speech difference indices, and quantized mean spatial speech direction indices using Golomb-Rice coding.
[0025] The spatial audio direction parameter can be a spherical coordinate azimuth value.
[0026] Averaging may involve converting spatial speech direction parameters from the spherical domain to Cartesian domain parameters, averaging the parameters in the Cartesian domain, and converting the averaged Cartesian domain parameters back to the spherical domain.
[0027] According to a third aspect, there is a device for spatial speech coding, comprising at least one processor and at least one memory containing computer program code, wherein the at least one memory and the computer program code are configured to cause the device to use at least one processor to quantize and index spatial speech direction parameters and form a quantized spatial speech direction index, wherein the spatial speech direction parameters are associated with time subframes of frequency subbands of speech frames, and to determine a quantized spatial speech difference index by calculating the difference between the quantized spatial speech direction index and the quantized average spatial speech direction index.
[0028] A computer program product stored on a medium can cause the device to perform the actions described herein.
[0029] The electronic device may include the apparatus as described herein.
[0030] The chipset may include the apparatus as described herein.
[0031] The embodiments of this application aim to address problems associated with the existing art.
[0032] For a deeper understanding of this application, the attached drawings are to be referenced as an example. [Brief explanation of the drawing]
[0033] [Figure 1] This diagram schematically illustrates a system of an apparatus suitable for implementing several embodiments. [Figure 2] This diagram schematically shows metadata encoders according to several embodiments. [Figure 3] Figure 2 shows a flowchart illustrating the operation of a metadata encoder as depicted in several embodiments. [Figure 4] A further flowchart of the operation of the metadata encoder, as shown in Figure 2, is provided, according to several embodiments. [Figure 5] This diagram schematically shows an exemplary device suitable for implementing the shown apparatus. [Modes for carrying out the invention]
[0034] The following further details suitable devices and possible mechanisms for providing effective spatial analysis derivation metadata parameters. In the following description, the multichannel system is described in relation to an embodiment of a multichannel microphone. However, as stated above, the input format may be any suitable input format, such as a multichannel loudspeaker, ambisonic (FOA / HOA), etc. Furthermore, the output of the exemplary system is a multichannel loudspeaker configuration. However, it is understood that the output may be rendered to the user via means other than a loudspeaker. Furthermore, the multichannel loudspeaker signal may be generalized to two or more reproducible audio signals. Such systems are currently being standardized by the 3GPP standardization body as Immersive Voice and Audio Service (IVAS). IVAS is intended to be an extension to the existing 3GPP Enhanced Voice Service (EVS) codec to facilitate immersive voice and audio services over existing and future mobile (cellular) and fixed-line networks. The application of IVAS can be to the delivery of immersive voice and audio services over 3GPP fourth-generation (4G) and fifth-generation (5G) networks. In addition, the IVAS codec, as an extension of EVS, can be used in storage-transfer applications where voice and speech content are encoded and stored in files for playback. It should be understood that IVAS can be used in conjunction with other voice and speech coding technologies that have the functionality to encode samples of voice and speech signals.
[0035] The metadata may consist of, at a minimum, spherical directions (elevation, azimuth), at least one energy ratio of the resulting direction, diffuse coherence, and direction-independent ambient coherence for each time-frequency (TF) block or tile considered, in other words, for each time / frequency subband. Overall, IVAS may have a number of different types of metadata parameters for each time-frequency (TF) tile. The types of spatial speech parameters that can constitute metadata for IVAS are shown in Table 1 below.
[0036] This data can be encoded by an encoder and transmitted (or stored) so that the spatial signal can be reconstructed in a decoder.
[0037] Furthermore, in some cases, metadata-assisted spatial audio (MASA) may support up to two directions per TF tile, which would require the parameters described above to be encoded and transmitted per direction per TF tile. This could potentially double the required bitrate according to Table 1 below.
[0038] [Table 1]
[0039] This data can be encoded by an encoder and transmitted (or stored) so that the spatial signal can be reconstructed in a decoder.
[0040] The bitrate allocated for metadata in practical immersive audio communication codecs can vary considerably. A typical overall operating bitrate for a codec may leave only 2–10 kbps for spatial metadata transmission / storage. However, some further embodiments may allow up to 30 kbps or more for spatial metadata transmission / storage. The encoding of directional parameters and energy ratio components has been examined previously, along with the encoding of coherence data. However, whatever the transmission / storage bitrate allocated for spatial metadata may be, it will always be necessary to represent these parameters using as few bits as possible, especially when TF tiles may support multiple directions corresponding to different sound sources within a spatial audio scene.
[0041] The concept described below is to quantize the spatial speech direction parameters (which may include azimuth and elevation values) for a speech frame by sequentially processing the spatial speech direction parameters across multiple subframes for each frequency subband.
[0042] Therefore, the present invention stems from the consideration that the bitrate required to transmit MASA data (or spatial metadata spatial speech parameters) can be reduced by quantizing the spatial speech direction parameters of the speech frame by using as few bits as possible to facilitate the transmission and storage of the encoded speech signal.
[0043] In this regard, Figure 1 shows exemplary apparatus and systems for carrying out embodiments of the present application. System 100 is shown to have an “analysis” portion 121 and a “synthesis” portion 131. The “analysis” portion 121 is the part from receiving a multi-channel signal to encoding metadata and downmix signals, and the “synthesis” portion 131 is the part from decoding the encoded metadata and downmix signals to presenting the reproduced signal (for example, in the form of a multi-channel loudspeaker).
[0044] The input to system 100 and the “analysis” section 121 is a multi-channel signal 102. While the following example describes a microphone channel signal input, in other embodiments, any suitable input (or composite multi-channel) format may be implemented. For example, in some embodiments, the spatial analyzer and spatial analysis may be performed outside the encoder. For example, in some embodiments, spatial metadata associated with the speech signal may be provided to the encoder as a separate bitstream. In some embodiments, the spatial metadata may be provided as a set of spatial (directional) index values. These are examples of metadata-based speech input formats.
[0045] The multi-channel signal is passed to the transport signal generator 103 and the analysis processor 105.
[0046] In some embodiments, the transport signal generator 103 is configured to receive a multi-channel signal, generate a suitable transport signal containing a specified number of channels, and output the transport signal 104. For example, the transport signal generator 103 may be configured to generate a two-channel audio downmix of the multi-channel signal. The specified number of channels can be any suitable number of channels. In some embodiments, the transport signal generator is configured to otherwise select the input audio signal into a specified number of channels, or combine them, for example by a beamforming technique, and output these as the transport signal.
[0047] Depending on the embodiment, the transport signal generator 103 is optional, and the multi-channel signal is passed to the encoder 107 in the same manner as the transport signal without being processed.
[0048] In some embodiments, the analysis processor 105 is also configured to receive a multichannel signal, analyze the signal, and create metadata 106 associated with the multichannel signal and, therefore, with the transport signal 104. The analysis processor 105 may be configured to generate metadata for each time-frequency analysis interval that may include a directional parameter 108, an energy ratio parameter 110 (including a directional to total energy ratio and a diffuse to total energy ratio), and a coherence parameter 112. The directional, energy ratio, and coherence parameters may, in some embodiments, be considered spatial speech parameters. In other words, spatial speech parameters include parameters intended to characterize the sound field produced / captured by the multichannel signal (or generally two or more speech signals).
[0049] Depending on the embodiment, the generated parameters may differ for each frequency band. For example, in band X, all parameters may be generated and transmitted, while in band Y, only one of the parameters may be generated and transmitted, and in band Z, no parameters may be generated or transmitted. An example of this is that for some frequency bands, such as the highest band, parameters may not be required for perceptual reasons. The transport signal 104 and metadata 106 may be passed to the encoder 107.
[0050] The encoder 107 may include an audio encoder core 109 configured to receive a transport (e.g., downmix) signal 104 and generate a preferred encoding of these audio signals. Depending on the embodiment, the encoder 107 may be a computer (running preferred software stored in memory and on at least one processor), or alternatively, a specific device utilizing, for example, an FPGA or ASIC. Encoding may be performed using any preferred method. The encoder 107 may further include a metadata encoder / quantizer 111 configured to receive metadata and output the information in an encoded or compressed form. Depending on the embodiment, the encoder 107 may further interleave the metadata before transmission or storage, as shown by the dashed lines in Figure 1, or multiplex it into a single data stream, or embed it within the encoded downmix signal. Multiplexing may be performed using any preferred method.
[0051] On the decoder side, the received or acquired data (stream) may be received by the decoder / demultiplexer 133. The decoder / demultiplexer 133 may multiplex the encoded stream and pass the audio encoded stream to a transport extractor 135 configured to decode the audio signal and obtain a transport signal. Similarly, the decoder / demultiplexer 133 may include a metadata extractor 137 configured to receive encoded metadata and generate metadata. Depending on the embodiment, the decoder / demultiplexer 133 may be a computer (running preferred software stored in memory and on at least one processor), or alternatively, a specific device utilizing, for example, an FPGA or ASIC.
[0052] The decoded metadata and transported audio signal can be passed to the synthesis processor 139.
[0053] The “synthesis” portion 131 of system 100 further illustrates a synthesis processor 139 configured to receive transport and metadata and, based on the transport signals and metadata, to reproduce synthesized spatial speech in the form of multi-channel signals 110 (which may be in multi-channel loudspeaker format, or, depending on the embodiment, any preferred output format such as binaural or ambisonic signals, depending on the use case) in any preferred format.
[0054] Therefore, to summarize, first, the system (analysis part) is configured to receive multi-channel audio signals.
[0055] Next, the system (analysis portion) is configured to generate a suitable transport audio signal and spatial audio parameters as metadata (for example, by selecting or downmixing some of the audio signal channels).
[0056] Next, the system is configured to encode the transport signals and metadata for storage / transmission.
[0057] After this, the system can store / transmit the encoded transport signal and metadata.
[0058] The system can acquire / receive encoded transport signals and metadata.
[0059] Next, the system is configured to extract the transport signal and metadata from the encoded transport signal and metadata parameters, for example, by multiplexing and decoding the encoded transport signal and metadata parameters.
[0060] The system (synthesis component) is configured to synthesize an output multi-channel audio signal based on the extracted transport audio signal and metadata.
[0061] With respect to Figure 2, exemplary analysis processors 105 and metadata encoders / quantizers 111 (as shown in Figure 1) according to several embodiments will be described in further detail.
[0062] Figures 1 and 2 show the metadata encoder / quantizer 111 and the analysis processor 105 coupled together. However, it should be understood that in some embodiments, these two processing entities do not need to be closely coupled, so that the analysis processor 105 can reside on a different device from the metadata encoder / quantizer 111. Thus, the device containing the metadata encoder / quantizer 111 may be given transport signals and metadata streams for processing and encoding independently of the ingestion and analysis processes.
[0063] Depending on the embodiment, the analysis processor 105 includes a time-frequency domain transformer 201.
[0064] In some embodiments, the time-frequency domain converter 201 is configured to receive a multi-channel signal 102 and apply a suitable time-frequency domain transformation, such as a Short Time Fourier Transform (STFT), to convert the input time-domain signal into a suitable time-frequency signal. These time-frequency signals can then be passed to a spatial analyzer 203.
[0065] Therefore, for example, the time-frequency signal 202 can be expressed in time-frequency domain form by the following equation. s i (b,n), Here, b is the frequency bin exponent, n is the time-frequency block (frame) exponent, and i is the channel exponent. In another equation, n can be considered as a time exponent with a lower sampling rate than that of the original time-domain signal. These frequency bins can be grouped into subbands, where one or more of the bins are grouped into subbands with a bandwidth exponent k=0,...,K-1. Each subband k is the lowest bin b k,low and the highest bin b k,high It has a subband of b k,low ~b k,high It encompasses all bins. The width of the subband can approximate any suitable distribution, such as the Equivalent Rectangular Bandwidth (ERB) scale or the Burke scale.
[0066] Therefore, a time-frequency (TF) tile (or block) is a specific subband within a subframe of a frame.
[0067] It can be understood that the number of bits required to represent spatial speech parameters may depend at least in part on the TF (Time-Frequency) tile resolution (i.e., the number of TF subframes or tiles). For example, a 20ms speech frame may be divided into four 5ms time-domain subframes, each time-domain subframe may have up to 24 frequency subbands divided in the frequency domain according to the Burke scale, an approximation thereof, or any other preferred division. In this particular example, the speech frame may be divided into four time-domain subframes with 96 TF subframes / tile, in other words, 24 frequency subbands. Thus, the number of bits required to represent spatial speech parameters for a speech frame may depend on the TF tile resolution. For example, if each TF tile were encoded according to the distribution in Table 1 above, then each TF tile would require 64 bits (for one source direction per TF tile) and 104 bits (for two source directions per TF tile, considering parameters independent of source direction).
[0068] In the embodiment, the analysis processor 105 may include a spatial analyzer 203. The spatial analyzer 203 may be configured to receive time-frequency signals 202 and estimate direction parameters 108 based on these signals. The direction parameters may be determined based on arbitrary voice-based “direction” determination.
[0069] For example, in some embodiments, the spatial analyzer 203 is configured to estimate the direction of a sound source using two or more signal inputs.
[0070] Therefore, the spatial analyzer 203 may be configured to provide at least one azimuth and elevation angle (spatial speech direction parameter) for each time-frequency block within the frame of the speech signal, expressed as azimuth angle φ(k,n) and elevation angle θ(k,n). The spatial speech direction parameter 108 for the time subframe may also be passed to the spatial parameter set encoder 207.
[0071] The spatial analyzer 203 may also be configured to determine an energy ratio parameter 110. The energy ratio can be thought of as a determination of the energy of an audio signal that can be assumed to be arriving from a particular direction. The directivity-to-total energy ratio r(k,n) can be estimated, for example, using a stability index for directivity estimation, or using an arbitrary correlation index, or any other suitable method for obtaining the ratio parameter. Each directivity-to-total energy ratio corresponds to a specific spatial direction and describes how much of the total energy comes from that particular spatial direction compared to the total energy. This value can also be expressed separately for each time-frequency tile. The spatial direction parameter and the directivity-to-total energy ratio describe how much of the total energy for each time-frequency tile comes from a particular direction. Generally, the spatial direction parameter can also be thought of as the direction of arrival (DOA).
[0072] In this embodiment, the directivity-to-total energy ratio parameter can be estimated based on the normalized cross-correlation parameter cor'(k,n) between microphone pairs in band k, where the value of the cross-correlation parameter is between -1 and 1. The directivity-to-total energy ratio parameter r(k,n) is calculated by comparing the normalized cross-correlation parameter with the diffuse field normalized cross-correlation parameter (cor). D By comparing with (k,n),
[0073]
number
[0074] The spatial analyzer 203 may be further configured to determine a number of coherence parameters 112, which may include ambient coherence (γ(k,n)) and diffusion coherence (ζ(k,n)), both of which are analyzed in the time-frequency domain.
[0075] The term "audio source" may relate to the dominant direction of the propagating sound wave, which may encompass the actual direction of the sound source.
[0076] Therefore, for each subband k, there exists a set of spatial speech parameters associated with subband k and subframe n. In this case, each subband k and subframe n (in other words, a TF tile) may have the following spatial speech parameters associated with each sound source direction: at least one azimuth angle and elevation angle expressed as azimuth φ(k,n) and elevation angle θ(k,n), as well as diffuse coherence (ζ(k,n)) and directivity-to-gross energy ratio parameter r(k,n). Naturally, if there is more than one direction for each TF tile, then the TF tile may have each of the above-listed parameters associated with each sound source direction. In addition, the set of spatial speech parameters may also include ambient coherence (γ(k,n)). The parameter also includes the diffuse-to-gross energy ratio r diff (k,n) may also be included.
[0077] In this embodiment, the diffusion to total energy ratio r diff (k,n) is the energy ratio of omnidirectional acoustics across the periphery, and typically there is a single diffuse-to-gross energy ratio (and ambient coherence (γ(k,n)) for each TF tile. The diffuse-to-gross energy ratio can be thought of as the energy ratio remaining when the directional-to-gross energy ratio (for each direction) is subtracted from 1. Hereafter, the above parameters may be referred to as a set of spatial acoustic parameters (spatial acoustic parameter set) for a particular TF tile.
[0078] The spatial parameter set encoder 207 may be configured to quantize the direction parameter 108 in addition to the energy ratio parameter 110 and the coherence parameter 112.
[0079] The quantization of the direction parameters 108 (such as the azimuth angle φ(k,n) and the elevation angle θ(k,n)) can be based on an arrangement of spheres that forms a spherical lattice arranged within an annulus on the "surface" sphere defined by a look-up table defined by a determined quantization resolution. In other words, the spherical lattice uses the idea of covering a sphere with smaller spheres and considering the centers of the smaller spheres as points that define a lattice in directions of approximately equal distance. Thus, the smaller spheres define a cone or solid angle around a center point that can be indexed according to any suitable indexing algorithm. Next, the direction parameters 108 of the azimuth angle φ(k,n) and the elevation angle θ(k,n) can be associated with points, and the spherical lattice uses a vector distance metric to provide quantization indices to the spherical lattice. Such a spherical quantization scheme can be found in Patent Application Publication, International Publication Nos. WO 2019 / 091575 and WO 2019 / 129350. Alternatively, the direction parameters 108 of the azimuth angle φ(k,n) and the elevation angle θ(k,n) can be quantized according to any suitable linear or non-linear quantization means.
[0080] Thus, the result of quantizing the spatial audio direction parameters 108 of the azimuth angle φ(k,n) and the elevation angle θ(k,n) by the spatial parameter set encoder 207 is at least one azimuth quantization index I φ (k,n) and at least one elevation quantization index I θ (k,n).
[0081] FIG. 3 shows a computer software or hardware-implementable process for encoding spatial audio direction parameters (such as azimuth and elevation values) for sub-frames of a frequency band.
[0082] In some embodiments, a scheme for encoding and quantizing the spatial speech direction parameter 108 may include an initial step of finding the average spatial speech parameter over subframes (n=0:N-1) of a particular subband k. In some embodiments, this initial step may include finding the average spatial speech parameter for subframes of a first subband k=0.
[0083] With respect to the directional parameters 108 of azimuth angle φ(k,n) and elevation angle θ(k,n), this step can be performed by first finding the average azimuth angle value Avgφ(k) and average elevation angle value Avgθ(k) over N subframes of the frame, also for the subband k(k=0).
[0084] Each average sound direction parameter can be calculated by first calculating the average in Cartesian coordinates, and then converting the average Cartesian coordinates to average spherical coordinates.
[0085] In other words, the calculation of the average azimuth and elevation values for subband k involves the X-axis component x(k,n)=r(k,n)cosθ(k,n)cosφ(k,n) It has as such, where the average X-axis for the subband k is,
number
number
number
[0086] Each Cartesian coordinate can be weighted by the respective directional-to-total energy ratio parameter r(k,n) associated with the TF tile.
[0087] In other embodiments, weighting may not be performed for each Cartesian coordinate.
[0088] Next, the mean azimuth and elevation values can be determined by taking the mean Cartesian coordinate values mentioned above for the subband k and converting them to the original spherical region.
[0089] In this embodiment, this transformation may be performed using the following equation.
number
[0090] The above processing step for determining the average spatial speech direction parameter for a subframe of the frequency band is shown as processing step 301 in Figure 3.
[0091] Next, the average azimuth and elevation values for subband k are quantized as outlined above, and the quantized index I is the average audio direction index across subframes for subband k. avgφ (k) and I avgθ (k) can be given.
[0092] The processing step of quantizing and indexing the average spatial speech direction parameter for subframes of the frequency band is shown as processing step 303 in Figure 3.
[0093] In this embodiment, I avgφ (k) and I avgθ The above calculation for (k) can be performed for the first subband in the frame, i.e., subband k=0, Iavgφ (0) and I avgθ It results in (0).
[0094] Next, the coding process for the speech direction parameter (as performed by the spatial parameter set encoder 207) may determine the speech direction difference index for each subframe (n=0:N-1) in the first subband (k=0). Here, the speech direction difference index for subframe n may take the form of determining the difference between the speech direction quantization index for subframe n and the average speech direction index (as determined above) for subframes across the frame. In the case of azimuth values, this routine may take the following form: Subframe n 0 to N-1 I diffφ (0,n)=I φ (0,n)-I avgφ (0)
[0095] The same step can be applied to the elevation angle value as well. That is, Subframe n 0 to N-1 I diffθ (0,n)=I θ (0,n)-I avgθ (0)
[0096] The processing step of quantizing and indexing the spatial speech direction parameter for each subband of the frequency band is shown as processing step 305 in Figure 3. The processing step of determining the average spatial speech direction difference index for the subframes of the frequency band is shown as processing step 307 in Figure 3.
[0097] Depending on the embodiment, the coding process for the audio direction parameter of the first subband may also include a further processing step (not shown in Figure 3) of processing the audio direction difference index for the subband (k=0) so that all values are positive.
[0098] In these embodiments, the processing of the audio direction difference index into a series of positive values can be performed by the following C code. for (i = 0; i < len; i++) { if (dif_idx[i] < 0) { dif_idx[i] = -2 * dif_idx[i]; } else if (dif_idx[i] > 0) { dif_idx[i] = dif_idx[i] * 2 - 1; } else { dif_idx[i] = 0; } }
[0099] Next, to facilitate entropy-based coding (of the speech directionality indices across subframes in the first frequency band), a set of speech directionality indices can be rearranged in either ascending or descending order of magnitude.
[0100] These processing steps for the audio direction difference index involve the azimuth and / or elevation parameter I for the subband k, in this case for the first subband k=0. diffφ (0,n), I diffθ (0,n) can be performed individually for n=0:N-1.
[0101] As described above, the rearranged audio direction difference index for subframes in the first frequency subband can be encoded using entropy coding, such as Golomb-Rice coding. This coding may depend on the number of bits available for encoding the audio direction for the frame.
[0102] Other subbands (e.g., subbands k=1:K-1) may employ different approaches for encoding speech direction parameters across their respective subbands. For simplicity, subbands on which these calculations may be performed will henceforth be denoted as subband k (where k≠0).
[0103] Figure 4 shows a computer software or hardware-implementable process for encoding spatial speech direction parameters (such as azimuth and elevation values) for subframes in frequency bands other than the first frequency band within an audio frame.
[0104] This approach first measures the spatial speech parameter mean exponent for subband k=0 as determined by processing step 303, Figure 3, I avgφ (0) and I avgθ This involves taking (0) and using this value to determine the spatial speech parameter direction difference index of the first subframe (for subband k≠0) and the average index for subband k=0. In other words, the spatial speech direction difference index I diffφ (k,0), I diffθ (k,0) is I for the azimuth and elevation values, respectively. diffθ (k,0)=I θ (k,0)-I avgθ (0), and I diffφ (k,0)=I φ (k,0)-I avgφ This can be found by determining (0).
[0105] Here, I θ (k,0) and I φ (k,0) is the quantization exponent for the azimuth and elevation values for the first subframe of the k-th frequency band.
[0106] With respect to Figure 4, step 401 shows this processing step of quantizing and indexing the spatial speech direction parameter for the first subframe of subband k, where k is not the first subband in the speech frame. Furthermore, processing step 403 is the spatial speech direction difference index (I) corresponding to the spatial speech direction parameter for the first subframe of subband k. diffφ (k,0), I diffθ This shows the step of determining (k,0).
[0107] Continuing from this step, the average spatial voice direction parameter for this first subframe (n=0) can be started to be the value of the actual spatial voice direction parameter for the subframe (n=0). With respect to the azimuth and elevation values, this can be expressed as Avgφ(k,0)=φ(k,0) and Avgθ(k,0)=θ(k,0). Next, the average spatial voice direction parameter for this first subframe (n=0) can be quantized to give the average exponent of the spatial voice direction parameter for the first subframe with respect to the azimuth and elevation terms, which is I avgφ (k,0) and I avgθ It can be represented as (k,0).
[0108] In some embodiments, the average spatial speech direction parameter for this first subframe (n=0) can be averaged with the corresponding average spatial speech direction parameter from previous speech frames. This can be performed as a weighted average, where the weighting is such that the average spatial speech direction parameter from the current speech frame is given priority. In this case, a weight w (less than 0.5) may be applied to the average spatial speech direction parameter from previous speech frames, and a weight of 1-w may be applied to the average spatial speech direction parameter from the current frame. In embodiments, the averaging calculation may be performed in the Cartesian coordinate domain as outlined above.
[0109] The step of determining the average spatial speech direction parameter for the first subframe of subband k is shown as processing step 405 in Figure 4. The step of quantizing and indexing the average spatial speech direction parameter for the first subframe of subband k is shown as processing step 407 in Figure 4.
[0110] Figure 4 also shows a path 402 that displays the spatial speech direction difference index for the first subframe of the k-th subband as the output from the spatial speech direction coding process.
[0111] Following this step, the average spatial speech parameter index for the first subframe (n=0) (for subband k≠0) can be used in determining the spatial speech parameter direction difference index for a further subframe (n=1) within the same subband k. This step is shown by path 404. Processing step 417 shows setting the average spatial speech parameter index for the first subframe (n=0) for use as the average spatial speech parameter index for the next subsequent subframe (n=1).
[0112] Processing for further subframes (in this case, n=1) can take the following steps:
[0113] 1. Determine the spatial speech parameter directional difference index. In other words, I diffθ (k,1)=I θ (k,1)-I avgθ (k,0) and I diffφ (k,1)=I φ (k,1)-I avgφ By determining (k,0), the spatial speech direction difference index I diffφ (k,1), I diffθ (k,1) can be found. Here, I θ (k,1) and I φ (k,1) are the quantization indices for the azimuth and elevation values for the second subframe of the k-th frequency band, respectively. With respect to Figure 4, this step is represented by processing steps 409 and 411.
[0114] 2. Next, the average spatial audio direction parameter for a further subframe (n=1) can be determined by calculating the average of the actual spatial audio direction parameter for the subframe (n=1) and the actual spatial audio direction parameter for the previous subframe (n=0). With respect to the azimuth and elevation values, this can be expressed as Avgφ(k,1)=(φ(k,1)+φ(k,0)) / 2 and Avgθ(k,0)=(θ(k,1)+θ(k,0)) / 2. In embodiments, the averaging calculation may be performed in the Cartesian coordinate domain as outlined above. With respect to Figure 4, this step is represented by processing step 413.
[0115] 3. Next, the average spatial audio direction parameter for further subframes (n=1) can be quantized to give the average exponent of the spatial audio direction parameter for further subframes (n=1). With respect to the terms of azimuth and elevation values, this is I avgφ (k,1) and I avgθ This can be expressed as (k,1). With respect to Figure 4, the quantization and indexing of the average spatial speech direction parameter are shown as processing step 415.
[0116] Next, the spatial speech parameter mean index for a further subframe (n=1) (for subband k≠0) can be used in determining the spatial speech parameter direction difference index for yet another subframe (n=2) within the same subband k. This step is shown in Figure 4 as path 406 coupled with processing step 417.
[0117] Next, processing steps 1-3 are repeated for further subframes (e.g., n=2). This is as follows:
[0118] 1(2).I diffθ (k,2)=I θ (k,2)-I avgθ (k,1) and I diffφ (k,2)=I φ (k,2)-I avgφBy determining (k,1), the spatial speech parameter directional difference index for further subframes can be determined. Here, I θ (k,2) and I φ (k,2) are the quantization indices for the azimuth and elevation values for the third subframe of the k-th frequency band, respectively.
[0119] 2(2). Next, the average spatial audio direction parameter for further subframes (n=2) can be determined by calculating the average of the actual spatial audio direction parameter for subframe (n=2) and the actual spatial audio direction parameter for the previous subframe (n=0,1). With respect to the azimuth and elevation values, this can be expressed as Avgφ(k,2)=(φ(k,2)+(φ(k,1)+φ(k,0)) / 3 and Avgθ(k,2)=(θ(k,2)+θ(k,1)+θ(k,0)) / 3. In embodiments, the averaging calculation may be performed in the Cartesian coordinate domain as described above.
[0120] 3(2). Next, the average spatial audio direction parameter for further subframes (n=2) can be quantized to give the average exponent of the spatial audio direction parameter for further subframes (n=2). With respect to the terms of azimuth and elevation values, this is I avgφ (k,2) and I avgθ This can be expressed as (k,2). Next, the spatial speech parameter mean index for a further subframe (n=2) (for subband k≠0) can be used in determining the spatial speech parameter directional difference index for yet another subframe (n=3) within the same subband k.
[0121] It should be understood that steps 1, 2, and 3 can be repeated until the spatial speech parameter directional difference index is determined for all subframes within the frequency band k (where k ≠ 0). The process of repeating the above steps for further subframes within the frequency subband k is shown in Figure 4 by the return path 408.
[0122] Once the spatial speech parameter directional difference index is determined for at least some (or all) of the subframes of frequency band k, the spatial speech parameter directional difference index can be processed using the C code described above (not shown in Figure 4) so that all values are positive. Next, to facilitate entropy-based coding (of the speech directional difference indices across subframes of the k-th frequency band), the set of speech directional difference indices can be rearranged in either ascending or descending order of magnitude.
[0123] The output spatial audio direction difference index for each processed subframe is shown in Figure 4 as path 410.
[0124] In the embodiment, the spatial speech parameter directional difference index associated with a subframe of a subband can be coded using Golomb-Rice coding. Thus, with respect to azimuth and elevation values, the azimuth difference index for a subframe of a subband can be coded according to Golomb-Rice coding, and the elevation difference index for a subframe of a subband can also be coded according to Golomb-Rice coding. Golomb coding of the azimuth difference index and the elevation difference index can each be performed for each subband of the speech frame.
[0125] In addition, the average spatial speech parameter exponent value of the first subband (k=0) can also be Golomb-Rice coded. For example, I avgθ (0) and I avgφ (0) can also be Golomb-Rice coded.
[0126] Therefore, the spatial speech parameter stream encoded for each speech frame is the entropy-encoded average spatial speech parameter exponent value of the first subband (k=0), e.g., Golomb-Rice encoded I avgθ (0) and I avgφ (0) It should be understood that this may include entropy-encoded spatial speech parameter directional difference indices associated with each subframe of each subband of the speech frame. For example, Golomb-Rice encoded elevation and azimuth difference indices I for k=0:M-1 and n=0:N-1. diffφ (k,n), Idiffθ (k,n)
[0127] It should be understood that the above coding steps for spatial speech direction parameters, as outlined in Figures 3 and 4, can be performed for either the azimuth value in the speech frame, the elevation value in the speech frame, or both the azimuth and elevation values in the frame.
[0128] The steps described above for encoding spatial speech direction parameters, as summarized in Figure 3, may be referred to as a fixed-average encoding method for encoding spatial speech direction parameters. The steps described above for encoding spatial speech direction parameters, as summarized in Figure 4, may be referred to as an adaptive-average encoding method for encoding spatial speech direction parameters. Therefore, the method steps (fixed-average encoding method) as summarized from Figure 3 can be developed as an independent method. In addition, the method steps of Figure 4 (adaptive-average encoding method) for encoding spatial speech direction parameters can also be configured as an independent method for encoding spatial speech direction parameters for subframes of a subband. In such a case, the average spatial speech direction parameter index used as input to Figure 4 can be derived as the average of at least the spatial speech direction parameters of subframes from the previous frequency subband.
[0129] Depending on the embodiment, the method used to encode the spatial audio parameters for a subframe of a frequency subband can be selected between a fixed mean encoding method or an adaptive mean encoding method. The criterion by which one of the two methods (fixed mean or adaptive mean) can be selected may depend on the variance of the average azimuth angle value for each subband. The metric by which the selection can be made may be based on the ratio of the variance of the spatial audio direction parameters over a subframe of the subband to the average of the spatial audio direction parameters for the audio frame. Next, this metric can be compared to a threshold, such that if the calculated metric is less than the threshold, then, at this time, the fixed mean method can be used to encode the spatial audio direction parameters associated with the subframe of the frequency subband. Conversely, if the calculated metric is greater than or equal to the threshold, then, at this time, the adaptive mean method can be used to encode the spatial audio direction parameters associated with the subframe of the frequency subband.
[0130] The above method of encoding the spatial audio direction parameters of an audio frame as described above and presented in FIGS. 3 and 4 can be incorporated into a more general encoding framework that includes a number of different spatial audio direction parameter encoding mechanisms. The selection of the encoding mechanism (for encoding the spatial audio direction parameters of an audio frame) can be determined on a per audio frame basis and may depend on the bit allocation for that purpose.
[0131] The general framework can have the following pseudo-code structure.
[0132] The input to the general framework can include the quantized spatial audio direction parameters (azimuth and elevation), as well as the number of bits allowed (bits_allowed). 1. If bits_EC1 < bits_allowed, then use the process EC1 to encode the parameters a. Encode the quantized direction parameters by method EC1 2. Otherwise a. Use bandwise encoding EC2 (with a potential reduction in quantization resolution) b. If bits_EC2 < bits_allowed, then i. Encoding using EC2 c. Otherwise i. Reduce the quantization resolution ii. Use EC3 d. End if 3. End if
[0133] Here, EC1 refers to a method of encoding spatial audio direction parameters using the difference index as outlined above and presented in FIGS. 3 and 4. Methods EC2 and EC3 can refer to different methods for encoding spatial audio direction parameters. For example, EC2 can refer to a method of encoding azimuth and elevation values as described in International Application PCT / FI2020 / 050578, and EC3 can refer to a method of encoding azimuth and elevation values as described in the international application published as International Publication No. 2020 / 070377.
[0134] The metadata coder / quantizer 111 may also include an energy ratio parameter encoder that receives energy ratio parameters for each TF tile and is configured to perform a suitable compression and encoding scheme.
[0135] Similarly, the metadata coder / quantizer 111 may also include a coherence encoder that receives the ambient coherence value γ and the diffuse coherence value ζ and is configured to determine a suitable encoding for compressing the ambient and diffuse coherence values.
[0136] The encoded direction, energy ratio, and coherence values can be passed to a combiner. The combiner may be configured to receive the encoded (or quantized / compressed) direction parameter, energy ratio parameter, and coherence parameter, combine them, and generate a suitable output (e.g., a metadata bitstream that can be combined with the transport signal or transmitted or stored separately from the transport signal).
[0137] In some embodiments, the encoded data stream is passed to a decoder / demultiplexer 133. The decoder / demultiplexer 133 multiplexes and separates the encoded quantized spatial speech parameter sets for the frame and passes them to a metadata extractor 137, and the decoder / demultiplexer 133 may also, in some embodiments, extract the transport speech signal to a transport extractor for decoding and extraction.
[0138] Next, the decoded spatial audio parameters can be passed to the synthesis processor 139 to form the decoded metadata output from the metadata extractor 137 and to form the multichannel signal 110.
[0139] Figure 5 shows an exemplary electronic device that may be used as an analysis or synthesis device. The device may be any suitable electronic device or apparatus. For example, depending on the embodiment, device 1400 may be a mobile device, user equipment, tablet computer, computer, audio playback device, etc.
[0140] Depending on the embodiment, the device 1400 comprises at least one processor or central processing unit 1407. The processor 1407 can be configured to execute various program codes, such as those described herein.
[0141] In some embodiments, device 1400 includes a memory 1411. In some embodiments, at least one processor 1407 is coupled to the memory 1411. The memory 1411 can be any preferred storage means. In some embodiments, the memory 1411 includes a program code section for storing program code that can be implemented on the processor 1407. Furthermore, in some embodiments, the memory 1411 may further include a storage data section for storing data, for example, data that has been processed or is to be processed according to embodiments as described herein. The implementable program code stored in the program code section and the data stored in the storage data section can be retrieved by the processor 1407 via the memory-processor coupling at any time as needed.
[0142] In some embodiments, device 1400 includes a user interface 1405. In some embodiments, the user interface 1405 may be coupled to a processor 1407. In some embodiments, the processor 1407 can control the operation of the user interface 1405 and receive input from the user interface 1405. In some embodiments, the user interface 1405 may allow a user to input commands to device 1400, for example, via a keypad. In some embodiments, the user interface 1405 may allow a user to obtain information from device 1400. For example, the user interface 1405 may include a display configured to show information from device 1400 to the user. In some embodiments, the user interface 1405 may include a touchscreen or touch interface having the ability to both allow information to be input to device 1400 and to further display the information to the user of device 1400. In some embodiments, the user interface 1405 may be a user interface for communicating with a position determination device as described herein.
[0143] In some embodiments, the device 1400 includes an input / output port 1409. The input / output port 1409 includes a transceiver, in some embodiments. In such embodiments, the transceiver is coupled to the processor 1407 and can be configured, for example, to enable communication with other devices or electronic devices via a wireless communication network. The transceiver, or any preferred transceiver or transmitter and / or receiver means, can be configured, in some embodiments, to communicate with other electronic devices or devices via wiring or wired coupling.
[0144] The transceiver can communicate with further devices by any suitable known communication protocol. For example, depending on the embodiment, the transceiver may use a suitable universal mobile telecommunications system (UMTS) protocol, a wireless local area network (WLAN) protocol such as IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth, or an infrared data communication pathway (IRDA).
[0145] The transceiver input / output port 1409 may be configured to receive signals and, depending on the embodiment, determine parameters as described herein by using a processor 1407 that executes a preferred code. Furthermore, the device may generate preferred downmix signals and parameter outputs to be transmitted to the synthesis device.
[0146] In some embodiments, device 1400 may be employed as at least part of a synthesis device. Therefore, the input / output port 1409 may be configured to generate a preferred audio signal format output by using a processor 1407 that receives a downmix signal determined in an acquisition or processing device as described herein, and, in some embodiments, receives parameters and executes a preferred code. The input / output port 1409 may be coupled, for example, to a multi-channel speaker system and / or headphones or similar devices, to any preferred audio output.
[0147] In general, various embodiments of the present invention may be implemented in the form of hardware or dedicated circuitry, software, logic, or any combination thereof. For example, some embodiments may be implemented in hardware form, while others may be implemented in the form of firmware or software that can be executed by a controller, microprocessor, or other computing device. However, the present invention is not limited to these. Various embodiments of the present invention may be illustrated and described using block diagrams, flowcharts, or any other graphical representation, but it should be understood that these blocks, apparatus, systems, techniques, or methods described herein may, in non-limiting examples, be implemented in the form of hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers, or other computing devices, or any combination thereof.
[0148] Embodiments of the present invention may be implemented, for example, by computer software or hardware executable by the data processor of a mobile device within a processor entity, or by a combination of software and hardware. Furthermore, in this regard, it should be noted that any block of the logic flow shown in the figure may represent a program step, or an interconnected set of logic circuits, blocks, and functions, or a combination of program steps and logic circuits, blocks, and functions. The software may be stored on a physical medium such as a memory chip, or a memory block implemented within the processor, a magnetic medium such as a hard disk or floppy disk, or an optical medium such as a DVD, its data variants, or a CD.
[0149] Memory can be of any type suitable for the local technical environment and can be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. Data processors can be of any type suitable for the local technical environment and, in non-limiting examples, may include one or more of the following: general-purpose computers, dedicated computers, microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), gate-level circuits, and processors based on multi-core processor architectures.
[0150] Embodiments of the present invention can be implemented in various components, such as integrated circuit modules. Integrated circuit design is generally a highly automated process. Complex and powerful software tools are available to translate logic-level designs into ready-to-etch and formed semiconductor circuit designs on semiconductor substrates.
[0151] The program can route wires and determine the placement of components on a semiconductor chip using well-established design rules and a library of pre-stored design modules. Once the design for the semiconductor circuit is complete, the resulting design can be transmitted to a semiconductor manufacturing facility or "fab" for production in a standardized electronic format.
[0152] The above description, using exemplary non-limiting examples, provides a complete and informative description of exemplary embodiments of the invention. However, those skilled in the art will see various modifications and adaptations in consideration of the above description when read in conjunction with the accompanying drawings and claims. However, all such modifications and similar modifications of the teachings of the invention remain within the scope of the invention as defined in the accompanying claims. [Explanation of Symbols]
[0153] 100 Systems 102 Multichannel signal 103 Transport signal generator 104 Transport signal 105 Analytical Processor 106 Metadata 107 encoder 108 Directional Parameters 109 Voice Encoder Core 110 Energy ratio parameters 111 Metadata Encoder / Quantizer 112 Coherence Parameters 121 Analysis part 131 Synthesis part 133 Decoder / Demultiplexer 135 Transfer extractor 137 Metadata Extractor 139 Synthetic Processors 201 Time-Time Frequency Domain Converter 202-hour frequency signal 203 Spatial Analyzer 207 Spatial Parameter Set Encoder 1400 devices 1405 User Interface 1407 Central Processing Unit, Processor 1409 Input / Output Ports 1411 memory
Claims
1. A device for spatial speech coding, Means for quantizing and indexing the spatial speech direction parameter of a first time subframe of a frequency subband of a speech frame, and for providing a quantized spatial speech direction index for the first time subframe, Means for determining a quantized spatial speech direction difference index for a first time subframe by calculating the difference between the quantized spatial speech direction index for the first time subframe and the quantized average spatial speech direction index of a previous frequency subband of the speech frame, Means for determining the average spatial speech direction parameter for the first time subframe by taking the average of the average spatial speech direction parameter for the corresponding frequency subband of the previous speech frame and the spatial speech direction parameter of the first time subframe, and by weighting the average, Means for quantizing and indexing the average spatial speech direction parameter for the first time subframe, and for providing a quantized average spatial speech direction index to be used in the second time subframe of the frequency subband, Means for quantizing and indexing the spatial speech direction parameter of the second time subframe and for providing a quantized spatial speech direction index for the second time subframe, Means for determining a quantized spatial speech direction difference index for a second time subframe by calculating the difference between the quantized spatial speech direction index for the second time subframe and the quantized average spatial speech direction index used in the second time subframe, Means for determining an average spatial speech direction parameter for a second time subframe by averaging the spatial speech direction parameter of the first time subframe and the spatial speech direction parameter of the second time subframe, Means for quantizing and indexing the average spatial speech direction parameter for the second time subframe, and for providing a quantized average spatial speech direction index to be used in the third time subframe of the frequency subband, A device equipped with the following features.
2. The means for averaging the average spatial speech direction parameter for the corresponding frequency subband of a previous speech frame weighted by the spatial speech direction parameter of the first time subframe, Means for weighting the average spatial speech direction parameter for the corresponding frequency subband of the aforementioned prior speech frame with a first weight, Means for weighting the spatial speech direction parameter for the first time subframe with a second weight, Means for averaging the weighted average spatial speech direction parameter for the corresponding frequency subband of the preceding speech frame and the second weighted spatial speech direction parameter for the first time subframe to obtain the average spatial speech direction parameter for the first time subframe, The apparatus according to claim 1, comprising:
3. The apparatus according to claim 1 or 2, further comprising means for encoding the quantized spatial speech direction index for the first and second time subframes of the frequency subband of the speech frame, and the quantized average spatial speech direction index for the first frequency subband of the speech frame, using Golomb-Rice coding.
4. The apparatus according to any one of claims 1 to 3, wherein the spatial sound direction parameter is a spherical coordinate azimuth angle value.
5. The means for averaging is, Converting the aforementioned spatial speech direction parameter from a spherical domain to a Cartesian domain parameter, Averaging the aforementioned Cartesian domain parameters, Converting the averaged Cartesian domain parameters to the spherical domain, The apparatus according to any one of claims 1 to 4, which is performed in the Cartesian domain.
6. A device for spatial speech coding, The spatial speech direction parameter of a first time subframe of a frequency subband of a speech frame is quantized and indexed, and a quantized spatial speech direction index is given for the first time subframe. The quantized spatial directional difference index for the first time subframe is determined by calculating the difference between the quantized spatial directional index for the first time subframe and the quantized average spatial directional index of the previous frequency subband of the speech frame, Determining the average spatial speech direction parameter for the first time subframe by taking the average of the average spatial speech direction parameter for the corresponding frequency subband of the previous speech frame and the spatial speech direction parameter for the first time subframe, and weighting the average, The average spatial speech direction parameter for the first time subframe is quantized and indexed to obtain a quantized average spatial speech direction index to be used in the second time subframe of the frequency subband, The spatial speech direction parameter of the second time subframe is quantized and indexed, and a quantized spatial speech direction index is given for the second time subframe, The quantized spatial speech direction difference index for the second time subframe is determined by calculating the difference between the quantized spatial speech direction index for the second time subframe and the quantized average spatial speech direction index used in the second time subframe, The average spatial speech direction parameter for the second time subframe is determined by averaging the spatial speech direction parameter of the first time subframe and the spatial speech direction parameter of the second time subframe, The mean spatial directional parameter for the second time subframe is quantized and indexed to obtain a quantized mean spatial directional index to be used in the third time subframe of the frequency subband, A method that includes this.
7. Average the average spatial speech direction parameter for the corresponding frequency subband of the previous speech frame, weighted by the spatial speech direction parameter of the first time subframe, Weighting the average spatial speech direction parameter for the corresponding frequency subband of the aforementioned prior speech frame with a first weight, Weighting the spatial speech direction parameter for the first time subframe with a second weight, Averaging the weighted average spatial speech direction parameter for the corresponding frequency subband of the preceding speech frame and the second weighted spatial speech direction parameter for the first time subframe to obtain the average spatial speech direction parameter for the first time subframe, The method according to claim 6, including the method described in claim 6.
8. The method according to claim 6 or 7, further comprising encoding the quantized spatial speech direction index for the first and second time subframes of the frequency subband of the speech frame, and the quantized average spatial speech direction index for the first frequency subband of the speech frame, using Golomb-Rice coding.
9. The method according to any one of claims 6 to 8, wherein the spatial sound direction parameter is a spherical coordinate azimuth angle value.
10. The above averaging, Converting the aforementioned spatial speech direction parameter from a spherical domain to a Cartesian domain parameter, Averaging the aforementioned Cartesian domain parameters, Converting the averaged Cartesian domain parameters to the spherical domain, The method according to any one of claims 6 to 9, which is performed in the Cartesian domain.
11. at least, The spatial speech direction parameter of a first time subframe of a frequency subband of a speech frame is quantized and indexed, and a quantized spatial speech direction index is given for the first time subframe. The quantized spatial directional difference index for the first time subframe is determined by calculating the difference between the quantized spatial directional index for the first time subframe and the quantized average spatial directional index of the previous frequency subband of the speech frame, Determining the average spatial speech direction parameter for the first time subframe by taking the average of the average spatial speech direction parameter for the corresponding frequency subband of the previous speech frame and the spatial speech direction parameter for the first time subframe, and weighting the average, The average spatial speech direction parameter for the first time subframe is quantized and indexed to obtain a quantized average spatial speech direction index to be used in the second time subframe of the frequency subband, The spatial speech direction parameter of the second time subframe is quantized and indexed, and a quantized spatial speech direction index is given for the second time subframe, The quantized spatial speech direction difference index for the second time subframe is determined by calculating the difference between the quantized spatial speech direction index for the second time subframe and the quantized average spatial speech direction index used in the second time subframe, The average spatial speech direction parameter for the second time subframe is determined by averaging the spatial speech direction parameter of the first time subframe and the spatial speech direction parameter of the second time subframe, The mean spatial directional parameter for the second time subframe is quantized and indexed to obtain a quantized mean spatial directional index to be used in the third time subframe of the frequency subband, Setting the quantized average spatial speech direction index for the second time subframe as the quantized average spatial speech direction index for the third time subframe of the frequency subband, A computer program that stores instructions for executing a command.
Citation Information
Patent Citations
Method and apparatus for decoding audio signal, and system for processing audio signal
JP2011209741A
Apparatus and method for encoding or decoding directional audio coding parameters using quantization and entropy coding
WO2019097018A1
Determination of the significance of spatial audio parameters and associated encoding
WO2020193865A1