Decoding device, decoding method, program, and encoding device

By dividing the audio signal into envelope components and flattened waveforms, and encoding them with different numbers of bits, the problem of audio signal auditory quality and sense of presence at low bit rates in existing technologies is solved, achieving good sound quality and a sense of presence at live events at low bit rates.

CN121464480APending Publication Date: 2026-02-03SONY GROUP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480044149.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-07-05
Filing Date
2024-06-18
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Existing coding schemes struggle to maintain the auditory quality of audio signals at low bit rates, especially due to the distortion of the envelope component, which causes auditory problems. Furthermore, they are difficult to transmit and reconstruct without sacrificing the sense of presence of a live event.

Method used

The audio signal is divided into an envelope component and a flattened waveform, which are encoded with different numbers of bits. The envelope component is encoded with fewer bits, while the flattened waveform is encoded with more bits. The frequency band division signal is synthesized during decoding to generate the content signal.

Benefits of technology

It maintains good sound quality and a sense of presence even at low bit rates, and can dynamically adjust the number of encoded bits according to the content type and network status, improving encoding efficiency while maintaining the freedom to change timbre and atmosphere.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121464480A_ABST
    Figure CN121464480A_ABST
Patent Text Reader

Abstract

The present disclosure pertains to a decoding device, a decoding method, a program, and an encoding device that can maintain good sound quality even at a low bit rate. A demultiplexing unit separates, for each frequency band, first encoded data obtained by encoding an envelope component of a content signal by frequency conversion, and second encoded data obtained by encoding an envelope component of the content signal by frequency conversion from an encoded bitstream of the content signal. The second encoded data is obtained by encoding the planarized waveform of the content signal with a number of bits different from the number of bits for the envelope component. A band decoding unit of the present disclosure generates a band division signal by combining an envelope component decoded from first encoded data and a planarized waveform decoded from second encoded data. A band synthesis unit of the present disclosure generates a content signal by synthesizing band division signals of respective bands. For example, the technique according to the present disclosure can be applied to an audio signal transmission system that transmits an encoded bitstream.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to a decoding device, a decoding method, a program, and an encoding device, and particularly relates to a decoding device, a decoding method, a program, and an encoding device capable of maintaining good sound quality even at a low bit rate. BACKGROUND

[0002] In recent years, the number of held remote live events has increased with the expansion of network transmission bandwidth. The term remote live refers to real-time distribution of only video and audio recordings of a performance by a performer or a performer and an audience on site at a live event site for an entertainment event such as music or a play to an audience outside the live event site (remote audience).

[0003] In distribution such as a remote live, a high-quality audio signal needs to be simultaneously transmitted to a large number of terminals of remote audiences without losing the sense of presence at the live event site. However, existing waveform coding schemes have limitations in compression performance and it is difficult to significantly reduce the bit rate while maintaining the quality. On the other hand, according to parametric coding in which audio is modeled and synthesized as parameters, even if the bit rate can be significantly reduced, the target of processing is limited to speech and the quality of general acoustic signals is degraded, making it difficult to perform transmission and reconstruction without losing the sense of presence at the live event site.

[0004] In order to improve coding efficiency, there is a method of extracting and coding an envelope component of a signal waveform or spectrum. For example, PTL 1 discloses a technique of outputting an index by using a codebook for vector quantization of envelope information for coding.

[0005] PRIOR ART DOCUMENTS

[0006] PATENT LITERATURE

[0007] Patent Literature 1: Patent Application Publication No. 9-146593 SUMMARY

[0008] PROBLEMS TO BE SOLVED BY THE INVENTION

[0009] Meanwhile, it has recently been found that an envelope component of an audio signal significantly affects hearing. However, in existing coding schemes, the envelope component tends to be distorted at a low bit rate, resulting in a problem in hearing.

[0010] The present disclosure is made in view of such circumstances, and aims to bring about satisfactory sound quality even at a low bit rate.

[0011] SOLUTION TO PROBLEM

[0012] The decoding device according to the first aspect of the present disclosure is a decoding device including: a demultiplexing unit configured to separate, for each frequency band, first encoded data and second encoded data from an encoded bitstream of a content signal, the first encoded data being obtained by encoding an envelope component of the content signal via a frequency transform, the second encoded data being obtained by encoding a flattened waveform of the content signal using a different number of bits than a number of bits used for the envelope component; a band decoding unit configured to generate a band division signal by synthesizing the envelope component decoded from the first encoded data and the flattened waveform decoded from the second encoded data; and a band synthesizing unit configured to generate the content signal by synthesizing the band division signals of the respective frequency bands.

[0013] The decoding method according to the first aspect of the present disclosure is a decoding method including: separating, for each frequency band, first encoded data and second encoded data from an encoded bitstream of a content signal by a decoding device, the first encoded data being obtained by encoding an envelope component of the content signal via a frequency transform, the second encoded data being obtained by encoding a flattened waveform of the content signal using a different number of bits than a number of bits used for the envelope component; generating, by the decoding device, a band division signal by synthesizing the envelope component decoded from the first encoded data and the flattened waveform decoded from the second encoded data; and generating, by the decoding device, the content signal by synthesizing the band division signals of the respective frequency bands.

[0014] The program according to the first aspect of the present disclosure is a program causing a computer to execute the following processing: separating, for each frequency band, first encoded data and second encoded data from an encoded bitstream of a content signal, the first encoded data being obtained by encoding an envelope component of the content signal via a frequency transform, the second encoded data being obtained by encoding a flattened waveform of the content signal using a different number of bits than a number of bits used for the envelope component; generating a band division signal by synthesizing the envelope component decoded from the first encoded data and the flattened waveform decoded from the second encoded data; and generating the content signal by synthesizing the band division signals of the respective frequency bands.

[0015] The encoding device according to the second aspect of the present disclosure is an encoding device including: a band division unit configured to divide a content signal into band division signals of respective frequency bands; a band encoding unit configured to separate the band division signals into an envelope component and a flattened waveform, encode the envelope component via a frequency transform, and encode the flattened waveform using a different number of bits than a number of bits used for the envelope component; and a multiplexing unit configured to generate an encoded bitstream by multiplexing first encoded data obtained by encoding the envelope component and second encoded data obtained by encoding the flattened waveform.

[0016] According to a first aspect of the present application, for each frequency band, first encoded data obtained by encoding an envelope component of a content signal by frequency transform and second encoded data obtained by encoding a flattened waveform of the content signal using a different number of bits from that used for the envelope component are separated from an encoded bit stream of the content signal, a frequency band division signal is generated by synthesizing the envelope component decoded from the first encoded data and the flattened waveform decoded from the second encoded data, and a content signal is generated by synthesizing the frequency band division signals of the respective frequency bands.

[0017] According to a second aspect of the present application, a content signal is divided into frequency band division signals of respective frequency bands, the frequency band division signals are divided into an envelope component and a flattened waveform, the envelope component is encoded by frequency transform, the flattened waveform is encoded using a different number of bits from that used for the envelope component, an encoded bit stream is generated by multiplexing first encoded data and second encoded data, the first encoded data is obtained by encoding the envelope component, and the second encoded data is obtained by encoding the flattened waveform. BRIEF DESCRIPTION OF DRAWINGS

[0018] Figure 1 is a graph for describing a relationship between an envelope component and a timbre of an audio signal.

[0019] Figure 2 is a diagram showing a configuration example of an audio signal transmission system according to an embodiment of the present disclosure.

[0020] Figure 3 is a block diagram showing a functional configuration example of an encoding apparatus.

[0021] Figure 4 is a diagram for describing separation of an envelope component and a flattened waveform.

[0022] Figure 5 is a diagram showing an example of parameters according to a type of content.

[0023] Figure 6 is a diagram showing an example of a transmission format of an encoded bit stream.

[0024] Figure 7 is a flowchart for describing a flow of an encoding process of an audio signal.

[0025] Figure 8 is a block diagram showing a functional configuration example of a decoding apparatus.

[0026] Figure 9 is a block diagram showing a configuration example of a synthesis parameter determination unit.

[0027] Figure 10 This is a diagram illustrating an embodiment of parameters based on the type of content.

[0028] Figure 11 This is a block diagram illustrating a configuration embodiment of the envelope weighting unit.

[0029] Figure 12 This is a diagram illustrating an embodiment of weighting information.

[0030] Figure 13 This is a block diagram illustrating another configuration embodiment of the envelope weighting unit.

[0031] Figure 14 This is a diagram illustrating an embodiment of parameter training for a neural network.

[0032] Figure 15 It is a flowchart used to describe the decoding process of audio signals.

[0033] Figure 16 This is a block diagram illustrating another functional configuration embodiment of the decoding device.

[0034] Figure 17 This is a block diagram illustrating a configuration embodiment of the envelope processing unit.

[0035] Figure 18 This is a block diagram illustrating another configuration embodiment of the synthesis parameter determination unit.

[0036] Figure 19 This is a block diagram illustrating yet another configuration embodiment of the synthesis parameter determination unit.

[0037] Figure 20 This is a block diagram illustrating another functional configuration embodiment of the decoding device.

[0038] Figure 21 This is a block diagram illustrating another functional configuration embodiment of the encoding device.

[0039] Figure 22 This is a diagram illustrating an embodiment of the editing screen of a music production / editing tool.

[0040] Figure 23 This is a diagram illustrating an example of parameters corresponding to timbre.

[0041] Figure 24 This is a block diagram illustrating an embodiment of a computer hardware configuration. Detailed Implementation

[0042] The modes of implementing this disclosure (hereinafter referred to as implementation methods) will be described below. The descriptions will be given in the following order.

[0043] 1. Related technologies and their problems

[0044] 2. The relationship between the envelope component of an audio signal and its timbre.

[0045] 3. Overview of the technology according to the present invention and the audio signal transmission system

[0046] 4. Configuration and operation of the encoding device

[0047] 5. Configuration and operation of the decoding device

[0048] 6. Modifications of the decoding device

[0049] 7. Variations of encoding devices and music production / editing tools

[0050] 8. Summary

[0051] 9. Examples of Computer Hardware Configuration

[0052] 1. Related technologies and their problems

[0053] background

[0054] In recent years, with the expansion of network transmission bandwidth, the number of remote live events has been increasing. The term remote live event refers to a live event, such as a music or drama performance, in which video and audio recordings of the performance by the performers or audience at the event location are distributed in real time to an audience outside the event venue (remote audience).

[0055] For example, one known system displays video reflecting the actions of other remote spectators, even when they join from outside the event venue, to enhance the sense of participation or integration with the performers and other audience members during the event. Another system operates in which pre-selected remote spectators use individual cameras and microphones to record video and audio, transmit the data in real-time to the event venue, and display video of their facial expressions and movements on a monitor at the event venue, while simultaneously outputting their sound from speakers, allowing them to cheer from outside the event venue.

[0056] In events such as remote live broadcasts, it is essential to transmit high-quality audio signals simultaneously to a large number of remote audiences without sacrificing the immersive experience of the event. Furthermore, to relay responses and cheers from remote audiences back to the live event location, it is crucial that a large number of users be able to transmit data simultaneously and bidirectionally with low latency. Recently, transmitting audio data alongside high-quality video data such as 4K video has required transmitting massive amounts of data. However, even with expanded bandwidth in networks such as fiber optics and 5G, the immersive experience is severely diminished when thousands to tens of thousands of people participate simultaneously due to network congestion, insufficient bandwidth leading to connection difficulties, audio interruptions, long latency, and very low sound quality.

[0057] To avoid this degradation in transmission quality, it is necessary to significantly reduce the amount of information transmitted without sacrificing the original sense of presence and quality of the content. To achieve this reduction, numerous audio coding / compression techniques (codecs) have been researched. Recent examples of audio codecs include MPEG4 High Efficiency Advanced Audio Coding (HE-AAC) (ISO / IEC 14496-3) and the 3GPP Enhanced Voice Services (3GPP EVS).

[0058] Problems with existing coding schemes

[0059] However, existing waveform coding schemes have limitations in compression performance and struggle to significantly reduce bit rate while maintaining quality. On the other hand, parametric coding, which uses audio modeling and synthesis as parameters, can significantly reduce bit rate, but the processing target is limited to speech, and the quality of the general acoustic signal deteriorates, making it difficult to transmit and reconstruct without sacrificing the sense of presence at the live event location.

[0060] Furthermore, to improve coding efficiency, there are methods for extracting and encoding the envelope components of a signal waveform or spectrum. For example, PTL 1 discloses a technique for outputting an index using a codebook for vector quantization used to encode the envelope information.

[0061] Recent findings indicate that the envelope component of an audio signal significantly affects hearing. However, in existing coding schemes, the quality of the envelope component cannot be directly controlled, leading to interference at low bit rates and consequently, auditory problems.

[0062] Furthermore, existing encoding schemes are not designed to reproduce the subjective sense of presence or the way sound is perceived (as they are greatly influenced by envelope components), and the decoded audio signal cannot adequately reproduce the sense of presence conveying the atmosphere of a live event location. Additionally, in existing technologies, distributors and viewers cannot easily change the timbre of the transmitted atmosphere according to the environment during distribution and reproduction.

[0063] 2. The relationship between the envelope component of an audio signal and its timbre.

[0064] Studies have shown that the envelope component of a waveform is important in the reproduction of audio signals.

[0065] For example, Masayuki Takada's "Calculation Methods of Sound Quality Metrics and Its Application," the Journal of the Acoustical Society of Japan, 2019, Vol.75, No.10, pp.582 to 589, discloses a method for calculating sound quality evaluation indices based on the fact that "roughness," which indicates the coarseness of sound, or "fluctuation intensity," which indicates the fluctuation of sound, strongly influences the subjective perception of sound, such as roughness or fluctuation. Furthermore, it has been disclosed that sound roughness occurs at modulation frequencies from 15 Hz to 300 Hz, increases from 30 Hz to 150 Hz, and reaches its maximum level, particularly at approximately 70 Hz, and that fluctuation intensity occurs at modulation frequencies of 20 Hz or lower, particularly reaching its maximum level, particularly at 4 Hz to 5 Hz.

[0066] In their paper "Modeling auditory processing of amplitude modulation. I. Detection and masking with narrow-band carriers," published in the Journal of the Acoustical Society of America, vol. 102, 5Pt 1 (1997): 2892-905, Dau, T, et al. proposed a model to explain the results of such psychoacoustic experiments. Here, it is assumed that the envelope component (amplitude modulation component) of the waveform after frequency band division by an "auditory filter bank" in the human auditory system is further analyzed by a "modulation filter bank."

[0067] Therefore, the envelope component can be considered to play an important role in the reproduction of audio signals. Thus, by extracting the envelope component of the audio signal and efficiently transmitting and reproducing it, a sense of presence can be effectively reproduced with less transmitted information. Furthermore, encoding efficiency can be improved by encoding the components needed to optimize the subjective perception of sound with high precision, and by coarsely quantizing or removing unimportant envelope components based on the content or type of sound.

[0068] As mentioned above, the modulation frequency component (envelope component) plays an important role in the reproduction of audio signals, and if an unexpected component is generated at a modulation frequency of about 30 to 150 Hz, this greatly affects the "roughness," for example, the roughness of the sound will be perceived as unpleasant.

[0069] Based on this knowledge, for example, S. van de Par, S. Disch, A. Niedermeier, and B. Edler disclosed a method in their paper "Informed postprocessing for audio roughness removal for low-bitrate audio coders" published in AES Convention Paper 10534, (Oct. 2021): suppressing modulation frequency components that are generated during encoding but not present in the original sound by processing them in the spectral domain (without extracting the envelope component). As mentioned above, in the prior art, only "unnecessary modulation frequency components" are suppressed, and there is no intention to actively change the timbre by processing the modulation frequency components.

[0070] On the other hand, for example, such as Figure 1 As shown, the signal x(t) from the sound source is band-divided and envelope-separated to extract the envelope component, and some of the envelope component is processed for envelope synthesis and band synthesis, thereby altering the signal to have an appropriate timbre according to individual preferences and the type of sound source. For example, the perception of fluctuation can be altered by processing the envelope component below 20 Hz that affects "fluctuation intensity." Alternatively, roughness can be suppressed by reducing the envelope component around 30 Hz to 150 Hz that affects "roughness," thus changing the sound into a calmer, more listenable sound, or aggression, sharpness, and further dryness can be enhanced by increasing the envelope component at 200 Hz or higher.

[0071] In this way, in addition to extracting and efficiently transmitting the envelope component of the audio signal, the subjective perception of the sound can be easily altered during distribution or reproduction by processing some of the extracted envelope components according to the content or type of sound.

[0072] 3. Overview of the technology according to the present invention and the audio signal transmission system

[0073] Figure 2 This is a diagram illustrating a configuration embodiment of an audio signal transmission system according to an embodiment of the present disclosure.

[0074] Figure 2The audio signal transmission system 1 shown is configured to include an encoding device 100 and a decoding device 200.

[0075] For example, the audio signal transmission system 1 can be configured as a system for remote live streaming. In this case, the encoding device 100 is configured as a server (distribution device) on the site of the event, and the decoding device 200 is configured as a playback terminal on the remote audience side.

[0076] In the audio signal transmission system 1, the encoding device 100 separates the frequency band division information obtained by dividing the input audio signal according to each frequency band into an envelope component and a flattened waveform excluding the envelope component. Considering the importance of the envelope component to auditory perception, the encoding device 100 encodes and quantizes the envelope component through frequency transformation, and encodes the flattened waveform using fewer bits than the number of bits for the envelope component. In this way, the encoded bitstream obtained by encoding the audio signal is transmitted to the decoding device 200.

[0077] The decoding device 200 decodes the encoded data of the envelope component and the encoded data of the flattened waveform for each frequency band from the encoded bit stream transmitted by the encoding device 100. The decoding device 200 generates a frequency band division signal by synthesizing the decoded envelope component and the flattened waveform, and outputs an audio signal by synthesizing the frequency band division signals of each frequency band.

[0078] The encoding device 100 and the decoding device 200 will be described in detail below.

[0079] 4. Configuration and operation of the encoding device

[0080] Configuration of encoding device

[0081] Figure 3 This is a block diagram illustrating an example of the functional configuration of the encoding device 100.

[0082] like Figure 3 As shown, the encoding device 100 includes a frequency band division unit 110, a frequency band encoding unit 120, a content type acquisition unit 130, a network status detection unit 140, and an encoding / multiplexing unit 150.

[0083] The band-division unit 110 uses a band-division filter to divide the audio signal input to the encoding device 100 into a predetermined K bands (e.g., K=16). For example, a quadrature mirror filter (QMF) or a polyphase quadrature filter (PQF) can be used as the bandpass filter. The band-division signal for each band division is provided to the band-coding unit 120. Note that although the audio signal is primarily the processing target in this embodiment, it is not limited to this, and content signals including video signals, haptic signals, and other types of signals can also be used as the processing target.

[0084] The frequency band coding unit 120 has a frequency band coding unit 120A and a frequency band coding unit 120B for coding signals that are divided into frequency bands for each frequency band. For example, the frequency band coding unit 120 is provided with a frequency band coding unit 120A corresponding to K' frequency bands out of K frequency bands and a frequency band coding unit 120B corresponding to (K-K') frequency bands.

[0085] The band coding unit 120A separates the envelope component and the flattened waveform from the band-divided signal, then encodes the envelope component through frequency transformation, and encodes the flattened waveform with a different number of bits than the envelope component (more specifically, a lower number of bits than the envelope component). On the other hand, the band coding unit 120B encodes the band-divided signal without changing (without separating the envelope component) the signal through frequency transformation. Note that the flattened waveform can be encoded with a higher number of bits than the envelope component, rather than a lower number of bits.

[0086] Band coding unit 120A and band coding unit 120B do not need to be configured to correspond to consecutive frequency bands respectively. Each frequency band is associated with either band partitioning unit 110 or band coding unit 120A, which only needs to be switched by band coding unit 120B.

[0087] In other words, the frequency band division unit 110 can determine an encoding method for each frequency band, which involves separating the envelope component and flattened waveform from the frequency band division signal and encoding the envelope component and flattened waveform, or encoding the frequency band division signal as is.

[0088] Furthermore, the frequency band associated with the frequency band coding unit 120A can be switched according to the network transmission status and the processing load of the playback terminal (decoding device 200). For example, since the processing load of synthesizing the separated envelope components and flattening the waveform is relatively large, the processing load of the playback terminal can be reduced by reducing the number of frequency bands associated with the frequency band coding unit 120A when the resource load of the playback terminal is large.

[0089] In this case, the frequency band division unit 110 can determine the frequency band to be associated with the frequency band coding unit 120A based on the transmission status of the network detected by the network status detection unit 140.

[0090] The frequency band coding unit 120A includes an envelope separation unit 161, a frequency conversion unit 162, an envelope quantization unit 163, and a parameterization unit 164.

[0091] Envelope separation unit 161 separates the frequency band division signal of each frequency band into an envelope component and components other than the envelope component. Specifically, as shown in... Figure 4 As shown, the envelope separation unit 161 detects the envelope component of a time signal (band-divided signal) whose amplitude varies with the time axis, and separates the time signal into the envelope component and other components (flattened waveform). While the method for separating the envelope component is not particularly limited, methods such as using the Hilbert transform, squaring the signal waveform, and then passing it through a low-pass filter are commonly used. The separated envelope component is supplied to the frequency conversion unit 162, and the flattened waveform is supplied to the parameterization unit 164.

[0092] Frequency conversion unit 162 converts the envelope components from envelope separation unit 161 into envelope spectrum coefficients through frequency conversion, and provides the envelope spectrum coefficients to envelope quantization unit 163. For frequency conversion, modified discrete cosine transform (MDCT) is usually used, but discrete cosine transform (DCT), discrete Fourier transform (DFT) filter bank, cosine modulation filter bank, etc. can be used (transform coding is equivalent to filter bank).

[0093] Here, the content type acquisition unit 130 acquires the type of content or sound as the content type of the audio signal. The content type is acquired from metadata such as content, or is automatically detected as information such as program type or category information (music, news, movie, sports, etc.) for video content, category (classical, pop, jazz, etc.) or instrument type for music content. These content types can be selected through user input. Furthermore, when the content to be encoded is object data associated with each sound source of the audio, only the instrument type or object priority of the object needs to be acquired as the content type.

[0094] In addition, the network status detection unit 140 obtains the transmission bit rate, latency, etc. from the system as the network transmission status, and obtains the load processing status from the playback terminal (decoding device 200).

[0095] The envelope quantization unit 163 quantizes the envelope spectrum coefficients from the frequency conversion unit 162 using a predetermined number of bits based on the content type obtained by the content type acquisition unit 130 and the network transmission status detected by the network status detection unit 140, and outputs the obtained quantized spectrum to the encoding / multiplexing unit 150. The envelope spectrum coefficients are normalized using the maximum value in the quantization band, and quantized using a predetermined number of bits.

[0096] At this point, bits are dynamically allocated based on the type of content or sound obtained as content type, so that the quantization accuracy of important frequency bands (hereinafter also referred to as "modulation bands") in the envelope component is high, and the quantization accuracy of less important modulation bands is low.

[0097] For example, when the content type is healing music, ambient sounds, etc., the sound type (timbre and atmosphere) only needs to be calm, relaxing, or smooth, such as... Figure 5 As shown in the diagram. In this case, the modulation band from 5 Hz to 20 Hz is considered important, and for this band, the quantization bandwidth is narrowed and the number of bits to be allocated is increased.

[0098] Additionally, when the content type is classical, jazz, etc., the sound type (timbre and atmosphere) only needs to be calm and smooth. In this case, the modulation band from 30Hz to 150Hz is considered important, and for this band, the quantization bandwidth becomes narrower and the number of bits to be allocated increases.

[0099] Furthermore, when the content type is pop music, rock music, movie sound, etc., the sound type (timbre and atmosphere) only needs to be sharp, clear, aggressive, dry, etc. In this case, the modulation frequency band above 200Hz is considered less important, and the quantization bandwidth is increased, while the number of bits allocated to the frequency band is reduced.

[0100] Furthermore, bits can be dynamically allocated not only based on the type of content or sound, but also based on the network's transmission status or object priority. For example, when the network's transmission status is insufficient, the number of quantization bits for modulation bands determined to be less important to the content type can be reduced, or no bits can be allocated, so that information about the band is not transmitted. Additionally, the processing load status of the playback terminal can be transmitted to the server (encoding device 100), and the bands whose information is not transmitted can be determined based on the processing load status. Furthermore, for objects with low priority, the transmission band can be suppressed by reducing the number of bits allocated to that band and modulation bands with low importance to the content, or by not allocating any bits, thus preventing information transmission.

[0101] The parameterization unit 164 performs parameter encoding on the flattened waveform from the envelope separation unit 161 to express the flattened waveform with predetermined parameters. For example, in parameter encoding, multiple signal waveform patterns are prepared in advance as a table, and the flattened waveform is replaced with the index number of the most recent signal waveform pattern. Furthermore, when the flattened waveform is close to a sine wave or a polyphonic waveform, the flattened waveform can be transformed into sine wave parameters such as the frequency, gain, and phase of several sine waves. The flattening parameters obtained from the parameter encoding are output to the encoding / multiplexing unit 150.

[0102] On the other hand, the frequency band coding unit 120B is configured by the frequency conversion unit 171 and the quantization unit 172.

[0103] The frequency conversion unit 171 converts the frequency band division signal of each frequency band into spectral coefficients through frequency conversion, and provides the spectral coefficients to the quantization unit 172.

[0104] The quantization unit 172 quantizes the spectral coefficients from the frequency conversion unit 171 using a predetermined number of bits and outputs the obtained quantized spectrum to the encoding / multiplexing unit 150.

[0105] Then, the encoding / multiplexing unit 150 transforms the quantized spectra of the band coding units 120A (envelope quantization unit 163) and the quantization spectra of the band coding units 120B (quantization unit 172) from each frequency band into entropy codes such as Huffman codes. Furthermore, the encoding / multiplexing unit 150 multiplexes the transformed entropy codes together with the flattening parameters of the band coding units 120A (parameterization unit 164) from each frequency band in a predetermined format to generate a coded bitstream.

[0106] Transmission format of encoded bit stream

[0107] Figure 6 A diagram illustrating a warp knitting example.

[0108] exist Figure 6 In the embodiments described, processing related to the syntax of the transmission format is illustrated. Figure 6 In the table, bold text indicates the syntax recorded in the bitstream, and this syntax is recorded using the number of bits described in the right column of the table. Descriptions other than bold text are auxiliary descriptions, such as processing related to the number of syntax or loops.

[0109] Num_band is the number of transmission frequency bands, and it is determined by... Figure 3The number of frequency band division signals encoded and transmitted to the decoding device 200 from the frequency band division signals divided by the frequency band division unit 110. For example, if the frequency band division signal is divided into 16, then Num_band is represented by 4 bits. The number of transmission frequency bands is controlled according to the transmission status of the network detected by the network status detection unit 140, and when the transmission status is poor, the number of transmission frequency bands can be reduced to save transmission frequency bands.

[0110] `isEnvelopeSpectral` is a 1-bit flag that indicates true if the spectral data for band k is the spectrum of the envelope component, and false otherwise. That is, `isEnvelopeSpectral` indicates whether the processing of band k includes the separation and synthesis of the envelope component. When `isEnvelopeSpectral` is true, the subsequent spectrum represents the envelope spectrum; when `isEnvelopeSpectral` is false, the subsequent spectrum represents the spectrum of the band-divided signal itself.

[0111] Num_qu[k] is the number of quantization bands transmitted in band k. A quantization band is a unit used to quantize the spectrum, and one quantization band contains several spectra. Figure 6 In this embodiment, Num_qu[k] is represented by 4 bits. By varying the number of quantization bands based on content type and network transmission conditions, it is possible to reduce the amount of information to be transmitted while maintaining the modulation frequency components that are subjectively important to the content. For example, based on... Figure 5 The parameters of the important frequency bands for each content described herein are determined based on the network transmission conditions to include modulation frequency components that are important to the timbre and content.

[0112] Note that the algorithm used to determine the quantization band based on isEnvelopeSpectral (true / false) is switched. This is because the appropriate quantization bandwidth differs between the spectrum of the envelope component and the spectrum of the band-divided signal itself.

[0113] `scale_factor_index[k][q]` is the index of the scaling factor in the quantized frequency band q of frequency band k. The scaling factor is the coefficient used to normalize the spectrum, and the index represents the exponent part of a power of 2. For example, if the number of bits is 5, it can represent 2^k. 5 =32 levels of normalization.

[0114] `quant_nbits[k][q]` represents the value {(number of quantization bits) - 1} applied to frequency band k and quantization band q (because 0 bits are not used). By varying the number of quantization bits per spectrum according to the content type and network transmission conditions, the amount of information to be transmitted can be effectively reduced while maintaining the modulation frequency components that are subjectively important to the content. Similar to the number of quantization bands mentioned above, it can be based on... Figure 5 The parameters shown allocate more bits to the frequency components and modulation frequency components that are important to the content.

[0115] The underscore-indicated Num_spec(k, q) is a function used to obtain the quantization bandwidth (the number of spectra included) in frequency band k and quantization band q. The quantization bandwidth for each frequency band is predefined.

[0116] `code_spec[k][q]` is the spectral code encoded by Huffman isentropically. Since this code is of variable length, the number of bits is variable. `Huff_decode()`, indicated by the underscore, is the function used to decode the Huffman code using a specified table to obtain the quantized spectrum. While an example of encoding the spectral coefficients one by one is shown here for simplicity, the same applies to multidimensional Huffman codes where multiple spectra are grouped.

[0117] The spectral coefficients Spec[k][n] of band k, indicated by an underscore, are obtained by inverse quantization, which is performed by dividing the decoded quantized spectrum by the quantization precision resolution, and then multiplying the result by a scaling factor to inversely normalize it. As shown above, if isEnvelopeSpectral is true, the spectral coefficients to be obtained are the spectrum of the envelope component; if isEnvelopeSpectral is false, the spectral coefficients to be obtained are the spectrum of the band-divided signal itself.

[0118] If isEnvelopeSpectral is true, information about the flattened waveform (carrier signal) that has been further parameterized is recorded. wave_type[k] indicates the type of flattened waveform for frequency band k. Here, 0 represents a general waveform such as noise or a composite signal. Numbers 1 to 3 represent multitone waveforms consisting of a representative number of sine waves. When wave_type[k] represents a general waveform, wave_index[k] represents the index of the waveform pattern, and wave_gain[k] represents its gain index. The flattened waveform is generated using the waveform pattern and gain corresponding to this index. Furthermore, when wave_type[k] represents multitone waveforms, the frequency index tone_freq, gain index tone_gain, and phase index tone_phase of the tone are recorded repeatedly by the number of tones.

[0119] Based on the above transmission format, the encoded bit stream generated in the encoding device 100 is transmitted to the decoding device 200.

[0120] Operation of the encoding device

[0121] Next, we will refer to Figure 7 The flowchart describes the encoding process of the audio signal by the encoding device 100. Figure 7 The process will mainly describe the encoding process for the frequency bands where isEnvelopeSpectral is True, but the encoding process for the frequency bands where isEnvelopeSpectral is False will also be performed in parallel.

[0122] In step S101, the frequency band division unit 110 divides the audio signal input to the encoding device 100 into frequency band division signals for each frequency band.

[0123] In step S102, the frequency band coding unit 120A (envelope separation unit 161) separates the frequency band division signal of each frequency band into an envelope component and a flattened waveform.

[0124] In step S103, the frequency band coding unit 120A (frequency conversion unit 162 and envelope quantization unit 163) obtains the quantized spectrum by encoding the envelope components. Specifically, the frequency conversion unit 162 transforms the envelope components into envelope spectrum coefficients through frequency conversion, and the envelope quantization unit 163 quantizes the envelope spectrum coefficients to obtain the quantized spectrum.

[0125] In step S104, the frequency band coding unit 120A (parameterization unit 164) encodes the flattened waveform through parameter coding to obtain flattening parameters.

[0126] In step S105, the encoding / multiplexing unit 150 generates an encoded bitstream by multiplexing the quantization spectrum and flattening parameters of each frequency band.

[0127] Based on the above configuration and processing, the frequency band division signal is separated into an envelope component and other components. The envelope component, which is important for timbre and sonic atmosphere, is encoded with high precision through frequency transformation, while the other components are encoded with a smaller number of bits. Therefore, even with a small amount of information transmitted, reconstruction can be performed without losing the original timbre, sonic atmosphere, or even the sense of presence. That is, good sound quality can be maintained even at low bit rates.

[0128] 5. Configuration and operation of the decoding device

[0129] Configuration of decoding device

[0130] Figure 8This is a block diagram illustrating a functional configuration embodiment of the decoding device 200.

[0131] like Figure 8 As shown, the decoding device 200 includes a decoding / demultiplexing unit 210, a frequency band decoding unit 220, a synthesis parameter determination unit 230, and a frequency band synthesis unit 240.

[0132] The decoding / demultiplexing unit 210 performs decoding processing on the encoded bitstream of the audio signal input to the decoding device 200.

[0133] Specifically, the decoding / demultiplexing unit 210 separates the quantization spectrum (first encoded data) and flattening parameters (second encoded data) from the encoded bitstream of the audio signal for each frequency band. The quantization spectrum (first encoded data) is obtained by encoding the envelope component of the audio signal through frequency transformation, and the flattening parameters (second encoded data) are obtained by encoding the flattened waveform of the audio signal with fewer bits than the envelope component (parameter encoding). Furthermore, the decoding / demultiplexing unit 210 separates the spectrum (third encoded data) from the encoded bitstream of the audio signal for each frequency band, obtained by frequency transformation of the band division signal itself.

[0134] At this point, the decoding / demultiplexing unit 210 determines, for each frequency band, whether to separate the quantization spectrum and flattening parameters or only the quantization spectrum, based on the flag (isEnvelopeSpectral) contained in the coded bitstream. That is, if isEnvelopeSpectral is true, the quantization spectrum and flattening parameters are separated from the coded bitstream of that frequency band; if isEnvelopeSpectral is false, only the quantization spectrum (the spectrum obtained by frequency transformation of the band-divided signal itself) is separated from the coded bitstream of that frequency band. The quantization spectrum (flattening parameters) separated for each frequency band is provided to the band decoding unit 220.

[0135] The band decoding unit 220 includes a band decoding unit 220A and a band decoding unit 220B for decoding the quantization spectrum (flattening parameters) of each band. For example, the band decoding unit 220 is provided with a band decoding unit 220A corresponding to K' bands out of K bands and a band decoding unit 220B corresponding to (K-K') bands.

[0136] For frequency bands where isEnvelopeSpectral is true, the frequency band decoding unit 220A generates a frequency band division signal by synthesizing the envelope component decoded from the quantization spectrum and the flattened waveform decoded from the flattening parameters. On the other hand, for frequency bands where isEnvelopeSpectral is false, the frequency band decoding unit 220B outputs the frequency band division signal decoded from the quantization spectrum.

[0137] Band decoding units 220A and 220B do not need to be configured to correspond to consecutive frequency bands. Note that which frequency band is associated with in band decoding units 220A and 220B only needs to be switched based on isEnvelopeSpectral contained in the coded bitstream.

[0138] The frequency band decoding unit 220A includes an inverse quantization unit 251, an envelope weighting unit 252, an inverse frequency conversion unit 253, a flattening waveform generation unit 254, and an envelope synthesis unit 255.

[0139] The inverse quantization unit 251 performs a predetermined inverse quantization process on the quantized spectrum from the decoding / demultiplexing unit 210, such as restoring the quantization index to the spectrum or multiplying the quantization index by the inverse gain of the normalization process, thereby obtaining N envelope spectral coefficients. Here, N is the number of frequency elements when the envelope components are frequency-transformed. The envelope spectral coefficients are provided to the envelope weighting unit 252.

[0140] Envelope weighting unit 252 performs weighting processing related to the envelope spectral coefficients of each frequency band based on weighting information (synthesis parameters) provided by synthesis parameter determination unit 230. Specifically, envelope weighting unit 252 multiplies the envelope spectral coefficients by weighting coefficients different for each frequency slot based on the weighting information from synthesis parameter determination unit 230, thereby obtaining weighted envelope spectral coefficients. The obtained weighted envelope spectral coefficients are provided to inverse frequency transformation unit 253. Synthesis parameter determination unit 230 determines the weighting information used to set the weighting coefficients in the weighting processing by using any of the pre-prepared parameters. The following will refer to... Figure 9 Details of the synthesis parameter determination unit 230 and weight information are described.

[0141] The inverse frequency transformation unit 253 recovers the weighted envelope spectrum coefficients from the envelope weighting unit 252 into the time-domain signal of the envelope component through inverse frequency transformation, and provides the time-domain signal to the envelope synthesis unit 255.

[0142] Then, the flattening waveform generation unit 254 generates a signal waveform pattern corresponding to the parameter index by generating the flattening parameters from the decoding / demultiplexing unit 210, and generates a flattened waveform by synthesizing multitone from sine wave parameters. The obtained flattened waveform is provided to the envelope synthesis unit 255.

[0143] Envelope synthesis unit 255 generates a band-division signal by synthesizing the envelope component from inverse frequency conversion unit 253 and the flattened waveform from flattened waveform generation unit 254. The band-division signal is generated by performing amplitude modulation on the envelope component of the flattened waveform as a carrier. The generated band-division signal is output to band synthesis unit 240.

[0144] On the other hand, the frequency band decoding unit 220B includes an inverse quantization unit 261 and an inverse frequency conversion unit 262.

[0145] The inverse quantization unit 261 performs a predetermined inverse quantization process on the quantized spectrum from the decoding / demultiplexing unit 210 to obtain spectral coefficients, and supplies the spectral coefficients to the inverse frequency conversion unit 262.

[0146] The inverse frequency conversion unit 262 restores the spectral coefficients from the inverse quantization unit 261 into a frequency band division signal through inverse frequency conversion, and outputs the frequency band division signal to the frequency band synthesis unit 240.

[0147] Then, the frequency band synthesis unit 240 combines the frequency band division signals from the frequency bands of the frequency band decoding unit 220 by using a synthesis filter to generate an audio signal of a time-domain waveform signal.

[0148] Configuration of the synthesis parameter determination unit

[0149] Here, we will refer to Figure 9 The configuration of the synthesis parameter determination unit 230 is described.

[0150] like Figure 9 As shown, the synthesis parameter determination unit 230 includes a content type acquisition unit 271, preset synthesis parameters 272, and an envelope synthesis control unit 273.

[0151] Content type acquisition unit 271 has the same as Figure 7 The content type acquisition unit 130 performs the same function as the audio signal content type acquisition unit, and acquires the type of content or sound as the content type of the audio signal. The content type is acquired from metadata such as content, or is automatically detected as information, such as the program type of video content (music, news, movie, sports, etc.) or the category of music content (classical, pop, jazz, etc.) or instrument type, etc. These content types can be selected by user input. Furthermore, when the content to be encoded is data of an audio object (audio object data), for example, only the instrument type or object priority of the audio object needs to be acquired as the content type.

[0152] Preset synthesis parameter 272 is a preset parameter (parameter set) used to weight the envelope component to have the timbre and sound atmosphere corresponding to the sound or type of the content obtained as a content type.

[0153] For example, when the content type is healing music, ambient sounds, etc., such as Figure 10 As explained, the sound type (timbre and atmosphere) only needs to be calm, relaxed, or smooth. Correspondingly, parameters are prepared for the modulation bands across the entire frequency band to reduce (-6dB) by 5 to 20 Hz.

[0154] Additionally, when the content type is classical, jazz, etc., the sound type (timbre and atmosphere) only needs to be calm and smooth. As a corresponding parameter, parameters are prepared for reducing (-6dB) the modulation band from 30Hz to 150Hz for the frequency band of 100Hz or higher.

[0155] Furthermore, when the content type is pop music, rock music, movie sound effects, etc., the sound type (timbre and atmosphere) only needs to be sharp, clear, aggressive, dry, etc. As corresponding parameters, parameters are prepared for increasing (increasing +6dB) the modulation frequency band from 200Hz to 1kHz in the 1kHz to 3.5kHz frequency band.

[0156] return Figure 9 As described, the envelope synthesis control unit 273 selects the optimal parameters for the content type (category, instrument type, timbre, etc.) obtained by the content type acquisition unit 271 from the preset synthesis parameters 272. The selected parameters are provided as weighting information to the envelope weighting unit 252.

[0157] Detailed configuration of the envelope weighting unit

[0158] Here, the detailed configuration of the envelope weighting unit 252 will be described further.

[0159] Figure 11 This is a block diagram illustrating a configuration embodiment of the envelope weighting unit 252.

[0160] The weighting coefficient setting unit 281 sets weighting coefficients for each envelope spectrum coefficient based on the weighting information from the synthesis parameter determination unit 230. The weighting unit 282 then assigns weighting coefficients to the N envelope spectrum coefficients X in frequency band k. k,n (n=0, 1, ..., N-1) multiplied by different weighting coefficients W for each frequency unit k,n This outputs N weighted envelope spectral coefficients Y. k,n (n=0, 1, ..., N-1).

[0161] Since the above processing is performed on all frequency bands k, the weight information from the synthesis parameter determination unit 230 is as follows: Figure 12 The two-dimensional table shown. Figure 12The weight information shown is presented in a table with Num_band rows and N columns. As shown above, Num_band represents the number of transmission frequency bands.

[0162] It should be noted that, although in this embodiment the envelope spectral coefficients are multiplied by a simple weighting coefficient W k、n However, it can perform the weighted summation of N envelope spectral coefficients. That is, for frequency band k, there are N×N weighting coefficients W k, n, m Used to perform the following processing.

[0163] Y k,n =Σm[W k,n,m ×X k,n ] (n, m=0, 1,..., N-1)

[0164] Furthermore, the weighted processing of the envelope spectrum coefficients can be performed using a neural network.

[0165] Figure 13 This is a block diagram illustrating another configuration embodiment of the envelope weighting unit 252.

[0166] The weighting coefficient setting unit 291 sets the weighting coefficients of the network unit 292 based on the weight information from the synthesis parameter determination unit 230. The network unit 292 uses a neural network to set the weighting coefficients X of the N envelope spectral coefficients in frequency band k. k,n (n=0, 1, ..., N-1) Perform weighted processing to output N weighted envelope spectral coefficients Y. k,n (n=0, 1, ..., N-1).

[0167] The parameters of the neural network are learned using signals from a sound source that are manually edited to have the desired timbre and desired sonic atmosphere as reference signals.

[0168] For example, it can be in the case of Figure 14 The parameters are learned in the process shown. First, based on... Figure 3 The encoding device 100 in the middle is configured similarly to perform frequency band division and envelope separation on the signal x(t) of the sound source to be processed, thereby obtaining the envelope component Ev. x (k, f) (where k represents the index of the frequency band and f represents the index of the envelope spectral coefficients). Envelope component Ev x (k, f) is input into a network such as a deep neural network (DNN), and the resulting output is determined by Ev x′ (k, f) represents the network configuration. The network configuration here can be arbitrary.

[0169] On the other hand, the envelope component obtained by similarly performing frequency banding and envelope separation on the signal y(t) of a reference sound source with the desired timbre is represented as Ev. y (k, f).

[0170] Then, the parameters of the network are learned so that the above network outputs Ev x′ (k, f) and the envelope component Ev of the reference sound source y (k, f) becomes as close as possible. For example, the mean squared error (MSE) of the two components is minimized as a loss function.

[0171] Furthermore, generative adversarial networks (GANs) can be used in conjunction with this approach. These networks employ discriminators to distinguish whether the input signal is from a reference sound source or generated by the network (generator). In addition to the learning methods mentioned above, reinforcement learning, unsupervised learning, and other techniques can also be employed.

[0172] Operation of the decoding device

[0173] Next, we will refer to Figure 15 The flowchart describes the decoding process of the audio signal by the decoding device 200. Figure 15 The main description will focus on the decoding process for the frequency bands where isEnvelopeSpectral is True, but the decoding process for the frequency bands where isEnvelopeSpectral is False will also be performed in parallel.

[0174] In step S201, the decoding / demultiplexing unit 210 separates the quantization spectrum and flattening parameters for each frequency band from the encoded bitstream of the audio signal input to the decoding device 200.

[0175] In step S202, the band decoding unit 220A (inverse quantization unit 251, envelope weighting unit 252, and inverse frequency transformation unit 253) decodes the quantized spectrum separated from the coded bitstream to obtain the envelope component. Specifically, the inverse quantization unit 251 performs inverse quantization on the quantized spectrum to obtain envelope spectrum coefficients, the envelope weighting unit 252 performs weighting processing on the envelope spectrum coefficients, and the inverse frequency transformation unit 253 transforms the weighted envelope spectrum coefficients into envelope components through inverse frequency transformation.

[0176] In step S203, the frequency band decoding unit 220A (flattening waveform generation unit 254) decodes the flattening parameters separated from the encoded bit stream to obtain the flattened waveform.

[0177] In step S204, the frequency band decoding unit 220A (envelope synthesis unit 255) generates a frequency band division signal by synthesizing envelope components and flattened waveforms.

[0178] Then, in step S205, the frequency band synthesis unit 240 generates an audio signal by synthesizing the frequency band division signals of each frequency band.

[0179] Based on the above configuration and processing, since the decoded bitstream is encoded in such a way that the envelope component, which is important for timbre and sonic atmosphere, is encoded with high precision through frequency transformation, while other components are encoded with fewer bits, reconstruction can be performed without loss of the original timbre and sonic atmosphere, and further without loss of presence, even with a small amount of information transmitted. In other words, good sound quality can be maintained even at low bit rates.

[0180] 6. Modifications of the decoding device

[0181] A variation of the decoding device 200 described above will be described.

[0182] Weighting of envelope components after inverse frequency transform

[0183] Figure 16 This is a block diagram illustrating another functional configuration embodiment of the decoding device 200.

[0184] Figure 16 Decoding device 200 and Figure 8 The difference of the decoding device 200 is that the envelope weighting unit 252 is not provided in the decoding unit 220A, and an envelope processing unit 311 is newly provided between the inverse frequency conversion unit 253 and the envelope synthesis unit 255.

[0185] Right now, Figure 16 The decoding device 200 (band decoding unit 220A) does not directly perform weighting processing on the N envelope spectrum coefficients obtained through decoding and inverse quantization. Instead, it converts the envelope spectrum coefficients into envelope components through inverse frequency transformation. Then, the envelope processing unit 311 performs weighting processing on the envelope components within the desired (specific) frequency range. In this way, the envelope processing unit 311 also performs weighting processing related to the envelope spectrum coefficients of each frequency band based on the weighting information provided by the synthesis parameter determination unit 230.

[0186] In this case, processing of the modulation frequency components more suited to human auditory characteristics can be performed without relying on the limitations of the number of band-division filters and bandwidth, such as MDCT. For example, the envelope component can be unequally divided using a constant Q filter bank to simulate a “modulation filter bank,” and then some of the divided signals can be processed. A “modulation filter bank” is described in “Modeling auditory processing of amplitude modulation. I. Detection and masking with narrow-band carriers” by Dau, T et al., published in Journal of the Acoustical Society of America vol. 102, 5Pt 1 (1997): 2892-905. Note that when the signal input to the encoding device 100 is a video signal or a tactile signal, processing suitable for human visual or tactile characteristics can be performed on the decoding device 200 side.

[0187] Figure 17 This is a block diagram illustrating a configuration embodiment of the envelope processing unit 311.

[0188] The weighting coefficient setting unit 321 sets the weighting coefficients for each frequency band based on the weighting information from the synthesis parameter determination unit 230. The analysis filter bank unit 322 analyzes the envelope component x of frequency band k. k Divided into M frequency bands k, m (m = 0, 1, ..., M-1). Weighting unit 323 weights the M frequency band signals u in frequency band k. k, m Multiplied by the weighting coefficient w of each frequency band division k, m The synthesis filter bank unit 324 resynthesizes the weighted M frequency band signals u. k, m To output the processed envelope component y k .

[0189] Note that the above Figure 8 The operation of the envelope processing unit 311 and the envelope weighting unit 252 is based on the premise that the frequency band is envelope-separated (isEnvelopeSpectral is true). That is, since the decoding device 200 cannot apply envelope processing to frequency bands that are not envelope-separated, it is expected that the encoding device 100 will perform envelope separation on frequency bands whose envelope components can be processed during decoding.

[0190] In this case, for frequency bands that are not envelope-separated (isEnvelopeSpectral is false), information about the frequency band other than the weight information from the synthesis parameter determination unit 230 is ignored.

[0191] When processing the envelope component of a frequency band that has not been envelope-separated, such as in the processing of the coding device 100, the envelope component can be processed after envelope separation. More specifically, it is only necessary to... Figure 3 When an envelope separation unit 161 or a frequency conversion unit 162 is set in the subsequent stage of the frequency band decoding unit 220B (inverse frequency conversion unit 262), the envelope component (envelope spectrum coefficients) and flattened waveform are obtained, and the envelope weighting unit 252 ( Figure 8 ) or envelope processing unit 311 ( Figure 16 The envelope synthesis unit 255 processes the data and then resynthesizes and outputs the data.

[0192] Variation of the synthesis parameter determination unit

[0193] Figure 18 This is a block diagram illustrating another configuration example of the synthesis parameter determination unit 230 included in the decoding device 200.

[0194] Figure 18 Synthesis parameter determination unit 230 and Figure 9 The difference between the synthesis parameter determination unit 230 and the content type acquisition unit 271 is that a user configuration acquisition unit 331 is set up instead of a content type acquisition unit 271.

[0195] User configuration acquisition unit 331 acquires user configuration information such as the user's auditory characteristics, user preferences, and device attributes of the hearing device worn by the user as the content to be viewed.

[0196] For example, the user profile acquisition unit 331 acquires the user's preferred tone, the device characteristics of the headphones being used, and (if the user is hearing impaired) their auditory characteristics and hearing aid device characteristics from a pre-recorded user profile, and outputs the acquired information as configuration information to the envelope synthesis control unit 273. Furthermore, the pre-recorded user profile can be updated by optimizing it individually for each user, or new data can be added.

[0197] The envelope synthesis control unit 273 determines the weighting information by selecting the parameters that best match the configuration information of the user configuration acquisition unit 331 from the parameters prepared in advance as preset synthesis parameters 272. The parameters for each frequency band are prepared as preset synthesis parameters 272 to obtain the best sound based on the user's preferred timbre, auditory characteristics (especially in the case of hearing-impaired individuals), characteristics of the device to be used, etc.

[0198] For example, regardless of whether a user prefers a less harsh tone, when the sound sounds harsh due to increased sensitivity of the modulation band related to harshness caused by the characteristics of the user's equipment, the tone can be improved by reducing the intensity of the modulation frequency component of the band.

[0199] Furthermore, based on the findings disclosed in the paper "Measurement of Temporal Resolution in Presbyacusis" published by Okamoto Yasuhide, Kanzaki Akira, Nukano Ayako, Nakaichi Kenji, Morimoto Ryuji, Harada Kouta, Kubota Eri, and Ogawa Iku in AUDIOLGY JAPAN, 2014, Vol. 57, No. 6, pp. 694-702—that older adults with age-related hearing loss often have higher modulation thresholds in the time modulation transfer function (TMTF) than those with normal hearing—it is possible to alleviate the difficulties in perceiving speech and music caused by hearing loss by enhancing the intensity of modulation frequency components that have higher modulation thresholds compared to those with normal hearing.

[0200] Furthermore, for users who have difficulty understanding speech due to hearing loss, the findings of Drulllman et al. in their paper "Effect of reducing slow temporal modulations on speech reception" published in *The Journal of the Acoustical Society of America*, Vol. 95, No. 5, Pt. 1 (1994): 2670-80—namely, "the modulation frequency of the speech amplitude envelope is 8 to 10 Hz"—can be correspondingly enhanced in these frequency ranges. More specifically, for example, for the frequency band around 100 to 2000 Hz (corresponding to the frequency band of human speech), the envelope spectrum corresponding to the modulation band of 8 to 10 Hz can be increased by +6 dB.

[0201] Figure 19 This is a block diagram illustrating another configuration embodiment of the synthesis parameter determination unit 230 included in the decoding device 200.

[0202] Figure 19 Synthesis parameter determination unit 230 and Figure 9 The difference between the synthesis parameter determination unit 230 and the content type acquisition unit 271 is that the terminal information acquisition unit 341 is provided instead of the content type acquisition unit 271.

[0203] The terminal information acquisition unit 341 acquires terminal information of the playback terminal (decoding device 200), transmission status (network status) of the network transmitting the encoded bit stream, etc.

[0204] For example, the terminal information acquisition unit 341 acquires the processing resources and load status of its own device, such as the remaining battery power, CPU utilization and memory consumption of the decoding device 200, and outputs the processing resources and load status as terminal information to the envelope synthesis control unit 273.

[0205] The envelope synthesis control unit 273 determines the weight information by selecting the optimal parameters for the terminal information of the terminal information acquisition unit 341 from the parameters prepared in advance as preset synthesis parameters 272.

[0206] When the remaining battery power or processing resources of the playback terminal are insufficient, by ignoring the envelope spectral coefficients that have little impact on the main timbre, a portion of the computational processing in the inverse frequency transformation unit 253 and the envelope synthesis unit 255 can be omitted in the subsequent stage of the envelope weighting unit 252. As an embodiment that further reduces the amount of processing, by ignoring 1 / 2 or 3 / 4 of the envelope spectral coefficients on the high-frequency band side as described above, the transformation size of the inverse frequency transformation (such as inverse MDCT) is reduced to 1 / 2 or 1 / 4 of the original size to perform the transformation, thereby allowing the transformation of the downsampled envelope component to be performed with less computation. In this case, the envelope synthesis unit 255 can perform envelope synthesis processing on the downsampled envelope component as is.

[0207] Additionally, when network conditions are poor, resulting in strange or unusual sounds due to packet loss or transmission errors, these sounds can be transformed into more palatable tones using appropriate preset parameters. For example, by changing the sound to a more aggressive or harsh tone, like rock music, the noise becomes less noticeable.

[0208] Furthermore, the envelope synthesis control unit 273 can control the band decoding unit 220A and the band synthesis unit 240 for each band based on terminal information. When the processing load is high, the processing load can be significantly reduced by not performing all decoding and synthesis processing on the predetermined high-frequency bands that have been divided into bands. Optionally, the load status of the playback terminal can be notified to the encoding device 100, and the server can switch the band used for encoding and transmission according to the load status of the playback terminal. In this case, the network status detection unit 140 ( Figure 3 It only needs to receive and detect the load status of the reproducing terminal.

[0209] Configuration for decoding audio object encoded data

[0210] The decoding device 200 can decode audio object encoded data, including audio object data and audio object metadata, into an encoded bitstream.

[0211] Figure 20 This is a block diagram illustrating another functional configuration embodiment of the decoding device 200.

[0212] Figure 20 Decoding device 200 and Figure 8 and Figure 16 The difference in the decoding device 200 is that a rendering unit 351 is newly provided in the post-stage of each of the band decoding units 220A and 220B set for each band. Note that in Figure 20 In the decoding device 200, with Figure 8 and Figure 16 The components of the decoding device 200 are represented by the same reference numerals, and their descriptions will be largely omitted.

[0213] However, in addition to the functions described above, the decoding / demultiplexing unit 210 decodes audio object data and object metadata from the encoded bitstream. The object metadata is provided to the rendering unit 351 for rendering and is also provided to the synthesis parameter determination unit 230.

[0214] Rendering unit 351 sets up for each frequency band generated by the frequency band division, performs rendering processing on the audio object for each frequency band, and outputs the rendered audio signal as a multi-channel signal. The rendered audio signals of each frequency band are synthesized by frequency band synthesis unit 240 to obtain the time-domain waveform signal of each channel.

[0215] In addition, the content type acquisition unit 271 of the synthesis parameter determination unit 230 ( Figure 9 The system obtains the object information of the audio object from the object metadata decoded by the decoding / demultiplexing unit 210. The object information includes the content category (attribute information) to which the audio object belongs, the sound type (instrument, sound, ambient sound, etc.), the instrument type, the timbre information (parameters of timbre elements such as timbre type, roughness, sharpness, etc.), the object priority (priority information), and the propagation degree of the audio object (propagation information).

[0216] The envelope synthesis control unit 273 of the synthesis parameter determination unit 230 ( Figure 9 The weighting information is determined by selecting the best preset parameters based on the object information obtained by the content type acquisition unit 271. Since audio objects are usually classified according to instrument and sound type, parameters that provide the best timbre are selected based on the category, sound type, instrument type, etc. included in the object information. In addition, when the object metadata contains timbre information (parameters of timbre elements, such as timbre type, roughness, and sharpness), a parameter set close to each timbre element is selected based on the timbre information.

[0217] Furthermore, the parameter set can be selected based on object priority. For example, when an object has a high priority, its impact on the overall sound reproduction is considered significant, and therefore a parameter set that particularly emphasizes the timbre components based on the type or category of the audio object's sound is selected. Conversely, when an object has a low priority, only the envelope components that are important for the category are retained, while other envelope components are set to zero, thereby further reducing the processing load during decoding.

[0218] In this way, the optimal parameters are selected based on the object information and output as weight information.

[0219] In object encoding, signals from sound sources (such as musical instruments and singing) located at the same position are encoded into audio objects for each sound source. Audience cheers, shouts, etc., without a definite location of the sound source can be represented by virtual objects. As the object-encoded data is decoded and reproduced, the rendering unit 351 decodes the audio signals of multiple audio objects included in the encoded bitstream to recover the time-domain waveform signal. Subsequently, the rendering unit 351 maps the sound in three-dimensional space based on the position information, gain information, etc., of the audio objects contained in the object metadata, and maps the sound to multiple channel signals according to the reproduction environment (rendering processing).

[0220] For example, in the case of loudspeaker reproduction, the spatial location of the sound image is reproduced using a method called Vector-Based Amplitude Translation (VBAP), based on the arrangement and number of loudspeakers. The audio signal of each audio object is mapped to a channel signal based on the location information of the audio objects included in the object metadata. In the case of headphone reproduction, the room's reverberation impulse response (e.g., the so-called Head-Related Transfer Function (HRTF) or Room Impulse Response (RIR)) is applied so that when viewing an audio object, the sound image is located at the position of each audio object, producing stereo sound (also known as a virtualizer). Furthermore, Binaural Chamber Impulse Response (BRIR), High-Order High-Fidelity Stereo (HOA), and others can be used for rendering processing.

[0221] Typically, this rendering process is performed after all the audio waveforms of the desired audio object have been decoded. That is, a decoding device 200 must be provided for each audio object, or the rendering process must be performed after repeating the decoding process as many times as the number of audio objects. Furthermore, different frequency bands may require different rendering processes. For example, in the case of speakers, because the characteristics and proper arrangement of devices such as woofers, midrange speakers, and tweeters differ in each frequency band, it is desirable to perform different rendering processes for each frequency band (this also applies to headphones).

[0222] As mentioned above, due to the high processing load of decoding and reproducing audio objects, Figure 20The decoding device 200 performs reproduction processing on each frequency band in the pre-stage of the frequency band synthesis unit 240. With this configuration, when rendering processing is performed differently for each frequency, it is not necessary to perform frequency band division processing again in the subsequent stage, thus reducing the amount of processing.

[0223] Note that in Figure 20 In this embodiment, the rendering unit 351 is positioned after the band decoding units 220A and 220B for each frequency band. While not limited thereto, considering that the envelope component greatly aids in the perceptual localization of the sound image, the rendering unit 351 may be positioned after the band decoding unit 220A (…). Figure 8 The inverse frequency transformation unit 253 and the envelope synthesis unit 255 are located within the inverse frequency transformation unit 253 and 262. In addition, the rendering unit 351 can be set in the stage before the inverse frequency transformation units 253 and 262, and can perform rendering processing on the envelope spectrum coefficients (spectral coefficients) in the frequency domain (multiplied by a different gain for each frequency).

[0224] Note that for flattened waveforms, it is assumed that the contribution of flattened waveforms to the perceptual localization of the sound image is insignificant (especially in the low-frequency range). Therefore, rendering processing can be performed in a simple way: by summing the flattened waveforms of each object without changing them and mixing the summed flattened waveforms into each channel signal, or by adjusting the gain parameter such as VBAP when parametrically encoding the flattened waveforms.

[0225] 7. Variations of encoding devices and music production / editing tools

[0226] Modifications of encoding devices

[0227] Figure 21 This is a block diagram illustrating another functional configuration embodiment of the encoding device 100.

[0228] Figure 21 The encoding device 100 and Figure 3 The difference in the encoding device 100 is that the envelope correction unit 411 is disposed between the frequency conversion unit 162 and the envelope quantization unit 163 in the band coding unit 120A, and a synthesis parameter determination unit 420 is newly provided. Note that in Figure 21 In the encoding device 100, with Figure 3 The components of the encoding device 100 are indicated by the same reference numerals, and their description will be substantially omitted. Additionally, in Figure 21 In, not included Figure 3 The encoding device 100 includes a content type acquisition unit 130 and a network status detection unit 140.

[0229] In this manner, the envelope correction unit 411 performs weighted processing related to the envelope spectral coefficients of each frequency band based on the weighting information provided by the synthesis parameter determination unit 420. Specifically, the envelope correction unit 411 obtains weighted envelope spectral coefficients by multiplying the envelope spectral coefficients from the frequency conversion unit 162 by different weighting coefficients for each frequency slot based on the weighting information from the synthesis parameter determination unit 420. The obtained weighted envelope spectral coefficients are provided to the envelope quantization unit 163.

[0230] The synthesis parameter determination unit 420 includes a timbre selection unit 431, preset synthesis parameters 432, and an envelope synthesis control unit 433.

[0231] Note that the preset synthesis parameters 432 and the envelope synthesis control unit 433 have parameters respectively contained in... Figure 9 The preset synthesis parameter 272 in the synthesis parameter determination unit 230 and the envelope synthesis control unit 273 have basically the same functions.

[0232] The timbre selection unit 431 acquires the timbre information of the content based on the input from the predetermined graphical user interface (GUI), and outputs the acquired timbre information to the envelope synthesis control unit 433.

[0233] The envelope synthesis control unit 433 selects the parameter that is best for the timbre information of the timbre selection unit 431 from the parameters prepared in advance as preset synthesis parameters 432.

[0234] As configured above Figure 21 The encoding device 100 has two types of operating modes (i.e., editing mode and archiving mode) as operating modes.

[0235] Edit mode is used when the encoding device 100 is invoked from the editing interface (GUI) of the music production / editing tool described below, and the editor converts the desired signal waveform into the desired timbre.

[0236] In edit mode, envelope correction unit 411 multiplies each component of the envelope spectrum coefficients from frequency conversion unit 162 by a weighting coefficient based on weighting information from synthesis parameter determination unit 420, and outputs the weighted envelope spectrum coefficients.

[0237] On the other hand, the timbre selection unit 431 of the synthesis parameter determination unit 420 acquires the timbre type of the content input from the editing interface and provides the acquired timbre type to the envelope synthesis control unit 433. Based on the timbre information from the timbre selection unit 431, the envelope synthesis control unit 433 selects the parameter that is optimal for the timbre information from the pre-prepared preset synthesis parameters 432 and outputs the selected parameter as weight information.

[0238] In this way, Figure 21 The encoding device 100 can convert the desired signal into a sound with the desired timbre or atmosphere during encoding.

[0239] For example, when performing automatic archiving on content, an archiving mode is used. In the signal of the content to be archived, envelope components (modulation frequency components) with abnormally strong or rapidly changing timbre cause unpleasant sounds, even without clipping. Therefore, during archiving, abnormalities affecting the timbre of the envelope components are detected, and if an abnormality is detected, the envelope components are corrected to have the optimal timbre.

[0240] In archive mode, envelope correction unit 411 determines whether the envelope spectrum coefficients from frequency conversion unit 162 are within a predetermined range. When the envelope spectrum coefficients exceed a predetermined threshold or are outside the range, envelope correction unit 411 corrects the envelope spectrum coefficients to optimal coefficients by multiplying them by a predetermined weighting coefficient based on the weighting information from synthesis parameter determination unit 420, and outputs the corrected coefficients as weighted envelope spectrum coefficients. The weighted envelope spectrum coefficients are then provided to envelope quantization unit 163, and subsequently, [further processing is performed]. Figure 3 The encoding device 100 processes the same process.

[0241] On the other hand, the timbre selection unit 431 of the synthesis parameter determination unit 420 acquires the timbre type of the content manually input or automatically detected from the audio signal, and provides the acquired timbre type to the envelope synthesis control unit 433. Based on the timbre information from the timbre selection unit 431, the envelope synthesis control unit 433 selects the parameter that is optimal for the timbre information from the pre-prepared preset synthesis parameters 432, and outputs the selected parameter as weight information.

[0242] In this way, according to Figure 21 The encoding device 100, during the archiving process, simultaneously detects anomalies in the envelope component that affect the timbre, and if an anomaly is found, corrects the envelope component to an envelope component with the best timbre.

[0243] Editing screen of music production / editing tools

[0244] Figure 22 This is a simplified diagram showing an example of the editing screen of the aforementioned music production / editing tool.

[0245] Music production / editing tools include Figure 21 The encoding device 100 corrects or transforms the timbre of the selected signal waveform.

[0246] exist Figure 22In the middle, the main editing screen 500 on the left is a screen used to display the signal waveform of each track or object of the audio data being edited. When the editor selects a predetermined waveform portion WP in track (object) 02 as the track (object) to be edited and presses the change timbre button 501, the timbre selection dialog box 510 in the upper right corner of the figure is displayed.

[0247] In the timbre selection dialog box 510, the editor selects the target timbre to be emphasized or reduced. Here, you can choose fluctuation, sharpness, or roughness as the timbre.

[0248] Furthermore, instead of the tone selection dialog box 510, the tone processing dialog box 520 in the lower right corner of the diagram can be displayed. In this case, the editor can use the sliders to adjust the intensity of the aforementioned fluctuation, sharpness, and roughness. The intensity of each item has a positive and negative direction, and when the intensity is 0, the corresponding tone does not change.

[0249] When the OK button 511 is pressed in the tone selection dialog box 510 or the OK button 521 is pressed in the tone processing dialog box 520, the focus returns to the main editing screen 500. In the main editing screen 500, the selected waveform portion WP of the audio data is processed to emphasize or reduce the tone selected and adjusted in the tone selection dialog box 510 or the tone processing dialog box 520.

[0250] That is, in Figure 21 In the encoding device 100, for the selected timbre, parameters corresponding to the timbre and content are obtained from the preset synthesis parameters 432. The envelope correction unit 411 corrects and processes the timbre by changing the weighting coefficients of the envelope components (a method for processing envelope spectral coefficients) using these parameters. Note that when the intensity is 0, the corresponding envelope spectral coefficients are not processed. Although the range of envelope spectral coefficients processed according to the selected timbre is not defined, Figure 23 The diagram illustrates the corresponding frequency band and modulation frequency band.

[0251] like Figure 23 As explained, when it is desired to change the fluctuation, it is only necessary to process the envelope spectral coefficients corresponding to the modulation band from 5 to 20 Hz based on the intensity adjusted for the entire frequency band in the timbre processing dialog box 520.

[0252] Additionally, when it is desired to change the fluctuations, for frequency bands of 100Hz or higher, it is only necessary to process the envelope spectral coefficients corresponding to the modulation frequency bands from 30Hz to 150Hz based on the intensity adjusted in the timbre processing dialog box 520.

[0253] Furthermore, when it is desired to change the sharpness, it is only necessary to process the envelope spectral coefficients corresponding to the modulation band from 200kHz to 1kHz based on the intensity adjusted in the timbre processing dialog box 520 for the band from 1kHz to 3.5kHz.

[0254] When the above processing is completed, the editor confirms the processed audio data by pressing the play button 502, and stores the processed audio data by pressing the OK button 503. 8. Summary of the Invention

[0256] According to the encoding device 100 of this disclosure ( Figure 1 The audio signal is separated into an envelope component and other components. The envelope component, which is important for timbre and sonic atmosphere, is frequency-transformed and encoded with high precision. The other components are encoded with fewer bits. Therefore, even with a small amount of information transmitted, the original atmosphere, timbre and sense of presence of the sound can be reconstructed without loss.

[0257] Furthermore, since a large number of bits are allocated to the quantization of the envelope component, which is important for the reproduction of timbre, based on the timbre and atmosphere of the sound corresponding to the sound category, even if the signal is transmitted with a small amount of information, the signal can be reproduced as having the best timbre for the sound category.

[0258] Furthermore, according to the encoding device 100 of this disclosure ( Figure 1 ), transmission format ( Figure 6 ) and decoding device 200 ( Figure 8 Because the configuration is dynamically switched for each frequency band between a configuration in which the frequency band division signal is divided into envelope components and components other than envelope components and encoded, and a configuration in which the envelope components are not separated and are directly frequency-converted and encoded, it is possible to reproduce the envelope components of the frequency bands that are important for the reproduced timbre, and reproduce the signal with the best timbre, while suppressing the processing load according to conditions such as the processing load of the reproduction terminal.

[0259] The decoding device according to the present invention ( Figure 8 , Figure 16 and Figure 20 This allows for easy modification of the timbre and sonic atmosphere of the audio signal during playback to achieve the desired characteristics. Specifically, content can be reproduced with optimal timbre based on its type (such as video and audio). Furthermore, even under conditions of poor network transmission or heavy processing load on the playback terminal, sound with optimal timbre, characterized by fewer dropouts and less sound quality degradation, can be heard. Additionally, it can dynamically switch to the timbre optimal for the user's hearing loss characteristics and viewing / hearing device.

[0260] Furthermore, by using a trained neural network to process the envelope component by multiplying it by weighting coefficients, it is easier to transform it into an audio signal that is closer to the desired timbre without tuning the weighting coefficients of the timbre transformation.

[0261] The encoding device 100 according to the present invention ( Figure 21 ) and music production / editing tools ( Figure 22 The editing screen effectively facilitates the transformation into a sound with the desired timbre and atmosphere during music production or editing, thus reducing the time and effort spent on music content creation. Furthermore, because parts causing abnormal sounds or unpleasant sensations can be automatically detected and corrected during archiving to achieve optimal timbre, the workload of archiving can be reduced and the quality of archived content can be improved.

[0262] 9. Computer Configuration Examples

[0263] The above series of processes can be performed by hardware or software. In the case where the series of processes are performed by software, the program constituting the software is installed from the program recording medium into a computer, general-purpose personal computer, or similar device embedded in dedicated hardware.

[0264] Figure 24 This is a block diagram illustrating an example of the hardware configuration of a computer executing the series of processes described above. The encoding device 100 and the decoding device 200 are, for example, composed of components similar to those in... Figure 24 The configuration shown is the configuration of the computer 800.

[0265] CPU 801, Read-Only Memory (ROM) 802 and Random Access Memory (RAM) 803 are interconnected via bus 804.

[0266] The input / output interface 805 is also connected to the bus 804. Input units 806, including a keyboard and mouse, and output units 807, including a display and speakers, are connected to the input / output interface 805. Furthermore, storage units 808, including hard disks and non-volatile memory, communication units 809, including network interfaces, and drivers 810 that drive removable media 811 are also connected to the input / output interface 805.

[0267] In the computer 800 configured as described above, for example, the CPU 801 loads the program stored in the storage unit 808 into the RAM 803 via the input / output interface 805 and the bus 804 and executes the program, thereby performing the series of processes described above.

[0268] For example, the program executed by CPU 801 is recorded on removable medium 811, or provided via wired or wireless transmission media such as local area network, Internet, and digital broadcasting, and installed in storage unit 808.

[0269] Note that the program executed by the computer 800 may be a program for performing processing sequentially in the order described in this specification, or it may be a program for performing processing in parallel or at necessary time intervals (such as when an execution call is performed).

[0270] Note that in this specification, "system" means a group of multiple components (devices, modules (parts), etc.), and it is irrelevant whether all components are housed in the same housing. Therefore, a system refers to multiple devices housed in separate housings and connected via a network, as well as a device in which multiple modules are housed in one housing.

[0271] It should be noted that the effects described in this specification are merely examples and are not limiting, and may produce other effects.

[0272] It should be noted that the embodiments of this disclosure are not limited to the above-described embodiments, and various modifications can be made without departing from the spirit of this disclosure.

[0273] For example, embodiments of this disclosure may employ a cloud computing configuration in which multiple devices perform the processing of a function in a shared and collaborative manner via a network.

[0274] Furthermore, each step described in the above flowchart can be performed by a single device or can be shared and performed by multiple devices.

[0275] Furthermore, when a single step includes multiple processes, the multiple processes contained in a single step can be executed by a single device, or they can be shared and executed by multiple devices.

[0276] The effects described in this specification are merely examples and are not limiting. Other effects may exist.

[0277] Furthermore, the technology according to this disclosure may have the following configurations. (1)

[0279] A decoding device, comprising: The demultiplexing unit is configured to separate first coded data and second coded data from the coded bitstream of the content signal for each frequency band. The first coded data is obtained by encoding the envelope component of the content signal via frequency transformation, and the second coded data is obtained by encoding the flattened waveform of the content signal using a number of bits different from the number of bits used for the envelope component. The frequency band decoding unit is configured to generate a frequency band division signal by synthesizing the envelope component decoded from the first coded data and the flattened waveform decoded from the second coded data; and The frequency band synthesis unit is configured to generate content signals by synthesizing the frequency band division signals of each frequency band. (2)

[0281] According to the decoding device described in (1), the content signal is at least one of an audio signal, a video signal, and a tactile signal. (3)

[0283] According to the decoding device in (1), The first encoded data is a spectrum obtained by frequency transformation of the envelope component, and The second encoded data is the parameters obtained by parametric encoding the flattened waveform. (4)

[0285] According to the decoding device described in (1) or (2), The demultiplexing unit extracts third coded data from the coded bitstream for each frequency band. This third coded data is the spectrum obtained by frequency transformation of the band-divided signal. The frequency band decoding unit outputs the frequency band division signal decoded from the third encoded data. (5)

[0287] According to the decoding apparatus of (4), the demultiplexing unit determines, based on the flags contained in the coded bitstream, whether to separate the first coded data and the second coded data for each frequency band, or to separate only the third coded data. (6)

[0289] According to any one of (1) to (5), the decoding device wherein the frequency band decoding unit performs weighted processing on the first coded data of each frequency band. (7)

[0291] According to the decoding device of (6), the frequency band decoding unit multiplies each frequency element of the first coded data of each frequency band by a different weighting coefficient. (8)

[0293] According to the decoding device described in (7), the band decoding unit uses a trained neural network for weighted processing. (9)

[0295] According to any one of (1) to (5), the band decoding unit performs weighted processing on a specific frequency range of the envelope component decoded from the first coded data of each band. (10)

[0297] According to the decoding apparatus described in (9), the band decoding unit multiplies the band signal obtained by dividing the envelope component by the first filter bank by a weighting coefficient, and then resynthesizes the band signal by the second filter bank. (11)

[0299] The decoding apparatus according to any one of (1) to (10) further includes a weight information determination unit, which determines weight information for setting weighting coefficients in weighting processing related to the first coded data of each frequency band by using pre-prepared arbitrary parameters. (12)

[0301] According to the decoding device of (11), the weight information determination unit selects the parameters based on at least one of the content type of the content signal, the user's configuration information, the terminal information of the decoding device, and the state of the network that sends the encoded bit stream.

[0302] (12-1)

[0303] According to the decoding device described in (12), the content type includes at least one of the following: program type of video, category information of video or music, instrument type and timbre information.

[0304] (12-2)

[0305] According to the decoding device described in (12), the configuration information includes at least one of the user's auditory characteristics, the user's preferences, and the device characteristics of the hearing device worn by the user.

[0306] (12-3)

[0307] According to the decoding device of (12), the terminal information includes at least one of the following: remaining battery power, CPU usage, and memory consumption. (13)

[0309] According to the decoding device described in (11), The encoded bitstream is audio object encoded data containing audio object data and audio object metadata. It further includes a rendering unit configured to render audio objects for each frequency band segmented signal. (14)

[0311] According to the decoding device of (13), the metadata includes at least one of the attribute information, location information and priority information of the audio object. (15)

[0313] According to the decoding apparatus described in (13), the rendering process is a rendering process using at least one of vector-based amplitude shift (VBAP), head-related transfer function (HRTF), room impulse response (RIR), and high-order high-fidelity stereo (HOA). (16)

[0315] According to the decoding device of (13), the weight information determination unit selects parameters based on the metadata selection parameters. (17)

[0317] The decoding device according to any one of (13) to (16), wherein the rendering unit performs rendering processing based on metadata. (18)

[0319] A decoding method, comprising: The decoding device separates first coded data and second coded data from the coded bitstream of the content signal for each frequency band. The first coded data is obtained by encoding the envelope component of the content signal via frequency transformation, and the second coded data is obtained by encoding the flattened waveform of the content signal using a number of bits different from the number of bits used for the envelope component. The decoding device generates a frequency band division signal by synthesizing the envelope component decoded from the first coded data and the flattened waveform decoded from the second coded data; and The content signal is generated by synthesizing the frequency band division signals of each frequency band through a decoding device. (19)

[0321] A program that causes a computer to perform the following processes: For each frequency band, first coded data and second coded data are separated from the coded bitstream of the content signal. The first coded data is obtained by encoding the envelope component of the content signal via frequency transformation, and the second coded data is obtained by encoding the flattened waveform of the content signal using a different number of bits than the number of bits used for the envelope component. A frequency band division signal is generated by synthesizing the envelope component decoded from the first coded data and the flattened waveform decoded from the second coded data; and Content signals are generated by synthesizing the frequency band division signals of each frequency band. (20)

[0323] An encoding device, comprising: The frequency band division unit divides the content signal into frequency band division signals for each frequency band; The frequency band coding unit separates the frequency band-divided signal into an envelope component and a flattened waveform. It encodes the envelope component through frequency transformation and encodes the flattened waveform using a different number of bits than those used for the envelope component. The multiplexing unit generates an encoded bitstream by multiplexing first encoded data and second encoded data. The first encoded data is obtained by encoding the envelope component, and the second encoded data is obtained by encoding the flattened waveform. (twenty one)

[0325] According to the encoding device described in (20), The first encoded data is a spectrum obtained by frequency transformation of the envelope component, and The second encoded data is the parameters obtained by parametric encoding the flattened waveform. (twenty two)

[0327] According to the encoding device of (21), the parameters include either a predetermined signal waveform pattern or a sine wave parameter. (twenty three)

[0329] According to the encoding device described in (21), Among them, the frequency band coding unit encodes the frequency band division signal without changing it through frequency transformation, and The multiplexing unit generates a coded bit stream by multiplexing the first coded data, the second coded data, and the third coded data. The third coded data is the spectrum obtained by encoding the frequency band division signal through frequency transformation. (twenty four)

[0331] According to the encoding device of (23), the frequency band division unit determines the encoding method for each frequency band, and the encoding method is to encode the frequency band division signal after separating the frequency band division signal into envelope components and flattened waveforms, or to encode the frequency band division signal without changing it. (25)

[0333] According to the encoding apparatus described in (23), the frequency band division unit determines the encoding method based on the resource information of the content signal reproduction terminal or the network state in which the encoded bit stream is transmitted. (26)

[0335] According to the encoding apparatus of (21), the frequency band encoding unit quantizes the first encoded data obtained by encoding the envelope component using the number of quantization bits or the quantization bandwidth based on at least one of the content type of the content signal and the network state of the transmitted encoded bit stream. (27)

[0337] According to the encoding device of (26), the content type includes at least one of the content type, the sound type of the content, and the priority level of the audio object. (28)

[0339] According to the coding device of (21), the frequency band coding unit performs weighted processing on the first coded data of each frequency band. (29)

[0341] The encoding apparatus according to (28) further includes a weight information determination unit, which determines the weight information for setting the weighting coefficients in the weighting process by using any one of the pre-prepared parameters. (30)

[0343] According to the encoding device of (29), the weight information determination unit selects parameters from the parameters based on the input from the graphical user interface (GUI). (31)

[0345] According to the coding device of (28), the frequency band coding unit performs weighting processing based on the changes of specific frequency components in the first coding data of each frequency band.

[0346] Reference number list

[0347] 1 Audio signal transmission system; 100 Encoding device; 110 Frequency band division unit; 120, 120A, 120B Frequency band encoding units; 130 Content type acquisition unit; 140 Network status detection unit; 150 Encoding / multiplexing unit; 161 Envelope separation unit; 162 Frequency conversion unit; 163 Envelope quantization unit; 164 Parameterization unit; 171 Frequency conversion unit; 172 Quantization unit. 200 Decoding device, 210 Decoding / demultiplexing unit, 220, 220A, 220B Bandwidth decoding unit, 230 Synthesis parameter determination unit, 240 Bandwidth synthesis unit, 251 Inverse quantization unit, 252 Envelope weighting unit, 253 Inverse frequency transformation unit, 254 Flattened waveform generation unit, 255 Envelope synthesis unit, 261 Inverse quantization unit, 262 Inverse frequency transformation unit, 271 Content type acquisition unit, 272 Preset synthesis parameters, 273 Envelope synthesis control unit, 311 Envelope processing unit, 331 User configuration acquisition unit, 341 Terminal information acquisition unit, 351 Rendering unit, 411 Envelope correction unit, 420 Synthesis parameter determination unit, 431: Timbre selection unit, 432: Preset synthesis parameters, 433: Envelope synthesis control unit.

Claims

1. A decoding device, comprising: The demultiplexing unit is configured to separate first coded data and second coded data from the coded bitstream of the content signal for each frequency band. The first coded data is obtained by encoding the envelope component of the content signal via frequency transformation, and the second coded data is obtained by encoding the flattened waveform of the content signal using a number of bits different from the number of bits used for the envelope component. The frequency band decoding unit is configured to generate a frequency band division signal by synthesizing the envelope component decoded from the first coded data and the flattened waveform decoded from the second coded data; as well as The frequency band synthesis unit is configured to generate the content signal by synthesizing the frequency band division signals of each frequency band.

2. The decoding device according to claim 1, wherein, The content signal is at least one of audio signals, video signals, and tactile signals.

3. The decoding device according to claim 1, in, The first encoded data is a spectrum obtained by frequency transformation of the envelope components, and The second encoded data are parameters obtained by parametric encoding the flattened waveform.

4. The decoding device according to claim 1, in, The demultiplexing unit extracts third coded data from the coded bitstream for each frequency band. This third coded data is a spectrum obtained by frequency transformation of the frequency band division signal. The frequency band decoding unit outputs the frequency band division signal decoded from the third encoded data.

5. The decoding device according to claim 4, wherein, The demultiplexing unit determines, based on the flags contained in the coded bitstream, whether to separate the first coded data and the second coded data for each frequency band, or to separate only the third coded data.

6. The decoding device according to claim 1, wherein, The frequency band decoding unit performs weighted processing on the first encoded data of each frequency band.

7. The decoding apparatus according to claim 6, wherein, The frequency band decoding unit multiplies each frequency element of the first coded data in each frequency band by a different weighting coefficient.

8. The decoding apparatus according to claim 7, wherein, The frequency band decoding unit uses a trained neural network to perform the weighted processing.

9. The decoding device according to claim 1, wherein, The frequency band decoding unit performs weighted processing on a specific frequency range of the envelope component decoded from the first coded data of each frequency band.

10. The decoding apparatus according to claim 9, wherein, The frequency band decoding unit multiplies the frequency band signal obtained by dividing the envelope component by the first filter group by a weighting coefficient, and then resynthesizes the frequency band signal through the second filter group.

11. The decoding apparatus according to claim 1, further comprising: The weight information determination unit is configured to determine weight information for setting weighting coefficients in the weighting process associated with the first coded data of each frequency band by using pre-prepared arbitrary parameters.

12. The decoding apparatus according to claim 11, wherein, The weight information determination unit selects the parameters based on at least one of the following: the content type of the content signal, the user's configuration information, the terminal information of the decoding device, and the status of the network that sends the encoded bit stream.

13. The decoding apparatus according to claim 11, in, The encoded bitstream is encoded audio object data containing audio object data and metadata of the audio objects; and It further includes a rendering unit configured to perform audio object rendering processing on the frequency band division signal of each frequency band.

14. The decoding apparatus according to claim 13, wherein, The metadata includes at least one of the following: attribute information, location information, and priority information of the audio object.

15. The decoding apparatus according to claim 13, wherein, The rendering process uses at least one of vector-based amplitude translation (VBAP), head-related transfer function (HRTF), room impulse response (RIR), and high-order high-fidelity stereo (HOA).

16. The decoding apparatus according to claim 13, wherein, The weight information determination unit selects the parameters from the parameters based on the metadata.

17. The decoding apparatus according to claim 13, wherein, The rendering unit performs the rendering process based on the metadata.

18. A decoding method, comprising: The decoding device separates first coded data and second coded data from the coded bitstream of the content signal for each frequency band. The first coded data is obtained by encoding the envelope component of the content signal via frequency transformation, and the second coded data is obtained by encoding the flattened waveform of the content signal using a number of bits different from the number of bits used for the envelope component. The decoding device generates a frequency band division signal by synthesizing the envelope component decoded from the first encoded data and the flattened waveform decoded from the second encoded data; and The decoding device generates the content signal by synthesizing the frequency band division signals of each frequency band.

19. A program that causes a computer to perform the following processes: For each frequency band, first coded data and second coded data are separated from the coded bitstream of the content signal, wherein, The first encoded data is obtained by encoding the envelope component of the content signal via frequency transformation, and the second encoded data is obtained by encoding the flattened waveform of the content signal using a number of bits different from the number of bits used for the envelope component. A frequency band division signal is generated by synthesizing the envelope component decoded from the first coded data and the flattened waveform decoded from the second coded data; as well as The content signal is generated by synthesizing the frequency band division signals of each frequency band.

20. An encoding device, comprising: The frequency band division unit is configured to divide the content signal into frequency band division signals for each frequency band; A frequency band coding unit is configured to separate the frequency band division signal into an envelope component and a flattened waveform, encode the envelope component by frequency transformation, and encode the flattened waveform using a number of bits different from the number of bits used for the envelope component. as well as The multiplexing unit is configured to generate an coded bitstream by multiplexing first coded data and second coded data, wherein the first coded data is obtained by encoding the envelope component and the second coded data is obtained by encoding the flattened waveform.