Audio reconstruction method and device using machine learning
Through machine learning technology, the high cost and copyright issues of high-quality audio transmission are solved through machine learning technology, and the efficient reconstruction and quality improvement of low-quality audio is achieved.
Patent Information
- Application Number
- CN201780095363.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2017-10-24
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2037-10-24
AI Technical Summary
The prior art requires high bandwidth and high cost when transmitting and storing high-quality audio, and has copyright problems. At the same time, low-quality audio lacks high-quality records and is difficult to effectively reconstruct.
Through machine learning methods, high-quality audio signals are reconstructed using decoded parameters and audio codec information, including the use of machine learning models to determine and correct the characteristics of decoded parameters, application of bandwidth expansion techniques and steady-state/transient signal recognition.
Improves the effect of low-quality audio reconstruction into high-quality audio, reduces bandwidth requirements, reduces copyright risks, and improves the quality of audio signals.
Smart Images

Figure CN111164682B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an audio reconstruction method and apparatus, and more particularly, to an audio reconstruction method and apparatus that provide improved sound quality by reconstructing decoded parameters or audio signals obtained from a bitstream using machine learning. Background Art
[0002] Audio codec technologies capable of transmitting, reproducing, and storing high-quality audio content have been developed, and current ultra-high sound quality technologies enable the transmission, reproduction, and storage of audio with a resolution of 24 bits / 192 kHz. A 24-bit / 192 kHz resolution means that the original audio is sampled at 192 kHz, and the sampled signal can be represented in 2^24 states using 24 bits.
[0003] However, high-bandwidth data transmission may be required to transmit high-quality audio content. In addition, high-quality audio content has a high service price and requires a high-quality audio codec, which may cause copyright issues. Furthermore, although high-quality audio services have recently started to be provided, there may be no audio recorded in high quality. Therefore, there is an increasing need for technologies for reconstructing low-quality audio into high-quality. To reconstruct low-quality audio into high-quality, artificial intelligence (AI) can be used.
[0004] An AI system is a computer system capable of achieving human-level intelligence and refers to a system in which a machine can autonomously learn, make decisions, and become more intelligent, unlike existing rule-based intelligent systems. The recognition rate can be improved in proportion to the iteration of the AI system and user preferences can be understood more accurately. Therefore, existing rule-based intelligent systems are gradually being replaced by AI systems based on deep learning.
[0005] AI technology includes machine learning (or deep learning) and elemental technologies using machine learning. Machine learning refers to algorithmic technologies for automatically classifying / learning the characteristics of input data, and elemental technologies refer to technologies that mimic the functions of the human brain (e.g., recognition and decision-making) by using machine learning algorithms such as deep learning, including technical fields such as language understanding, visual understanding, reasoning / prediction, knowledge representation, and operation control.
[0006] Examples of various fields to which AI technology can be applied are described below. Language understanding refers to technologies for recognizing and applying / processing human languages / characters, including natural language processing, machine translation, dialogue systems, queries and responses, speech recognition / synthesis, etc. Visual understanding refers to technologies for recognizing and processing objects like human vision, including object recognition, object tracking, image search, human recognition, scene understanding, spatial understanding, image enhancement, etc. Reasoning / prediction refers to a technology for determining information and making logical inferences and predictions, including knowledge / probability-based reasoning, optimized prediction, preference-based planning, recommendation, etc. Knowledge representation refers to a technology for automatically processing human experience information into knowledge data, including knowledge construction (data generation / data classification), knowledge management (data utilization), etc. Operation control refers to technologies for controlling the autonomous driving of vehicles and the movement of robots, and includes motion control (e.g., navigation, collision avoidance, and driving control), manipulation control (e.g., action control), etc.
[0007] According to the present disclosure, machine learning is performed using original audio and various decoding parameters of an audio codec to obtain reconstructed decoding parameters. According to the present disclosure, the reconstructed decoding parameters can be used to reconstruct higher-quality audio. Summary of the Invention
[0008] Solution to the Problem
[0009] A method and device for reconstructing decoding parameters or an audio signal obtained from a bitstream using machine learning are provided.
[0010] According to an aspect of the present disclosure, an audio reconstruction method includes: obtaining a plurality of decoding parameters of a current frame by decoding a bitstream; determining characteristics of a second parameter included in the plurality of decoding parameters and associated with a first parameter based on the first parameter; obtaining a reconstructed second parameter by applying a machine learning model to at least one of the plurality of decoding parameters, the second parameter, and the characteristics of the second parameter; and decoding an audio signal based on the reconstructed second parameter.
[0011] Decoding the audio signal may include: obtaining a corrected second parameter by correcting the reconstructed second parameter based on the characteristics of the second parameter; and decoding the audio signal based on the corrected second parameter.
[0012] Determining the characteristics of the second parameter may include: determining a range of the second parameter based on the first parameter, and wherein obtaining the corrected second parameter includes: when the reconstructed second parameter is not within the range, obtaining a value within the range that is closest to the reconstructed second parameter as the corrected second parameter.
[0013] Determining the characteristics of the second parameter may include: determining the characteristics of the second parameter by using a machine learning model pre-trained based on at least one of the first parameter and the second parameter.
[0014] Obtaining the reconstructed second parameter may include: determining a plurality of candidates for the second parameter based on the characteristics of the second parameter; and selecting one candidate from the plurality of candidates for the second parameter based on the machine learning model.
[0015] Obtaining the reconstructed second parameter may further include: obtaining the reconstructed second parameter of the current frame based on at least one decoding parameter among a plurality of decoding parameters of a previous frame.
[0016] The machine learning model may be generated by machine learning an original audio signal and at least one decoding parameter among the plurality of decoding parameters.
[0017] According to another aspect of the present disclosure, an audio reconstruction method includes: obtaining a plurality of decoding parameters of a current frame by decoding a bitstream; decoding an audio signal based on the plurality of decoding parameters; selecting one machine learning model from a plurality of machine learning models based on the decoded audio signal and at least one decoding parameter among the plurality of decoding parameters; and reconstructing the decoded audio signal by using the selected machine learning model.
[0018] The machine learning model may be generated by machine learning the decoded audio signal and the original audio signal
[0019] Selecting the machine learning model may include: determining a start frequency of bandwidth expansion based on at least one decoding parameter among the plurality of decoding parameters; and selecting a machine learning model of the decoded audio signal based on the start frequency and the frequency of the decoded audio signal.
[0020] Selecting the machine learning model may include: obtaining a gain of the current frame based on at least one decoding parameter among the plurality of decoding parameters; obtaining an average value of the gains of the current frame and a frame adjacent to the current frame; when a difference between the gain of the current frame and the average value of the gains is greater than a threshold, selecting a machine learning model for a transient signal; when the difference between the gain of the current frame and the average value of the gains is less than the threshold, determining whether a window type included in the plurality of decoding parameters is indicated as short; when the window type is indicated as short, selecting the machine learning model for the transient signal; and when the window type is not indicated as short, selecting a machine learning model for a steady-state signal.
[0021] According to another aspect of the present disclosure, an audio reconstruction device includes: a memory that stores a received bitstream; and at least one processor configured to obtain a plurality of decoding parameters for a current frame by decoding the bitstream, determine a characteristic of a second parameter included in the plurality of decoding parameters and associated with the first parameter based on a first parameter included in the plurality of decoding parameters, obtain a reconstructed second parameter by applying a machine learning model to at least one decoding parameter, the second parameter, and the characteristic of the second parameter among the plurality of decoding parameters, and decode an audio signal based on the reconstructed second parameter.
[0022] The at least one processor may be further configured to obtain a corrected second parameter by correcting the reconstructed second parameter based on the characteristic of the second parameter, and decode an audio signal based on the corrected second parameter.
[0023] The at least one processor may be further configured to determine the characteristic of the second parameter by using a machine learning model pre-trained based on at least one of the first parameter and the second parameter.
[0024] The at least one processor may be further configured to obtain the reconstructed second parameter by determining a plurality of candidates for the second parameter based on the characteristic of the second parameter and selecting one candidate from the plurality of candidates for the second parameter based on the machine learning model.
[0025] The at least one processor may be further configured to obtain the reconstructed second parameter for the current frame further based on at least one decoding parameter among a plurality of decoding parameters of a previous frame.
[0026] The machine learning model may be generated by machine learning an original audio signal and at least one decoding parameter among the plurality of decoding parameters.
[0027] According to another aspect of the present disclosure, an audio reconstruction device includes: a memory that stores a received bitstream; and at least one processor configured to obtain a plurality of decoding parameters for a current frame by decoding the bitstream, decode an audio signal based on the plurality of decoding parameters, select a machine learning model from a plurality of machine learning models based on the decoded audio signal and at least one decoding parameter among the plurality of decoding parameters, and reconstruct the decoded audio signal by using the selected machine learning model.
[0028] According to another aspect of the present invention, a computer-readable recording medium having recorded thereon a computer program for performing the above method. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 is a block diagram of an audio reconstruction device according to an embodiment.
[0030] Figure 2 is a block diagram of an audio reconstruction device according to an embodiment.
[0031] Figure 3 is a flowchart of an audio reconstruction method according to an embodiment.
[0032] Figure 4 is a block diagram for describing machine learning according to an embodiment.
[0033] Figure 5 shows a prediction of the characteristics of decoding parameters according to an embodiment.
[0034] Figure 6 shows a prediction of the characteristics of decoding parameters according to an embodiment.
[0035] Figure 7 is a flowchart of an audio reconstruction method according to an embodiment.
[0036] Figure 8 shows decoding parameters according to an embodiment.
[0037] Figure 9 shows a change in decoding parameters according to an embodiment.
[0038] Figure 10 shows a change in decoding parameters in the case of an increase in the number of bits according to an embodiment.
[0039] Figure 11 shows a change in decoding parameters according to an embodiment.
[0040] Figure 12 is a block diagram of an audio reconstruction device according to an embodiment.
[0041] Figure 13 is a flowchart of an audio reconstruction method according to an embodiment.
[0042] Figure 14 is a flowchart of an audio reconstruction method according to an embodiment.
[0043] Figure 15 is a flowchart of an audio reconstruction method according to an embodiment. Detailed Description
[0044] One or more embodiments of the present disclosure and their implementation methods can be more easily understood by referring to the following detailed description and the accompanying drawings. However, the present disclosure may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the present disclosure to those of ordinary skill in the art.
[0045] The terms used in this specification will now be briefly described before the detailed description of the present disclosure.
[0046] Although as many of the terms used herein as possible have been selected from currently widely used general terms in consideration of the functions obtained according to the present disclosure, these terms may be replaced by other terms for the purpose of one of ordinary skill in the art, custom, or the emergence of new technologies. In certain cases, terms arbitrarily selected by the applicant may be used, and in such cases, the meanings of these terms may be described in the relevant parts of the present disclosure. Therefore, it should be noted that the terms used herein are to be interpreted based on their actual meanings and the entire content of this specification, rather than merely based on the names of the terms.
[0047] As used herein, the singular forms "a", "an", and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. Additionally, unless the context clearly indicates otherwise, the plural forms are also intended to include the singular form.
[0048] It will be understood that when used herein, the terms "comprises", "comprising", "includes", and / or "including" specify the presence of the stated element(s), but do not preclude the presence or addition of one or more other elements.
[0049] As used herein, the term "unit" represents a software or hardware element and performs a specific function. However, a "unit" is not limited to software or hardware. The "unit" may be formed in an addressable storage medium or may be formed to operate one or more processors. Thus, for example, a "unit" may include elements (such as software elements, object-oriented software elements, class elements, and task elements), processes, functions, attributes, programs, subroutines, program code segments, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, or variables. The functions provided by the elements and "units" may be combined into a smaller number of elements and "units" or may be divided into additional elements and "units".
[0050] According to an embodiment of the present disclosure, a "unit" may be implemented as a processor and a memory. The term "processor" should be interpreted broadly to cover general-purpose processors, central processing units (CPUs), microprocessors, digital signal processors (DSPs), controllers, microcontrollers, state machines, etc. In some cases, a "processor" may refer to an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a field-programmable gate array (FPGA), etc. The term "processor" may refer to a combination of processing devices, such as a combination of a DSP and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors and a DSP core, or any other such configured combination.
[0051] The term "memory" should be interpreted broadly to cover any electronic component capable of storing electronic information. The term memory may refer to various types of processor-readable media, such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, magnetic or optical data storage, and registers. When a processor can read information from a memory and / or record information on the memory, the memory is said to be in electronic communication with the processor. Memory integrated with a processor is in electronic communication with the processor.
[0052] Hereinafter, embodiments of the present disclosure will be described in detail by referring to the accompanying drawings. In the drawings, parts irrelevant to the embodiments of the present disclosure are not shown for clarity.
[0053] High-quality audio content requires high service prices and high-quality audio codecs, which may cause copyright issues. In addition, high-quality audio services have recently started, yet there may be no audio recorded in high quality. Therefore, there is an increasing need for techniques for reconstructing audio encoded in low quality into high quality. One of the methods available for reconstructing audio encoded in low quality into high quality is a method using machine learning. Now, reference will be made to Figures 1 to 15 Describe a method for improving the quality of decoded audio by using decoding parameters of a codec and machine learning.
[0054] Figure 1 is a block diagram of an audio reconstruction device 100 according to an embodiment.
[0055] The audio reconstruction device 100 may include a receiver 110 and a decoder 120. The receiver 110 may receive a bitstream. The decoder 120 may output an audio signal decoded based on the received bitstream. Now, reference will be made to Figure 2 Describe the audio reconstruction device 100 in detail.
[0056] Figure 2 It is a block diagram of an audio reconstruction device 100 according to an embodiment.
[0057] The audio reconstruction device 100 may include a codec information extractor 210 and at least one decoder. The codec information extractor 210 may equivalently correspond to Figure 1 the receiver 110. The at least one decoder may include at least one of a first decoder 221, a second decoder 222, and an Nth decoder 223. At least one of the first decoder 221, the second decoder 222, and the Nth decoder 223 may equivalently correspond to Figure 1 the decoder 120.
[0058] The codec information extractor 210 may receive a bitstream. The bitstream may be generated by an encoding device. The encoding device may encode and compress the original audio into the bitstream. The codec information extractor 210 may receive the bitstream from the encoding device or a storage medium through wired or wireless communication. The codec information extractor 210 may store the bitstream in a memory. The codec information extractor 210 may extract various types of information from the bitstream. The various types of information may include codec information. The codec information may include information about the technology used to encode the original audio. The technology used to encode the original audio may include, for example, MPEG Layer-3 (MP3), Advanced Audio Coding (AAC), or High Efficiency AAC (HE-AAC) technology. The codec information extractor 210 may select a decoder from the at least one decoder based on the codec information.
[0059] The at least one decoder may include a first decoder 221, a second decoder 222, and an Nth decoder 223. The decoder selected by the codec information extractor 210 from the at least one decoder may decode an audio signal based on the bitstream. For ease of explanation, the Nth decoder 223 will now be described. The first decoder 221 and the second decoder 222 may have a structure similar to that of the Nth decoder 223.
[0060] The Nth decoder 223 may include an audio signal decoder 230. The audio signal decoder 230 may include a lossless decoder 231, an inverse quantizer 232, a stereo signal reconstructor 233, and an inverse converter 234.
[0061] The lossless decoder 231 can receive a bitstream. The lossless decoder 231 can decode the bitstream and output at least one decoded parameter. The lossless decoder 231 can decode the bitstream without losing information. The inverse quantizer 232 can receive at least one decoded parameter from the lossless decoder 231. The inverse quantizer 232 can inverse-quantize at least one decoded parameter. The inverse-quantized decoded parameter can be a mono signal. The stereo signal reconstructor 233 can reconstruct a stereo signal based on the inverse-quantized decoded parameter. The inverse converter 234 can convert the stereo signal in the frequency domain and output the decoded audio signal in the time domain.
[0062] The decoded parameter can include at least one of a spectral bin, a scale factor gain, a global gain, spectral data, and a window type. The decoded parameter can be a parameter used by a codec such as an MP3, AAC, or HE-AAC codec. However, the decoded parameter is not limited to a specific codec, and decoded parameters called by different names can perform similar functions. The decoded parameter can be sent in units of frames. A frame is a unit divided from the original audio signal in the time domain.
[0063] The spectral bin can correspond to a signal amplitude depending on the frequency in the frequency domain.
[0064] The scale factor gain and the global gain are values for scaling the spectral bin. For multiple frequency bands included in a frame, the scale factor can have different values.
[0065] For all frequency bands in a frame, the global gain can have the same value. The audio reconstruction device 100 can obtain an audio signal in the frequency domain by multiplying the spectral bin by the scale factor gain and the global gain.
[0066] The spectral data is information indicating the characteristics of the spectral bin. The spectral data can indicate the sign of the spectral bin. The spectral data can indicate whether the value of the spectral bin is 0.
[0067] The window type can indicate the characteristics of the original audio signal. The window type can correspond to a time period used to convert the original audio signal in the time domain to the frequency domain. When the original audio signal is a steady-state signal with little change, the window type may indicate "long". When the original audio signal is a transient signal with large changes, the window type may indicate "short".
[0068] The Nth decoder 223 may include at least one of a parameter characteristic determiner 240 and a parameter reconstructor 250. The parameter characteristic determiner 240 may receive at least one decoding parameter and determine the characteristics of the at least one decoding parameter. The parameter characteristic determiner 240 may use machine learning to determine the characteristics of the at least one decoding parameter. The parameter characteristic determiner 240 may use a first decoding parameter included in the at least one decoding parameter to determine the characteristics of a second decoding parameter included in the at least one decoding parameter. The parameter characteristic determiner 240 may output the at least one decoding parameter and the characteristics of the decoding parameter to the parameter reconstructor 250. The parameter characteristic determiner 240 will be described in detail below with reference to Figures 4 to 6 describe the parameter characteristic determiner 240 in detail.
[0069] According to an embodiment of the present disclosure, the parameter reconstructor 250 may receive the at least one decoding parameter from the lossless decoder 231. The parameter reconstructor 250 may reconstruct the at least one decoding parameter. The parameter reconstructor 250 may use a machine learning model to reconstruct the at least one decoding parameter. The audio signal decoder 230 may output a decoded audio signal close to the original audio based on the reconstructed at least one decoding parameter.
[0070] According to another embodiment of the present disclosure, the parameter reconstructor 250 may receive the at least one decoding parameter and the characteristics of the decoding parameter from the parameter characteristic determiner 240. The parameter reconstructor 250 may output the reconstructed parameter by applying a machine learning model to the at least one decoding parameter and the characteristics of the decoding parameter. The parameter reconstructor 250 may output the reconstructed parameter by applying a machine learning model to the at least one decoding parameter. The parameter reconstructor 250 may correct the reconstructed parameter based on the parameter characteristics. The parameter reconstructor 250 may output the corrected parameter. The audio signal decoder 230 may output a decoded audio signal close to the original audio based on the corrected parameter.
[0071] The parameter reconstructor 250 may output at least one of the reconstructed at least one decoding parameter and the corrected parameter to the parameter characteristic determiner 240 or the parameter reconstructor 250. At least one of the parameter characteristic determiner 240 and the parameter reconstructor 250 may receive at least one of the at least one decoding parameter and the corrected parameter of the previous frame. The parameter characteristic determiner 240 may output the parameter characteristics of the current frame based on at least one of the at least one decoding parameter and the corrected parameter of the previous frame. The parameter reconstructor 250 may obtain the reconstructed parameter of the current frame based on at least one of the at least one decoding parameter and the corrected parameter of the previous frame.
[0072] Now, the parameter characteristic determiner 240 and the parameter reconstructor 250 will be described in detail with reference to Figures 3 to 11 describe the parameter characteristic determiner 240 and the parameter reconstructor 250 in detail.
[0073] Figure 3It is a flowchart of an audio reconstruction method according to an embodiment.
[0074] In operation 310, the audio reconstruction device 100 may obtain a plurality of decoding parameters of the current frame by decoding the bitstream. In operation 320, the audio reconstruction device 100 may determine the characteristics of the second parameter. In operation 330, the audio reconstruction device 100 may obtain the reconstructed second parameter by using a machine learning model. In operation 340, the audio reconstruction device 100 may decode the audio signal based on the reconstructed second parameter.
[0075] The audio reconstruction device 100 may obtain a plurality of decoding parameters of the current frame by decoding the bitstream (operation 310). The lossless decoder 231 may obtain a plurality of decoding parameters by decoding the bitstream. The lossless decoder 231 may output the decoding parameters to the inverse quantizer 232, the parameter characteristic determiner 240, or the parameter reconstructor 250. The audio reconstruction device 100 may determine where to output the decoding parameters by analyzing the decoding parameters. According to an embodiment of the present disclosure, the audio reconstruction device 100 may determine where to output the decoding parameters based on a predetermined rule. However, the present disclosure is not limited thereto, and the bitstream may include information on where the decoding parameters need to be output. The audio reconstruction device 100 may determine where to output the decoding parameters based on the information included in the bitstream.
[0076] When high audio quality can be ensured without modifying at least one of the plurality of decoding parameters, the audio reconstruction device 100 may not modify at least one of the decoding parameters. The lossless decoder 231 may output at least one of the decoding parameters to the inverse quantizer 232. At least one of the decoding parameters does not pass through the parameter characteristic determiner 240 or the parameter reconstructor 250 and thus may not be modified. The audio reconstruction device 100 does not use the parameter characteristic determiner 240 and the parameter reconstructor 250 for some decoding parameters and thus may effectively use computing resources.
[0077] According to an embodiment of the present disclosure, the audio reconstruction device 100 may determine to modify at least one of the decoding parameters. The lossless decoder 231 may output at least one of the decoding parameters to the parameter reconstructor 250. The audio reconstruction device 100 may obtain the reconstructed decoding parameters based on the decoding parameters by using a machine learning model. The audio reconstruction device 100 may decode the audio signal based on the reconstructed decoding parameters. The audio reconstruction device 100 may provide an audio signal with improved quality based on the reconstructed decoding parameters. The machine learning model will be described in detail below with reference to Figure 4 Describe the machine learning model in detail.
[0078] According to another embodiment of the present disclosure, the audio reconstruction device 100 may determine to modify a plurality of decoding parameters. The lossless decoder 231 may output the plurality of decoding parameters to the parameter characteristic determiner 240.
[0079] The parameter characteristic determiner 240 may determine the characteristic of a second parameter included in the plurality of decoding parameters based on a first parameter included in the plurality of decoding parameters (operation 320). The second parameter may be associated with the first parameter. The first parameter may directly or indirectly represent the characteristic of the second parameter. For example, the first parameter may include at least one of a scale factor gain, a global gain, spectral data, and a window type of the second parameter.
[0080] The first parameter may be a parameter adjacent to the second parameter. The first parameter may be a parameter included in the same frequency band or frame as the second parameter. The first parameter may be a parameter included in a frequency band or frame adjacent to the frequency band or frame including the second parameter.
[0081] Although the first parameter and the second parameter are described as different parameters for ease of explanation herein, the first parameter may be the same as the second parameter. That is, the parameter characteristic determiner 240 may determine the characteristic of the second parameter based on the second parameter.
[0082] The parameter reconstructor 250 may obtain a reconstructed second parameter by applying a machine learning model to at least one of the plurality of decoding parameters, the second parameter, and the characteristic of the second parameter (operation 330). The audio reconstruction device 100 may decode the audio signal based on the reconstructed second parameter (operation 340). Decoding the audio signal based on the second parameter reconstructed by applying the machine learning model may provide excellent quality. Now, the machine learning model will be described with reference to Figure 4 is described in detail.
[0083] Figure 4 is a block diagram for describing machine learning according to an embodiment.
[0084] The data learner 410 and the data applicator 420 may operate at different times. For example, the data learner 410 may operate earlier than the data applicator 420. The parameter characteristic determiner 240 and the parameter reconstructor 250 may include at least one of the data learner 410 and the data applicator 420.
[0085] Refer to Figure 4 According to an embodiment, the data learner 410 may include a data acquirer 411, a preprocessor 412, and a machine learner 413. The process in which the data learner 410 receives the input data 431 and outputs the machine learning model 432 may be referred to as a training process.
[0086] The data acquirer 411 can receive input data 431. The input data 431 can include at least one of a raw audio signal and decoding parameters. The raw audio signal can be a high-quality recorded audio signal. The raw audio signal can be represented in the frequency domain or the time domain. The decoding parameters can correspond to the result of encoding the raw audio signal. When encoding the raw audio signal, some information may be lost. That is, compared with the raw audio signal, the audio signal decoded based on multiple decoding parameters may have lower quality.
[0087] The preprocessor 412 can preprocess the input data 431 for learning. The preprocessor 412 can process the input data 431 into a preset format in such a way that the machine learner 413 described below can use the input data 431. When the raw audio signal and multiple decoding parameters have different formats, the raw audio signal or multiple decoding parameters can be converted to the format of the other. For example, when the raw audio signal and multiple decoding parameters are related to different codecs, the codec information of the raw audio signal and multiple decoding parameters can be modified to achieve compatibility between them. When the raw audio signal and multiple decoding parameters are represented in different domains, they can be modified to be represented in the same domain.
[0088] The preprocessor 412 can select the data required for learning from the input data 431. The selected data can be provided to the machine learner 413. The preprocessor 412 can select the data required for learning from the preprocessed data according to a preset criterion. The preprocessor 412 can select data according to a preset criterion through the learning of the machine learner 413 described below. Since a large amount of input data requires a long data processing time, the efficiency of data processing can be improved when a part of the input data 431 is selected.
[0089] The machine learner 413 can output a machine learning model 432 based on the selected input data. The selected input data can include at least one of multiple decoding parameters and a raw audio signal. The machine learning model 432 can include a criterion for reconstructing at least one parameter from multiple decoding parameters. The machine learner 413 can learn to minimize the difference between the raw audio signal and the audio signal decoded based on the reconstructed decoding parameters. The machine learner 413 can learn a criterion for selecting a part of the input data 431 to reconstruct at least one parameter from multiple decoding parameters.
[0090] The machine learning unit 413 can learn a machine learning model 432 by using input data 431. In this case, the machine learning model 432 can be a pre-trained model. For example, the machine learning model 432 can be a model pre-trained by receiving default training data (e.g., at least one decoding parameter). The default training data can be the initial data for constructing the pre-trained model.
[0091] For example, the machine learning model 432 can be selected in consideration of the application field of the machine learning model 432, the learning purpose, or the computing performance of the device. The machine learning model 432 can be, for example, a neural network-based model. For example, the machine learning model 432 can use a deep neural network (DNN) model, a recurrent neural network (RNN) model, or a bidirectional recurrent deep neural network (BRDNN) model, but is not limited thereto.
[0092] According to various embodiments, when there are multiple pre-built machine learning models, the machine learning unit 413 can determine a machine learning model highly relevant to the input data 431 or the default training data as the machine learning model to be trained. In this case, the input data 431 or the default training data can be pre-classified according to the data type, and the machine learning models can be pre-built according to the data type. For example, the input data 431 or the default training data can be pre-classified according to various criteria (e.g., the region where the data is generated, the time when the data is generated, the data size, the data type, the data generator, the object type in the data, and the data format).
[0093] The machine learning unit 413 can train the machine learning model 432 by using, for example, a learning algorithm (e.g., error backpropagation or gradient descent).
[0094] The machine learning unit 413 can train the machine learning model 432 by using the input data 431 as an input value through, for example, supervised learning. The machine learning unit 413 can train the machine learning model 432 through, for example, unsupervised learning for finding the criteria for making decisions based on the data type required for making decisions autonomously. The machine learning unit 413 can train the machine learning model 432 through, for example, reinforcement learning by using feedback on whether the result of the decision made through learning is correct.
[0095] The machine learning unit 413 can perform machine learning by using Equation 1 and Equation 2.
[0096] [Equation 1]
[0097]
[0098] [Equation 2]
[0099] y = softmax(evidencei)
[0100] In Equations 1 and 2, x represents the selected input data for the machine learning model, y represents the probability of each candidate, i represents the index of the candidate, j represents the index of the selected input data for the machine learning model, W represents the weight matrix of the input data, and b represents the deflection parameter.
[0101] The machine learner 413 can obtain prediction data by using arbitrary weights W and arbitrary deflection parameters b. The prediction data can be the reconstructed decoding parameters. The machine learner 413 can calculate the cost y. The cost can be the difference between the actual data and the prediction data. For example, the cost can be the difference between the data related to the original audio signal and the data related to the reconstructed decoding parameters. The machine learner 413 can update the weights W and the deflection parameter b to minimize the cost.
[0102] The machine learner 413 can obtain the weights and deflection parameters corresponding to the minimum cost. The machine learner 413 can represent the weights and deflection parameters corresponding to the minimum cost in a matrix. The machine learner 413 can obtain the machine learning model 432 by using at least one of the weights and deflection parameters corresponding to the minimum cost. The machine learning model 432 can correspond to the weight matrix and the deflection parameter matrix.
[0103] When training the machine learning model 432, the machine learner 413 can store the trained machine learning model 432. In this case, the machine learner 413 can store the trained machine learning model 432 in the memory of the data learner 410. The machine learner 413 can store the trained machine learning model 432 in the memory of the data applicator 420 described below. Optionally, the machine learner 413 can store the trained machine learning model 432 in the memory of an electronic device or a server connected to a wired or wireless network.
[0104] In this case, the memory storing the trained machine learning model 432 can also store, for example, commands or data related to at least one other component of the electronic device. The memory can store software and / or programs. The programs can include, for example, a kernel, middleware, an application programming interface (API), and / or an application (or “app”).
[0105] A model evaluator (not shown) can input evaluation data into the machine learning model 432 and, when the result output using the evaluation data does not meet a specific criterion, request the machine learner 413 to relearn. In this case, the evaluation data can be preset data for evaluating the machine learning model 432.
[0106] For example, when the number or proportion of the evaluation data corresponding to inaccurate results in the results of the machine learning model 432 trained based on the evaluation data exceeds a preset threshold, the model evaluator may evaluate that the specific criteria are not met. For example, when the specific criteria are defined as a ratio of 2%, and when the trained machine learning model 432 outputs incorrect results for more than 20 pieces of evaluation data out of a total of 1000 pieces of evaluation data, the model evaluator may evaluate that the trained machine learning model 432 is not appropriate.
[0107] When there are multiple trained machine learning models, the model evaluator may evaluate whether each trained machine learning model meets the specific criteria and determine the model that meets the specific criteria as the final machine learning model. In this case, when multiple machine learning models meet the specific criteria, the model evaluator may determine any machine learning model or a certain number of machine learning models preset in the order of evaluation scores as the final machine learning model 432.
[0108] At least one of the data acquirer 411, the preprocessor 412, the machine learner 413, and the model evaluator in the data learner 410 may be fabricated in the form of at least one hardware chip and installed in an electronic device. For example, at least one of the data acquirer 411, the preprocessor 412, the machine learner 413, and the model evaluator may be fabricated in the form of a dedicated hardware chip for AI or as part of a general-purpose processor (e.g., a central processing unit (CPU) or an application processor) or a dedicated graphics processor (e.g., a graphics processing unit (GPU)), and installed in various electronic devices.
[0109] The data acquirer 411, the preprocessor 412, the machine learner 413, and the model evaluator may be installed in one electronic device or separately installed in different electronic devices. For example, some of the data acquirer 411, the preprocessor 412, the machine learner 413, and the model evaluator may be included in the electronic device, while others may be included in the server.
[0110] At least one of the data acquirer 411, the preprocessor 412, the machine learner 413, and the model evaluator may be implemented as a software module. When at least one of the data acquirer 411, the preprocessor 412, the machine learner 413, and the model evaluator is implemented as a software module (or a program module including instructions), the software module may be stored in a non-transitory computer-readable medium. In this case, at least one software module may be provided by an operating system (OS) or a specific application. Optionally, some of the at least one software module may be provided by the OS, while others may be provided by the specific application.
[0111] Reference Figure 4, according to the embodiment, the data applicator 420 may include a data acquirer 421, a preprocessor 422, and a result provider 423. The process in which the data applicator 420 receives the input data 441 and the machine learning model 432 and outputs the output data 442 may be referred to as a testing process.
[0112] The data acquirer 421 may acquire the input data 441. The input data 441 may include at least one decoding parameter for decoding an audio signal. The preprocessor 422 may preprocess the input data 441 to make it available. The preprocessor 422 may process the input data 441 into a preset format in such a way that the result provider 423 described below can use the input data 441.
[0113] The preprocessor 422 may select the data to be used by the result provider 423 from the preprocessed input data. The preprocessor 422 may select at least one decoding parameter for improving the quality of the audio signal from the preprocessed input data. The selected data may be provided to the result provider 423. The preprocessor 422 may select a part or all of the preprocessed input data according to a preset criterion for improving the quality of the audio signal. The preprocessor 422 may select the data according to a criterion preset through the learning of the machine learner 413.
[0114] The result provider 423 may output the output data 442 by applying the data selected by the preprocessor 422 to the machine learning model 432. The output data 442 may be a reconstructed decoding parameter for providing an improved sound quality. The audio reconstruction device 100 may output a decoded audio signal close to the original audio signal based on the reconstructed decoding parameter.
[0115] The result provider 423 may provide the output data 442 to the preprocessor 422. The preprocessor 422 may preprocess the output data 442 and provide it to the result provider 423. For example, the output data 442 may be the reconstructed decoding parameter of the previous frame. The result provider 423 may provide the output data 442 of the previous frame to the preprocessor 422. The preprocessor 422 may provide the selected decoding parameter of the current frame and the reconstructed decoding parameter of the previous frame to the result provider 423. The result provider 423 may generate the output data 442 of the current frame by reflecting not only the reconstructed decoding parameter of the current frame but also the information about the previous frame. The output data 442 of the current frame may include at least one of the reconstructed decoding parameter of the current frame and the corrected decoding parameter. The audio reconstruction device 100 may provide higher-quality audio based on the output data 442 of the current frame.
[0116] A model updater (not shown) may control the update of the machine learning model 432 based on an evaluation of the output data 442 provided by the result provider 423. For example, the model updater may request the machine learning engine 413 to update the machine learning model 432 by providing the output data 442 provided by the result provider 423 to the machine learning engine 413.
[0117] At least one of the data acquirer 421, the pre-processor 422, the result provider 423, and the model updater in the data applicator 420 may be fabricated in the form of at least one hardware chip and installed in an electronic device. For example, at least one of the data acquirer 421, the pre-processor 422, the result provider 423, and the model updater may be fabricated in the form of a dedicated hardware chip for AI or as part of a general-purpose processor (e.g., CPU or application processor) or a dedicated graphics processor (e.g., GPU), and installed in various electronic devices.
[0118] The data acquirer 421, the pre-processor 422, the result provider 423, and the model updater may be installed in one electronic device or separately installed in different electronic devices. For example, some of the data acquirer 421, the pre-processor 422, the result provider 423, and the model updater may be included in an electronic device, while others may be included in a server.
[0119] At least one of the data acquirer 421, the pre-processor 422, the result provider 423, and the model updater may be implemented as a software module. When at least one of the data acquirer 421, the pre-processor 422, the result provider 423, and the model updater is implemented as a software module (or a program module including instructions), the software module may be stored in a non-transitory computer-readable medium. In this case, the OS or a certain application may provide at least one software module. Optionally, some of the at least one software module may be provided by the OS, while others may be provided by a specific application.
[0120] Now, reference will be made to Figures 5 to 11 a detailed description Figure 1 of the operation of the audio reconstruction device 100 and Figure 4 the data learner 410 and the data applicator 420.
[0121] Figure 5 shows a prediction of the characteristics of the decoding parameters according to an embodiment.
[0122] The parameter characteristic determiner 240 may determine the characteristics of the decoding parameters. The audio reconstruction device 100 does not need to process parameters that do not satisfy the characteristics of the decoding parameters, so the computational amount can be reduced. Compared with the input decoding parameters, the audio reconstruction device 100 can prevent the output of reconstructed decoding parameters with poor quality.
[0123] The graph 510 shows the signal amplitude depending on the frequency in the frame. The multiple decoding parameters obtained by the audio reconstruction device 100 based on the bitstream may include the signal amplitude depending on the frequency. For example, the signal amplitude may correspond to a spectral bin.
[0124] The multiple decoding parameters may include a first parameter and a second parameter. The parameter characteristic determiner 240 may determine the characteristics of the second parameter based on the first parameter. The first parameter may be a parameter adjacent to the second parameter. The audio reconstruction device 100 may determine the characteristics of the second parameter based on the trend of the first parameter. The characteristics of the second parameter may include the range of the second parameter.
[0125] According to an embodiment of the present disclosure, the second parameter may be the signal amplitude 513 at the frequency f3. The first parameter may include the signal amplitudes 511, 512, 514, and 515 corresponding to the frequencies f1, f2, f4, and f5. The audio reconstruction device 100 may determine that the signal amplitudes 511, 512, 514, and 515 corresponding to the first parameter are in an upward trend. Therefore, the audio reconstruction device 100 may determine that the range of the signal amplitude 513 corresponding to the second parameter is between the signal amplitudes 512 and 514.
[0126] Figure 2 The parameter characteristic determiner 240 may include Figure 4 a data learner 410. The machine learning model 432 may be pre-trained by the data learner 410.
[0127] For example, the data learner 410 of the parameter characteristic determiner 240 may receive information corresponding to the original audio signal. The information corresponding to the original audio signal may be the original audio signal itself or the information obtained by high-quality encoding of the original audio signal. The data learner 410 of the parameter characteristic determiner 240 may receive the decoding parameters. The parameters received by the data learner 410 of the parameter characteristic determiner 240 may correspond to at least one frame. The data learner 410 of the parameter characteristic determiner 240 may output the machine learning model 432 based on the operations of the data acquirer 411, the pre-processor 412, and the machine learner 413. The machine learning model 432 of the data learner 410 may be a machine learning model for determining the characteristics of the second parameter based on the first parameter. For example, the machine learning model 432 may be specified as the weight for each of at least one first parameter.
[0128] The parameter characteristic determiner 240 may include Figure 4 a data applicator 420. The parameter characteristic determiner 240 may determine the characteristics of a second parameter based on at least one of a first parameter and a second parameter. The parameter characteristic determiner 240 may use a pre-trained machine learning model to determine the characteristics of the second parameter.
[0129] For example, the data applicator 420 of the parameter characteristic determiner 240 may receive at least one of the first parameter and the second parameter included in a plurality of decoded parameters of a current frame. The data applicator 420 of the parameter characteristic determiner 240 may receive the machine learning model 432 from the data learner 410 of the parameter characteristic determiner 240. The data applicator 420 of the parameter characteristic determiner 240 may determine the characteristics of the second parameter based on the operations of the data acquirer 421, the pre-processor 422, and the result provider 423. For example, the data applicator 420 of the parameter characteristic determiner 240 may determine the characteristics of the second parameter by applying the machine learning model 432 to at least one of the first parameter and the second parameter.
[0130] According to another embodiment of the present disclosure, the audio reconstruction device 100 may provide high-bitrate audio by reconstructing a second parameter not included in the bitstream. The second parameter may be the signal amplitude at the frequency f0. The bitstream may not contain information about the signal amplitude at the frequency f0. The audio reconstruction device 100 may estimate the signal characteristics at the frequency f0 based on the first parameter. The first parameter may include signal amplitudes 511, 512, 513, 514, and 515 corresponding to frequencies f1, f2, f3, f4, and f5. The audio reconstruction device 100 may determine that the signal amplitudes 511, 512, 513, 514, and 515 corresponding to the first parameter are in an upward trend. Therefore, the audio reconstruction device 100 may determine that the range of the signal amplitude corresponding to the second parameter is between the signal amplitudes 514 and 515. The audio reconstruction device 100 may include Figure 4 at least one of a data learner 410 and a data applicator 420. A detailed description of the operations of the data learner 410 or the data applicator 420 is provided above, and thus will not be repeated here.
[0131] Referring to graph 520, the second parameter can be the signal amplitude 523 at frequency f3. The first parameter can include signal amplitudes 521, 522, 524, and 525 corresponding to frequencies f1, f2, f4, and f5. The audio reconstruction device 100 can determine that the signal amplitudes 521, 522, 524, and 525 corresponding to the first parameter are in a trend of rising first and then falling. Since the signal amplitude 524 corresponding to frequency f4 is greater than the signal amplitude 522 corresponding to frequency f2, the audio reconstruction device 100 can determine that the range of the signal amplitude 523 corresponding to the second parameter is greater than or equal to the signal amplitude 524.
[0132] Referring to graph 530, the second parameter can be the signal amplitude 533 at frequency f3. The first parameter can include signal amplitudes 531, 532, 534, and 535 corresponding to frequencies f1, f2, f4, and f5. The audio reconstruction device 100 can determine that the signal amplitudes 531, 532, 534, and 535 corresponding to the first parameter are in a trend of falling first and then rising. Since the signal amplitude 534 corresponding to frequency f4 is less than the signal amplitude 532 corresponding to frequency f2, the audio reconstruction device 100 can determine that the range of the signal amplitude 533 corresponding to the second parameter is less than or equal to the signal amplitude 534.
[0133] Referring to graph 540, the second parameter can be the signal amplitude 543 at frequency f3. The first parameter can include signal amplitudes 541, 542, 544, and 545 corresponding to frequencies f1, f2, f4, and f5. The audio reconstruction device 100 can determine that the signal amplitudes 541, 542, 544, and 545 corresponding to the first parameter are in a falling trend. The audio reconstruction device 100 can determine that the range of the signal amplitude 543 corresponding to the second parameter is between the signal amplitudes 542 and 544.
[0134] Figure 6 Shows a prediction of the characteristics of the decoding parameters according to an embodiment.
[0135] The audio reconstruction device 100 can use multiple frames to determine the characteristics of the decoding parameters in a frame. The audio reconstruction device 100 can use the previous frame of the frame to determine the characteristics of the decoding parameters in the frame. For example, the audio reconstruction device 100 can use at least one decoding parameter included in frame n - 2, frame n - 1 610, frame n 620, or frame n + 1 630 to determine the characteristics of at least one decoding parameter included in frame n + 1 630.
[0136] The audio reconstruction device 100 may obtain decoding parameters from a bitstream. The audio reconstruction device 100 may obtain graphs 640, 650, and 660 based on the decoding parameters in multiple frames. Graph 640 shows the decoding parameters in frame n - 1 610 in the frequency domain. The decoding parameters shown in graph 640 may represent a frequency-dependent signal amplitude. Graph 650 shows the signal amplitude depending on the frequency in frame n 620 in the frequency domain. Graph 660 shows the signal amplitude depending on the frequency in frame n + 1 630 in the frequency domain. The audio reconstruction device 100 may determine the characteristics of the signal amplitude included in graph 660 based on at least one of the signal amplitudes included in graphs 640, 650, and 660.
[0137] According to an embodiment of the present disclosure, the audio reconstruction device 100 may determine the characteristics of the signal amplitude 662 included in graph 660 based on at least one of the signal amplitudes included in graphs 640, 650, and 660. The audio reconstruction device 100 may examine the trends of the signal amplitudes 641, 642, and 643 in graph 640. The audio reconstruction device 100 may examine the trends of the signal amplitudes 651, 652, and 653 in graph 650. The trends may rise and then fall near f3. The audio reconstruction device 100 may determine the trend of graph 660 based on graphs 640 and 650. The audio reconstruction device 100 may determine that the signal amplitude 662 is greater than or equal to the signal amplitudes 661 and 663.
[0138] According to another embodiment of the present disclosure, the audio reconstruction device 100 may determine the characteristics of the signal amplitude at f0 in graph 660 based on at least one of the signal amplitudes included in graphs 640, 650, and 660. The audio reconstruction device 100 may examine the trend of the signal amplitude in graph 640. The audio reconstruction device 100 may examine the trend of the signal amplitude in graph 650. The trend may fall near f0. The audio reconstruction device 100 may determine the trend of graph 660 based on graphs 640 and 650. The audio reconstruction device 100 may determine that the signal amplitude at f0 is less than or equal to the signal amplitude at f4 and greater than or equal to the signal amplitude at f5. The audio reconstruction device 100 may include Figure 4 at least one of a data learner 410 and a data applicator 420. A detailed description of the operations of the data learner 410 or the data applicator 420 is provided above, and thus will not be repeated here.
[0139] According to an embodiment of the present disclosure, the audio reconstruction device 100 may use a previous frame of a frame to determine characteristics of decoding parameters included in the frame. The audio reconstruction device 100 may determine characteristics of frequency-dependent signals in the current frame based on frequency-dependent signals in the previous frame. The audio reconstruction device 100 may determine characteristics of the frequency-dependent decoding parameters in the current frame based on, for example, a distribution range, an average value, a median value, a minimum value, a maximum value, a deviation, or a sign of the frequency-dependent signals in the previous frame.
[0140] For example, the audio reconstruction device 100 may determine characteristics of the signal amplitude 662 included in the graph 660 based on at least one signal amplitude included in the graphs 640 and 650. The audio reconstruction device 100 may determine characteristics of the signal amplitude 662 at the frequency f3 in the graph 660 based on the signal amplitude 642 at the frequency f3 in the graph 640 and the signal amplitude 652 at the frequency f3 in the graph 650. The characteristics of the signal amplitude 662 may be based on, for example, a distribution range, an average value, a median value, a minimum value, a maximum value, a deviation, or a sign of the signal amplitudes 642 and 652.
[0141] According to an embodiment of the present disclosure, the audio reconstruction device 100 may obtain decoding parameters from a bitstream. The decoding parameters may include a second parameter. Characteristics of the second parameter may be determined based on a predetermined parameter that is not a decoding parameter.
[0142] For example, a quantization step may not be included in the decoding parameters. The second parameter may correspond to a signal amplitude depending on a frequency in a frame. The signal amplitude may correspond to a spectral bin. The audio reconstruction device 100 may determine a range of the spectral bin based on the quantization step. The quantization step is determined as a range of the signal amplitude of the spectral bin. The quantization step may vary with frequency. The quantization step may be dense in the audible frequency region. The quantization step may be sparse in a region other than the audible frequency region. Therefore, when the frequency corresponding to the spectral bin is known, the quantization step may be determined. The range of the spectral bin may be determined based on the quantization step.
[0143] According to another embodiment of the present disclosure, the audio reconstruction device 100 may obtain decoding parameters from a bitstream. The decoding parameters may include a first parameter and a second parameter. Characteristics of the second parameter may be determined based on the first parameter. The characteristics of the second parameter may include a range of the second parameter.
[0144] For example, the first parameter may include a scaling factor and a masking threshold. The quantization step size can be determined based on the scaling factor and the masking threshold. As described above, the scaling factor is a value used to scale the spectral bins. For the multiple frequency bands included in a frame, the scaling factor may have different values. When there is noise called a masker, the masking threshold is the minimum amplitude of the current signal that makes the current signal audible. The masking threshold may vary depending on the frequency and the type of masker. When the masker is close to the current signal in frequency, the masking threshold can be increased.
[0145] For example, the current signal may appear at f0, while the masker signal may appear at f1 close to f0. The masking threshold at f0 can be determined based on the masker at f1. When the amplitude of the current signal at f0 is less than the masking threshold, the current signal may be an inaudible sound. Therefore, the audio reconstruction device 100 may ignore the current signal at f0 during the encoding or decoding process. Otherwise, when the amplitude of the current signal at f0 is greater than the masking threshold, the current signal may be an audible sound. Therefore, the audio reconstruction device 100 may not ignore the current signal at f0 during the encoding or decoding process.
[0146] The audio reconstruction device 100 may set the quantization step size in the scaling factor and the masking threshold to a smaller value. The audio reconstruction device 100 may determine the range of the spectral bins based on the quantization step size.
[0147] Figure 7 is a flowchart of an audio reconstruction method according to an embodiment.
[0148] In operation 710, the audio reconstruction device 100 may obtain a plurality of decoding parameters for the current frame of the audio signal to be decoded by decoding the bitstream. In operation 720, the audio reconstruction device 100 may determine the characteristics of a second parameter included in the plurality of decoding parameters based on a first parameter included in the plurality of decoding parameters. In operation 730, the audio reconstruction device 100 may obtain the reconstructed second parameter by using a machine learning model based on at least one of the plurality of decoding parameters. In operation 740, the audio reconstruction device 100 may obtain the corrected second parameter by correcting the second parameter based on the characteristics of the second parameter. In operation 750, the audio reconstruction device 100 may decode the audio signal based on the corrected second parameter.
[0149] Operation 710 and operation 750 may be performed by the audio signal decoder 230. Operation 720 may be performed by the parameter characteristic determiner 240. Operations 730 and 740 may be performed by the parameter reconstructor 250.
[0150] Return to refer to according to an embodiment of the present disclosure Figure 3, the data learner 410 and the data applicator 420 of the parameter reconstructor 250 may receive the characteristics of the second parameter as input. That is, the parameter reconstructor 250 may perform machine learning based on the characteristics of the second parameter. The data learner 410 of the parameter reconstructor 250 may output the machine learning model 432 by reflecting the characteristics of the second parameter. The data applicator 420 of the parameter reconstructor 250 may output the output data 442 by reflecting the characteristics of the second parameter.
[0151] Referring to another embodiment according to the present disclosure Figure 7 , the data learner 410 and the data applicator 420 of the parameter reconstructor 250 may not receive the characteristics of the second parameter as input. That is, the parameter reconstructor 250 may perform machine learning only based on the decoded parameter without based on the characteristics of the second parameter. The data learner 410 of the parameter reconstructor 250 may output the machine learning model 432 without reflecting the characteristics of the second parameter. The data applicator 420 of the parameter reconstructor 250 may output the output data 442 without reflecting the characteristics of the second parameter.
[0152] The output data 442 may be the reconstructed second parameter. The parameter reconstructor 250 may determine whether the reconstructed second parameter satisfies the characteristics of the second parameter. When the reconstructed second parameter satisfies the characteristics of the second parameter, the parameter reconstructor 250 may output the reconstructed parameter to the audio signal decoder 230. When the reconstructed second parameter does not satisfy the characteristics of the second parameter, the parameter reconstructor 250 may obtain the corrected second parameter by correcting the reconstructed second parameter based on the characteristics of the second parameter. The parameter reconstructor 250 may output the corrected parameter to the audio signal decoder 230.
[0153] For example, the characteristics of the second parameter may include the range of the second parameter. The audio reconstruction device 100 may determine the range of the second parameter based on the first parameter. When the reconstructed second parameter is not within the range of the second parameter, the audio reconstruction device 100 may obtain the value closest to the reconstructed second parameter within the range as the corrected second parameter. Now referring to Figure 8 A detailed description thereof will be provided.
[0154] Figure 8 Shows the decoded parameter according to an embodiment.
[0155] The curve graph 800 shows the signal amplitude depending on the frequency of the original audio signal. The curve graph 800 may correspond to a frame of the original audio signal. The original audio signal is represented by a curve 805 with a continuous waveform. The original audio signal may be sampled at frequencies f1, f2, f3, and f4. The magnitudes of the original audio signal at frequencies f1, f2, f3, and f4 may be represented by points 801, 802, 803, and 804. The original audio signal may be encoded. The audio reconstruction device 100 may generate decoding parameters by decoding the encoded original audio signal.
[0156] The curve graph 810 shows the signal amplitude depending on the frequency. The dashed line 815 shown in the curve graph 810 may correspond to the original audio signal. The points 811, 812, 813, and 814 shown in the curve graph 810 may correspond to the decoding parameters. The decoding parameters may be output from the lossless decoder 231 of the audio reconstruction device 100. At least one of the original audio signal and the decoding parameters may be scaled and displayed in the curve graph 810.
[0157] As shown in the curve graph 810, the dashed line 815 may be different from the points 811, 812, 813, and 814. The difference between the dashed line 815 and the points 811, 812, 813, and 814 may be due to errors caused during the encoding and decoding of the original audio signal.
[0158] The audio reconstruction device 100 may determine the characteristics of the decoding parameters corresponding to the points 811, 812, 813, and 814. The audio reconstruction device 100 may use a machine learning model to determine the characteristics of the decoding parameters. A detailed description of determining the characteristics of the decoding parameters is provided above regarding Figure 5 and Figure 6 and will not be elaborated here. The decoding parameters may be spectral bins. The characteristics of the decoding parameters may include the range of the spectral bins.
[0159] The range of the spectral bins determined by the audio reconstruction device 100 is shown in the curve graph 830. That is, the arrow 835 indicates that the point 831 corresponds to the available range of the spectral bin. The arrow 836 indicates that the point 832 corresponds to the available range of the spectral bin. The arrow 837 indicates that the point 833 corresponds to the available range of the spectral bin. The arrow 838 indicates that the point 834 corresponds to the available range of the spectral bin.
[0160] The audio reconstruction device 100 may determine the signal characteristics at f0 between f2 and f3. The audio reconstruction device 100 may not receive the decoding parameters at f0. The audio reconstruction device 100 may determine the characteristics of the decoding parameters at f0 based on the decoding parameters related to f0.
[0161] For example, the audio reconstruction device 100 may not receive information about the amplitude of the spectral bin at f0. The audio reconstruction device 100 may determine the range of the signal amplitude at f0 by using the spectral bins of frequencies adjacent to f0 and the spectral bins of frames adjacent to the current frame. As described above regarding Figure 5 and Figure 6 provides its detailed description and will not be elaborated here.
[0162] The audio reconstruction device 100 may reconstruct decoding parameters. The audio reconstruction device 100 may use a machine learning model. To reconstruct the decoding parameters, the audio reconstruction device 100 may apply at least one decoding parameter and the characteristics of the decoding parameter to the machine learning model.
[0163] The decoding parameters reconstructed by the audio reconstruction device 100 are shown in the graph 850. The points 851, 852, 853, 854, and 855 represent the reconstructed decoding parameters. Compared with the decoding parameters before reconstruction, the reconstructed decoding parameters may have a larger error. For example, although the point 834 corresponding to the spectral bin in the graph 830 is close to the original audio signal, the point 854 corresponding to the spectral bin in the graph 850 may be far from the original audio signal represented by the dashed line 860.
[0164] The audio reconstruction device 100 may correct the decoding parameters. The audio reconstruction device 100 may determine whether the decoding parameter is within the available range of the decoding parameter. When the decoding parameter is not within the available range of the decoding parameter, the audio reconstruction device 100 may correct the decoding parameter. The corrected decoding parameter may be within the available range of the decoding parameter.
[0165] For example, the graph 870 shows the corrected spectral bin. The points 871, 872, 873, and 875 corresponding to the spectral bin may be within the available range of the spectral bin. However, the point 874 corresponding to the spectral bin may not be within the available range 878 of the spectral bin. When the reconstructed spectral bin is not within the available range 878 of the spectral bin, the audio reconstruction device 100 may obtain the value closest to the reconstructed spectral bin within the range 878 as the corrected spectral bin. When the value of the point 874 corresponding to the reconstructed spectral bin is greater than the maximum value of the range 878, the audio reconstruction device 100 may obtain the maximum value of the range 878 as the point 880 corresponding to the corrected spectral bin. That is, the audio reconstruction device 100 may correct the point 874 corresponding to the reconstructed spectral bin to the point 880. The point 880 may correspond to the corrected spectral bin.
[0166] The audio reconstruction device 100 can decode an audio signal based on the corrected decoding parameters. By using the point 875 corresponding to the reconstructed spectral bin at the frequency f0, the sampling rate of the audio signal can be improved. By using the point 880 corresponding to the reconstructed spectral bin at the frequency f4, the amplitude of the audio signal can be accurately represented. Since the corrected decoding parameters are close to the original audio signal in the frequency domain, the decoded audio signal can be close to the original audio signal.
[0167] Figure 9 Shows the change of the decoding parameters according to an embodiment.
[0168] The graph 910 corresponds to Figure 8 the graph 810 of. The graph 910 shows the signal amplitude depending on the frequency. The dashed line 915 shown in the graph 910 can correspond to the original audio signal. The points 911, 912, 913, and 914 shown in the graph 910 can correspond to the decoding parameters. At least one of the original audio signal and the decoding parameters can be scaled and shown in the graph 910.
[0169] The audio reconstruction device 100 can determine the characteristics of the decoding parameters corresponding to the points 911, 912, 913, and 914. The audio reconstruction device 100 can use a machine learning model to determine the characteristics of the decoding parameters. The above regarding Figure 5 and Figure 6 provides a detailed description of determining the characteristics of the decoding parameters, which will not be repeated here. The decoding parameters can be spectral bins. The characteristics of the decoding parameters can include the range of the spectral bins. The range of the spectral bins determined by the audio reconstruction device 100 is shown in the graph 930.
[0170] The audio reconstruction device 100 can determine candidates for fine-tuning each spectral bin. The audio reconstruction device 100 can represent the spectral bins by using multiple bits. The audio reconstruction device 100 can finely represent the spectral bins proportionally to the number of bits used to represent the spectral bins. The audio reconstruction device 100 can increase the number of bits used to represent the spectral bins to fine-tune the spectral bins. Now, the case of increasing the number of bits used to represent the spectral bins will be described with reference to Figure 10 Describe the case of increasing the number of bits used to represent the spectral bins.
[0171] Figure 10 Shows the change of the decoding parameters in the case of increasing the number of bits according to an embodiment.
[0172] Referring to curve graph 1000, the audio reconstruction device 100 can use two bits to represent the quantized decoding parameters. In this case, the audio reconstruction device 100 can represent the quantized decoding parameters by using "00", "01", "10", and "11". That is, the audio reconstruction device 100 can represent four magnitudes of the decoding parameters. The audio reconstruction device 100 can assign the minimum value of the decoding parameters to "00". The audio reconstruction device 100 can assign the maximum value of the decoding parameters to "11".
[0173] The magnitude of the decoding parameter received by the audio reconstruction device 100 can correspond to point 1020. The magnitude of the decoding parameter can be "01". However, the actual magnitude of the decoding parameter before being quantized can correspond to star 1011, 1012, or 1013. When the actual magnitude of the decoding parameter corresponds to star 1011, the error range can correspond to arrow 1031. When the actual magnitude of the decoding parameter corresponds to star 1012, the error range can correspond to arrow 1032. When the actual magnitude of the decoding parameter corresponds to star 1013, the error range can correspond to arrow 1033.
[0174] Referring to curve graph 1050, the audio reconstruction device 100 can use three bits to represent the quantized decoding parameters. In this case, the audio reconstruction device 100 can represent the quantized decoding parameters by using "000", "001", "010", "011", "100", "101", "110", and "111". That is, the audio reconstruction device 100 can represent eight magnitudes of the decoding parameters. The audio reconstruction device 100 can assign the minimum value of the decoding parameters to "000". The audio reconstruction device 100 can assign the maximum value of the decoding parameters to "111".
[0175] The magnitudes of the decoding parameters received by the audio reconstruction device 100 can correspond to points 1071, 1072, and 1073. The magnitudes of the decoding parameters can be "001", "010", and "011". The actual magnitudes of the decoding parameters can correspond to stars 1061, 1062, and 1063. When the actual magnitude of the decoding parameter corresponds to star 1061, the error range can correspond to arrow 1081. When the actual magnitude of the decoding parameter corresponds to star 1062, the error range can correspond to arrow 1082. When the actual magnitude of the decoding parameter corresponds to star 1063, the error range can correspond to arrow 1083.
[0176] When comparing curve graph 1000 with curve graph 1050, the error of the decoding parameter in curve graph 1050 is smaller than the error of the decoding parameter in curve graph 1000. As Figure 10 shown, the decoding parameter can be accurately represented in proportion to the number of bits used by the audio reconstruction device 100 to represent the decoding parameter.
[0177] Return reference Figure 9 Figure 9 , the audio reconstruction device 100 can determine candidates for fine-tuning each decoding parameter. Referring to the reference curve 950, the audio reconstruction device 100 can additionally use one bit to represent the decoding parameter. The audio reconstruction device 100 can determine candidates 951, 952, and 953 corresponding to the decoding parameter 931 of the curve 930. The audio reconstruction device 100 can use the characteristics of the decoding parameter to determine the candidates 951, 952, and 953 of the decoding parameter. For example, the characteristics of the decoding parameter can include the range 954 of the decoding parameter. The candidates 951, 952, and 953 can be within the range 954 of the decoding parameter.
[0178] The audio reconstruction device 100 can select one of the candidates 951, 952, and 953 of the decoding parameter based on a machine learning model. The audio reconstruction device 100 can include at least one of a data learner 410 and a data applicator 420. The audio reconstruction device 100 can select a decoding parameter by applying at least one of the decoding parameter of the current frame and the decoding parameter of the previous frame to the machine learning model. The machine learning model can be pre-trained.
[0179] The decoding parameter can include a first parameter and a second parameter. The audio reconstruction device 100 can use the first parameter associated with the second parameter to select one of multiple candidates of the second parameter.
[0180] Referring to the reference curve 960, the audio reconstruction device 100 can obtain the selected decoding parameter 961. The audio reconstruction device 100 can obtain an audio signal decoded based on the selected decoding parameter 961.
[0181] Referring to the reference curve 970, the audio reconstruction device 100 can additionally use two bits to represent the decoding parameter. The audio reconstruction device 100 can determine candidates 971, 972, 973, 974, and 975 corresponding to the decoding parameter 931 of the curve 930. Compared with the candidates 951, 952, and 953 of the curve 950, the candidates 971, 972, 973, 974, and 975 have more detailed values. Compared with the case of using one bit, the audio reconstruction device 100 can reconstruct a more accurate decoding parameter by using two bits. The audio reconstruction device 100 can use the characteristics of the decoding parameter to determine the candidates 971, 972, 973, 974, and 975 of the decoding parameter. For example, the characteristics of the decoding parameter can include the range 976 of the decoding parameter. The candidates 971, 972, 973, 974, and 975 can be within the range 976 of the decoding parameter.
[0182] The audio reconstruction device 100 can select one of the candidates 971, 972, 973, 974, and 975 of the decoding parameters based on a machine learning model. The audio reconstruction device 100 can select a decoding parameter by applying at least one of the decoding parameters of the current frame and the decoding parameters of the previous frame to the machine learning model. The decoding parameters can include a first parameter and a second parameter. The audio reconstruction device 100 can use the first parameter associated with the second parameter to select one of the multiple candidates of the second parameter.
[0183] Referring to the graph 980, the audio reconstruction device 100 can obtain the selected decoding parameter 981. Compared with the selected decoding parameter 961 of the graph 960, the selected decoding parameter 981 can have a more accurate value. Compared with the selected decoding parameter 961, the selected decoding parameter 981 can be closer to the original audio signal. The audio reconstruction device 100 can obtain an audio signal decoded based on the selected decoding parameter 981.
[0184] Figure 11 Shows the change of the decoding parameter according to an embodiment.
[0185] The audio reconstruction device 100 can receive a bitstream. The audio reconstruction device 100 can obtain decoding parameters based on the bitstream. The audio reconstruction device 100 can determine the characteristics of the decoding parameters. The characteristics of the decoding parameters can include a symbol. When the amplitude of the decoding parameter is 0, the amplitude 0 can be used as the characteristic of the decoding parameter.
[0186] For example, the decoding parameter can be spectral data. The spectral data can indicate the symbol of the spectral bin. The spectral data can indicate whether the value of the spectral bin is 0. The spectral data can be included in the bitstream. The audio reconstruction device 100 can generate spectral data based on the bitstream.
[0187] The decoding parameters can include a first parameter and a second parameter. The audio reconstruction device 100 can determine the characteristics of the second parameter based on the first parameter. The first parameter can be spectral data. The second parameter can be a spectral bin.
[0188] The graph 1110 shows the amplitude of the decoding parameter depending on the frequency. The decoding parameter can be a spectral bin. The decoding parameter can have various symbols. For example, the decoding parameter 1111 can have a negative sign. The decoding parameter 1113 can have a positive sign. The audio reconstruction device 100 can determine the symbols of the decoding parameters 1111 and 1113 as the characteristics of the decoding parameters 1111 and 1113. The amplitude of the decoding parameter 1112 can be 0. The audio reconstruction device 100 can determine the amplitude of 0 as the characteristic of the decoding parameter 1112.
[0189] According to an embodiment of the present disclosure, the audio reconstruction device 100 may determine reconstructed decoding parameters by applying decoding parameters to a machine learning model. The graph 1130 shows the magnitudes of the reconstructed decoding parameters depending on the frequency. The audio reconstruction device 100 may obtain the reconstructed decoding parameters 1131, 1132, and 1133 by reconstructing the decoding parameters 1111, 1112, and 1113. However, the reconstructed decoding parameters 1131 and 1133 may have signs different from those of the decoding parameters 1111 and 1113. Different from the decoding parameter 1112, the reconstructed decoding parameter 1132 may have a non-zero value.
[0190] The audio reconstruction device 100 may obtain corrected decoding parameters by correcting the reconstructed decoding parameters based on the characteristics of the decoding parameters. The audio reconstruction device 100 may correct the reconstructed decoding parameters based on the signs of the decoding parameters. Referring to the graph 1150, the audio reconstruction device 100 may obtain the corrected decoding parameters 1151 and 1153 by correcting the signs of the reconstructed decoding parameters 1131 and 1133. The audio reconstruction device 100 may obtain the corrected decoding parameter 1152 by correcting the magnitude of the reconstructed decoding parameter 1132 to 0.
[0191] According to another embodiment of the present disclosure, the audio reconstruction device 100 may obtain reconstructed decoding parameters by applying a machine learning model to the decoding parameters and the characteristics of the decoding parameters. That is, the audio reconstruction device 100 may obtain the reconstructed parameters shown in the graph 1150 based on the decoding parameters shown in the graph 1110.
[0192] Figure 12 is a block diagram of the audio reconstruction device 100 according to an embodiment.
[0193] The audio reconstruction device 100 may include a codec information extractor 1210, an audio signal decoder 1220, a bitstream analyzer 1230, a reconstruction method selector 1240, and at least one reconstructor.
[0194] The codec information extractor 1210 may equivalently correspond to Figure 1 the receiver 110 of Figure 2 The codec information extractor 1210 may equivalently correspond to the codec information extractor 210 of
[0195] The codec information extractor 1210 may receive a bitstream and determine the technology used to encode the bitstream. The technology used to encode the original audio may include, for example, MP3, AAC, or HE-AAC technology. Figure 2The audio signal decoder 230. The audio signal decoder 1220 may include a lossless decoder, an inverse quantizer, a stereo signal reconstructor, and an inverse converter. The audio signal decoder 1220 may output a reconstructed audio signal based on the codec information received from the codec information extractor 1210.
[0196] The bitstream analyzer 1230 may obtain decoding parameters for the current frame based on the bitstream. The bitstream analyzer 1230 may check the characteristics of the reconstructed audio signal based on the decoding parameters. The bitstream analyzer 1230 may send information about the signal characteristics to the reconstruction method selector 1240.
[0197] For example, the decoding parameters may include at least one of spectral bins, scale factor gains, global gains, window types, buffer levels, time-domain noise shaping (TNS) information, and perceptual noise substitution (PNS) information.
[0198] The spectral bins may correspond to signal amplitudes depending on frequencies in the frequency domain. The audio encoding device may send accurate spectral bins only for the frequency range sensitive to humans to reduce data. For high-frequency or low-frequency regions that humans cannot hear, the audio encoding device may not send any spectral bins or may send inaccurate spectral bins. The audio reconstruction device 100 may apply bandwidth extension techniques to regions for which no spectral bins are sent. The bitstream analyzer 1230 may determine the frequency regions where the spectral bins are accurately sent and the frequency regions where the spectral bins are inaccurately sent by analyzing the spectral bins. The bitstream analyzer 1230 may send information about the frequencies to the reconstruction method selector 1240.
[0199] For example, bandwidth extension techniques can generally be applied to high-frequency regions. The bitstream analyzer 1230 may determine the minimum frequency value of the frequency region where the spectral bins are inaccurately sent as the starting frequency. The bitstream analyzer 1230 may determine that bandwidth extension techniques need to be applied starting from the starting frequency. The bitstream analyzer 1230 may send information about the starting frequency to the reconstruction method selector 1240.
[0200] The scale factor gains and global gains are values used to scale the spectral bins. The bitstream analyzer 1230 may obtain the characteristics of the reconstructed audio signal by analyzing the scale factor gains and global gains. For example, when the scale factor gains and global gains of the current frame change rapidly, the bitstream analyzer 1230 may determine that the current frame is a transient signal. When the scale factor gains and global gains of the frame are almost unchanged, the bitstream analyzer 1230 may determine that the frame is a steady-state signal. The bitstream analyzer 1230 may send information indicating whether the frame is a steady-state signal or a transient signal to the reconstruction method selector 1240.
[0201] The window type can correspond to a time period used to convert an original audio signal in the time domain into the frequency domain. When the window type of the current frame indicates "long", the bitstream analyzer 1230 can determine that the current frame is a steady-state signal. When the window type of the current frame indicates "short", the bitstream analyzer 1230 can determine that the current frame is a transient signal. The bitstream analyzer 1230 can send information indicating whether the frame is a steady-state signal or a transient signal to the reconstruction method selector 1240.
[0202] The buffer level is information about the available number of bits remaining after encoding a frame. The buffer level is used to encode data using variable bit rate (VBR). When the frame of the original audio is a steady-state signal with little change, the audio encoding device can encode the original audio using a small number of bits. However, when the frame of the original audio is a transient signal with large changes, the audio encoding device can encode the original audio using a large number of bits. The audio encoding device can use the available bits remaining after encoding the steady-state signal to encode the transient signal. That is, a high buffer level for the current frame means that the current frame is a steady-state signal. A low buffer level for the current frame means that the current frame is a transient signal. The bitstream analyzer 1230 can send information indicating whether the frame is a steady-state signal or a transient signal to the reconstruction method selector 1240.
[0203] The TNS information is information for reducing pre-echo. The starting position of an attack signal in the time domain can be found based on the TNS information. The attack signal represents a suddenly loud sound. Since the starting position of the attack signal can be found based on the TNS information, the bitstream analyzer 1230 can determine that the signal before the starting position is a steady-state signal. The bitstream analyzer 1230 can determine that the signal after the starting position is a transient signal.
[0204] The PNS information is information about the part where holes are generated in the frequency domain. A hole refers to a part for which spectral bins are not sent to save bits in the bitstream and are filled with arbitrary noise during the decoding process. The bitstream analyzer 1230 can send information about the position of the hole to the reconstruction method selector 1240.
[0205] The reconstruction method selector 1240 can receive the characteristics of the decoded audio signal and the decoded parameters. The reconstruction method selector 1240 can select a method for reconstructing the decoded audio signal. The decoded audio signal can be reconstructed by one of at least one reconstructor based on the selection of the reconstruction method selector 1240.
[0206] At least one reconstructor may include, for example, a first reconstructor 1250, a second reconstructor 1260, and an Nth reconstructor. At least one of the first reconstructor 1250, the second reconstructor 1260, and the Nth reconstructor may use a machine learning model. The machine learning model may be a model generated by machine learning at least one of an original audio signal, a decoded audio signal, and decoded parameters. At least one of the first reconstructor 1250, the second reconstructor 1260, and the Nth reconstructor may include a data acquirer 1251, a preprocessor 1252, and a result provider 1253. At least one of the first reconstructor 1250, the second reconstructor 1260, and the Nth reconstructor may include Figure 4 a data learner 410. At least one of the first reconstructor 1250, the second reconstructor 1260, and the Nth reconstructor may receive at least one of a decoded audio signal and decoded parameters as an input.
[0207] According to an embodiment of the present disclosure, the characteristics of the decoded parameters may include information about a frequency region for which spectral bins are accurately transmitted and a frequency region for which spectral bins are inaccurately transmitted. For a frequency region of the current frame for which spectral bins are accurately transmitted, the reconstruction method selector 1240 may determine a reconstructed decoded audio signal based on at least one of the decoded parameters and the decoded audio signal. The reconstruction method selector 1240 may determine to reconstruct the decoded audio signal by using the first reconstructor 1250. The first reconstructor 1250 may output a reconstructed audio signal by using a machine learning model.
[0208] For a frequency region of the current frame for which spectral bins are not accurately transmitted, the reconstruction method selector 1240 may determine to reconstruct the audio signal by using a bandwidth extension technique. The bandwidth extension technique includes spectral band replication (SBR). The reconstruction method selector 1240 may determine to reconstruct the decoded audio signal by using the second reconstructor 1260. The second reconstructor 1260 may output a reconstructed audio signal by using a bandwidth extension technique improved by a machine learning model.
[0209] According to another embodiment of the present disclosure, the characteristics of the decoded parameters may include information indicating whether a frame is a steady-state signal or a transient signal. When the frame is a steady-state signal, the reconstruction method selector 1240 may use the first reconstructor 1250 for the steady-state signal. When the frame is a transient signal, the reconstruction method selector 1240 may use the second reconstructor 1260 for the transient signal. The first reconstructor 1250 or the second reconstructor 1260 may output a reconstructed audio signal.
[0210] According to another embodiment of the present disclosure, the characteristics of the decoded parameters may include information about the position of the holes. For an audio signal decoded using a signal that does not correspond to the position of the holes, the reconstruction method selector 1240 may determine a reconstructed decoded audio signal based on the decoded parameters and the decoded audio signal. The reconstruction method selector 1240 may determine to reconstruct the decoded audio signal by using the first reconstructor 1250. The first reconstructor 1250 may output a reconstructed audio signal by using a machine learning model. For an audio signal decoded using a signal that corresponds to the position of the holes, the reconstruction method selector 1240 may use the second reconstructor 1260 for the signal corresponding to the position of the holes. The second reconstructor 1260 may output a reconstructed audio signal by using a machine learning model.
[0211] Since the method for reconstructing the decoded audio signal can be selected by the reconstruction method selector 1240 based on the characteristics of the audio signal, the audio reconstruction device 100 can effectively reconstruct the audio signal.
[0212] Figure 13 is a flowchart of an audio reconstruction method according to an embodiment.
[0213] In operation 1310, the audio reconstruction device 100 decodes the bitstream to obtain a plurality of decoded parameters for the current frame. In operation 1320, the audio reconstruction device 100 decodes the audio signal based on the plurality of decoded parameters. In operation 1330, the audio reconstruction device 100 selects one machine learning model from a plurality of machine learning models based on at least one of the decoded audio signal and the plurality of decoded parameters. In operation 1340, the audio reconstruction device 100 reconstructs the decoded audio signal by using the selected machine learning model.
[0214] Based on Figure 13 the audio reconstruction device 100 and based on Figure 3 the audio reconstruction device 100 can both improve the quality of the decoded audio signal. The audio reconstruction device 100 based on Figure 13 is less dependent on the decoded parameters and thus can achieve higher generality.
[0215] Now, the operation of the audio reconstruction device 100 will be described in detail with reference to Figure 14 and Figure 15 The operation of the audio reconstruction device 100 will be described in detail.
[0216] Figure 14 is a flowchart of an audio reconstruction method according to an embodiment.
[0217] The codec information extractor 1210 may receive the bitstream. The audio signal decoder 1220 may output an audio signal decoded based on the bitstream.
[0218] The bitstream analyzer 1230 can obtain the characteristics of decoding parameters based on the bitstream. For example, the bitstream analyzer 1230 can determine the starting frequency of bandwidth extension (operation 1410) based on at least one of multiple decoding parameters.
[0219] Referring to the reference curve graph 1460, the audio encoding device can accurately transmit spectral bins for the frequency region below the frequency f. However, for the frequency region above the frequency f and inaudible to humans, the audio encoding device may not transmit spectral bins or may transmit spectral bins poorly. The codec information extractor 1210 can determine the starting frequency f of bandwidth extension based on the spectral bins. The codec information extractor 1210 can output the information about the starting frequency f of bandwidth extension to the reconstruction method selector 1240.
[0220] The reconstruction method selector 1240 can select a machine learning model for the decoded audio signal based on the starting frequency f and the frequency of the decoded audio signal. The reconstruction method selector 1240 can compare the frequency of the decoded audio signal with the starting frequency f (operation 1420). The reconstruction method selector 1240 can select a reconstruction method based on the comparison.
[0221] When the frequency of the decoded audio signal is lower than the starting frequency f, the reconstruction method selector 1240 can select a certain machine learning model. A specific machine learning model can be pre-trained based on the decoded audio signal and the original audio signal. The audio reconstruction device 100 can reconstruct the decoded audio signal by using the machine learning model (operation 1430).
[0222] When the frequency of the decoded audio signal is higher than the starting frequency f, the reconstruction method selector 1240 can reconstruct the decoded audio signal by using bandwidth extension technology. For example, the reconstruction method selector 1240 can select a machine learning model to which the bandwidth extension technology is applied. A machine learning model can be pre-trained by using at least one of the parameters related to the bandwidth extension technology, the decoded audio signal, and the original audio signal. The audio reconstruction device 100 can reconstruct the decoded audio signal by using the machine learning model to which the bandwidth extension technology is applied (operation 1440).
[0223] Figure 15 is a flowchart of an audio reconstruction method according to an embodiment.
[0224] The codec information extractor 1210 can receive the bitstream. The audio signal decoder 1220 can output an audio signal decoded based on the bitstream.
[0225] The bitstream analyzer 1230 may obtain characteristics of decoding parameters based on the bitstream. For example, the bitstream analyzer 1230 may obtain the gain A of the current frame (operation 1510) based on at least one of a plurality of decoding parameters. The bitstream analyzer 1230 may obtain the average value of the gains of the current frame and the frame adjacent to the current frame (operation 1520).
[0226] The reconstruction method selector 1240 may compare the difference between the gain of the current frame and the average value of the gains with a threshold (operation 1530). When the difference between the gain of the current frame and the average value of the gains is greater than the threshold, the reconstruction method selector 1240 may select a machine learning model for transient signals. The audio reconstruction device 100 may reconstruct the decoded audio signal by using the machine learning model for transient signals (operation 1550).
[0227] When the difference between the gain of the current frame and the average value of the gains is less than the threshold, the reconstruction method selector 1240 may determine whether the window type included in the plurality of decoding parameters is indicated as short (operation 1540). When the window type is indicated as short, the reconstruction method selector 1240 may select a machine learning model for transient signals (operation 1550). When the window type is not indicated as short, the reconstruction method selector 1240 may select a machine learning model for steady-state signals. The audio reconstruction device 100 may reconstruct the decoded audio signal by using the machine learning model for steady-state signals (operation 1560).
[0228] The machine learning model for transient signals may perform machine learning based on the original audio signal classified as a transient signal and the decoded audio signal. The machine learning model for steady-state signals may perform machine learning based on the original audio signal classified as a steady-state signal and the decoded audio signal. Since steady-state signals and transient signals have different characteristics, and the audio reconstruction device 100 learns steady-state signals and transient signals separately, the decoded audio signal can be effectively reconstructed.
[0229] The present disclosure has been specifically shown and described with reference to various embodiments of the present disclosure. Those of ordinary skill in the art will understand that various changes may be made in form and detail without departing from the scope of the present disclosure. Therefore, the above embodiments should be considered only in a descriptive sense and not for purposes of limitation. The scope of the present disclosure is not defined by the above description, but by the appended claims, and all variations derived from the scope defined by the claims and their equivalents will be construed as being included within the scope of the present disclosure.
[0230] The above embodiments of the present disclosure can be written as a computer program and can be implemented in a general-purpose digital computer that executes the program by using a computer-readable recording medium. Examples of the computer-readable recording medium include magnetic storage media (e.g., ROM, floppy disks, and hard disks) and optical recording media (e.g., CD-ROMs and DVDs).
Claims
1. An audio signal reconstruction method, the method comprising: Obtaining a signal amplitude according to a frequency in a current frame by decoding a bitstream including an audio signal; Determining a range of the signal amplitude at the first frequency based on signal amplitudes corresponding to frequencies adjacent to the first frequency; Obtaining a reconstructed signal amplitude at the first frequency by applying a machine learning model to the range of the signal amplitude at the first frequency and the signal amplitude according to the frequency; And Decoding the audio signal of the current frame from the bitstream based on the reconstructed signal amplitude at the first frequency.
2. The audio signal reconstruction method according to claim 1, wherein, Decoding the audio signal includes: Obtaining a corrected signal amplitude at the first frequency by correcting the reconstructed signal amplitude at the first frequency to be within the range of the signal amplitude at the first frequency; and Decoding the audio signal based on the corrected signal amplitude.
3. The audio signal reconstruction method according to claim 2, Among them, Obtaining the corrected signal amplitude includes: when the reconstructed signal amplitude is not within the range, obtaining a value closest to the reconstructed signal amplitude within the range as the corrected signal amplitude.
4. The audio signal reconstruction method according to claim 1, wherein, Determining the range of the signal amplitude at the first frequency includes: determining the range of the signal amplitude at the first frequency by using a machine learning model pre-trained based on the signal amplitude according to the frequency.
5. The audio signal reconstruction method according to claim 1, wherein, Obtaining the reconstructed signal amplitude at the first frequency includes: Determining candidates for the signal amplitude at the first frequency based on the range; and Selecting one candidate from the candidates for the signal amplitude at the first frequency based on the machine learning model.
6. The audio signal reconstruction method according to claim 1, wherein, Obtaining the reconstructed signal amplitude at the first frequency includes: further obtaining the reconstructed signal amplitude at the first frequency of the current frame based on at least one decoding parameter among a plurality of decoding parameters of a previous frame.
7. The audio signal reconstruction method according to claim 1, wherein, The machine learning model is generated by machine learning an original audio signal and at least one signal amplitude according to the frequency.
8. The audio signal reconstruction method according to claim 1, the method further comprising: Selecting a machine learning model from a plurality of machine learning models based on the decoded audio signal and at least one signal amplitude according to the frequency; And Reconstructing the decoded audio signal by using the selected machine learning model.
9. The audio signal reconstruction method according to claim 8, wherein, Selecting the machine learning model includes: Determining a start frequency of bandwidth expansion based on at least one signal amplitude according to the frequency; and Selecting the machine learning model of the decoded audio signal based on the start frequency and the frequency of the decoded audio signal.
10. The audio signal reconstruction method according to claim 8, wherein, Selecting the machine learning model includes: Obtaining a gain of the current frame based on at least one signal amplitude according to the frequency; Obtaining an average value of the gains of the current frame and a frame adjacent to the current frame; When the difference between the gain of the current frame and the average value of the gain is greater than a threshold, select the machine learning model for the transient signal; When the difference between the gain of the current frame and the average value of the gain is less than the threshold, determine whether the window type included in the signal amplitude according to the frequency indicates short; When the window type indicates short, select the machine learning model for the transient signal; and When the window type does not indicate short, select the machine learning model for the steady-state signal.
11. An audio signal reconstruction device, the device comprising: a memory that stores the received bitstream; and at least one processor configured to: obtain the signal amplitude according to the frequency in the current frame by decoding the bitstream including the audio signal, determine the range of the signal amplitude at the first frequency based on the signal amplitude corresponding to the frequency adjacent to the first frequency, obtain the reconstructed signal amplitude at the first frequency by applying a machine learning model to the range of the signal amplitude at the first frequency and the signal amplitude according to the frequency, and decode the audio signal of the current frame from the bitstream based on the reconstructed signal amplitude at the first frequency.
12. The audio signal reconstruction device according to claim 11, wherein, The at least one processor is further configured to obtain the corrected signal amplitude at the first frequency by correcting the reconstructed signal amplitude at the first frequency to be within the range of the signal amplitude at the first frequency, and decode the audio signal based on the corrected signal amplitude.
13. The audio signal reconstruction device according to claim 12, wherein, The at least one processor is further configured to determine the range of the signal amplitude at the first frequency by using a machine learning model pre-trained based on the signal amplitude according to the frequency.
14. The audio signal reconstruction device according to claim 11, wherein, The at least one processor is further configured to obtain the reconstructed signal amplitude at the first frequency by determining candidates for the signal amplitude at the first frequency based on the range and selecting one candidate from the candidates for the signal amplitude at the first frequency based on the machine learning model.
15. The audio signal reconstruction device according to claim 11, wherein, The at least one processor is further configured to further obtain the reconstructed signal amplitude at the first frequency of the current frame based on at least one decoding parameter among the plurality of decoding parameters of the previous frame.
16. The audio signal reconstruction device according to claim 11, wherein, The machine learning model is generated by machine learning the original audio signal and at least one signal amplitude according to the frequency.
17. The audio signal reconstruction device according to claim 11, wherein, The at least one processor is further configured to select a machine learning model from a plurality of machine learning models based on the decoded audio signal and at least one signal amplitude according to the frequency, and reconstruct the decoded audio signal by using the selected machine learning model.
18. A computer-readable recording medium having recorded thereon a computer program for performing the method according to claim 1.
Citation Information
Patent Citations
Predictive vector quantization techniques in a higher order ambisonics (HOA) framework
US20160093308A1
Multimode coding of speech-like and non-speech-like signals
WO2009114656A1