Method for processing audio signal using psychoacoustic model and electronic device for performing same

The audio signal processing method leverages a psychoacoustic model and neural networks to optimize quantization levels for each subband, addressing inefficiencies in existing methods by maintaining high audio quality through imperceptible quantization noise.

WO2026063616A1PCT designated stage Publication Date: 2026-03-26SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

Existing audio signal processing methods fail to efficiently compress audio signals while maintaining high audio quality by not adequately utilizing human auditory characteristics.

Method used

An audio signal processing method that uses a psychoacoustic model to determine masking thresholds for each subband, combined with neural network-based encoders and quantizers, to optimize quantization levels and minimize perceptible quantization noise.

Benefits of technology

The method achieves high compression efficiency with minimal degradation in sound quality by applying different quantization levels based on human auditory perception, ensuring that quantization noise is imperceptible.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025010611_26032026_PF_FP_ABST
    Figure KR2025010611_26032026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed are an audio signal processing method using a psychoacoustic model and an electronic device for performing same. The audio signal processing method may comprise: an operation of receiving an input audio signal; an operation of converting the input audio signal into an audio signal in the frequency domain; an operation of dividing the audio signal in the frequency domain into audio signals of a plurality of subbands; an operation of determining a masking threshold corresponding to each of the subbands by using a psychoacoustic model that takes the audio signals of the subbands as inputs; an operation of acquiring audio vector data corresponding to each of the subbands by using neural network–based encoders; and an operation of generating a compressed audio signal by quantizing the audio vector data corresponding to each of the subbands on the basis of the masking threshold corresponding to each of the subbands.
Need to check novelty before this filing date? Find Prior Art

Description

Audio signal processing method using a psychoacoustic model and electronic device for performing the same

[0001] The present disclosure relates to an audio signal processing method using a psychoacoustic model and an electronic device for performing the same.

[0002] Electronic devices can provide functions related to audio signal processing. For example, electronic devices can provide call functions for collecting and transmitting audio signals, recording functions for recording audio signals, and audio output functions for outputting audio signals. Electronic devices can output audio through external audio output devices, such as earphones and headphones, or through audio output modules built into the electronic device. Electronic devices can perform signal processing, such as data compression, to reduce the data size of audio signals.

[0003] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. None of the foregoing is to be claimed as prior art related to the present disclosure, nor is it to be used to determine prior art.

[0004] An audio signal processing method according to an embodiment may include: receiving an input audio signal; converting the input audio signal into a frequency domain audio signal; dividing the frequency domain audio signal into audio signals of a plurality of subbands; determining a masking threshold value corresponding to each of the subbands using a psychoacoustic model that takes the audio signals of the subbands as inputs; obtaining audio vector data corresponding to each of the subbands using neural network-based encoders; and generating a compressed audio signal by quantizing the audio vector data corresponding to each of the subbands based on the masking threshold value corresponding to each of the subbands.

[0005] An electronic device according to an embodiment may include one or more memories for storing instructions and one or more processors. When the instructions are executed individually or collectively by one or more processors, the electronic device may generate a compressed audio signal by converting an input audio signal into a frequency domain audio signal, dividing the frequency domain audio signal into audio signals of a plurality of subbands, determining a masking threshold value corresponding to each of the subbands using a psychoacoustic model that takes the audio signals of the subbands as input, obtaining audio vector data corresponding to each of the subbands using neural network-based encoders, and quantizing the audio vector data corresponding to each of the subbands based on the masking threshold value corresponding to each of the subbands.

[0006] FIG. 1 is a block diagram illustrating an exemplary configuration of an electronic device according to various embodiments.

[0007] FIGS. 2 and FIGS. 3 are flowcharts for explaining the operations of an audio signal processing method according to various embodiments.

[0008] FIG. 4 is a diagram illustrating operations for processing audio signals using a psychoacoustic model according to various embodiments.

[0009] FIGS. 5A, 5B, and 5C are drawings illustrating the characteristics and masking threshold values ​​of a psychoacoustic model according to various embodiments.

[0010] FIG. 6 is a diagram illustrating how the quantization processing of a residual vector quantization (RVQ) is controlled based on quantization control values ​​according to various embodiments.

[0011] FIG. 7 is a diagram illustrating an example of performing audio signal processing by dividing an input audio signal into four subbands according to various embodiments.

[0012] FIG. 8 is a diagram illustrating the interaction between an electronic device performing audio signal processing according to various embodiments and another electronic device.

[0013] Hereinafter, embodiments will be described in detail with reference to the attached drawings. In the description with reference to the attached drawings, identical components are given the same reference numeral regardless of the drawing number, and redundant descriptions thereof will be omitted.

[0014] FIG. 1 is a block diagram illustrating an exemplary configuration of an electronic device according to various embodiments.

[0015] FIG. 1 is a block diagram of an electronic device (101) in a network environment (100) according to various embodiments. Referring to FIG. 1, in the network environment (100), the electronic device (101) may communicate with another electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or may communicate with at least one of another electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through the server (108).

[0016] According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (160), audio module (170), sensor (176), interface (177), connection terminal (178), haptic module (179), camera (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor (176), camera (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)).

[0017] The processor (120) may be implemented as one or more IC (integrated circuit (or circuitry)) chips and may perform various data processing operations. The processor (120) may include at least one electrical circuit and may process instructions (or programs (140), data, etc.) stored in memory (130) individually or collectively in a distributed manner. When the instructions are executed by the processor (1210), they may control the electronic device (101) to perform one or more operations of the electronic device (101) described in this disclosure. The processor (120) may include a processor assembly comprising one or more processing circuits. The processor (120) may include any operative processing circuit to control the performance and operations of one or more components of the electronic device (101) (e.g., memory (130), display module (160), camera (180), communication module (190), and / or sensor (176)).

[0018] The processor (120) can control at least one other component (e.g., hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., program (140)) and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store instructions or data received from other components (e.g., sensor (176) or communication module (190)) in volatile memory (132), process the instructions or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134).

[0019] According to one embodiment, the processor (120) may include one or more processors, and the operations of the electronic device (101) described in this disclosure may be performed by one processor or by a combination of multiple processors. In this disclosure, "processor" may include a processing circuit or a plurality of processors. For example, as used in this disclosure including in the claims, the term "processor" may include various processing circuits including one or more processors, wherein one or more processors may be configured to perform various functions described in this disclosure in a distributed manner, individually and / or collectively. Where in this disclosure "processor," "at least one processor," and "one or more processors" are described as being configured to perform a plurality of functions, these terms include, but are not limited to, situations where, for example, one processor performs some of the cited functions and another processor performs other of the cited functions, and situations where a single processor can perform all of the cited functions. Additionally, one or more processors may include, for example, a combination of processors performing various cited / disclosed functions in a distributed manner. One or more processors can execute instructions to achieve or perform various functions.

[0020] According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor). If the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.

[0021] The auxiliary processor (123) may control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera (180) or communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or through a separate server (e.g., server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include multiple artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.

[0022] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor (176)). The data may include, for example, input data or output data for software (e.g., program (140)) and related instructions. The memory (130) may include volatile memory (132) or non-volatile memory (134).

[0023] In one embodiment, the memory (130) may include one or more memories. Instructions for controlling the processor (120) to perform operations of the electronic device (101) described in this disclosure may be stored in one memory or may be divided and stored in multiple memories.

[0024] The program (140) may be stored as software in memory (130). The program (140) may include, for example, an operating system (142), middleware (144), or an application (146).

[0025] The input module (150) can receive instructions or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0026] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.

[0027] The display module (160) can visually provide information to an external (e.g., user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.

[0028] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or another electronic device (e.g., electronic device (102)) (e.g., speaker or headphones) that is directly or wirelessly connected to the electronic device (101).

[0029] The sensor (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor (176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor. For example, the sensor (176) may include an inertial measurement unit (IMU).

[0030] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to another electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0031] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to another electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0032] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that can be perceived by the user through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.

[0033] The camera (180) can capture still images and video. According to one embodiment, the camera (180) may include one or more lenses, one or more image sensors, one or more image signal processors, or one or more flashes.

[0034] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) may be implemented, for example, as at least part of a power management integrated circuit (PMIC).

[0035] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0036] A communication module (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and another electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication circuits. The communication module (190) may include one or more communication processors (CP) that operate independently of a processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (194) (e.g., LAN (local region network) communication module, or power line communication module). These communication modules can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data relation)) or a second network (199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).

[0037] The wireless communication module (192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (192) can support high-frequency bands (e.g., mmWave band) to achieve high data transmission rates, for example. The wireless communication module (192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), other electronic devices (e.g., electronic device (104)), or network systems (e.g., second network (199)).

[0038] An antenna module (197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (198) or a second network (199), may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).

[0039] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.

[0040] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., instructions or data) with each other.

[0041] According to one embodiment, instructions or data may be transmitted or received between an electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199).

[0042] Each of the external electronic devices, such as the electronic device (102, 104) and the server (108), may be of the same or different type as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more external electronic devices, such as the electronic device (102, 104) or the server (108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the request may perform at least part of the requested function or service, or additional functions or services related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may provide the result as is or additionally processed as at least part of the response to the request.

[0043] In one embodiment, the electronic device (101) can compress an audio signal using a neural network-based audio codec. The neural network-based audio codec may also be referred to as a 'neural audio codec'. The electronic device (101) can further optimize the quantization processing of the audio signal using a psychoacoustic model described in this disclosure (e.g., the psychoacoustic model (440) of FIG. 4, the psychoacoustic model (720) of FIG. 7). By determining the optimal quantization level for each subband of the audio signal using the psychoacoustic model, the electronic device (101) can provide compression performance with high audio quality relative to compression efficiency (e.g., minimizing quantization noise). The audio signal compressed by the electronic device (101) can be transmitted to another electronic device (102, 104) (e.g., the electronic device (800) of FIG. 8) through a network (e.g., the first network (198) or the second network (199). Below, the processing of the audio signal by the electronic device (101) will be described in detail.

[0044] FIGS. 2 and FIGS. 3 are flowcharts for explaining the operations of an audio signal processing method according to various embodiments.

[0045] FIG. 2 is a flowchart illustrating operations for compressing an audio signal among audio signal processing methods. Operations for compressing an audio signal may be performed by an electronic device described in this disclosure (e.g., the electronic device (101) of FIG. 1). Referring to FIG. 2, in operation (210), the electronic device may receive an input audio signal. The input audio signal may include an audio signal input (or transmitted) to the electronic device for audio processing and / or an audio signal acquired by the electronic device (via a microphone). The input audio signal may be, for example, an analog audio signal converted into a digital audio signal through a pulse code modulation (PCM) process.

[0046] In operation (220), the electronic device can convert the input audio signal into an audio signal in the frequency domain. The input audio signal may be an audio signal in the time domain in which the audio signal is recorded over time. The electronic device can convert the audio signal in the time domain into an audio signal in the frequency domain using, for example, a transformation method such as FFT (fast Fourier transform), STFT (short time Fourier transform), or MDCT (modified discrete cosine transform). The electronic device can convert the input audio signal into an audio signal in the frequency domain to process the input audio signal by dividing it into frequency bands.

[0047] In operation (230), the electronic device can divide the audio signal in the frequency domain into audio signals of multiple sub-bands. The electronic device can, for example, divide the entire frequency band of the audio signal in the frequency domain into defined sub-bands.

[0048] In operation (240), the electronic device may determine a masking threshold corresponding to each of the subbands using a psychoacoustic model (e.g., the psychoacoustic model (440) of FIG. 4 or the psychoacoustic model (720) of FIG. 7) that takes the audio signals of the subbands as input. The psychoacoustic model is a model that mimics the way the human auditory system perceives sound. The psychoacoustic model can help improve the compression performance of the audio signal by using human auditory characteristics to remove sound information that humans do not actually perceive. The psychoacoustic model can provide a masking threshold corresponding to the magnitude of the audio signal at which the listener does not perceive quantization noise generated by the quantization process that reduces the data size of the audio signal. The masking threshold may represent a threshold of sound that is estimated to be imperceptible to humans.

[0049] Psychoacoustic models can output masking thresholds for each frequency band of an input audio signal based on information regarding the human minimum audibility threshold and the masking effect caused by signal amplitudes across frequency bands of audio signals in subbands. The minimum audibility threshold information represents the absolute hearing threshold, which is the minimum sound level at which the human ear can perceive sound in a specific frequency band. Sounds quieter than the absolute hearing threshold may not be perceived by the listener. The masking effect refers to the phenomenon where quiet sounds below a specific threshold are masked by loud sounds at a particular frequency. The listener may not perceive these masked quiet sounds. If quantization noise generated during the quantization process (corresponding to the difference between the original audio signal and the quantized audio signal) falls below the corresponding threshold, the quantization noise may be masked and become inaudible to the listener. Psychoacoustic models can provide such thresholds (masking thresholds) for the entire frequency band. The electronic device can obtain a masking threshold for the entire frequency band of the input audio signal from the psychoacoustic model and determine a masking threshold for each of the subbands based on the obtained masking threshold for the entire frequency band. The masking threshold for the entire frequency band may depend on the frequency band-specific signal magnitudes of the audio signals of the subbands input to the psychoacoustic model.

[0050] In operation (250), the electronic device may acquire audio vector data corresponding to each of the subbands using neural network-based encoders (e.g., encoders of FIG. 4 (452, 454) or encoders of FIG. 7 (732, 734, 736, 738)). Audio signals of different subbands are input to each of the encoders, and each of the encoders may output audio vector data corresponding to different subbands. Audio signals of different subbands may be converted into audio vector data downsampled through subband-specific encoders. Each encoder may encode the input audio signal of a specific subband into a latent space of a smaller dimension than the original audio signal, and may output audio vector data containing latent vector values ​​as a result of encoding.

[0051] In one embodiment, encoders may be implemented as a neural network model (e.g., a deep neural network). A neural network model may refer to a model in which artificial neurons (or nodes) forming a network through synaptic connections change the strength of the synaptic connections through training or machine learning to possess problem-solving capabilities. The artificial neurons of a neural network model may include a combination of weights and / or biases, and the neural network may include one or more layers formed of multiple artificial neurons. The encoders may be, for example, encoders included in an autoencoder model.

[0052] According to the embodiment, the operation (240) and the operation (250) may be performed in parallel, and the operation (250) may be performed before the operation (240). It should not be interpreted as being limited by the order of execution of the operations shown in FIG. 2.

[0053] In operation (260), the electronic device can generate a compressed audio signal by quantizing audio vector data corresponding to each subband based on a masking threshold value corresponding to each subband. The electronic device can determine a quantization control value to be applied to quantizers for each subband (e.g., residual vector quantizers (462, 464) of FIG. 4 or residual vector quantizers (740, 750, 760, 770) of FIG. 7) based on a masking threshold value corresponding to each subband. The following describes an embodiment in which a residual vector quantizer is used as a quantizer, but the scope of the embodiment is not limited thereto. For example, a group vector quantizer (GVQ) and / or a group residual vector quantizer (GRVQ) may be used as a quantizer.

[0054] According to one embodiment, the electronic device can determine a quantization control value applied to each residual vector quantizer so that the data size of the compressed audio signal is less than or equal to the target data size. For example, the electronic device can determine a quantization control value such that the quantization error is minimized for each subband. The electronic device can adjust the quantization stage (or quantization level) applied to each residual vector quantizer based on a masking threshold value corresponding to each subband. In one embodiment, the electronic device can determine a quantization control value such that the quantization error for each subband is placed below the masking threshold of each subband. Based on the determined quantization control value for each subband, the electronic device can control the quantization processing by each residual vector quantizer corresponding to each subband. The audio vector data of each subband can be quantized by the residual vector quantizer to match the target bit size.

[0055] In subbands with a large masking threshold, it is highly likely that the listener will not perceive the quantization error even if the quantization error is somewhat large. For such subbands, even if a small number of bits are allocated by increasing the compression ratio, the sound quality may not be significantly affected. In the case of subbands with a small masking threshold, the sound quality is significantly affected; therefore, for subbands with a small masking threshold, it is necessary to allocate a large number of bits to the corresponding subband by lowering the compression ratio so that the quantization error is placed below the masking threshold. In one embodiment, the electronic device may determine a quantization control value with a high compression ratio for the audio vector data of the corresponding subband if the masking threshold determined for the subband is large, and determine a quantization control value with a low compression ratio for the audio vector data of the corresponding subband if the masking threshold determined for the subband is small.

[0056] According to one embodiment, each of the residual vector quantizers can control the quantization stage of the audio vector data input to each residual vector quantizer based on a quantization control value applied to each residual vector quantizer. The residual vector quantizers can convert audio vector data into quantized audio vector data through quantization processing. A compression effect in which the data size is reduced occurs through quantization processing, and the quantization stage applied to the residual vector quantizer can determine the extent to which the audio vector data is compressed. Accordingly, the data size of the audio vector data input to each residual vector quantizer is larger than the data size of the quantized audio vector data, and the data size of the quantized audio vector data can be determined according to the quantization stage applied to the residual vector quantizer. Depending on the quantization control value, the number of bits allocated for the compression of the audio vector data input to the residual vector quantizer or the number of quantization stages applied to the quantization of the audio vector data can be determined. For example, if the quantization control value corresponds to the number of bits allocated to each subband, the number of quantization stages applied during the quantization process in the residual vector quantizer can be adjusted according to the number of bits represented by the quantization control value. The number of quantization stages applied during the quantization process can be determined so that the data size of the quantized audio vector data is less than or equal to the corresponding number of bits. If the quantization control value corresponds to the number of quantization stages activated for each subband, the residual vector quantizer can perform the quantization process according to the number of quantization stages represented by the quantization control value. An embodiment in which the data size of the quantized audio vector data is determined according to the quantization stage is described in more detail in the description related to Figure 6 below.

[0057] In one embodiment, it is assumed that residual vector quantizers include a first residual vector quantizer and a second residual vector quantizer. The electronic device may determine a first quantization control value that causes the first residual vector quantizer, which quantizes first audio vector data corresponding to a first subband, to perform quantization processing with a first quantization stage determined for a first masking threshold value corresponding to the first subband. The electronic device may determine, for example, among candidate quantization control values ​​corresponding to different quantization stages, a candidate quantization control value that minimizes the quantization error value generated by the quantization processing of the first residual vector quantizer as the first quantization control value. The electronic device may determine a second quantization control value that causes the second residual vector quantizer, which quantizes second audio vector data corresponding to a second subband, to perform quantization processing with a second quantization stage determined for a second masking threshold value corresponding to the second subband. The electronic device can determine, for example, among candidate quantization control values ​​corresponding to different quantization stages, a candidate quantization control value that minimizes the quantization error value generated by the second residual vector quantizer quantization processing as the second quantization control value.

[0058] A first residual vector quantizer can generate first quantized audio vector data by quantizing first audio vector data into a first quantization stage according to a first quantization control value, and a second residual vector quantizer can generate second quantized audio vector data by quantizing second audio vector data into a second quantization stage according to a second quantization control value. The first audio vector data can be transmitted from a first encoder that encodes an audio signal of a first subband, and the second audio vector data can be transmitted from a second encoder that encodes an audio signal of a second subband. In one embodiment, if the first masking threshold is greater than the second masking threshold, the data size (e.g., number of bits) of the first quantized audio vector data generated by the first residual vector quantizer may be smaller than the data size of the second quantized audio vector data generated by the second residual vector quantizer.

[0059] In one embodiment, an electronic device may transmit a compressed audio signal to another electronic device (e.g., the electronic device (102, 104) of FIG. 2, the electronic device (800) of FIG. 8). The compressed audio signal may be transmitted to the other electronic device via a wireless network in the form of a bitstream. The other electronic device may receive the compressed audio signal and perform operations to restore the compressed audio signal according to the audio signal processing method described in FIG. 3 below.

[0060] As described above, the electronic device performs quantization using a masking threshold obtained through a psychoacoustic model, thereby setting the quantization noise generated by quantization to fall below the masking threshold for information that the listener cannot actually hear. By setting a high compression ratio for information that the listener cannot actually hear and a low compression ratio for information that the listener can actually hear, the electronic device can increase the compression efficiency of the audio signal while maintaining sound quality similar to the original sound. In addition, the electronic device can minimize the degradation of the audio signal quality caused by quantization by determining a masking threshold according to the characteristics of the audio signal for each subband and determining the quantization level applied to each residual vector quantizer according to the determined masking threshold for each subband.

[0061] FIG. 3 is a flowchart illustrating operations for decompressing a compressed audio signal to obtain a restored audio signal among audio signal processing methods. The operations for obtaining the restored audio signal may be performed by an electronic device described in this disclosure (e.g., the electronic device (102, 104) of FIG. 2, the electronic device (800) of FIG. 8). Referring to FIG. 3, in operation (310), the electronic device may receive the compressed audio signal. For example, the electronic device may receive the compressed audio signal via a wireless network.

[0062] In operation (320), the electronic device may obtain restored audio data corresponding to each of the subbands using neural network-based decoders (e.g., decoders (472, 474) of FIG. 4 or decoders (782, 784, 785, 788) of FIG. 7). The decoders may be, for example, decoders included in an autoencoder model. Each decoder may correspond to each encoder described in FIG. 2. In one embodiment, it is assumed that the encoders include a first encoder and a second encoder, and the decoders include a first decoder and a second decoder. If the first encoder encodes the audio signal of the first subband and outputs audio vector data, the first decoder corresponding to the first encoder may decode the audio vector data for the first subband and restore the audio signal of the first subband. If the second encoder encodes the audio signal of the second subband and outputs audio vector data, the second decoder corresponding to the second encoder can decode the audio vector data for the second subband and restore the audio signal of the second subband.

[0063] In operation (330), the electronic device can convert the audio data restored by subband into an audio signal in the time domain. In one embodiment, the electronic device can combine the audio data restored by subband to generate an audio signal restored in the frequency domain, and convert the audio signal restored in the frequency domain into a restored audio signal in the time domain using an inverse transform method of IFFT (inverse fast Fourier transform), ISTFT (inverse short time Fourier transform), or IMDCT (inverse modified discrete cosine transform).

[0064] An output audio signal is generated through the above operations, and the output audio signal can be provided to a listener through a speaker.

[0065] FIG. 4 is a diagram illustrating operations for processing audio signals using a psychoacoustic model according to various embodiments.

[0066] Referring to FIG. 4, an audio signal compression architecture comprising encoders (452, 454), a psychoacoustic model (440), a quantization controller (445), and residual vector quantizers (462, 464), and an audio signal restoration architecture comprising decoders (472, 474) are illustrated. In one embodiment, the architecture of a neural audio codec of an autoencoder model may include an audio signal compression architecture and an audio signal restoration architecture. Operations performed in the audio signal compression architecture may be performed by an audio encoding device (402) (e.g., the electronic device (101) of FIG. 1). Operations performed in the audio signal restoration architecture may be performed by an audio decoding device (404) (e.g., the electronic device (102, 104) of FIG. 2, the electronic device (800) of FIG. 8).

[0067] According to one embodiment, an input audio signal (410) in the time domain, which is subject to compression of the audio signal, may be input. The input audio signal (410) in the time domain may be converted (420) into an audio signal in the frequency domain through a conversion method, for example, FFT, STFT, or MDCT. The audio signal in the frequency domain may be divided into audio signals of multiple subbands through frequency band division (430). Audio signals of different subbands may be obtained through frequency band division (430).

[0068] The audio signals of the subbands can be input to encoders (452, 454) corresponding to each subband. Each encoder (452, 454) can perform encoding to output audio vector data (e.g., latent vector values) corresponding to each subband. For example, the audio signal of the first subband is input to the first encoder (452), and the first encoder (452) can output latent vector values ​​for the first subband. The audio signal of the second subband is input to the second encoder (454), and the second encoder (454) can output latent vector values ​​for the second subband. Each of the encoders (452, 454) can extract important features of the input audio signal of the subband and compress the extracted features into latent vector values ​​having a data size smaller than that of the input audio signal. The structures of the neural network models of the encoders (452, 454) may be the same or different from each other. If the structures are the same, the encoders (452, 454) may have different parameters of the neural network models (e.g., weights or biases). The encoders (452, 454) may be trained using unsupervised learning or self-supervised learning methods.

[0069] Additionally, audio signals of the subbands are input into a psychoacoustic model (440), and a masking threshold value for each frequency band of the input audio signal can be obtained from the psychoacoustic model (440). A quantization controller (445) can determine a masking threshold value for each subband based on the masking threshold value for each frequency band obtained from the psychoacoustic model (440). A quantization controller (445) can determine a quantization control value to be applied to the residual vector quantizers (462, 464) for each subband based on the masking threshold value corresponding to each subband. A quantization controller (445) can adjust the quantization stage (or quantization level) applied to each of the residual vector quantizers (462, 464) based on the masking threshold value corresponding to each subband. The quantization controller (445) can determine a quantization control value applied to each residual vector quantizer (462, 464) so ​​that the data size of the compressed audio signal is less than or equal to the target data size. In one embodiment, the quantization controller (445) can determine a quantization control value such that, for audio vector data by subband, a large number of bits are allocated to sensitive subbands that have a large impact on sound quality (corresponding to subbands with a small masking threshold), and a small number of bits are allocated to subbands that have a relatively small impact on sound quality (corresponding to subbands with a large masking threshold).

[0070] Residual vector quantizers (462, 464) can convert audio vector data into quantized audio vector data through quantization processing. A compression effect in which the data size is reduced occurs through quantization processing, and the compression rate for each subband can be determined according to the quantization control value applied to each of the residual vector quantizers (462, 464). In one embodiment, the quantization control value can be determined such that if the masking threshold value determined for the subband is large, the compression rate is determined to be high, and if the masking threshold value determined for the subband is small, the compression rate is determined to be low. A high compression rate for a specific subband indicates that the original sound of that subband is relatively less preserved in the compressed audio signal, and a low compression rate for a specific subband indicates that the original sound of that subband is relatively more preserved in the compressed audio signal.

[0071] In one embodiment, audio vector data of subbands quantized by residual vector quantizers (462, 464) can be concatenated to generate a compressed audio signal.

[0072] The above audio signal compression architecture can effectively improve compression ratio and sound quality by using a psychoacoustic model (440) in a neural network model-based audio codec to effectively control quantization errors by utilizing human auditory characteristics and a masking effect on the input audio signal (410) that changes in real time.

[0073] In one embodiment, the audio encoding device (402) may transmit the compressed audio signal to an audio decoding device (404) that performs the restoration of the audio signal. The audio encoding device (402) may transmit the compressed audio signal to the audio decoding device (404), for example, in the form of a bitstream. The bitstream may be transmitted through a communication network (e.g., a wired network or a wireless network).

[0074] The audio decoding device (404) receives a compressed audio signal from the audio encoding device (402) and can input the compressed audio signal to different decoders (472, 474) for each subband. Audio data restored for each subband can be output from the decoders (472, 474). For example, a latent vector value compressed (or quantized) for a first subband can be input to a first decoder (472), and the first decoder (472) can perform decoding on the latent vector value compressed for the first subband to output an audio signal restored for the first subband. A latent vector value compressed for a second subband can be input to a second decoder (474), and the second decoder (474) can perform decoding on the latent vector value compressed for the second subband to output an audio signal restored for the second subband. The structures of the neural network models of the decoders (472, 474) may be the same or different from each other. If the structures are the same, the decoders (472, 474) may have different parameters of the neural network models (e.g., weights or biases). The decoders (472, 474) may be trained using unsupervised learning or self-supervised learning methods.

[0075] The audio data restored for each subband is an audio signal in the frequency domain, and the audio data restored for each subband can be converted into a restored audio signal (490) in the time domain through an inverse transform (480) such as IFFT, ISTFT, or IMDCT, for example. The restored audio signal (490) can be provided to a listener through a speaker.

[0076] FIGS. 5A, 5B, and 5C are drawings illustrating the characteristics and masking threshold values ​​of a psychoacoustic model according to various embodiments.

[0077] FIG. 5a is a diagram illustrating the minimum audible threshold information used by a psychoacoustic model to determine a masking threshold. In FIG. 5a, the graph (510) represents the minimum audible threshold information by frequency band (e.g., kHz). The minimum audible threshold information by frequency band may include information on the absolute auditory threshold, which is the minimum sound level (e.g., dB) for each frequency band at which humans can perceive sound. In the case of an audio signal (522) in the first frequency band, since the signal level is greater than the absolute auditory threshold, the audio signal (522) corresponds to an audio signal that humans can perceive. In contrast, in the case of an audio signal (524) in the second frequency band, even if the signal level is equal to the signal level of the audio signal (522), since the signal level is smaller than the absolute auditory threshold, the audio signal (524) corresponds to an audio signal that humans cannot perceive.

[0078] FIG. 5b is a diagram illustrating the masking effect used by a psychoacoustic model to determine a masking threshold. The masking effect represents a phenomenon in which small sounds below a specific threshold value in the surrounding area are obscured by a loud sound of a specific frequency. In FIG. 5b, the graph (530) represents information on the minimum audible threshold for each frequency band. Before an audio signal (540) with a large signal size occurs in a specific frequency band, the signal sizes of the audio signals (552, 554, 556) are all greater than the absolute auditory threshold, so the listener can hear all of the corresponding audio signals (552, 554, 556). However, when an audio signal (540) with a large signal size occurs in a specific frequency band, the absolute auditory threshold in the vicinity of that specific frequency band changes to a masking threshold (545), and as a result, the signal sizes of the audio signals (552, 554) are below the masking threshold (545), so they become audio signals that the listener cannot hear. In this case as well, since the signal magnitude of the audio signal (556) is greater than the masking threshold (545), the audio signal (556) can be heard by the listener.

[0079] FIG. 5c is a diagram illustrating the determination of masking threshold values ​​for each subband by considering the minimum audible limit information and the masking effect on the audio signal input in real time. In FIG. 5c, graph (570) represents the input audio signal (audio signal of the frequency band), and graph (560) represents the masking threshold value for the entire frequency band determined by a psychoacoustic model (e.g., the psychoacoustic model (440) of FIG. 4 or the psychoacoustic model (720) of FIG. 7).

[0080] An electronic device according to one embodiment (e.g., the electronic device (1010) of FIG. 1) can divide an audio signal in a frequency band into a plurality of subbands (582, 584, 586, 588) and determine a masking threshold value (590) corresponding to each subband (582, 584, 586, 588). The electronic device can determine a quantization control value such that the quantization error for each of the subbands (582, 584, 586, 588) lies below the masking threshold value (590) of each subband (582, 584, 586, 588). The electronic device can adjust the quantization stage (or quantization level) applied to each of the residual vector quantizers (e.g., the residual vector quantizers (462, 464) of FIG. 4 or the residual vector quantizers (740, 750, 760, 770) of FIG. 7) based on a masking threshold value (590) corresponding to each of the subbands (582, 584, 586, 588). The electronic device can determine a quantization control value such that for each subband (582, 584, 586, 588), the quantization error lies below the masking threshold value (590) of each subband (582, 584, 586, 588). The electronic device can determine a quantization control value to be applied to residual vector quantizers based on a masking threshold value (590) determined for each subband (582, 584, 586, 588). In the illustrated embodiment, since the masking threshold value for subband (586) among the subbands (582, 584, 586, 588) is the smallest, the electronic device can determine a quantization control value so that quantization proceeds with the largest number of bits (low compression rate) for the audio vector data of subband (586). Since the masking threshold value for subband (588) is the largest, the electronic device can determine a quantization control value so that quantization proceeds with the smallest number of bits (high compression rate) for the audio vector data of subband (588).

[0081] FIG. 6 is a diagram illustrating that the quantization processing of a residual vector quantizer is controlled based on quantization control values ​​according to various embodiments.

[0082] Referring to FIG. 6, an encoder (610) that performs encoding on an audio signal of a specific frequency band and outputs audio vector data, a quantization controller (445) that generates a quantization control value for controlling the quantization processing of a residual vector quantizer (620) based on a masking threshold value, a residual vector quantizer (620) that generates quantized audio vector data by performing quantization processing on the audio vector data output from the encoder (610) based on the quantization control value, and a decoder (630) that performs decoding on the quantized audio vector data for a specific subband and outputs a restored audio signal are shown. The encoder (610), the residual vector quantizer (620), and the decoder (630) may correspond to the encoder (452 ​​or 454), the residual vector quantizer (462 or 464), and the decoder (472 or 474) of FIG. 4, respectively.

[0083] In one embodiment, the residual vector quantizer (620) may include a plurality of vector quantizers (622, 624, 626, 628). Hereinafter, it is assumed that the residual vector quantizer (620) includes four vector quantizers: a first vector quantizer (622), a second vector quantizer (624), a third vector quantizer (262), and a fourth vector quantizer (628). However, the scope of the embodiment is not limited thereto, and the residual vector quantizer (620) may include two or more vector quantizers.

[0084] In one embodiment, the vector quantizers (622, 624, 626, 628) may each correspond to a different quantization stage (or quantization level). The quantization stage may be activated stepwise. As the quantization stage increases, quantization may be performed continuously. A quantization control value generated by the quantization controller (445) may determine the number of quantization stages to be activated among the quantization stages of the vector quantizers (622, 624, 626, 628).

[0085] For example, (1) when only the first quantization stage by the first vector quantizer (622) is activated by the quantization control value, quantization is performed on the audio vector data transmitted from the encoder (610) to generate the first quantized vector data Q1. The first quantized vector data Q1 can be output as quantized audio vector data from the residual vector quantizer (620). (2) When the first quantization stage by the first vector quantizer (622) and the second quantization stage by the second vector quantizer (624) are activated by the quantization control value, in addition to generating the first quantized vector data Q1, quantization is performed on the first residual signal by the second vector quantizer (624) to generate the second quantized vector data Q2. The first residual signal may correspond to the residual signal between the signal before quantization is performed by the first vector quantizer (622) and the signal after quantization is performed. The first quantized vector data Q1 and the second quantized vector data Q2 may be combined and output as quantized audio vector data from the residual vector quantizer (620). (3) When the first quantization stage by the first vector quantizer (622), the second quantization stage by the second vector quantizer (624), and the third quantization stage by the third vector quantizer (626) are activated by the quantization control value, in addition to the generation of the first quantized vector data Q1 and the generation of the second quantized vector data Q2, quantization may be performed on the second residual signal by the third vector quantizer (626) to generate the third quantized vector data Q3. The second residual signal can correspond to the residual signal between the signal before being quantized by the second vector quantizer (624) and the signal after being quantized.The first quantized vector data Q1, the second quantized vector data Q2, and the third quantized vector data Q3 can be combined and output as quantized audio vector data from the residual vector quantizer (620). (4) When the first quantization stage by the first vector quantizer (622), the second quantization stage by the second vector quantizer (624), the third quantization stage by the third vector quantizer (626), and the fourth quantization stage by the fourth vector quantizer (628) are all activated by the quantization control value, in addition to the generation of the first quantized vector data Q1, the generation of the second quantized vector data Q2, and the generation of the third quantized vector data Q3, quantization is performed on the third residual signal by the fourth vector quantizer (628) to generate the fourth quantized vector data Q4. The third residual signal may correspond to the residual signal between the signal before quantization and the signal after quantization by the third vector quantizer (626). The first quantized vector data Q1, the second quantized vector data Q2, the third quantized vector data Q3, and the fourth quantized vector data Q4 may be combined and output as quantized audio vector data from the residual vector quantizer (620). Among the above cases, the data size of the quantized audio vector data may increase in the order of (1), (2), (3), and (4). For example, the data size of the quantized audio vector data in the order of (1), (2), (3), and (4) may be 10 bits, 20 bits, 30 bits, and 40 bits. As the quantization stage increases, the data size allocated to the quantized audio vector data may increase, and the quantization error caused by quantization may decrease. Conversely, the smaller the quantization stage, the smaller the data size allocated to the quantized audio vector data can be, and the quantization error caused by quantization can be larger.

[0086] FIG. 7 is a diagram illustrating an example of performing audio signal processing by dividing an input audio signal into four subbands according to various embodiments.

[0087] Referring to FIG. 7, an audio signal compression architecture comprising encoders (732, 734, 736, 738) (e.g., encoder (452), encoder (454) of FIG. 4), a psychoacoustic model (720) (e.g., psychoacoustic model (440) of FIG. 4), a quantization controller (725) (e.g., quantization controller (445) of FIG. 4), and residual vector quantizers (740, 750, 760, 770) (e.g., residual vector quantizer (482, 484) of FIG. 4) and an audio signal restoration architecture comprising decoders (782, 784, 786, 788) (e.g., decoder (472, 474) of FIG. 4) are illustrated. Operations performed in the audio signal compression architecture may be performed by an audio encoding device (702) (e.g., the audio encoding device (402) of FIG. 4 or the electronic device (101) of FIG. 1). Operations performed in the audio signal restoration architecture may be performed by an audio decoding device (704) (e.g., the audio decoding device (404) of FIG. 4, the electronic device (102, 104) of FIG. 2, or the electronic device (800) of FIG. 8).

[0088] The input audio signal (705) in the time domain can be converted into an audio signal in the frequency domain (710) through a conversion method such as FFT, STFT, or MDCT. The audio signal in the frequency domain can be divided into four different subbands of audio signals through frequency band division (430).

[0089] The audio signals of the four subbands can be input to encoders (732, 734, 736, 738) corresponding to each subband. Each encoder (732, 734, 736, 738) can perform encoding to output audio vector data (e.g., latent vector values) corresponding to each of the four subbands.

[0090] According to one embodiment, audio signals of subbands are input to a psychoacoustic model (720), and a masking threshold value for each frequency band can be obtained from the psychoacoustic model (720). A quantization controller (725) can determine a masking threshold value for each of the four subbands based on the masking threshold value for each frequency band obtained from the psychoacoustic model (720). The quantization controller (725) can determine a quantization control value to be applied to residual vector quantizers (740, 750, 760, 770) for each of the four subbands based on the masking threshold value corresponding to each of the four subbands. The quantization controller (725) can adjust a quantization stage (or quantization level) applied to each of the residual vector quantizers based on the masking threshold value corresponding to each of the four subbands.

[0091] In one embodiment, the quantization controller (725) can determine a quantization control value by considering a given target bit size for the input audio signal (705). For example, assuming the given target bit size is 100 bits, the quantization controller (725) can determine a quantization control value such that the sum of the data sizes of the quantized audio vector data generated by the quantization processing of each of the residual vector quantizers (740, 750, 760, 770) is 100 bits or less.

[0092] Each of the residual vector quantizers (740, 750, 760, 770) can generate quantized audio vector data by performing quantization processing on audio vector data transmitted from each encoder (732, 734, 736, 738) based on a quantization control value transmitted from the quantization controller (725). Depending on the quantization control value, the quantization stage to be performed on each residual vector quantizer (740, 750, 760, 770) can be determined. For example, it is assumed that, given the target bit size of 100 bits above, the residual vector quantizer (740) is activated only up to the first quantization stage corresponding to 10 bits, the residual vector quantizer (750) is activated only up to the third quantization stage corresponding to 30 bits, the residual vector quantizer (760) is activated up to the fourth quantization stage corresponding to 40 bits, and the residual vector quantizer (770) is activated only up to the second quantization stage corresponding to 20 bits, and the quantization control values ​​for each residual vector quantizer (740, 750, 760, 770) are determined such that the residual vector quantizer (760) is activated only up to the second quantization stage corresponding to 20 bits. In this case, the audio signal of the subband for which the residual vector quantizer (760) is responsible for quantization may have a relatively low masking threshold value, which may indicate that the sound quality is relatively affected. On the other hand, the audio signal of the subband in which the residual vector quantizer (740) is responsible for quantization may have a high masking threshold value, which may indicate that it has relatively little effect on sound quality.

[0093] In one embodiment, audio vector data of subbands quantized by residual vector quantizers (740, 750, 760, 770) can be combined to generate a compressed audio signal. The audio encoding device (702) can transmit the compressed audio signal to an audio decoding device (704) that performs the restoration of the audio signal. The audio encoding device (702) can transmit the compressed audio signal to the audio decoding device (704), for example, in the form of a bitstream.

[0094] The audio decoding device (704) receives a compressed audio signal from the audio encoding device (702) and can input the compressed audio signal to different decoders (782, 784, 786, 788) for each of the four subbands. Audio data restored for the four subbands can be output from the decoders (782, 784, 786, 788). The audio data restored for each of the four subbands is an audio signal in the frequency domain, and the audio data restored for each subband can be converted into a time domain restored audio signal (795) through an inverse transform (790) such as, for example, IFFT, ISTFT, or IMDCT.

[0095] FIG. 8 is a diagram illustrating the interaction between an electronic device performing audio signal processing according to various embodiments and another electronic device.

[0096] Referring to FIG. 8, an electronic device (101) (e.g., the audio encoding device (402) of FIG. 4 or the audio encoding device (702) of FIG. 7) may perform operations of compressing an audio signal among the audio signal processing methods described in the present disclosure. An electronic device (800) (e.g., the electronic device (102, 104) of FIG. 1, the audio decoding device (404) of FIG. 4, or the audio decoding device (704) of FIG. 7) may perform operations of decompressing the compressed audio signal to obtain a restored audio signal. The electronic device (101) may be various electronic devices such as, for example, a mobile terminal, a terminal device, a smartphone, a personal computer (PC), a server, a laptop, a tablet PC, a VR (virtual reality) / AR (augmented reality) device, or a pad-type electronic device. The electronic device (800) may include, for example, an audio signal output device such as wireless earphones, a headset, and a wireless speaker device, a head-mounted device (HMD), a portable multimedia device, or a home appliance (e.g., a TV). A compressed audio signal may be transmitted from the electronic device (101) to the electronic device (800) via a network. The network may include a wired network of a cable network, a short-range wireless network, or a long-range wireless network. The short-range wireless network may include, for example, Bluetooth, wireless fidelity (WiFi), or infrared data association (IrDA), and the long-range wireless network may include a legacy cellular network, a 3G / 4G / 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN).

[0097] In one embodiment, the electronic device (101) may include one or more memories for storing instructions (e.g., memory (130) of FIG. 1) and one or more processors (e.g., processor (120) of FIG. 1). When the instructions stored in the memory are executed individually or collectively by one or more processors (120), the electronic device (101) may be able to perform operations such as compressing audio signals.

[0098] The electronic device (101) can convert an input audio signal into a frequency domain audio signal. The electronic device (101) can convert a time domain audio signal into a frequency domain audio signal using, for example, a conversion method of FFT, STFT, or MDCT. The electronic device (101) can divide the frequency domain audio signal into audio signals of multiple subbands. The electronic device (101) can determine a masking threshold value corresponding to each of the subbands using a psychoacoustic model (e.g., the psychoacoustic model (440) of FIG. 4 or the psychoacoustic model (720) of FIG. 7) that takes the audio signals of the subbands as input. The electronic device (101) can obtain audio vector data corresponding to each of the subbands using neural network-based encoders (e.g., the encoders (452, 454) of FIG. 4 or the encoders (732, 734, 736, 738) of FIG. 7). The electronic device (101) can generate a compressed audio signal by quantizing audio vector data corresponding to each of the subbands based on a masking threshold value corresponding to each of the subbands. The electronic device (101) can determine a quantization control value to be applied to quantizers for each of the subbands (e.g., residual vector quantizers (462, 464) of FIG. 4 or residual vector quantizers (740, 750, 760, 770) of FIG. 7) based on a masking threshold value corresponding to each of the subbands.

[0099] In one embodiment, the electronic device (101) can determine a quantization control value such that the quantization error is minimized for each subband. The electronic device can adjust the quantization stage (or quantization level) applied to each residual vector quantizer based on a masking threshold value corresponding to each subband. Each residual vector quantizer can adjust the quantization stage of the audio vector data input to each residual vector quantizer based on the quantization control value applied to each residual vector quantizer. The residual vector quantizers can convert audio vector data into quantized audio vector data through quantization processing.

[0100] The electronic device (10) can generate a compressed audio signal by combining audio vector data of subbands quantized by residual vector quantizers. In one embodiment, the electronic device (101) may further include a communication module (e.g., the communication module (190) of FIG. 1) for communicating with the electronic device (800). The electronic device (101) can transmit the compressed audio signal to the electronic device (800) through the communication module, thereby enabling the electronic device (800) to restore the compressed audio signal and generate a restored audio signal. The compressed audio signal may be transmitted to the electronic device (800) over a network in the form of a bitstream.

[0101] The electronic device (800) can receive a compressed audio signal. The electronic device (800) can obtain restored audio data corresponding to each of the subbands using neural network-based decoders (e.g., decoders (472, 474) of FIG. 4 or decoders (782, 784, 785, 788) of FIG. 7). The electronic device (800) can convert the restored audio data for each subband into a time-domain audio signal. In one embodiment, the electronic device (800) can combine the restored audio data for each subband to generate a restored audio signal in the frequency domain, and convert the restored audio signal in the frequency domain into a restored audio signal in the time domain using an inverse transform method of IFFT or IMDCT. The electronic device (800) can output the output audio signal generated through these operations through a speaker.

[0102] An audio signal processing method according to one embodiment may include: an operation of receiving an input audio signal (210); an operation of converting the input audio signal into a frequency domain audio signal (220); an operation of dividing the frequency domain audio signal into audio signals of a plurality of subbands (230); an operation of determining a masking threshold value corresponding to each of the subbands using a psychoacoustic model (440, 720) that takes the audio signals of the subbands as input (240); an operation of obtaining audio vector data corresponding to each of the subbands using neural network-based encoders (452, 454, 732, 734, 736, 738) (250); and an operation of generating a compressed audio signal by quantizing the audio vector data corresponding to each of the subbands based on the masking threshold value corresponding to each of the subbands (260).

[0103] The operation (260) of generating the compressed audio signal may include the operation of determining a quantization control value for each of the subbands based on a masking threshold value corresponding to each of the subbands, and the operation of controlling quantization processing by each of the residual vector quantizers (462, 464, 740, 750, 760, 770) corresponding to each of the subbands based on the determined quantization control value for each of the subbands.

[0104] Each of the above residual vector quantizers (462, 464, 740, 750, 760, 770) can control the quantization stage of the audio vector data input to each residual vector quantizer based on a quantization control value applied to each residual vector quantizer (462, 464, 740, 750, 760, 770).

[0105] According to the above quantization control value, the number of bits allocated to the compression of the audio vector data or the number of quantization stages applied to the quantization of the audio vector data may be determined.

[0106] The above residual vector quantizers (462, 464, 740, 750, 760, 770) can convert the input audio vector data into quantized audio vector data through the quantization process. The data size of the input audio vector data may be larger than the data size of the quantized audio vector data. The data size of the quantized audio vector data may be determined according to the quantization stage.

[0107] The operation of determining the above quantization control value may include: an operation of determining a first quantization control value such that a first residual vector quantizer that quantizes first audio vector data corresponding to a first subband performs quantization processing with a first quantization stage determined for a first masking threshold value corresponding to the first subband; and an operation of determining a second quantization control value such that a second residual vector quantizer that quantizes second audio vector data corresponding to a second subband performs quantization processing with a second quantization stage determined for a second masking threshold value corresponding to the second subband.

[0108] The first residual vector quantizer can generate first quantized audio vector data by quantizing the first audio vector data into a first quantization stage according to the first quantization control value. The second residual vector quantizer can generate second quantized audio vector data by quantizing the second audio vector data into a second quantization stage according to the second quantization control value. If the first masking threshold is greater than the second masking threshold, the data size of the first quantized audio vector data may be smaller than the data size of the second quantized audio vector data.

[0109] The operation of determining the first quantization control value may include determining, among candidate quantization control values ​​corresponding to different quantization stages, a candidate quantization control value that minimizes the quantization error value generated by the quantization processing of the first residual vector quantizer as the first quantization control value. The operation of determining the second quantization control value may include determining, among candidate quantization control values ​​corresponding to different quantization stages, a candidate quantization control value that minimizes the quantization error value generated by the quantization processing of the second residual vector quantizer as the second quantization control value.

[0110] The operation (240) of determining a masking threshold value corresponding to each of the above subbands may include the operation of obtaining a masking threshold value for the entire frequency band of the input audio signal from the psychoacoustic model (440, 720), and the operation of determining a masking threshold value for each of the above subbands based on the obtained masking threshold value for the entire frequency band.

[0111] The masking threshold for the entire frequency band may depend on the signal magnitude of the audio signals of the subbands input to the psychoacoustic model (440, 720) per frequency band.

[0112] The psychoacoustic model (440, 720) can output a masking threshold value for each frequency band of the input audio signal based on the minimum audible limit information of a person and the masking effect based on the signal magnitude of the audio signals of the subbands by frequency band.

[0113] Each of the encoders (452, 454, 732, 734, 736, 738) receives an audio signal of a different subband, and each of the encoders (452, 454, 732, 734, 736, 738) can output audio vector data corresponding to a different subband.

[0114] A computer-readable recording medium storing one or more computer programs according to one embodiment may include instructions for performing the audio signal processing method.

[0115] An electronic device (101) according to one embodiment may include one or more memories (130) for storing instructions and one or more processors (120). When the instructions are executed individually or collectively by the one or more processors (120), the electronic device (101) may generate a compressed audio signal by converting an input audio signal into a frequency domain audio signal, dividing the frequency domain audio signal into audio signals of a plurality of subbands, determining a masking threshold value corresponding to each of the subbands using a psychoacoustic model (440, 720) that takes the audio signals of the subbands as input, obtaining audio vector data corresponding to each of the subbands using neural network-based encoders (452, 454, 732, 734, 736, 738), and quantizing the audio vector data corresponding to each of the subbands based on the masking threshold value corresponding to each of the subbands.

[0116] When the above instructions are executed individually or collectively by the one or more processors (120), the electronic device (101) may determine a quantization control value for each of the subbands based on a masking threshold value corresponding to each of the subbands, and, based on the determined quantization control value for each of the subbands, control quantization processing by each of the residual vector quantizers (462, 464, 740, 750, 760, 770) corresponding to each of the subbands.

[0117] When the above instructions are executed individually or collectively by the one or more processors (120), the electronic device (101) may determine a first quantization control value that causes a first residual vector quantizer that quantizes first audio vector data corresponding to a first subband to perform quantization processing with a first quantization stage determined for a first masking threshold value corresponding to the first subband, and a second quantization control value that causes a second residual vector quantizer that quantizes second audio vector data corresponding to a second subband to perform quantization processing with a second quantization stage determined for a second masking threshold value corresponding to the second subband.

[0118] When the above instructions are executed individually or collectively by the one or more processors (120), the electronic device (101) may determine the first quantization control value among candidate quantization control values ​​corresponding to different quantization stages that minimizes the quantization error value generated by the quantization processing of the first residual vector quantizer, and determine the second quantization control value among candidate quantization control values ​​corresponding to different quantization stages that minimizes the quantization error value generated by the quantization processing of the second residual vector quantizer.

[0119] The above electronic device (101) may further include a communication module (190) for communicating with other electronic devices (102, 102, 800).

[0120] When the above instructions are executed individually or collectively by the one or more processors (120), the electronic device (101) may be made to transmit the compressed audio signal to the other electronic device (102, 104, 800) through the communication module (190), thereby allowing the other electronic device (102, 104, 800) to restore the compressed audio signal and generate a restored audio signal.

[0121] The other electronic devices (102, 104, 800) may include an audio signal output device.

[0122] The various embodiments of the present disclosure and the terms used therein are not intended to limit the technical features described in the present disclosure to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In the present disclosure, phrases such as “A or B,” “at least one of A and B,” “at least one of A or B,” “A, B or C,” “at least one of A, B and C,” and “at least one of A, B, or C” each may include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as “first,” “second,” or “first” or “second” may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as "coupled" or "connected" to another (e.g., 2nd) component, with or without the terms "functionally" or "communicationly," it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.

[0123] At least one of the operations described in the various embodiments of the present disclosure may be performed simultaneously or in parallel with other operations, and the order of the operations may be changed. Additionally, at least one of the operations may be omitted, and other operations may be performed additionally.

[0124] In the description of various embodiments of the present disclosure, the operation of “A transmits B to C” may include not only “A transmits B so that B is immediately delivered to C” but also “A transmits B, and D receives B in the intermediary and then D transmits B to C.” There may be one or more Ds that transmit B between A and C.

[0125] The term “module” as used in various embodiments of the present disclosure may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0126] Various embodiments of the present disclosure may be implemented as software comprising one or more instructions stored in a storage medium readable by a machine (e.g., the electronic device (101) of FIG. 1 or another electronic device (800) of FIG. 8). For example, a processor of the machine (e.g., the processor (120) of FIG. 1) may call at least one of the one or more instructions stored from the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.

[0127] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or instruct the processing unit independently or collectively. Software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, or computer storage medium or device so as to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computer systems and stored or executed in a distributed manner. Software and data may be stored on computer-readable recording media.

[0128] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created in a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0129] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

[0130] The embodiments described above may be implemented as hardware components, software components, and / or combinations of hardware and software components. For example, the devices, methods, and components described in the embodiments may be implemented using a general-purpose computer or a special-purpose computer, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include multiple processing elements and / or multiple types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. Additionally, other processing configurations, such as parallel processors, are also possible.

[0131] The hardware device described above may be configured to operate as one or more software modules to perform the operation of the embodiment, and vice versa.

[0132] Although the present disclosure has been illustrated and described with reference to various embodiments, it will be understood that the various embodiments are for illustrative purposes only and are not limiting. It will be further understood by those skilled in the art that various modifications of form and detail may be made without departing from the true spirit and full scope of the present disclosure, including the appended claims and their equivalents. Additionally, it will be understood that any embodiment(s) described in the present disclosure may be used in combination with any other embodiment(s) described in the present disclosure.

Claims

1. In an audio signal processing method, Operation of receiving an input audio signal (210); The operation of converting the above input audio signal into a frequency domain audio signal (220); The operation of dividing the audio signal in the above frequency domain into audio signals of a plurality of subbands (230); An operation (240) of determining a masking threshold value corresponding to each of the subbands using a psychoacoustic model (440; 720) that takes the audio signals of the subbands as input; An operation (250) of acquiring audio vector data corresponding to each of the subbands using neural network-based encoders (452; 454; 732; 734; 736; 738); and The operation (260) of generating a compressed audio signal by quantizing audio vector data corresponding to each of the subbands based on a masking threshold value corresponding to each of the subbands. An audio signal processing method including 2. In Paragraph 1, The operation (260) of generating the above compressed audio signal is, An operation of determining a quantization control value for each of the subbands based on a masking threshold value corresponding to each of the subbands; and Operation of controlling quantization processing by each of the residual vector quantizers (462; 464; 740; 750; 760; 770) corresponding to each of the subbands, based on the quantization control value for each of the subbands determined above. An audio signal processing method including 3. In Paragraph 2, Each of the above residual vector quantizers (462; 464; 740; 750; 760; 770) is, Based on the quantization control value applied to each residual vector quantizer (462; 464; 740; 750; 760; 770), the quantization stage of the audio vector data input to each residual vector quantizer is adjusted, and The number of bits allocated to the compression of the audio vector data is determined according to the above quantization control value, or the number of quantization stages applied to the quantization of the audio vector data is determined. Audio signal processing method.

4. In Paragraph 3, The above residual vector quantizers (462; 464; 740; 750; 760; 770) are, The input audio vector data is converted into quantized audio vector data through the above quantization processing, and The data size of the above-mentioned input audio vector data is, Larger than the data size of the above-mentioned quantized audio vector data, and The data size of the quantized audio vector data is determined according to the quantization stage. Audio signal processing method.

5. In any one of paragraphs 2 through 4, The operation of determining the above quantization control value is, An operation of determining a first quantization control value that causes a first residual vector quantizer, which quantizes first audio vector data corresponding to a first subband, to perform quantization processing with a first quantization stage determined for a first masking threshold value corresponding to the first subband; and An operation of determining a second quantization control value that causes a second residual vector quantizer, which quantizes second audio vector data corresponding to a second subband, to perform quantization processing with a second quantization stage determined for a second masking threshold value corresponding to the second subband. An audio signal processing method including 6. In Paragraph 5, The first residual vector quantizer generates first quantized audio vector data by quantizing the first audio vector data into a first quantization stage according to the first quantization control value, and The second residual vector quantizer generates second quantized audio vector data by quantizing the second audio vector data into a second quantization stage according to the second quantization control value, and If the first masking threshold is greater than the second masking threshold, the data size of the first quantized audio vector data is smaller than the data size of the second quantized audio vector data. Audio signal processing method.

7. In Paragraph 5 or 6, The operation of determining the first quantization control value above is, Among candidate quantization control values ​​corresponding to different quantization stages, the operation of determining a candidate quantization control value that minimizes the quantization error value generated by the quantization processing of the first residual vector quantizer as the first quantization control value is included. The operation of determining the above second quantization control value is, Among candidate quantization control values ​​corresponding to different quantization stages, the operation of determining a candidate quantization control value that minimizes the quantization error value generated by the quantization processing of the second residual vector quantizer as the second quantization control value, Audio signal processing method.

8. In any one of paragraphs 1 through 7, The operation (240) of determining a masking threshold value corresponding to each of the above subbands is, The operation of obtaining a masking threshold value for the entire frequency band of the input audio signal from the psychoacoustic model (440; 720); and Operation of determining a masking threshold value for each of the subbands based on the masking threshold value for the entire frequency band obtained above. Includes, The masking threshold value for the entire frequency band above is, It depends on the signal magnitudes of the audio signals of the subbands input to the psychoacoustic model (440; 720) according to frequency bands. Audio signal processing method.

9. In any one of paragraphs 1 through 8, The above psychoacoustic model (440; 720) is, Outputting a masking threshold value for each frequency band of the input audio signal based on information regarding the minimum audible limit of humans and the masking effect caused by the signal magnitude of the audio signals of the subbands according to frequency bands, Audio signal processing method.

10. In any one of paragraphs 1 through 9, Each of the encoders (452; 454; 732; 734; 736; 738) receives an audio signal of a different subband, and each of the encoders (452; 454; 732; 734; 736; 738) outputs audio vector data corresponding to a different subband. Audio signal processing method.

11. In an electronic device (101), One or more memories (130) for storing instructions; and One or more processors (120) Includes, When the above instructions are executed individually or collectively by the one or more processors (120), the electronic device (101) is enabled, Converts the input audio signal into a frequency domain audio signal, and The above-mentioned audio signal in the frequency domain is divided into audio signals of multiple subbands, and A masking threshold value corresponding to each of the subbands is determined using a psychoacoustic model (440; 720) that takes the audio signals of the subbands as input, and Audio vector data corresponding to each of the subbands is obtained using neural network-based encoders (452; 454; 732; 734; 736; 738), and Generating a compressed audio signal by quantizing audio vector data corresponding to each of the subbands based on a masking threshold value corresponding to each of the subbands. Electronic device (101).

12. In Paragraph 11, When the above instructions are executed individually or collectively by the one or more processors (120), the electronic device (101) is enabled, A quantization control value for each of the subbands is determined based on a masking threshold value corresponding to each of the subbands, and Based on the quantization control value for each of the determined subbands above, controlling the quantization processing by each of the residual vector quantizers (462; 464; 740; 750; 760; 770) corresponding to each of the subbands, Electronic device (101).

13. In Paragraph 12, When the above instructions are executed individually or collectively by the one or more processors (120), the electronic device (101) is enabled, A first residual vector quantizer that quantizes first audio vector data corresponding to a first subband determines a first quantization control value that causes quantization processing to be performed by a first quantization stage determined for a first masking threshold value corresponding to the first subband, and A second residual vector quantizer that quantizes second audio vector data corresponding to a second subband determines a second quantization control value that causes quantization processing to be performed by a second quantization stage determined for a second masking threshold value corresponding to the second subband. Electronic device (101).

14. In Paragraph 13, The first residual vector quantizer generates first quantized audio vector data by quantizing the first audio vector data into a first quantization stage according to the first quantization control value, and The second residual vector quantizer generates second quantized audio vector data by quantizing the second audio vector data into a second quantization stage according to the second quantization control value, and If the first masking threshold is greater than the second masking threshold, the data size of the first quantized audio vector data is smaller than the data size of the second quantized audio vector data. Electronic device (101).

15. In any one of paragraphs 12 through 14, The above residual vector quantizers (462; 464; 740; 750; 760; 770) are, Based on the quantization control value applied to each residual vector quantizer (462; 464; 740; 750; 760; 770), the quantization stage of the audio vector data input to each residual vector quantizer is adjusted, and The number of bits allocated to the compression of the audio vector data is determined according to the above quantization control value, or the number of quantization stages applied to the quantization of the audio vector data is determined. Electronic device (101).

Citation Information

Patent Citations

  • Audio data encoding apparatus and method

    KR100547113B1

  • A method and apparatus for processing an audio signal

    KR1020090122142A

  • A generative adversarial network-based wireless signal inpainting method with multiple discriminators for automatic modulation classification, and system thereof

    KR1020250126401A

  • Apparatus and method for encoding / decoding audio signal using filter bank

    KR102594160B1

  • System and method for providing charging station lighthouse service

    KR102670292B1