Electronic device and method for audio processing
By determining key and sub-object audio signals and compressing them for multi-channel output, the method addresses the challenge of rendering immersive object-based audio in devices with channel-based limitations, enhancing audio quality and reducing errors.
Patent Information
- Application Number
- PCT/KR2024/020699
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-28
- Filing Date
- 2024-12-19
- Publication Date
- 2025-07-03
AI Technical Summary
Existing audio processing technologies struggle to efficiently handle both channel-based and object-based audio signals, particularly in scenarios where devices can only render channel-based audio, leading to limitations in outputting immersive object-based audio.
A method and system for determining key and sub-object audio signals, compressing these signals, and generating bitstreams to enable rendering of multi-channel audio, even in devices that can only output channel-based audio, by prioritizing key object audio signals based on importance and spatial information.
Enhances audio output quality by improving sound quality and reducing compression errors, allowing devices to render immersive object-based audio effectively even when limited to channel-based rendering capabilities.
Smart Images

Figure KR2024020699_03072025_PF_FP_ABST
Abstract
Description
Electronic devices and methods for audio processing
[0001] Embodiments disclosed in this document relate to electronic devices and methods for audio processing.
[0002] Channel-based audio is audio in which multiple audio signals are output in a predefined channel layout. It is an audio signal output method in which it is predetermined which channel the multiple audio signals will be transmitted to.
[0003] Object-based audio is an audio implementation that allows users to experience immersive sound, the sound of objects moving in three-dimensional space. It is an audio signal output method in which the speakers to be output are determined based on the spatial information of the object, rather than a predefined channel layout.
[0004] A method for processing audio in an electronic device according to one embodiment of the present disclosure may include a step of determining at least one object audio signal among a plurality of object audio signals as a key object audio signal. The method may include a step of determining an object audio signal excluding the key object audio signal among the plurality of object audio signals as a sub-object audio signal. The method may include a step of obtaining a multi-channel audio signal using the key object audio signal and the sub-object audio signal. The method may include a step of generating a bitstream by compressing the key object audio signal and the multi-channel audio signal.
[0005] An electronic device according to one embodiment of the present disclosure may include a memory having one or more instructions stored therein, and at least one processor configured to execute one or more instructions stored in the memory. When the at least one processor executes the one or more instructions, the electronic device may determine at least one object audio signal among a plurality of object audio signals as a key object audio signal. When the at least one processor executes the one or more instructions, the electronic device may determine an object audio signal excluding the key object audio signal among the plurality of object audio signals as a sub-object audio signal. When the at least one processor executes the one or more instructions, the electronic device may obtain a multi-channel audio signal using the key object audio signal and the sub-object audio signal. The at least one processor may compress the key object audio signal and the multi-channel audio signal to generate a bitstream.
[0006] In one embodiment of the present disclosure, a computer-readable recording medium having recorded thereon a program for performing a method on a computer may be provided. The program recorded on the recording medium may include a step of determining at least one object audio signal among a plurality of object audio signals as a key object audio signal. The program recorded on the recording medium may include a step of determining an object audio signal excluding the key object audio signal among the plurality of object audio signals as a sub-object audio signal. The program recorded on the recording medium may include a step of obtaining a multi-channel audio signal using the key object audio signal and the sub-object audio signal. The program recorded on the recording medium may include a step of generating a bitstream by compressing the key object audio signal and the multi-channel audio signal.
[0007] A method for processing audio in an electronic device according to one embodiment of the present disclosure may include a step of decompressing a bitstream to obtain a key object audio signal and a first multi-channel audio signal in which the key object audio signal and the sub-object audio signal are rendered. The method may include a step of obtaining a second multi-channel audio signal that can be output from the electronic device based on the first multi-channel audio signal and information about a channel layout of the electronic device, if it is determined that the electronic device is not capable of outputting the key object audio signal. The method may include a step of obtaining a third multi-channel audio signal that can be output from the electronic device using the key object audio signal and the first multi-channel audio signal, if it is determined that the electronic device is capable of outputting the key object audio signal.
[0008] An electronic device according to one embodiment of the present disclosure may include a memory having one or more instructions stored therein, and at least one processor configured to execute one or more instructions stored in the memory. When the at least one processor executes the one or more instructions, the electronic device may decompress a bitstream to obtain a key object audio signal and a first multi-channel audio signal in which the key object audio signal and the sub-object audio signal are rendered. When the at least one processor executes the one or more instructions, the electronic device may obtain a second multi-channel audio signal that can be output from the electronic device based on the first multi-channel audio signal and information about a channel layout of the electronic device, if it is determined that the electronic device cannot output the key object audio signal. When the at least one processor executes the one or more instructions, the electronic device may obtain a third multi-channel audio signal that can be output from the electronic device using the key object audio signal and the first multi-channel audio signal.
[0009] In one embodiment of the present disclosure, a computer-readable recording medium having recorded thereon a program for performing a method on a computer may be provided. The program recorded on the recording medium may include a step of decompressing a bitstream to obtain a key object audio signal and a first multi-channel audio signal in which the key object audio signal and the sub-object audio signal are rendered. The program recorded on the recording medium may include a step of obtaining a second multi-channel audio signal that can be output from the electronic device based on the first multi-channel audio signal and information about a channel layout of the electronic device, if it is determined that the electronic device is not capable of outputting the key object audio signal. The program recorded on the recording medium may include a step of obtaining a third multi-channel audio signal that can be output from the electronic device using the key object audio signal and the first multi-channel audio signal, if it is determined that the electronic device is capable of outputting the key object audio signal.
[0010] FIG. 1 is a drawing for explaining an electronic device according to one embodiment of the present disclosure.
[0011] FIG. 2 is a flowchart illustrating an audio encoding method in an electronic device according to one embodiment of the present disclosure.
[0012] FIG. 3 is a diagram for explaining spatial information of an object corresponding to an object audio signal according to one embodiment of the present disclosure.
[0013] FIG. 4 is a flowchart illustrating a method for preprocessing a key object audio signal according to one embodiment of the present disclosure.
[0014] FIG. 5 is a diagram for explaining an audio encoding method according to one embodiment of the present disclosure.
[0015] FIG. 6 is a block diagram illustrating an electronic device performing audio encoding according to one embodiment of the present disclosure.
[0016] FIG. 7 is a flowchart illustrating an audio decoding method in an electronic device according to one embodiment of the present disclosure.
[0017] FIG. 8 is a flowchart illustrating a method for rendering a first multi-channel audio signal according to one embodiment of the present disclosure.
[0018] FIG. 9 is a block diagram illustrating an electronic device that performs audio decoding according to one embodiment of the present disclosure.
[0019] FIG. 10 is a flowchart illustrating an audio processing method in an electronic device according to one embodiment of the present disclosure.
[0020] FIG. 11 is a schematic block diagram of an electronic device that performs an audio processing method according to one embodiment of the present disclosure.
[0021] FIG. 12 is a specific block diagram of an electronic device according to one embodiment of the present disclosure.
[0022] Hereinafter, embodiments of the present disclosure are described in detail with reference to the attached drawings.
[0023] The present disclosure may be subject to various modifications and embodiments. Specific embodiments are illustrated in the drawings and described in detail herein. However, this is not intended to limit the embodiments of the present disclosure, and it should be understood that the present disclosure encompasses all modifications, equivalents, and alternatives falling within the spirit and technical scope of the various embodiments.
[0024] In describing the embodiments, detailed descriptions of related known technologies are omitted if they are deemed to unnecessarily obscure the gist of the present disclosure. Furthermore, numbers (e.g., "first," "second," etc.) used throughout the description of the specification are merely identifiers used to distinguish one component from another.
[0025] The terms used in the embodiments of this specification have been selected from widely used, current terms, taking into account the functions of the present disclosure. However, these terms may vary depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, the applicant may arbitrarily select terms, and in such cases, their meanings will be described in detail in the description of the relevant embodiments. Therefore, the terms used in this disclosure should not be defined simply as names of terms, but rather based on the meanings of the terms and the overall content of the present disclosure.
[0026] The scope of the present disclosure may be indicated by the claims that follow rather than the detailed description above. Various features mentioned in one claim category of the present disclosure (e.g., in a method claim) may also be claimed in another claim category (e.g., in a system claim). Furthermore, an embodiment of the present disclosure may include not only combinations of features specified in the appended claims, but also various combinations of individual features within the claims. The scope of the present disclosure should be interpreted to include all changes or modifications derived from the meaning and scope of the claims and their equivalent concepts.
[0027] In addition, components expressed as 'unit', 'module', etc. in the present disclosure may be two or more components combined into one component, or one component may be divided into two or more components with more detailed functions. These functions may be implemented by hardware or software, or a combination of hardware and software. In addition, each component described below may additionally perform some or all of the functions performed by other components in addition to its own main function, and of course, some of the main functions performed by each component may be dedicated and performed by other components.
[0028] Singular expressions may include plural expressions unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art described herein.
[0029] In the present disclosure, a processor is a component that controls a series of processes so that an electronic device operates according to the embodiments described below, and may be composed of one or more processors. One or more processors included in the processor may be circuitry such as a System on Chip (SoC), an Integrated Circuit (IC), etc. One or more processors included in the processor may be a general-purpose processor such as a Central Processing Unit (CPU), a Micro Processor Unit (MPU), an Application Processor (AP), a Digital Signal Processor (DSP), a graphics-only processor such as a Graphics Processing Unit (GPU), a Vision Processing Unit (VPU), an artificial intelligence-only processor such as a Neural Processing Unit (NPU), or a communication-only processor such as a Communication Processor (CP). When one or more processors included in the processor are artificial intelligence-only processors, the artificial intelligence-only processor may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
[0030] In the present disclosure, a processor may include various processing circuits and / or multiple processors. For example, the term “processor” as used herein, including in the claims, may include various processing circuits, including at least one processor. At least one processor, one or more processors may be configured to perform various functions described herein, individually and / or collectively, in a distributed fashion. As used herein, “processor,” “at least one processor,” and “one or more processors” may be configured to perform various functions. However, these terms encompass, without limitation, situations where one processor performs some of the functions and other processor(s) perform other parts of the functions, and situations where a single processor may perform all of the functions. Furthermore, the at least one processor may include a combination of processors that perform various functions of the disclosed functions in a distributed manner. At least one processor may execute program instructions to achieve or perform various functions.
[0031] It should be understood that the blocks and combinations of flowcharts in each of the flowcharts in this disclosure can be implemented by one or more computer programs containing computer-executable instructions. The one or more computer programs may be stored entirely in a single memory, or may be divided and stored across multiple different memories.
[0032] In this disclosure, all functions or operations described in this document may be processed by a single processor or a combination of processors. A single processor or a combination of processors may be a circuitry that performs processing, and may include circuitry such as an Application Processor (AP), a Communication Processor (CP), a Graphical Processing Unit (GPU), a Neural Processing Unit (NPU), a Microprocessor Unit (MPU), a System on Chip (SoC), an Integrated Chip (IC), etc.
[0033] The processor can write data to memory, read data stored in memory, and process data according to predefined operating rules or artificial intelligence models, particularly by executing a program or at least one instruction stored in memory. Accordingly, the processor can perform the operations described in the following embodiments, and operations described as performed by electronic devices or detailed components included in the electronic devices in the following embodiments can be considered to be performed by the processor, unless otherwise specified.
[0034] When a component is referred to as being "connected" or "connected" to another component in this disclosure, it should be understood that the component may be directly connected or connected to the other component, but may also be connected or connected via another component in between, unless otherwise specifically stated.
[0035] Throughout this disclosure, unless specifically stated otherwise, "or" is inclusive and not exclusive. Thus, unless explicitly stated otherwise or the context dictates otherwise, "A or B" can mean "A, B, or both." As used herein, the phrases "at least one of" or "one or more of" can mean that different combinations of one or more of the listed items can be used, or that only any one of the listed items is required. For example, "at least one of A, B, and C" can include any of the following combinations: A, B, C, A and B, A and C, B and C, or A and B and C.
[0036] It will be appreciated that each block of the flowchart drawings and combinations of the flowchart drawings can be performed by computer program instructions. These computer program instructions can be installed in a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing equipment, such that the instructions, when executed by the processor of the computer or other programmable data processing equipment, create a means for performing the functions described in the flowchart block(s). These computer program instructions can also be stored in a computer-available or computer-readable memory that can direct a computer or other programmable data processing equipment to perform the functions in a specific manner, such that the instructions stored in the computer-available or computer-readable memory can produce an article of manufacture that includes instruction means for performing the functions described in the flowchart block(s). Since the computer program instructions may be installed on a computer or other programmable data processing device, a series of operational steps may be performed on the computer or other programmable data processing device to create a computer-executable process, and the instructions that cause the computer or other programmable data processing device to perform the steps for performing the functions described in the flowchart block(s) may also provide steps for performing the functions described in the flowchart block(s).
[0037] Additionally, each block may represent a module, segment, or portion of code that contains one or more executable instructions for performing a specific logical function(s). It should also be noted that in some alternative implementation examples, the functions described in the blocks may occur out of order. For example, two blocks depicted in succession may actually be executed substantially concurrently, or the blocks may sometimes be executed in reverse order, depending on their respective functions.
[0038] Below, embodiments of the present disclosure are described in detail with reference to the attached drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In addition, for the purpose of clearly explaining the present disclosure in the drawings, parts irrelevant to the description are omitted, and similar parts are designated with similar reference numerals throughout the specification.
[0039] The terms used in this disclosure will be briefly explained, and an embodiment of the present invention will be specifically described.
[0040] The terms described below are defined based on their functions within the present invention and may vary depending on the intent or custom of the user or operator. Therefore, their definitions should be based on the contents of this specification.
[0041] In the present disclosure, the term "electronic device" may refer to any device that processes an audio signal. For example, the "electronic device" may perform at least one of a method for encoding or decoding an audio signal.
[0042] In the present disclosure, an 'audio signal' may mean a signal including information related to audio. For example, the 'audio signal' may be a 'channel-based audio signal' or an 'object-based audio signal', or may include both a 'channel-based audio signal' and an 'object-based audio signal'.
[0043] In the present disclosure, a 'channel-based audio signal' may refer to an audio signal through which it is predefined which channel a plurality of audio signals will be output through according to a channel layout. In the present disclosure, a 'channel audio signal' refers to a 'channel-based audio signal.'
[0044] In the present disclosure, 'channel layout' may mean a combination of at least one channel or a spatial arrangement of speakers connected to channels. For example, the channel layout may be an XYZ channel layout, where X may represent the number of surround channels, Y may represent the number of subwoofer channels, and Z may represent the number of height channels. Examples of the 'channel layout' include, but are not limited to, a 1.0.0 channel (mono channel) layout, a 2.0.0 channel (stereo channel) layout, a 5.1.0 channel layout, a 5.1.2 channel layout, a 5.1.4 channel layout, a 7.1.0 layout, a 7.1.2 layout, and a 3.1.2 channel layout.
[0045] In the present disclosure, 'rendering' may mean an operation of converting a plurality of audio signals into an audio signal that can be output from one or more speakers.
[0046] In the present disclosure, an 'object-based audio signal' may mean an audio signal rendered so that the speaker to be output is determined based on spatial information of the object, rather than a predefined channel layout. In the present disclosure, an 'object audio signal' means an 'object-based audio signal.'
[0047] In the present disclosure, a 'multi-channel audio signal' may mean an audio signal of n channels (where n is an integer greater than 2). The 'multi-channel audio signal' may mean a channel-based audio signal or an object-based audio signal, or an audio signal including both a channel-based audio signal and an object-based audio signal.
[0048] In the present disclosure, "compression" may refer to an operation of reducing the amount of data in an audio signal, thereby enabling more efficient storage or transmission of the audio signal. For example, "compression" may be performed by a codec or encoder, and the decompression operation may be performed by a codec or decoder.
[0049] In the present disclosure, 'mixing' means a signal processing operation of generating a new audio signal by adding the respective values obtained by multiplying each of a plurality of audio signals by their respective corresponding weights (i.e., mixing a plurality of audio signals).
[0050] In the present disclosure, 'de-mixing' may mean a signal processing operation of separating a specific audio signal from an audio signal in which various audio signals are mixed.
[0051] In the present disclosure, 'up-mixing' may mean a signal processing operation of converting an audio signal so that the number of channels of an audio signal output from an electronic device increases compared to the number of channels of an audio signal received from the electronic device.
[0052] In the present disclosure, 'down-mixing' may mean a signal processing operation of converting an audio signal so that the number of channels of an audio signal output from an electronic device is reduced compared to the number of channels of an audio signal received from an electronic device.
[0053] FIG. 1 is a drawing for explaining an electronic device (100) according to one embodiment of the present disclosure.
[0054] Referring to FIG. 1, the electronic device (100) can output a channel audio signal according to the channel layout.
[0055] A channel layout can consist of multiple channels. For example, a 5.1.2 channel layout can consist of five surround channels (SL, FL, SR, FR, and C channels), one subwoofer channel (LFE channel), and two height channels (HFL, HFR channels), for a total of eight channels.
[0056] The names of the multiple channels specified by the channel layout may vary, but for the sake of convenience of explanation, they will be unified. The names of the multiple channels may be named based on the spatial positions of each channel. For example, in a 5.1.2 channel layout, the channels located to the left of the user may be named the SL (Surround Left) channel and the FL (Front Left) channel, the channels located to the right of the user may be named the SR (Surround Right) channel and the FR (Front Right) channel, and the channel located in the center of the user may be named the C (Center) channel. In addition, the channel located to the upper left of the user may be named the HFL (Height Front Left) channel, and the channel located to the upper right of the user may be named the HFR (Height Front Right) channel. Meanwhile, since the subwoofer channel transmits low-frequency signals, it may be named the LFE (Low Frequency Effect) channel regardless of its position.
[0057] In one embodiment, the electronic device (100) may include a plurality of speakers connected to a plurality of channels. For example, in a 5.1.2 channel layout, the electronic device (100) may include a total of eight speakers, including five speakers connected to five surround channels, one speaker connected to one subwoofer channel, and two speakers connected to height channels.
[0058] In one embodiment, the spatial locations of multiple speakers connected to multiple channels can be specified by the channel layout. The spatial locations of the speakers may refer to the angles at which the speakers are positioned relative to the direction in which the user (10) views the electronic device (100). For example, the locations of the speakers according to the 5.1.2 channel layout may be predefined as shown in Table 1 below.
[0059]
[0060] In one embodiment, the electronic device (100) can output a channel audio signal using metadata of the channel audio signal.
[0061] Metadata of a channel audio signal may include various information. For example, metadata of a channel audio signal may include information related to the channel through which the channel audio signal is to be output. For example, information related to the channel may include the name of each channel according to the channel layout and the position of the speaker connected to each channel. For example, in the case of a 5.1.2 channel layout, information related to the channel in the metadata of the channel audio signal may include at least one of an SL channel, an FL channel, an SR channel, an FR channel, a C channel, an HFL channel, and an HFR channel as the channel name, the position of the speaker connected to each channel, etc.
[0062] In one embodiment, the electronic device (100) can render a channel audio signal using metadata of the channel audio signal and output the rendered channel audio signal. For example, if the electronic device (100) identifies information related to a channel included in the metadata of the channel audio signal and identifies a channel of the channel audio signal as an SL channel based on the information related to the identified channel, the electronic device (100) can determine the SL channel as a channel through which the channel audio signal is to be transmitted. Then, the electronic device (100) can render the channel audio signal so that it can be output from a speaker connected to the SL channel. Then, the electronic device (100) can transmit the rendered channel audio signal through the SL channel and control the speaker connected to the SL channel to output the rendered channel audio signal.
[0063] In one embodiment, the electronic device (100) can output an object audio signal using metadata of the object audio signal.
[0064] The metadata of an object audio signal may include various information. For example, the metadata of an object audio signal may include spatial information of an object corresponding to the object audio signal.
[0065] The object corresponding to an object audio signal may refer to various objects. For example, the object corresponding to an object audio signal may refer to a moving car, a flying bird, etc. These are examples and are not limiting.
[0066] An object's spatial information can include various information. For example, the object's spatial information can include the object's location, distance, direction of movement, and speed. These are examples and are not intended to be limiting.
[0067] In one embodiment, the electronic device (100) can render the object audio signal using the metadata of the object audio signal and output the rendered object audio signal. For example, the electronic device (100) can identify the location of the object from the metadata of the object audio signal. In one embodiment, the electronic device (100) can determine a channel of a location corresponding to the location of the identified object, and render the object audio signal so that it can be output to a speaker connected to the determined channel. In one embodiment, the electronic device (100) can control the speaker to output the rendered object audio signal.
[0068] In one embodiment, when the electronic device (100) can only render a channel audio signal and cannot render an object audio signal, the electronic device (100) cannot output an audio signal including an object audio signal. Even when the electronic device (100) can only render a channel audio signal, the object audio signal can be appropriately processed so that the audio signal can be output. A method for appropriately processing an object audio signal so that the audio signal can be output even when the electronic device (100) can only render a channel audio signal will be described below by dividing it into an audio encoding method in the mobile device (200) and an audio decoding method in the electronic device (300).
[0069] An exemplary embodiment of an electronic device (200) that performs audio encoding is described below with reference to FIG. 6, and an electronic device (300) that performs audio decoding is described below with reference to FIG. 9.
[0070] FIG. 2 is a flowchart for explaining an audio encoding method in an electronic device (200) according to one embodiment of the present disclosure.
[0071] Referring to FIG. 2, in step S202, the electronic device (200) may determine at least one object audio signal among a plurality of object audio signals as a key object audio signal.
[0072] In one embodiment, the electronic device (200) may determine at least one object audio signal among the plurality of object audio signals as a key object audio signal based on the importance of each of the plurality of object audio signals.
[0073] In one embodiment, the importance of each of the plurality of object audio signals may mean a value indicating the degree of importance of the plurality of objects corresponding to each of the plurality of object audio signals.
[0074] In one embodiment, the electronic device (200) can calculate the importance of each of the plurality of objects based on spatial information of each of the plurality of objects corresponding to the plurality of object audio signals.
[0075] In one embodiment, spatial information of each of the plurality of objects corresponding to the plurality of object audio signals may mean information included in metadata of each of the plurality of object audio signals.
[0076] In one embodiment, the spatial information of each of the plurality of objects corresponding to the plurality of object audio signals may include various information. For example, referring to FIG. 3, the spatial information of each of the plurality of objects may include at least one of the distance, elevation angle, or azimuth angle of the object.
[0077] In one embodiment, the spatial information of each of the plurality of objects may include information based on time-series data. For example, the spatial information of each of the plurality of objects corresponding to the plurality of object signals may include at least one of a plurality of distances acquired over time for each of the plurality of objects, a plurality of elevation angles acquired over time, a moving standard deviation of distances calculated from a plurality of azimuth angles acquired over time, a moving standard deviation of elevation angles, and a moving standard deviation of azimuth angles.
[0078] In one embodiment, the electronic device (200) may determine the mobility of each of the plurality of objects based on spatial information of each of the plurality of objects corresponding to each of the plurality of object audio signals, and may determine the importance of each of the plurality of object audio signals based on the mobility of each of the plurality of objects.
[0079] For example, the electronic device (200) can calculate the importance of each of the plurality of objects by weighting the standard deviation of the distance movement, the standard deviation of the elevation angle movement, and the standard deviation of the azimuth angle movement included in the spatial information of each object corresponding to the plurality of object signals as in Equation 1 below.
[0080]
[0081] represents the mobility of the produced object, Each represents the distance weight of the object, the elevation weight, and the azimuth weight, Each represents the standard deviation of the object's distance translation, the standard deviation of its elevation angle translation, and the standard deviation of its azimuth angle translation, respectively.
[0082] In one embodiment, the distance weight, elevation weight, and azimuth weight of an object may be determined differently depending on the type of object. For example, for an object for which linear movement is important, the distance weight of the object may be determined to have a higher value than the elevation weight and azimuth weight.
[0083] In one embodiment, the electronic device (200) may calculate the mobility of each of the plurality of objects corresponding to each of the plurality of object signals according to Equation 1, and determine the mobility of each of the plurality of objects as the importance of each of the plurality of object signals.
[0084] In one embodiment, the electronic device (200) may determine the importance of each of the plurality of object audio signals based on the mobility of each of the plurality of objects and the volume of each of the plurality of object audio signals. For example, the electronic device (200) may determine the importance of each of the plurality of objects by calculating a weighted average of the mobility of each of the plurality of object audio signals and the volume of each of the plurality of object audio signals, as shown in Equation 2 below.
[0085]
[0086] indicates the importance of the object, represents the volume of the object audio signal, can represent a weight greater than or equal to 0 and less than or equal to 1.
[0087] In one embodiment, can be determined differently depending on how important the object's mobility or volume is. For example, for objects where mobility is important, can be determined to have a value greater than 0.5.
[0088] In one embodiment, the electronic device (200) may determine a preset number of object audio signals having high importance among the importance of each of a plurality of object audio signals as key object audio signals. For example, among the plurality of object audio signals, P object audio signals having high importance calculated according to Equation 2 may be determined as key object audio signals. For example, among 122 object audio signals, 8 object audio signals having high importance calculated according to Equation 2 may be determined as key object audio signals.
[0089] Returning to FIG. 2, in step S204, the electronic device (200) can determine an object audio signal excluding a key object audio signal among a plurality of object audio signals as a sub-object audio signal.
[0090] In one embodiment, the electronic device (200) may generate a cluster of audio signals using a determined key object audio signal and one or more sub-object audio signals. For example, a pair of a key object audio signal and a sub-object audio signal in which a difference in mobility of each of a plurality of key object audio signals and a difference in mobility of each of the sub-object audio signals is less than or equal to a preset threshold value may be determined, and a cluster of audio signals may be generated by combining the determined pair of the key object audio signal and the sub-object audio signal.
[0091] In step S206, the electronic device (200) can obtain a multi-channel audio signal using the key object audio signal and the sub-object audio signal.
[0092] In one embodiment, a multi-channel audio signal obtained by rendering a key object audio signal and a sub-object audio signal may refer to a channel-based audio signal. For example, the electronic device (200) may render a key object audio signal and a sub-object audio signal to convert each of the key object audio signal and the sub-object audio signal into channel audio signals, and mix the converted channel audio signals to obtain a multi-channel audio signal. In one embodiment, the electronic device (200) may record information related to channels, including the name of each channel according to the channel layout and the position of a speaker connected to each channel, in metadata.
[0093] In one embodiment, the electronic device (200) may obtain a bed channel audio signal, which is a channel-based audio signal. For example, the bed channel audio signal may refer to a channel-based audio signal including background sound, sound effects, etc.
[0094] In one embodiment, the electronic device (200) can obtain a multi-channel audio signal using a bed channel audio signal, a key object audio signal, and a sub-object audio signal.
[0095] In one embodiment, the electronic device (200) can render a bed channel audio signal, a key object audio signal, and a sub-object audio signal. For example, the electronic device (200) can render a key object audio signal and a sub-object audio signal and convert them into channel audio signals, respectively.
[0096] In one embodiment, the electronic device (200) can obtain a multi-channel audio signal by mixing a rendered bed channel audio signal, a rendered key object audio signal, and a rendered sub-object audio signal. For example, the electronic device (200) can obtain a multi-channel audio signal, which is a channel audio signal, by mixing a rendered bed channel audio signal, a rendered key object audio signal, and a rendered sub-object audio signal.
[0097] In step S208, the electronic device (200) may compress the key object audio signal and the multi-channel audio signal to generate a bitstream. In one embodiment, the electronic device (200) may compress the key object audio signal and the multi-channel audio signal using a codec.
[0098] In one embodiment, when an audio signal needs to be compressed significantly, the electronic device (200) may generate a bitstream by compressing a key object audio signal and a multi-channel audio signal using at least one of OPUS and AAC (Advanced Audio Coding), which are codecs that perform lossy compression. Lossy compression may refer to a signal processing operation in which a portion of data is lost during the compression process.
[0099] In one embodiment, when it is necessary to compress the audio signal without loss, the electronic device (200) can generate a bitstream by compressing the key object audio signal and the multi-channel audio signal using FLAC (Free Lossless Audio Codec), which is a codec that performs lossless compression.
[0100] In one embodiment, when a multi-channel audio signal is compressed by a codec that performs lossy compression, some audio signals may be lost due to the lossy compression. In this case, in order to reduce the loss of the audio signal, pre-processing may be performed on the key object audio signal during the step of obtaining the multi-channel audio signal.
[0101] FIG. 4 is a flowchart illustrating a method for preprocessing a key object audio signal according to one embodiment of the present disclosure.
[0102] Referring to FIG. 4, in step S402, the electronic device (200) may compress the key object audio signal. For example, the electronic device (200) may compress the key object audio signal in the same manner as the method of compressing the multi-channel audio signal in step S208. In one embodiment, when the multi-channel audio signal is compressed using an audio codec that performs lossy compression, for example, the OPUS codec, in step S208, the key object audio signal may be compressed using the OPUS codec in step S402.
[0103] In step S404, the electronic device (200) can obtain a decompressed key object audio signal from the compressed key object audio signal. For example, the electronic device (200) can decompress the compressed key object audio signal using the OPUS codec used to compress the key object audio signal.
[0104] In step S406, the electronic device (200) may render the decompressed key object audio signal and sub-object audio signal to obtain a multi-channel audio signal. In one embodiment, the decompressed key object audio signal may refer to an audio signal that has been partially lost from the key object audio signal before compression is performed, since the decompressed key object audio signal has been subjected to lossy compression.
[0105] In one embodiment, the electronic device (200) can transmit the generated bitstream to an electronic device (300) that performs audio decoding. The audio decoding method performed in the electronic device (300) will be described below with reference to FIG. 7.
[0106] FIG. 5 is a block diagram illustrating an audio encoding method according to one embodiment of the present disclosure. Below, a description of the aforementioned operations is omitted for brevity.
[0107] Referring to FIG. 5, R audio signals may include N bed channel audio signals and L object audio signals. For example, 128 audio signals may include 6 bed channel audio signals and 122 object audio signals.
[0108] In step S502, the electronic device (200) can determine P key object audio signals among L object audio signals. For example, the electronic device (200) can determine 8 key object audio signals among 122 object audio signals.
[0109] In step S502, the electronic device (200) can determine Q sub-object audio signals excluding P key object audio signals among L object audio signals. For example, the electronic device (200) can determine 114 sub-object audio signals excluding 8 key object audio signals among 122 object audio signals.
[0110] In step S506, the electronic device (200) can obtain a multi-channel audio signal including M channels by using N bed channel audio signals, Q sub-object audio signals, and P key object audio signals. For example, the electronic device (200) can render 6 bed channel audio signals, 114 sub-object audio signals, and 8 key object audio signals. Then, the electronic device (200) can obtain a multi-channel audio signal including 8 channels by mixing the rendered 6 bed channel audio signals, the rendered 114 sub-object audio signals, and the rendered 8 key object audio signals.
[0111] In step S508, the electronic device (200) can generate a bitstream by compressing a multi-channel audio signal including P key object audio signals and M channels. For example, the electronic device (200) can obtain a bitstream by compressing a multi-channel audio signal including 8 key object audio signals and 8 channels.
[0112] According to the above-described embodiment, when calculating importance to determine a key object audio signal and compressing only the determined key object audio signal, the sound quality of the output audio signal can be improved by reducing the compression error when obtaining the third multi-channel audio signal described in FIG. 7.
[0113] FIG. 6 is a block diagram illustrating an electronic device performing audio encoding according to one embodiment of the present disclosure.
[0114] Referring to FIG. 6, the electronic device (200) may include a memory (210) and a processor (220). The electronic device (200) may be implemented as a device capable of audio processing, such as a server, TV, mobile phone, tablet PC, or laptop.
[0115] Although the memory (210) and the processor (220) are illustrated separately in FIG. 6, the memory (210) and the processor (220) may be implemented through a single hardware module (e.g., chip).
[0116] The memory (210) may store one or more instructions for processing an audio signal. In one embodiment, the memory (210) may store an audio signal and metadata of the audio signal. For example, the memory (210) may store a plurality of object audio signals and metadata of each of the plurality of object audio signals. For example, the memory (210) may store a bed channel audio signal and metadata of the bed channel audio signal.
[0117] A processor (220) configured to execute one or more instructions stored in memory (210) may be configured with multiple processors. In this case, it may be implemented as a combination of dedicated processors, or it may be implemented through a combination of multiple general-purpose processors, such as a CPU or GPU, and software.
[0118] In one embodiment, the processor (220) may determine at least one object audio signal among a plurality of object audio signals as a key object audio signal, determine an object audio signal excluding the key object audio signal among the plurality of object audio signals as a sub-object audio signal, obtain a multi-channel audio signal using the key object audio signal and the sub-object audio signal, and generate a bitstream by compressing the key object audio signal and the multi-channel audio signal.
[0119] In one embodiment, the processor (220) can render a key object audio signal and a sub-object audio signal. In one embodiment, the processor (220) can compress a key object audio signal, obtain a decompressed key object audio signal from the compressed key object audio signal, and render the decompressed key object audio signal and sub-object audio signal to obtain a multi-channel audio signal.
[0120] In one embodiment, the processor (220) may compress a key object audio signal using an encoder used to compress a multi-channel audio signal.
[0121] In one embodiment, metadata of each of the plurality of object audio signals includes spatial information of each object corresponding to the plurality of object audio signals, and the spatial information may include at least one of information about a distance, information about an elevation angle, and information about an azimuth angle of each object.
[0122] In one embodiment, the processor (220) may determine the mobility of each of the plurality of objects based on spatial information of each of the plurality of objects corresponding to each of the plurality of object audio signals, determine the importance of each of the plurality of object audio signals based on the mobility of each of the plurality of objects and the volume of each of the plurality of object audio signals, and determine a preset number of object audio signals having a high importance among the respective importances as key object audio signals.
[0123] FIG. 7 is a flowchart illustrating an audio decoding method in an electronic device (300) according to one embodiment of the present disclosure.
[0124] Referring to FIG. 7, in step S702, the electronic device (300) can decompress a bitstream to obtain a key object audio signal and obtain a first multi-channel audio signal in which the key object audio signal and the sub-object audio signal are rendered.
[0125] In one embodiment, a bitstream may be data in which an audio signal is compressed, and a bitstream may mean data received from an electronic device (200) or stored in a memory (310) of an electronic device (300).
[0126] In one embodiment, the key object audio signal may be the key object audio signal determined in step S204.
[0127] In one embodiment, the first multi-channel audio signal may be a multi-channel audio signal obtained in step S206. For example, the first multi-channel audio signal may be an audio signal converted into a channel audio signal by rendering a key object audio signal and a sub-object audio signal.
[0128] In step S704, the electronic device (300) may determine whether the electronic device (300) can output a key object audio signal. For example, if the electronic device (300) cannot render object audio, the electronic device (300) may determine that the electronic device (300) cannot output a key object audio signal. In one embodiment, if the electronic device (300) can render object audio, the electronic device (300) may determine that the electronic device (300) can output a key object audio signal.
[0129] If it is determined that the electronic device (300) cannot output a key object audio signal, a second multi-channel audio signal that can be output from the electronic device (300) can be obtained based on the first multi-channel audio signal and information about the channel layout of the electronic device (300) in step S706.
[0130] In one embodiment, the electronic device (300) can obtain a second multi-channel audio signal by converting a first multi-channel audio signal to correspond to the channel layout of the electronic device (300). For example, if the first multi-channel audio signal is a channel audio signal that can be output in a 5.1.2 channel layout and the channel layout of the electronic device (300) is a 5.1 channel layout, the electronic device (300) can downmix the first multi-channel audio signal to obtain a second multi-channel audio signal that is a channel audio signal that can be output in a 5.1 channel layout.
[0131] If it is determined that the electronic device (300) can output a key object audio signal, in step S708, the electronic device (300) can obtain a third multi-channel audio signal that can be output from the electronic device (300) using the key object audio signal and the first multi-channel audio signal.
[0132] In one embodiment, the electronic device (300) may render the key object audio signal and the first multi-channel audio signal, respectively, to obtain a third multi-channel audio signal that may be output from the electronic device (300). For example, the electronic device (300) may render the key object audio signal so that the key object audio signal may be output from a speaker. In one embodiment, the electronic device (300) may render the first multi-channel audio signal so that the first multi-channel audio signal may be mixed with the rendered key object audio signal.
[0133] FIG. 8 is a flowchart illustrating a method for rendering a first multi-channel audio signal according to one embodiment of the present disclosure.
[0134] Referring to FIG. 8, in step S802, the electronic device (300) may obtain a rendered sub-object audio signal from a first multi-channel audio signal. For example, the electronic device (300) may obtain a rendered sub-object audio signal by demixing the first multi-channel audio signal. The rendered sub-object audio signal obtained by demixing the first multi-channel audio signal may be a channel audio signal. For example, the electronic device (300) may simultaneously obtain a rendered bed channel audio signal and a rendered key object audio signal by demixing the first multi-channel audio signal in addition to the rendered sub-object audio signal by demixing the first multi-channel audio signal. The rendered bed channel audio signal, the rendered key object audio signal, and the rendered sub-object audio signal obtained by demixing the first multi-channel audio signal may be channel audio signals.
[0135] In step S804, the electronic device (300) can render the key object audio signal obtained in step S702. For example, the electronic device (300) can render the key object audio signal obtained in step S702 and convert it into a channel audio signal.
[0136] In step S806, the electronic device (300) can obtain a rendered first multi-channel audio signal using the rendered sub-object audio signal and the rendered key object audio signal. For example, the electronic device (300) can obtain a rendered first multi-channel audio signal, which is a channel audio signal, by mixing the rendered sub-object audio signal and the rendered key object audio signal.
[0137] In one embodiment, the electronic device (300) can obtain a third multi-channel audio signal using the rendered key object audio signal and the rendered first multi-channel audio signal. For example, the electronic device (300) can obtain a third multi-channel audio signal by mixing the rendered key object audio signal and the rendered first multi-channel audio signal. For example, the third multi-channel audio signal may mean an audio signal in which the object audio signal is rendered in a form that can be output from a speaker.
[0138] FIG. 9 is a block diagram illustrating an electronic device that performs audio decoding according to one embodiment of the present disclosure.
[0139] Referring to FIG. 9, the electronic device (300) may include a memory (310) and a processor (320). The electronic device (300) may be implemented as a device capable of audio processing, such as a server, TV, mobile phone, tablet PC, or laptop.
[0140] Although the memory (310) and the processor (320) are illustrated separately in FIG. 9, the memory (310) and the processor (320) may be implemented through a single hardware module (e.g., chip).
[0141] The memory (310) may store one or more instructions for processing an audio signal. In one embodiment, the memory (310) may store an audio signal and metadata of the audio signal. For example, the memory (310) may store a plurality of object audio signals and metadata of each of the plurality of object audio signals. For example, the memory (210) may store a bed channel audio signal and metadata of the bed channel audio signal.
[0142] A processor (320) configured to execute one or more instructions stored in memory (310) may be configured with multiple processors. In this case, it may be implemented as a combination of dedicated processors, or it may be implemented through a combination of multiple general-purpose processors, such as a CPU or GPU, and software.
[0143] In one embodiment, the processor (320) can decompress the bitstream to obtain a key object audio signal and a first multi-channel audio signal in which the key object audio signal and the sub-object audio signal are rendered.
[0144] In one embodiment, when the processor (320) determines that the electronic device (300) cannot output a key object audio signal, the processor (320) may obtain a second multi-channel audio signal that can be output from the electronic device (300) based on the first multi-channel audio signal and information about the channel layout of the electronic device (300).
[0145] In one embodiment, when the processor (320) determines that the decoding device (300) is capable of outputting a key object audio signal, the processor (320) may obtain a third multi-channel audio signal that can be output from the electronic device (300) using the key object audio signal and the first multi-channel audio signal.
[0146] In one embodiment, the processor (320) can obtain a second multi-channel audio signal by converting a first multi-channel audio signal to correspond to a channel layout.
[0147] In one embodiment, the processor (320) can render a key object audio signal and a first multi-channel audio signal.
[0148] In one embodiment, the processor (320) can obtain a third multi-channel audio signal using the rendered key object audio signal and the rendered first multi-channel audio signal.
[0149] In one embodiment, the processor (320) can obtain a rendered sub-object audio signal from a first multi-channel audio signal.
[0150] In one embodiment, the processor (320) can render a key object audio signal.
[0151] In one embodiment, the processor (320) can obtain a rendered first multi-channel audio signal using a rendered sub-object audio signal and a rendered key object audio signal.
[0152] The audio processing method in the electronic device (200) and the audio processing method in the electronic device (300) described above may be implemented in separate devices or in a single device. Hereinafter, a method for controlling an electronic device (100) in which the audio processing method in the electronic device (200) and the audio processing method in the electronic device (300) are implemented in a single device will be described.
[0153] FIG. 10 is a flowchart illustrating an audio processing method in an electronic device (100) according to one embodiment of the present disclosure. Hereinafter, a description of the aforementioned operations is omitted for brevity.
[0154] Referring to FIG. 10, in step S1002, the electronic device (100) can determine at least one object audio signal among a plurality of object audio signals as a key object audio signal.
[0155] In one embodiment, the electronic device (100) may determine at least one object audio signal among the plurality of object audio signals as a key object audio signal based on the importance of each of the plurality of object audio signals.
[0156] In step S1004, the electronic device (100) can determine an object audio signal excluding a key object audio signal among a plurality of object audio signals as a sub-object audio signal.
[0157] In step S1006, the electronic device (100) can obtain a multi-channel audio signal using a key object audio signal and a sub-object audio signal.
[0158] In step S1008, the electronic device (100) can compress the key object audio signal and the multi-channel audio signal to generate a bitstream.
[0159] In step S1010, the electronic device (100) can decompress the bitstream to obtain a key object audio signal and a first multi-channel audio signal in which the key object audio signal and the sub-object audio signal are rendered.
[0160] In step S1012, the electronic device (100) can determine whether the electronic device (100) can output a key object audio signal.
[0161] If it is determined that the electronic device (100) cannot output a key object audio signal, in step S1014, the electronic device (100) can obtain a second multi-channel audio signal that can be output from the electronic device (100) based on the first multi-channel audio signal and information about the channel layout of the electronic device.
[0162] If it is determined that the electronic device (100) can output a key object audio signal, in step S1016, the electronic device (100) can obtain a third multi-channel audio signal that can be output from the electronic device (100) using the key object audio signal and the first multi-channel audio signal.
[0163] FIG. 11 is a schematic block diagram of an electronic device (100) that performs an audio processing method according to one embodiment of the present disclosure.
[0164] Referring to FIG. 11, the electronic device (100) may include a memory (110) and at least one processor (120). The memory (110) and at least one processor (120) included in the electronic device (100) have the same configuration as the memory (210, 310) and the processor (220, 320) described above, and thus, redundant descriptions are omitted for brevity.
[0165] In one embodiment, at least one processor (120) may determine at least one object audio signal among a plurality of object audio signals as a key object audio signal, determine an object audio signal excluding the key object audio signal among the plurality of object audio signals as a sub-object audio signal, obtain a multi-channel audio signal using the key object audio signal and the sub-object audio signal, and generate a bitstream by compressing the key object audio signal and the multi-channel audio signal.
[0166] In one embodiment, at least one processor (120) can render a key object audio signal and a sub-object audio signal to obtain a multi-channel audio signal. In one embodiment, at least one processor (120) can compress a key object audio signal, obtain a decompressed key object audio signal from the compressed key object audio signal, and render the decompressed key object audio signal and the sub-object audio signal to obtain a multi-channel audio signal.
[0167] In one embodiment, at least one processor (120) can compress a key object audio signal using an encoder used to compress a multi-channel audio signal.
[0168] In one embodiment, metadata of each of the plurality of object audio signals includes spatial information of each object corresponding to the plurality of object audio signals, and the spatial information may include at least one of information about a distance, information about an elevation angle, and information about an azimuth angle of each object.
[0169] In one embodiment, at least one processor (120) may determine the mobility of each of the plurality of objects based on spatial information of each of the plurality of objects corresponding to each of the plurality of object audio signals, determine the importance of each of the plurality of object audio signals based on the mobility of each of the plurality of objects and the volume of each of the plurality of object audio signals, and determine a preset number of object audio signals having a high importance among the respective importances as key object audio signals.
[0170] In one embodiment, at least one processor (120) decompresses a bitstream to obtain a key object audio signal and a first multi-channel audio signal in which the key object audio signal and the sub-object audio signal are rendered, and if it is determined that the electronic device (100) cannot output the key object audio signal, it obtains a second multi-channel audio signal that can be output from the electronic device (100) based on the first multi-channel audio signal and information about a channel layout of the electronic device (100), and if it is determined that the electronic device (100) can output the key object audio signal, it obtains a third multi-channel audio signal that can be output from the electronic device (100) using the key object audio signal and the first multi-channel audio signal.
[0171] In one embodiment, at least one processor (120) can convert a first multi-channel audio signal to correspond to a channel layout to obtain a second multi-channel audio signal.
[0172] In one embodiment, at least one processor (120) can render a key object audio signal and a first multi-channel audio signal, and obtain a third multi-channel audio signal using the rendered key object audio signal and the rendered first multi-channel audio signal.
[0173] In one embodiment, at least one processor (120) can obtain a rendered sub-object audio signal from a first multi-channel audio signal, render a key object audio signal, and obtain a rendered first multi-channel audio signal using the rendered sub-object audio signal and the rendered key object audio signal.
[0174] FIG. 12 is a specific block diagram of an electronic device according to one embodiment of the present disclosure.
[0175] Referring to FIG. 12, the electronic device (1000) may, for example, constitute all or part of the electronic device (200) illustrated in FIG. 6, constitute all or part of the electronic device (300) illustrated in FIG. 9, and constitute all or part of the electronic device (100) illustrated in FIG. 11. Referring to FIG. 12, the electronic device (1000) may include a memory (110), at least one processor (120), a communication interface (130), a speaker (140), a user interface (150), and a display (160). However, not all of the illustrated components are essential components, and the electronic device (1000) may be implemented by more or fewer components than the illustrated components. Since the memory (110) and at least one processor (120) are the same as in FIG. 11, a repeated description will be omitted.
[0176] In one embodiment, the memory (110) may include a built-in memory (111). The built-in memory (111) may include, for example, at least one of a volatile memory (e.g., dynamic RAM (DRAM), static RAM (SRAM), synchronous dynamic RAM (SDRAM), etc.) or a non-volatile memory (e.g., one time programmable ROM (OTPROM), programmable ROM (PROM), erasable and programmable ROM (EPROM), electrically erasable and programmable ROM (EEPROM), mask ROM, flash ROM, NAND flash memory, NOR flash memory, etc.).
[0177] In one embodiment, at least one processor (120) may include one or more processors. In this case, one or more processors may be a general-purpose processor such as a CPU, an AP, a DSP (Digital Signal Processor), a graphics-only processor such as a GPU or a VPU (Vision Processing Unit), or an AI-only processor such as an NPU. If one or more processors are AI-only processors, the AI-only processors may be designed with a hardware structure specialized for processing a specific AI model.
[0178] The communication interface (130) can perform data transmission and reception with a base station or other device capable of communication connected to the electronic device (1000) through a network. For example, the communication interface (130) may include a short-range communication module (132) and a long-range communication module (134). The short-range communication module (short-range wireless communication interface) may be used when performing data transmission and reception with other devices capable of communication. For example, the short-range communication module (132) may include, but is not limited to, a Bluetooth communication module, a BLE (Bluetooth Low Energy) communication module, a near field communication interface, a WLAN (Wi-Fi) communication module, a Zigbee communication module, an IrDA (infrared Data Association) communication module, a WFD (Wi-Fi Direct) communication module, a UWB (Ultra Wideband) communication module, an Ant+ communication module, etc. The remote communication module (134) may include the Internet, a computer network (e.g., a LAN or WAN), and a mobile communication module. The mobile communication module transmits and receives wireless signals with at least one of a base station, an external terminal, and a server on a mobile communication network. Here, the wireless signals may include various types of data according to transmission and reception of voice call signals, video call signals, or text / multimedia messages. The mobile communication module may include, but is not limited to, a 3G module, a 4G module, an LTE module, a 5G module, a 6G module, an NB-IoT module, an LTE-M module, etc.
[0179] In one embodiment, the communication interface (130) can receive multiple object audio signals.
[0180] In one embodiment, the communication interface (130) can compress an audio signal and transmit the generated bitstream to another electronic device.
[0181] In one embodiment, the communication interface (130) can receive a bitstream transmitted from another electronic device.
[0182] The speaker (140) outputs sound. The speaker (140) can output various audio signals. In one embodiment, the speaker (140) can output channel audio signals and object audio signals rendered so that they can be output from the speaker (140), respectively, or simultaneously.
[0183] The user interface (150) may include an input interface (154) for receiving user input and an output interface (152) for outputting information. The output interface (152) is for outputting a video signal or an audio signal. When the display (160) and the touchpad are configured as a touchscreen in a layered structure, the display (160) may be used as an input interface (154) in addition to the output interface (152). In one embodiment, the output interface (152) may be connected to a speaker (140) to output an audio signal.
[0184] The user interface (150) may include, for example, a high-definition multimedia interface (HDMI) or a universal serial bus (USB). Additionally or alternatively, the user interface (150) may include, for example, a mobile high-definition link (MHL) interface, a secure digital (SD) card / multi-media card (MMC) interface, or an infrared data association (IrDA) standard interface.
[0185] The display (160) may include a display panel and a controller (not shown) that controls the display panel, and the display (160) may represent a display built into the electronic device (1000). The display panel may be implemented as various types of displays, such as an LCD (Liquid Crystal Display), an OLED (Organic Light Emitting Diodes) display, an AM-OLED (Active-Matrix Organic Light-Emitting Diode), a PDP (Plasma Display Panel), etc. The display panel may be implemented to be flexible, transparent, or wearable. The display (160) may be combined with a touch panel to be provided as a touch screen. For example, the touch screen may include an integrated module in which a display panel and a touch panel are combined in a laminated structure.
[0186] A method according to an embodiment of the present disclosure may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program commands, data files, data structures, etc., alone or in combination. The program commands recorded on the medium may be those specially designed and configured for the present disclosure or may be known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; magneto-optical media such as optical media such as CD-ROMs and DVDs; and hardware devices specially configured to store and execute program commands, such as ROMs, RAMs, and flash memories. Examples of program commands include not only machine language codes generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc.
[0187] Some embodiments of the present disclosure may also be implemented in the form of a recording medium containing computer-executable instructions, such as program modules, executed by a computer. Computer-readable media may be any available media that can be accessed by a computer, and include both volatile and nonvolatile media, removable and non-removable media. Furthermore, computer-readable media may include both computer storage media and communication media. Computer storage media includes both volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. Communication media typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave, or other transport mechanism, and includes any information delivery media. Furthermore, some embodiments of the present disclosure may also be implemented as a computer program or computer program product containing computer-executable instructions, such as a computer program that is executed by a computer.
[0188] A device-readable storage medium may be provided in the form of a non-transitory storage medium. Here, the term "non-transitory storage medium" simply means a tangible device that does not contain signals (e.g., electromagnetic waves). This term does not distinguish between cases where data is permanently stored in the storage medium and cases where data is temporarily stored. For example, a "non-transitory storage medium" may include a buffer in which data is temporarily stored.
[0189] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0190] In one embodiment, the audio processing method may include the step of rendering a key object audio signal and a sub-object audio signal. In one embodiment, the audio processing method may include the step of obtaining a decompressed key object audio signal from a compressed key object audio signal, and the step of rendering the decompressed key object audio signal and the sub-object audio signal to obtain a multi-channel audio signal.
[0191] In one embodiment, a bed channel audio signal, which is a channel-based audio signal, can be obtained. In one embodiment, an audio processing method may include a step of rendering a bed channel audio signal, a key object audio signal, and a sub-object audio signal. In one embodiment, an audio processing method may include a step of mixing the rendered bed channel audio signal, the rendered key object audio signal, and the rendered sub-object audio signal to obtain a multi-channel audio signal.
[0192] In one embodiment, an audio processing method may include compressing a key object audio signal using an encoder used to compress a multi-channel audio signal.
[0193] In one embodiment, metadata of each of the plurality of object audio signals includes spatial information of each object corresponding to the plurality of object audio signals, and the spatial information may include at least one of information about a distance, information about an elevation angle, and information about an azimuth angle of each object.
[0194] In one embodiment, an audio processing method may include a step of determining mobility of each of a plurality of objects based on spatial information of each of a plurality of objects corresponding to each of a plurality of object audio signals, a step of determining importance of each of a plurality of object audio signals based on the mobility of each of the plurality of objects and the volume of each of the plurality of object audio signals, and a step of determining a preset number of object audio signals having a high importance among the respective importances as key object audio signals.
[0195] In one embodiment, an audio processing method may include the steps of decompressing a bitstream to obtain a key object audio signal and a first multi-channel audio signal in which the key object audio signal and the sub-object audio signal are rendered; if it is determined that the electronic device cannot output the key object audio signal, the step of obtaining a second multi-channel audio signal that can be output from the electronic device based on the first multi-channel audio signal and information about a channel layout of the electronic device; and if it is determined that the electronic device can output the key object audio signal, the step of obtaining a third multi-channel audio signal that can be output from the electronic device using the key object audio signal and the first multi-channel audio signal.
[0196] In one embodiment, the audio processing method may include a step of converting a first multi-channel audio signal to correspond to a channel layout to obtain a second multi-channel audio signal.
[0197] In one embodiment, an audio processing method may include the steps of rendering a key object audio signal and a first multi-channel audio signal, and obtaining a third multi-channel audio signal using the rendered key object audio signal and the rendered first multi-channel audio signal.
[0198] In one embodiment, an audio processing method may include obtaining a rendered sub-object audio signal from a first multi-channel audio signal, rendering a key object audio signal, and obtaining a rendered first multi-channel audio signal using the rendered sub-object audio signal and the rendered key object audio signal.
Claims
1. A method for processing audio in an electronic device, A step (S202) of determining at least one object audio signal among a plurality of object audio signals as a key object audio signal; A step (S204) of determining an object audio signal excluding the key object audio signal among the plurality of object audio signals as a sub-object audio signal; A step (S206) of obtaining a multi-channel audio signal using the key object audio signal and the sub-object audio signal; and An audio processing method, comprising: a step (S208) of generating a bitstream by compressing the key object audio signal and the multi-channel audio signal.
2. In the first paragraph, the method further comprises a step of obtaining a bed channel audio signal, which is a channel-based audio signal, The step of obtaining the above multi-channel audio signal is: A step of rendering the bed channel audio signal, the key object audio signal and the sub object audio signal; and An audio processing method, comprising the step of obtaining the multi-channel audio signal by mixing the rendered bed channel audio signal, the rendered key object audio signal, and the rendered sub-object audio signal.
3. In any one of paragraphs 2 to 3, The step of rendering the above key object audio signal is: A step of compressing the above key object audio signal; A step of obtaining a decompressed key object audio signal of the compressed key object audio signal; and An audio processing method, comprising: a step of rendering the decompressed key object audio signal and the sub-object audio signal to obtain the multi-channel audio signal.
4. In paragraph 3, The step of compressing the above key object audio signal is: An audio processing method, comprising the step of compressing the key object audio signal using an encoder used to compress the multi-channel audio signal.
5. In any one of paragraphs 1 to 4, An audio processing method, wherein metadata of each of the plurality of object audio signals includes spatial information of each object corresponding to the plurality of object audio signals, and the spatial information includes at least one of information regarding a distance of each object, information regarding an elevation angle, and information regarding an azimuth angle.
6. In any one of paragraphs 1 to 5, The step of determining the above key object audio signal is: A step of determining the mobility of each of the plurality of objects based on spatial information of each of the plurality of objects corresponding to each of the plurality of object audio signals; A step of calculating the importance of each of the plurality of object audio signals based on the mobility of each of the plurality of objects and the volume of each of the plurality of object audio signals; and An audio processing method, comprising the step of determining a preset number of object audio signals having a high importance among the above-mentioned respective importance levels as the key object audio signals.
7. A method for processing audio in an electronic device, A step (S802) of decompressing a bitstream to obtain a key object audio signal and a first multi-channel audio signal in which the key object audio signal and the sub-object audio signal are rendered; If it is determined that the electronic device cannot output the key object audio signal, a step (S804) of obtaining a second multi-channel audio signal that can be output from the electronic device based on the first multi-channel audio signal and information about the channel layout of the electronic device; and An audio processing method, comprising: a step of obtaining a third multi-channel audio signal capable of being output from the electronic device by using the key object audio signal and the first multi-channel audio signal, when it is determined that the electronic device can output the key object audio signal (S806).
8. In paragraph 7, The step of obtaining the second multi-channel audio signal comprises: An audio processing method comprising: a step of converting the first multi-channel audio signal to correspond to the channel layout to obtain the second multi-channel audio signal.
9. In any one of paragraphs 7 to 8, The step of obtaining the third multi-channel audio signal comprises: a step of rendering the key object audio signal and the first multi-channel audio signal; and An audio processing method comprising: a step of obtaining the third multi-channel audio signal by using the rendered key object audio signal and the rendered first multi-channel audio signal.
10. In the 9th paragraph, the step of rendering the first multi-channel audio signal, A step of obtaining a rendered sub-object audio signal from the first multi-channel audio signal; a step of rendering the above key object audio signal; and An audio processing method, comprising: a step of obtaining the rendered first multi-channel audio signal by using the rendered sub-object audio signal and the rendered key object audio signal.
11. In an electronic device (100), A memory (110) in which one or more instructions are stored, and At least one processor (120) configured to execute one or more instructions stored in the memory, The electronic device (100) executes the one or more instructions by the at least one processor (120). determining at least one object audio signal among the plurality of object audio signals as a key object audio signal; Among the above multiple object audio signals, an object audio signal excluding the key object audio signal is determined as a sub-object audio signal, A multi-channel audio signal is obtained by using the above key object audio signal and the above sub-object audio signal, An electronic device that generates a bitstream by compressing the key object audio signal and the multi-channel audio signal.
12. In paragraph 11, The electronic device (100) executes the one or more instructions by the at least one processor (120), Obtain a bed channel audio signal, which is a channel-based audio signal, Rendering the bed channel audio signal, the key object audio signal and the sub object audio signal, An electronic device that obtains the multi-channel audio signal by mixing the rendered bed channel audio signal, the rendered key object audio signal, and the rendered sub object audio signal.
13. In any one of paragraphs 11 to 12, The electronic device (100) executes the one or more instructions by the at least one processor (120), Compress the above key object audio signal, Obtaining a decompressed key object audio signal of the compressed key object audio signal, An electronic device that renders the decompressed key object audio signal and the sub-object audio signal to obtain the multi-channel audio signal.
14. In paragraph 13, The electronic device (100) executes the one or more instructions by the at least one processor (120), An electronic device that compresses the key object audio signal using an encoder used to compress the multi-channel audio signal.
15. In any one of paragraphs 11 to 14, An electronic device, wherein metadata of each of the plurality of object audio signals includes spatial information of each object corresponding to the plurality of object audio signals, and wherein the spatial information includes at least one of information regarding a distance of each object, information regarding an elevation angle, and information regarding an azimuth angle.
Citation Information
Patent Citations
Head Tracking for Parametric Binaural Output Systems and Methods
JP6964703B2
Apparatus and method for generating audio output signals using object based metadata
KR101325402B1
Efficient coding of audio scenes comprising audio objects
KR101760248B1
Method and apparatus for rendering sound signal, and computer-readable recording medium
KR102574478B1
KR20230088400A