Encoding of multichannel audio signals
By coordinating the encoding mode selection of different channels in the audio encoder, the problem of encoding distortion of stereo or multi-channel signals in the prior art is solved, and better sound quality and spatial representation effects are achieved.
Patent Information
- Application Number
- CN202110304954.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2015-05-20
- Filing Date
- 2016-05-19
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2036-05-19
AI Technical Summary
In systems where stereo or multi-channel signals are required but available codecs do not include dedicated stereo modes, it is difficult for the prior art to effectively select and coordinate encoding modes of different channels, resulting in encoding distortion and unmasked effects.
By implementing a method in an audio encoder, multiple audio signal channels are obtained and the selection of encoding modes for these channels is coordinated or synchronized, ensuring that the encoding modes selected by different channels are coordinated, reducing encoding distortion and unmasked effects.
This method can effectively reduce the encoding distortion characteristics in different channels of stereo or multi-channel signals, improve the sound quality and spatial representation of the signal, and save computational complexity.
Smart Images

Figure CN113035212B_ABST
Abstract
Description
[0001] Division Explanation
[0002] This application is a divisional application of the invention patent application with the application date of May 19, 2016, the application number of 201680029059.0, and the invention title of "Encoding of Multichannel Audio Signals". Technical Field
[0003] The subject matter of the present disclosure relates to audio coding, and more particularly, to encoding stereo or multichannel signals using instances of two or more codecs including a number of codec modes. Background Art
[0004] Cellular communication networks are evolving towards higher data rates, improved capacity, and improved coverage. In the body of the 3rd Generation Partnership Project (3GPP) standards, a number of technologies have been developed and are currently being developed.
[0005] LTE (Long Term Evolution) is an example of a standardized technology. In LTE, an OFDM (Orthogonal Frequency Division Multiplexing)-based access technology is used for the downlink, while a single-carrier FDMA (SC-FDMA)-based access technology is used for the uplink. Resource allocation to wireless terminals (also referred to as user equipment, UEs) on both the downlink and the uplink is typically performed adaptively using fast scheduling considering the instantaneous traffic pattern and radio propagation characteristics of each wireless terminal. One type of data on LTE is, for example, audio data for voice conversations or streaming audio.
[0006] To improve the performance of low bitrate speech and audio coding, it is well-known to utilize prior knowledge about signal characteristics and employ signal modeling. In the case of more complex signals, several coding models or coding modes can be used for different signal types and different parts of the signal. Selecting an appropriate coding mode at any time is beneficial.
[0007] In systems where stereo or multichannel signals are to be sent and the available or preferred codec does not include a dedicated stereo mode, each channel of the signal can be encoded and sent with a currently available separate codec instance. This means that, for example, in the case of stereo with two channels, the codec runs once for the left channel and once for the right channel. Separate instances mean there is no coupling of left and right channel encoding. Encoding with "different instances" can be parallel, for example, preferably simultaneously, but can also be sequential. For the stereo case, the left / right representation and the mid / side representation can be considered as two channels of a stereo signal. Similarly, for the multichannel case, for different ways of encoding, the channels can be represented as they are presented or captured. When time-aligning the decoded signals at the earpiece, these signals can be used to render or reconstruct a stereo or multichannel signal. For the stereo case, this is typically referred to as dual-mono encoding.
[0008] In a typical scenario, each microphone can represent a channel that is encoded and played back by one speaker after decoding. However, virtual input channels can also be generated based on different combinations of microphone signals. For example, in the case of stereo, the mid / side representation is typically chosen instead of the left / right representation. In the simplest case, the mid signal is generated by adding the left and right channel signals, while the side signal is obtained by taking the difference. Conversely, at the decoder, a similar remapping can be done again, for example, from the mid / side representation to the left / right representation. The left signal (e.g., apart from a constant scale factor) can be obtained by adding the mid and side signals, and the right signal can be obtained by subtracting these signals. In general, there are corresponding mappings from N microphone signals to M virtual input channels that are encoded and from the M virtual output channels received by the decoder to K speakers. These mappings can be obtained by linear combinations of the individual input signals of the mapping, which can be expressed by mathematically multiplying the input signals with a mapping matrix.
[0009] Many newly developed codecs include multiple different coding modes and can select a coding mode based on the characteristics of the signal to be encoded / decoded. To select the best coding / decoding mode, the encoder and / or decoder can try all available modes in analysis-by-synthesis (also known as the closed-loop approach), or it can rely on a signal classifier that makes a decision on the coding mode based on signal analysis (also known as the open-loop decision). An example of a codec that includes different selectable coding modes can be a codec that includes both an ACELP (speech) coding strategy or mode and an MDCT (music) coding strategy or mode. Other important examples of the main coding modes are active signal coding and discontinuous transmission (DTX) schemes with comfort noise generation. For this case, a voice activity detector or a signal activity detector is typically used to select one of these coding modes. Other coding modes can be selected in response to the detected audio bandwidth. For example, if the input audio bandwidth is only narrowband (no signal energy above 4 kHz), a narrowband coding mode can be selected as compared to when the signal is broadband (signal energy up to 8 kHz), ultra-wideband (signal energy up to 16 kHz), or full-band (energy over the full audible spectrum). Another example of different coding modes is related to the bitrate used for encoding. A bitrate selector can select different bitrates for encoding based on the requirements of the audio input signal or the transmission network.
[0010] Generally, the main coding strategy accordingly includes multiple sub-strategies selected based on a signal classifier. Examples of such sub-strategies can be (when the main strategies are MDCT coding and ACELP coding) MDCT coding of noise-like signals and MDCT coding of harmonic signals, and / or different ACELP excitation representations.
[0011] Regarding audio signal classification, typical signal classes for speech signals are voiced and unvoiced speech utterances. For general audio signals, a distinction is typically made between speech, music, and potential background noise signals. SUMMARY OF THE INVENTION
[0012] According to a first aspect, there is provided a method for assisting in the selection of a coding mode for encoding a multi-channel audio signal, where different coding modes can be selected for different channels. The method is performed in an audio encoder and includes obtaining a plurality of audio signal channels and coordinating or synchronizing the selection of the coding modes for the obtained plurality of channels, where the coordination is based on the coding mode selected for one of the obtained channels or a group of channels among the obtained channels.
[0013] According to a second aspect, there is provided a device for assisting in the selection of an encoding mode for a multi-channel audio signal. The device includes a processor and a memory for storing instructions which, when executed by the processor, cause the device to: obtain a plurality of audio signal channels and coordinate or synchronize the selection of encoding modes for the obtained plurality of channels, wherein the coordination is based on an encoding mode selected for one of the obtained channels or a group of channels among the obtained channels.
[0014] According to a third aspect, there is provided a computer program for assisting in the selection of an audio encoding mode. The computer program includes computer program code which, when run on a device, causes the device to: obtain a plurality of audio signal channels and coordinate or synchronize the selection of encoding modes for the obtained plurality of channels, wherein the coordination is based on an encoding mode selected for one of the obtained channels or a group of channels among the obtained channels. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The drawings illustrate selected embodiments of the disclosed subject matter. In the drawings, like reference numerals represent like features.
[0016] Figure 1 is a schematic diagram of a cellular network in which embodiments proposed herein can be applied.
[0017] Figure 2 is a diagram showing a prior art solution with independent codecs for each channel and no mode synchronization.
[0018] Figure 3 is a diagram showing an exemplary mode determination structure within an example of an encoder according to the prior art.
[0019] Figure 4 shows a solution using an external mode determination unit that controls all encoder instances according to an embodiment.
[0020] Figure 5 shows an embodiment in which one codec is selected as the main encoder, i.e., the mode determination of this codec is imposed on all other encoders.
[0021] Figure 6 and Figure 7 is a flowchart of a method according to an embodiment.
[0022] Figures 8a - 8c is a schematic block diagram showing different implementations of an encoder according to an example embodiment.
[0023] Figure 9 is a schematic diagram showing some components of a wireless terminal.
[0024] Figure 10It is a schematic diagram showing some components of the transcoding node. Detailed implementation
[0025] The disclosed subject matter is described below with reference to various embodiments. These embodiments are presented as illustrative examples and are not to be construed as limiting the disclosed subject matter.
[0026] When using a codec with multiple coding strategies or modes having separate coding strategies or modes on two channels of a stereo signal or different channels of a multi-channel signal, different codec modes can be selected for different channels. This is because the mode determination of different instances of the codec is independent. An example scenario where different coding modes can be selected for different channels of a signal is, for example, a stereo signal captured by an AB microphone, where one channel is dominated by speech and the other channel is dominated by background music. In such a case, a codec including, for example, ACELP and MDCT coding modes may select the ACELP mode for the speech-dominated channel and the MDCT mode for the music-dominated channel. The characteristics or properties of the coding distortion generated by these two coding strategies may be quite different. For example, in one case, the characteristic of the coding distortion may be noise, while another characteristic caused by a different coding mode may be the pre-echo distortion sometimes observed in the MDCT coding mode. The presented signals with such different distortion characteristics can lead to an unmasking effect, that is, a distortion that is reasonably well masked when only one signal is provided to the listener becomes apparent or annoying when two signals with different distortion characteristics are provided to the listener simultaneously (for example, provided to the left and right ears respectively).
[0027] According to an embodiment of the proposed solution, the mode decision of different instances of a codec for encoding stereo or multichannel signals is coordinated. Coordination generally means that the mode decisions are synchronized, but it may also mean that these modes (although different) are selected such that the encoding distortion and the unmasking effect are minimized. To encode the different channels of a multichannel signal, the selection of the codec mode and possibly the codec sub - modes in different instances of the codec can be synchronized such that the same codec mode is selected for all channels, or at least such that for all channels of the multichannel signal, the codec instances select related codec modes with similar distortion characteristics. By synchronizing or coordinating the selection of the codec modes for the different channels of a multichannel signal, the characteristics of the encoding artifacts are similar for all channels. Thus, when the multichannel signal is reconstructed and played, there will be no unmasking effect or at least the unmasking is reduced. An embodiment of the solution may include a decision algorithm that determines or measures whether synchronization of the mode decision is necessary. For example, the algorithm can predict whether an unmasking effect as described above can or will occur in the different channels of the current multichannel signal. In the case of applying such an algorithm, the synchronization or coordination of the mode decision in different instances of the codec can be selectively activated, for example, only when the decision algorithm determines or indicates that this is necessary and / or beneficial.
[0028] By applying the embodiments related to synchronized or coordinated mode decision described herein, the deviation of the encoding distortion characteristics in different channels of a stereo or multichannel signal can be avoided or at least mitigated. This will advantageously improve the sound quality and the spatial representation of the signal. Additionally, embodiments of the solution are able to save computational complexity, for example, when only one mode decision needs to be made for all instances of the codec.
[0029] Figure 1 An exemplary network context is shown in Figure 1FIG. is a schematic view of a wireless network 8 to which embodiments presented herein can be applied. The wireless network 8 includes a core network 3 and one or more radio access nodes 1, where the radio access nodes 1 are in the form of evolved Node Bs (also referred to as eNodeBs or eNBs). The radio base stations 1 can also be in the form of Node Bs, BTSs (Base Transceiver Stations), and / or BSSs (Base Station Subsystems), etc. The radio base stations 1 provide radio connections to a plurality of wireless devices 2. The term wireless device is also referred to as a wireless communication device or a radio communication device, such as a UE, which is also referred to as, for example, a mobile terminal, a wireless terminal, a mobile station, a mobile phone, a cellular phone, a smart phone, and / or a target device. Other examples of different wireless devices include a laptop computer with wireless capabilities, a laptop embedded device (LEE), a laptop mounted device (LME), a USB dongle, a customer premise equipment (CPE), a modem, a personal digital assistant (PDA), a tablet computer (sometimes referred to as a surfboard with wireless capabilities, or simply a tablet), a device or UE with machine-to-machine (M2M) capabilities, a device-to-device (D2D) UE or wireless device, a file storage device equipped with a wireless interface (such as a printer or a file storage device), a machine type communication (MTC) device such as a sensor (e.g., a sensor equipped with a UE), just to mention a few examples.
[0030] As long as the principles described hereinafter apply, the wireless network 8 can conform to, for example, any one or a combination of LTE (Long Term Evolution), W-CDMA (Wideband Code Division Multiple Access), EDGE (Enhanced Data Rates for GSM (Global System for Mobile Communications) Evolution), GPRS (General Packet Radio Service), CDMA2000 (Code Division Multiple Access 2000), or any other current or future wireless network (such as Advanced LTE).
[0031] On the radio interface, uplink (UL) 4a communication from the wireless terminal 2 and downlink (DL) 4b communication to the wireless terminal 2 between the wireless terminal 2 and the radio base station 1 are performed. Due to the effects of fading, multipath propagation, interference, etc., the wireless radio interface quality for each wireless terminal 2 may vary over time and according to the location of the wireless terminal 2.
[0032] The radio base station 1 is also connected to the core network 3 to connect to central functions and external networks 7, such as the public switched telephone network (PSTN) and / or the Internet.
[0033] Audio data such as multi-channel signals can be encoded and decoded, for example, by a wireless terminal 2 and a transcoding node 5, where the transcoding node 5 is a network node arranged to perform transcoding of audio. The transcoding node 5 can be implemented, for example, in an MGW (Media Gateway), an SBG (Session Border Gateway) / BGF (Border Gateway Function), or an MRFP (Media Resource Function Processor). Thus, both the wireless terminal 2 and the transcoding node 5 are host devices including corresponding audio encoders and decoders. Obviously, the solution disclosed herein can be applied to any device or node that wishes to encode multi-channel audio signals.
[0034] The solution described herein relates at least to such a system where a multi-channel or stereo signal is encoded with one instance of the same codec for each channel, and each instance selects from a plurality of different operating modes related to MDCT and ACELP encoding. Figure 2 and Figure 3 An example of such a system in which embodiments applying the solution will benefit is depicted. Figure 2 The situation of the prior art is described, where each input audio channel is separately encoded by one instance of a codec. Figure 3 An example of a codec instance having multiple selectable coding modes (including a main mode and a sub-mode) is shown. Different modes can be selected according to signal characteristics, and different mode decision algorithms can be assumed herein to select the correct mode.
[0035] Figure 4 and Figure 5 Embodiments of the proposed solution are depicted. In Figure 4 , an external (i.e., outside the instance) mode decision algorithm controls the mode selection of all codec instances. In another embodiment or scenario, the external mode decision algorithm can detect or identify a group of channels that should be synchronized / coordinated. An example that may be meaningful is when there is a group of channels dominated by different source signals. Only a subset of the mode decisions can also be performed in the external mode decision unit, and some sub-modes can be determined only locally. For example, in a codec or device including multiple entities similar to the entity shown in Figure 3 , the main mode decision can be synchronized / coordinated, while the sub-mode decision can be performed locally. In Figure 5 , the mode decision algorithm (internal) from one of the codec instances is used to control all the codec instances, and the external unit selects the main codec instance, i.e., the codec instance whose mode decision should be imposed on other codec instances.
[0036] Figures 3 to 5The input of the determination module is all the channel signals or a subset thereof. The determination can involve identifying one or several main channels, for example, based on signal energy or other more complex criteria, such as the perceptual complexity or perceptual entropy of the signal which can be a metric for the coding requirements. The determination can also be based on certain combinations of the input channel signals. One possibility is that some channels are used to compensate for signal components in other channels (such as compensating for background noise), and these compensated channels will be used for the determination.
[0037] Referring to the embodiment where the main determination according to Figure 4 is external to the codec instance, even in the case of using only a single instance of the codec, it is important to include this embodiment as a specific embodiment, which allows encoding of only a single channel (mono) signal. In this specific embodiment, a separate stereo or multi-channel codec instance can be used to generate and transmit supplementary stereo or multi-channel coding information. For example, this can be the case when the stereo or multi-channel coding is parametric. In this embodiment, it is important that the mode determination of a single mono codec can be replaced / controlled by an external mode determination module.
[0038] According to at least some embodiments of the solution, in the case of encoding a stereo or other multi-channel signal using multiple instances of the same codec (e.g., in parallel), the codec or encoder mode determination of one encoder instance is applied to or imposed on other encoder instances.
[0039] Figures 6 - 7 are other embodiments
[0040] Next, embodiments related to a method for encoding a multi-channel audio signal (such as a stereo signal) will be described with reference to Figure 6 The method will be performed by a codec or encoder, for example, which includes multiple instances and in each instance includes multiple different selectable coding modes (such as ACELP and MDCT coding). Alternatively, it can be a codec device including multiple codecs or encoders, each of which includes multiple selectable coding modes. The encoder or codec can be configured to comply with one or more standards for audio coding. Figure 6The method shown includes obtaining multiple channels of an audio signal 601. The obtaining may include, for example, receiving audio signal channels from a microphone or some other entity, or retrieving them from a memory. The audio signal may be a stereo signal or include more than two channels. A multi-channel audio signal herein generally refers to an audio signal that includes more than one channel, i.e., at least two channels. The different channels obtained are provided to respective individual instances of an encoder (or separate encoders, depending on the terminology and / or implementation). The method further includes selecting 602 an encoding mode based on one or more channels, where the selected encoding mode will be used to encode at least a plurality of the channels obtained, i.e., not only for the one channel on which the encoding mode is based. The method further includes applying 603 the selected encoding mode to a plurality of the channels obtained (e.g., all channels or a subset). Alternatively, this may be described and / or implemented as: the method includes applying a selected encoding mode for one of the plurality of channels when encoding a plurality of the channels obtained. Alternatively, it may be described as controlling the encoding mode selection of multiple encoder instances based on an encoding mode selected by one of the encoder instances for one of the channels obtained. Alternatively, an embodiment may be described as encoding a plurality of channels in a multi-channel audio signal based on an encoding mode selection made according to (or for) one of the channels.
[0041] Reference will now be made to Figure 7 describe more detailed method embodiments. Figure 7The method shown includes obtaining multiple channels of an audio signal. As previously mentioned, the channels are provided to respective encoder instances for encoding. The method also includes determining 702 whether there is a risk of an unmasking effect or other unwanted effects for the obtained multiple channels, as previously mentioned, because different encoding modes are selected for different channels. Alternatively, action 702 can be described as determining whether it is necessary to coordinate the encoding mode selection for multiple instances encoding the multiple channels. This determination may involve, for example, determining whether different channels belong to different audio signal types (such as music or speech) or are dominated by them, where different types typically result in different encoding mode selections. If there is no risk or possibility of unwanted effects or artifacts (such as due to having different encoding mode selections), then there is no need to coordinate the encoding mode selection for different entities, and the encoding process can proceed according to a conventional process. However, if it is determined, for example, in action 702 that it is necessary to coordinate the encoding mode selection for different audio signal channels, such coordination should be carried out. The method may also include an optional action of determining 703 which channels actually need to be coordinated according to the encoding mode. This action may involve classifying the channels into different groups based on whether the channels belong to different audio signal types (such as music or speech) or are dominated by them. Then, the encoding mode selection made for encoding the channels classified into the first group can be controlled or coordinated 704 such that the encoding mode selected for the channels in the second group is also used for the first group. There may be more than two groups of signals. Then, the coordinated encoding mode selected for one channel or a group of channels can be used to encode the audio signal channels 705.
[0042] Exemplary embodiments
[0043] The above methods and techniques can be implemented in an encoder and / or decoder, which can be part of, for example, a communication device or other host device.
[0044] Encoder or codec, Figures 8a - 8c
[0045] In Figure 8a an encoder is shown in a general manner. The encoder is configured to encode an audio signal and supports encoding of multiple signals (such as multiple channels of a multi-channel audio signal) (e.g., parallel encoding by multiple instances of the encoder). The encoder may also include multiple different optional encoding modes, such as, for example, ACELP and MDCT encoding and their sub-modes as previously mentioned. The encoder may also be configured to encode other types of signals. Encoder 800 is configured to perform with reference to, for example, Figures 4 - 7At least one of the method embodiments described in any of them. The encoder 800 is associated with the same technical features, purposes, and advantages as the foregoing method embodiments. The decoder can be configured to conform to one or more standards for audio encoding / decoding. To avoid unnecessary repetition, the encoder will be briefly described.
[0046] The encoder can be implemented and / or described as follows:
[0047] The encoder 800 is configured to encode an audio signal including a plurality of channels. The encoder 800 includes a processing circuit or processing component 801 and a communication interface 802. The processing circuit 801 can be configured to, for example, enable the encoder 800 to obtain a plurality of channels of the audio signal and further coordinate or synchronize the selection of the encoding mode. The processing circuit 801 can also be configured to enable the encoder to apply the coordinated encoding mode to all or at least a plurality of the obtained plurality of channels. The communication interface 802, which can also be represented as, for example, an input / output (I / O) interface, includes an interface for sending data to and receiving data from other entities or modules.
[0048] As Figure 8b shown, the processing circuit 801 can include one or more processing components (such as a processor 803 (such as a CPU)) and a memory 804 for storing or holding instructions. Then, the memory will include instructions in the form of, for example, a computer program 805, which when executed by the processor 803, causes the encoder 800 to perform the above actions.
[0049] In Figure 8c an alternative embodiment of the processing circuit 801 is shown. The processing circuit here can include an obtaining unit 806, which is configured to enable the encoder 800 to obtain a plurality of channels of the audio signal. The processing circuit further includes a selection unit 807, which is configured to enable the encoder to select an encoding mode from a plurality of encoding modes based on one of the audio signal channels. The processing circuit can also include an application unit or control unit 808, which is configured to enable the encoder to apply the selected encoding mode to at least a plurality of channels. The processing circuit 801 can include more units, such as a determination unit 809, which is configured to enable the encoder to determine whether it is necessary to coordinate the selection of the encoding mode for the considered audio signal channels. The processing circuit also includes an encoding unit 810, which is configured to enable the encoder to actually encode the channels using the coordinated encoding mode. These latter units are shown in Figure 8c with a dashed outline to emphasize that they are optional compared to other units. These units can be combined according to needs or preferences to achieve a sufficient implementation.
[0050] The above encoder or codec can be configured for different method embodiments described herein.
[0051] It can be considered that the encoder 800 also includes other functions for performing conventional encoder functions when needed.
[0052] Figure 9 is a schematic diagram showing Figure 1 some components of the wireless terminal 2. The processor 70 is provided using any combination of one or more of a suitable central processing unit (CPU), multiprocessor, microcontroller, digital signal processor (DSP), application specific integrated circuit, etc. that can execute software instructions 76 stored in the memory 74 (and thus can be a computer program product). The processor 70 can execute the software instructions 76 to perform one or more embodiments of the method described above with reference to Figures 4 - 7 as described.
[0053] The memory 74 can be any combination of read - write memory (RAM) and read - only memory (ROM). The memory 74 can also include a persistent storage device, which can be, for example, any single one or combination of magnetic memory, optical memory, solid - state memory, or even remotely - mounted memory.
[0054] A data memory 72 is also provided for reading and / or storing data during the execution of software instructions in the processor 70. The data memory 72 can be any combination of read - write memory (RAM) and read - only memory (ROM).
[0055] The wireless terminal 2 also includes an I / O interface 72 for communicating with other external entities. The I / O interface 73 also includes a user interface, including a microphone, speaker, display, etc. Optionally, an external microphone and / or speaker / headphones can be connected to the wireless terminal.
[0056] The wireless terminal 2 also includes one or more transceivers 71, including analog and digital components and a suitable number of antennas 75, for wireless communication with Figure 1 the wireless terminal shown in
[0057] The wireless terminal 2 includes an audio encoder and an audio decoder. These can be implemented with software instructions 76, which can be executed by the processor 70 or using separate hardware (not shown).
[0058] To highlight the concepts presented herein, other components of the wireless terminal 2 are omitted.
[0059] Figure 10 is a schematic diagram showing Figure 1Schematic diagram of some components of transcoding node 5. A processor 80 is provided using any combination of one or more of a suitable central processing unit (CPU), multiprocessor, microcontroller, digital signal processor (DSP), application specific integrated circuit, etc. that can execute software instructions 86 stored in a memory 84 (and thus can be a computer program product). The processor 80 can be configured to execute the software instructions 86 to perform one or more embodiments of the method described above with reference to Figures 4 - 7 as described.
[0060] The memory 84 can be any combination of read-write memory (RAM) and read-only memory (ROM). The memory 84 can also include a persistent storage device, which can be, for example, any single one or combination of magnetic memory, optical memory, solid-state memory, or even remotely mounted memory.
[0061] A data memory 82 is also provided for reading and / or storing data during the execution of software instructions in the processor 80. The data memory 82 can be any combination of read-write memory (RAM) and read-only memory (ROM).
[0062] The transcoding node 5 further includes an I / O interface 83 for communicating with other external entities (such as Figure 1 wireless terminals) via a radio base station 1.
[0063] The transcoding node 5 includes an audio encoder and an audio decoder. These can be implemented with software instructions 86, and the software instructions 56 can be executed by the processor 80 or using separate hardware (not shown).
[0064] To highlight the concepts presented herein, other components of the transcoding node 5 are omitted.
[0065] The solution described herein also relates to a computer program product including a computer-readable medium. A computer program can be stored in the computer-readable medium, and the computer program can cause a processor to execute the method according to the embodiments described herein. The computer program product can be an optical disc such as a CD (compact disc), DVD (digital versatile disc), or Blu-ray disc. As described above, the computer program product can also be embodied in the memory of a device, such as Figure 8b the computer program product 804. The computer program can be stored in any manner suitable for the computer program product. The computer program product can be a removable solid-state memory, for example, a universal serial bus (USB) stick.
[0066] The solution described herein also relates to a carrier containing a computer program, which when executed on at least one processor causes the at least one processor to execute according to the embodiments described herein, for example. The carrier can be one of an electrical signal, an optical signal, a radio signal, or a computer-readable storage medium.
[0067] The following are certain exemplary embodiments that further illustrate various aspects of the subject matter of the present disclosure.
[0068] 1. A method for assisting in the selection of an audio coding mode, the method being executed in an audio encoder and comprising: obtaining a plurality of audio signal channels; and coordinating or synchronizing the selection of coding modes for a plurality of the obtained channels, wherein the coordination can be based on the coding mode selected for one of the obtained channels or a group of the obtained channels.
[0069] 2. The method according to embodiment 1, further comprising applying the coding mode selected for one of the obtained channels to encode the obtained plurality of channels.
[0070] 3. The method according to embodiment 1 or 2, further comprising determining whether coordination of the selection of coding modes is required and performing the coordination when required.
[0071] 4. The method according to any one of the preceding embodiments, further comprising determining which channels require coordination.
[0072] 5. The method according to any one of the preceding embodiments, further comprising encoding the audio signal channels according to the coordinated coding mode selection.
[0073] 6. A host device (2, 5) and / or encoder for assisting in the selection of an audio coding mode, the host device and / or encoder comprising: a processor (70, 80); and a memory (74, 84) storing instructions (76, 86), which when executed by the processor cause the host device (2, 5) and / or encoder to: obtain audio signal channels; and coordinate the selection of coding modes for the channels.
[0074] 7. The host device (2, 5) and / or encoder according to embodiment 6, further comprising instructions that, when executed by the processor, cause the host device (2, 5) and / or encoder to apply the coding mode selected for one of the obtained channels to encode the obtained plurality of channels.
[0075] 8. The host device (2, 5) and / or encoder according to embodiment 6 further comprises the following instructions: when executed by a processor, the instructions cause the host device (2, 5) and / or encoder to determine whether coordination of the selection of an encoding mode is required, and to perform coordination when required.
[0076] 9. The host device (2, 5) and / or encoder according to any one of embodiments 6 to 8, wherein the instructions for classifying an audio signal comprise the following instructions: when executed by the processor, the instructions cause the host device (2, 5) and / or encoder to determine which of the obtained sound channels require coordination.
[0077] 10. A computer program (66, 91) for assisting in the selection of an encoding mode for audio, the computer program comprising computer program code which, when running on a host device (2, 5) and / or encoder, causes the host device (2, 5) and / or encoder to: obtain audio signal channels; and coordinate the selection of an encoding mode for the channels.
[0078] 11. A computer program product comprising the computer program according to embodiment 10 and a computer-readable medium storing the computer program.
[0079] The steps, functions, processes, modules, units and / or blocks described herein can be implemented in hardware using any conventional technology, such as using discrete circuits or integrated circuit technology, including both general-purpose electronic circuits and application-specific circuits.
[0080] Specific examples include one or more suitably configured digital signal processors and other known electronic circuits, such as interconnected discrete logic gates for performing specialized functions, or application-specific integrated circuits (ASICs).
[0081] Alternatively, at least some of the above steps, functions, processes, modules, units and / or blocks can be implemented in software, such as a computer program executed by a suitable processing circuit comprising one or more processing units. Before and / or during the use of the computer program in a network node, the software can be carried by a carrier such as an electronic signal, an optical signal, a radio signal or a computer-readable storage medium. The above network node and index server can be implemented in a so-called cloud solution, which means that this implementation can be distributed, and thus the network node and index server can be so-called virtual nodes or virtual machines.
[0082] When executed by one or more processors, the flowcharts(s) presented herein may be considered computer flowcharts(s). The corresponding apparatus may be defined as a set of functional modules, where each step executed by the processor corresponds to a functional module. In this case, the functional modules are implemented as computer programs running on the processor.
[0083] Examples of processing circuitry include, but are not limited to: one or more microprocessors, one or more digital signal processors (DSPs), one or more central processing units (CPUs), and / or any suitable programmable logic circuitry, such as one or more field programmable gate arrays (FPGAs) or one or more programmable logic controllers (PLCs). That is, the units or modules in the arrangements at the different nodes described above may be implemented as a combination of analog or digital circuitry and / or one or more processors configured by software and / or firmware stored in a memory. One or more of these processors and other digital hardware may be included in a single application specific integrated circuit (ASIC), or several processors and various digital hardware may be distributed across several separate components, whether individually packaged or assembled as a system on a chip (SoC).
[0084] It should also be understood that the general processing capabilities of any conventional device or unit implementing the proposed technology can be reused. Existing software can also be reused, for example, by reprogramming existing software or adding new software components.
[0085] Merely by way of example, the above embodiments are presented, and it should be understood that the proposed technology is not limited thereto. Those skilled in the art will understand that various modifications, combinations, and changes can be made to this embodiment without departing from the scope of the present invention. Specifically, when technically feasible, different partial solutions in different embodiments can be combined in other configurations.
[0086] In some alternative embodiments, the functions / actions recorded in the boxes may occur in an order other than that shown in the flowchart. For example, depending on the functions / actions involved, two consecutive boxes shown may actually be executed substantially simultaneously, or the boxes may sometimes be executed in the reverse order. Additionally, the functions of a given module in the flowchart and / or block diagram may be separated into multiple boxes, and / or, the functions of two or more boxes in the flowchart and / or block diagram may be at least partially integrated. Finally, other boxes may be added / inserted between the boxes shown, and / or boxes / operations may be omitted, without departing from the scope of the subject matter of the present disclosure.
[0087] It should be understood that the selection of the interaction unit and the naming of the unit within the present disclosure are for illustrative purposes only, and the nodes suitable for performing any of the above methods can be configured in multiple alternative ways so as to be able to perform the proposed processing actions.
[0088] It should also be noted that the units described in the present disclosure should be considered as logical entities and do not necessarily have to be separate physical entities.
[0089] Although the subject matter of the present disclosure has been presented above with reference to various embodiments, it will be understood that various changes in form and detail may be made to the described embodiments without departing from the overall scope of the subject matter of the present disclosure.
Claims
1. A coding method for multi-channel audio signal coding, the coding method being executed by an audio encoder and comprising: Obtain (601) a plurality of audio signal channels; And Coordinate the selection (602) of coding modes for the obtained plurality of channels, wherein the coordination is based on a coding mode selected for one of the obtained channels, and wherein the coordination is selectively activated.
2. The coding method according to claim 1, further comprising applying (603) a coding mode selected for one of the obtained channels to coding the obtained multiple channels.
3. The coding method according to claim 1 or 2, further comprising determining which channels need to be coordinated.
4. The coding method according to claim 1 or 2, further comprising selecting a main codec instance, wherein the main codec instance imposes its mode determination on other codec instances.
5. The coding method according to claim 1 or 2, further comprising selecting to code the audio signal channels according to the coordinated coding mode.
6. An audio encoder (800) for coding a multi-channel audio signal, the audio encoder comprising: A processor (70, 80); And A memory (74, 84) storing instructions (76, 86), the instructions when executed by the processor cause the audio encoder to: Obtain a plurality of audio signal channels; And Coordinate the selection of coding modes for the obtained plurality of channels, wherein the coordination is based on a coding mode selected for one of the obtained channels, and wherein the coordination is selectively activated.
7. The audio encoder according to claim 6, further comprising instructions that, when executed by the processor, cause the audio encoder to apply a coding mode selected for one of the obtained channels to coding the obtained multiple channels.
8. The audio encoder according to claim 6 or 7, further comprising instructions that, when executed by the processor, cause the audio encoder to determine which of the obtained audio channels need to be coordinated.
9. The audio encoder according to claim 6 or 7, further comprising instructions that, when executed by the processor, cause the audio encoder to select to code the audio signal channels according to the coordinated coding mode.
10. The audio encoder according to claim 6 or 7, wherein, The audio encoder is included in a host device (2, 5).
11. The audio encoder according to claim 6 or 7, wherein, The audio encoder is included in a network node (5).
12. A computer-readable storage medium stores a computer program (805) for encoding a multi-channel audio signal. The computer program includes computer program code that, when executed by a processor (803) of an audio encoder (800), causes the audio encoder to: obtain a plurality of audio signal channels; and coordinate the selection of an encoding mode for the obtained plurality of channels, wherein the coordination is based on an encoding mode selected for one of the obtained channels, and wherein the coordination is selectively activated.
13. A computer program product (74) includes a computer program (805) for encoding a multi-channel audio signal. The computer program includes computer program code that, when executed by a processor (803) of an audio encoder (800), causes the audio encoder to: obtain a plurality of audio signal channels; and coordinate the selection of an encoding mode for the obtained plurality of channels, wherein the coordination is based on an encoding mode selected for one of the obtained channels, and wherein the coordination is selectively activated.
Citation Information
Patent Citations
Method, system and control server for processing audio
CN101471804A
Digital audio multi-channel coding method and system of DRA (Digital Recorder Analyzer) with low bit rate
CN101814289A