Receiver for decoding audio signal comprising switching device and corresponding method

By using a learnable model and switching device on the decoder side, the problem of selective decoding of mixed audio signals is solved, enabling flexible audio signal output and reducing the need for additional bitstreams, thereby improving the flexibility and efficiency of audio signal processing.

CN122070580APending Publication Date: 2026-05-19FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
Filing Date
2024-08-29
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing technologies cannot efficiently and selectively decode a specific audio signal or a mixture of audio signals in a mixed audio signal at the decoder side, and require additional bitstream transmission to achieve signal enhancement.

Method used

The decoder uses a learnable model and switching device to switch different decoding modes by controlling signals, enabling flexible decoding of mixed audio signals, selective output of single or multiple audio signals, and no additional bitstream transmission is required.

Benefits of technology

It enables flexible selective output on the decoder side, allowing users to choose between clean or noisy frequency signals as needed, improving the flexibility and efficiency of audio signal processing and reducing the need for additional bitstreams.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122070580A_ABST
    Figure CN122070580A_ABST
Patent Text Reader

Abstract

A receiver for decoding an encoded audio signal, the receiver comprising: a decoder configured to receive an encoded audio signal, the encoded audio signal comprising a mixture of a plurality of audio signals, and to decode the encoded audio signal by applying a learnable model to obtain an audio output signal, the learnable model is configured to decode the encoded audio signal such that the audio output signal comprises only one of the plurality of audio signals and such that the audio output signal comprises some or all of the plurality of audio signals; wherein the learnable model is configured to decode the audio signal in response to a control signal as a component of the decoder, such that the audio output signal comprises only one of the plurality of audio signals, or such that the audio output signal comprises some or all of the plurality of audio signals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to a receiver and a corresponding method for decoding audio signals. Preferred embodiments relate to a receiver for decoding audio signals having a "switch" or switching device located on the receiver or decoder side. Generally, embodiments of the present invention belong to the fields of joint neural coding and source separation, and the decoder side has control signals. Background Technology

[0002] Recent advances in neural speech coding [1], [2], [3] have demonstrated the powerful capabilities of end-to-end neural network methods in wideband speech coding at bit rates as low as 3.0 kbps. These studies tested the performance of these methods on clean and noisy speech, showcasing their robustness in a variety of scenarios.

[0003] For example, some neural encoders, or generalized trainable encoders, have the ability to distinguish between noisy and clean speech (i.e., so-called joint speech coding and enhancement). With this coding method, the decision to send a clean signal or a noisy signal is made, or must be made, at the transmitter side; that is, the transmitter encodes the noisy signal and sends it, or it performs speech enhancement before sending and sends the encoded, clean signal. Therefore, an improved method is needed. Summary of the Invention

[0004] The purpose of this invention is to provide a concept for efficiently encoding and decoding mixed signals (e.g., signal_1 + signal_2 + …, such as clean speech and noisy speech).

[0005] The subject matter of this invention solves this objective.

[0006] Embodiments of the present invention provide a receiver for decoding an encoded audio signal, the receiver comprising at least a decoder. The decoder is configured to receive an encoded audio signal comprising a mixture of multiple audio signals (e.g., clean speech and noisy speech). The decoder is also configured to apply a learnable model (such as a neural decoder) to decode the encoded audio signal to obtain an audio output signal. The learnable model is configured to decode the encoded audio signal such that the audio output signal comprises only one of the multiple audio signals, and such that the audio output signal comprises some or all of the multiple audio signals. In response to a control signal that is part of the decoder, the decoder decodes / outputs the audio signal such that the audio output signal comprises only one of the multiple audio signals, or such that the audio output signal comprises some or all of the multiple audio signals. In other words, this means that the decoder can choose at any time whether to decode the mixed signal or only one of the signals. According to an embodiment, the receiver provides the decoder with a control signal preprocessed by a switch, either as an input combined with the received bitstream or as a separate input. In other words, the control signal is independent of the encoder side and is given only on the decoder side.

[0007] According to the embodiment, the encoder simultaneously transmits the encoding of the clean signal and the noisy signal in a single bitstream, allowing the receiver to decode either the clean signal or the noisy signal without additional transmission overhead. Existing trainable encoders cannot achieve this.

[0008] An embodiment provides a receiver in which a learnable model or neural codec model is configured to switch between a first decoding mode and a second decoding mode in response to a control signal, wherein when operating in the first decoding mode, the audio output signal includes only one of a plurality of audio signals, and when operating in the second decoding mode, the audio output signal includes some or all of the plurality of audio signals. According to the embodiment, the switch is configured to receive the control signal and convert it into different representations for supplying to the decoder (thereby controlling it), i.e., providing the decoder with different representations as a response to the control signal.

[0009] According to an embodiment, the learnable model is a neural speech codec model, which includes Switch NESC in denoising or accurate reproduction mode, or a general speech enhancement module (codec), which may or may not be learning-based.

[0010] According to an embodiment, the learnable decoder is configured to operate in a first mode or a second mode in response to a control signal to decode an audio signal. The first mode and the second mode can be selected from the group described above. For example, the first mode enables the decoder to decode the audio signal such that the audio output signal includes only one of a plurality of audio signals, while the second mode enables the decoder to decode the audio signal such that the audio output signal includes some or all of the plurality of audio signals. For example, the decoder operating in the second mode could be a Switch_NESC decoder in exact reproduction mode, while the decoder operating in the first mode could be a Switch_NESC decoder in denoising mode.

[0011] Embodiments of the present invention are based on a novel principle / method of constructing a switch on the decoder side, which allows the end user on the receiver side to decide whether to play some or all of a plurality of audio signals, or only one audio signal, or, for example, a noisy version of encoded speech or a denoised version. This particular embodiment is primarily aimed at the case where the mixed signal is noisy speech and the desired separated signal is clean speech, i.e., joint encoding and enhancement of speech is desired. Conventional methods for solving such problems involve using a denoiser model on the encoder or decoder side, or jointly training the encoder and decoder to enhance the speech. However, the prior art does not cover methods on the decoder side for determining the more suitable codec for the current situation (input signal and / or user preference) by using a switch or control signal. Therefore, embodiments of the present invention have the advantage that selectable audio output signals can be generated on the decoder side, allowing the user or receiver to decide which generated audio output signal should be used. For example, the encoder outputs different versions (such as a "noisy" version and a "clean" version) in a single bitstream, and then the decoder side can decide whether to reproduce the "noisy" speech or the "clean" version. Therefore, no additional modules (additional speech enhancement systems) are required on the decoder side, nor is it necessary to transmit additional bitstreams (including mixed and separated signals).

[0012] According to an embodiment, the multiple audio signals include audio signals from multiple different audio sources (e.g., from different speakers or from different locations), or include background noise or music, voice, or a mixture of at least one of the foregoing elements.

[0013] According to a further embodiment, the learnable model is configured to decode the encoded audio signal such that the audio output signal includes audio signals from only one of the multiple audio sources, or includes a mixture of audio signals from some or all of the multiple audio sources.

[0014] According to an embodiment, the multiple audio signals include speech signals, and the learnable model / neural codec model is a neural speech codec model. For example, the mixing of audio signals includes a mixture of speech signals from multiple different audio sources (such as different speakers), or a mixture of background noise or music, voice, or at least one of the foregoing elements; and

[0015] The neural speech codec model is configured to decode the encoded audio signal such that the audio output signal includes a speech signal from only one of a plurality of audio sources, or a mixture of speech signals from some or all of the plurality of audio sources. Alternatively, the mixing of the audio signal includes a mixture of clean speech signal and noise signal (such as background noise), and

[0016] The neural speech codec model is configured to decode the encoded audio signal so that the audio output signal includes only clean speech signal or noisy speech signal, which includes a mixture of clean speech signal and noise signal.

[0017] According to embodiments, the decoder is configured to receive control signals from an application (such as a speech recognition application) or a user (such as a person listening to an audio signal), or to automatically generate control signals. Note that the pseudocode may describe a model for a specific use case. For example, control signals may be automatically generated based on signal-to-noise ratio (SNR) estimation.

[0018] Another embodiment provides a method for decoding an encoded audio signal, the core steps of which are:

[0019] • Receive encoded audio signals, which (e.g., bitstreams) comprise a mixture of multiple audio signals, and

[0020] • Apply a learnable model to decode the encoded audio signal to obtain the audio output signal (to be output by the receiver).

[0021] Here, the learnable model is configured to decode the encoded audio signal such that the audio output signal includes only one of a plurality of audio signals, or includes some or all of a plurality of audio signals; wherein, the learnable model is configured to decode the audio signal in response to a control signal that is part of the receiver, such that the audio output signal includes only one of a plurality of audio signals, or includes some or all of a plurality of audio signals. The receiver preprocesses the control signal (CS) via a switch and provides it to the decoder that is part of the receiver. Note that the learnable model is configured to decode the audio signal in response to a control signal that is part of the decoder, such that the audio output signal includes only one of a plurality of audio signals, or includes some or all of a plurality of audio signals.

[0022] According to an embodiment, the method can be implemented by a computer; therefore, another embodiment provides a computer program for performing the above-described method. Attached Figure Description

[0023] The embodiments of the present invention will now be discussed with reference to the accompanying drawings, wherein:

[0024] Figure 1 An exemplary block diagram of a decoder based on a basic implementation is shown;

[0025] Figure 2 An exemplary block diagram of the overall framework according to an embodiment is shown;

[0026] Figure 3 An exemplary block diagram of noise reduction according to an embodiment is shown;

[0027] Figure 4 An exemplary block diagram illustrating signal speaker quality improvement according to an embodiment is shown;

[0028] Figure 5 An exemplary block diagram illustrating an enhancement of a decoder operating in at least three modes according to an embodiment is shown. Detailed Implementation

[0029] Hereinafter, embodiments of the present invention will be discussed with reference to the accompanying drawings, wherein objects having the same or similar functions are given the same reference numerals, and their descriptions are interchangeable and mutually applicable.

[0030] Figure 1A receiver 1 is shown for decoding an encoded audio signal AS. The receiver 1 includes a decoder 10 configured to apply a learnable model or a neural codec model to provide an audio output signal OAS. For this purpose, the decoder 10 includes a learnable model (such as a neural codec model). It is configured to use different decoding modes 12a and 12b, for example, one mode in which it (10) outputs only one signal from the mixed signal, and another mode in which it outputs the entire mixed signal. In summary, the decoder may be a trained deep neural network (DNN) that decodes the received bitstream and is controlled by a control signal CS via a switch 16.

[0031] Furthermore, receiver 1 includes a switch 16 configured to, in response to a control signal CS, output / forward a signal decoded according to one of modes 12a and 12b, or switch between different modes 12a and 12b. This means that, according to embodiments, switch 16 can be positioned adjacent to decoder 10 in receiver 1 to switch the decoder between different modes 12a and 12b, or it can be arranged as part of a learnable model to switchably enable or use different modes 12a and 12b. For example, switch 16 receives the control signal CS from the user and configures decoder 10 accordingly. CS triggers mode m as the decoder's operating mode, and OAS is the result of decoding under mode m. According to embodiments, switch 16 can be logic circuitry, for example, for creating the control signal based on user input or other information available on the receiver side.

[0032] Switch 16 of receiver 1 receives control signals on the receiver side. This means that the control signals are either generated by decoder 10 itself or by the user using decoder 10.

[0033] Note that a bitstream (AS) is not explicitly composed of different components that encode different signals or mixtures of different signals. Instead, the bitstream interweaves information from individual signals, which allows information related to all individual signals and signal mixtures to be transmitted at a very low bit rate while maintaining the ability of the decoder to reproduce each individual signal or its mixture.

[0034] For example, different modes 12a and 12b, Switch_NESC_with / without_denoise_3k2bps (Neural End-to-End Speech Codec (= Robust, Scalable End-to-End Neural Speech Codec for High-Quality Wideband Speech Coding at 3kbps), can be used, with or without denoising: the neural speech codec model is trained to output both clean and noisy versions of the input bitstream. This is achieved by using a switch on the decoder side, thus requiring no additional bitrate, and the model can seamlessly switch between the two modes (noisy input -> clean or noisy output).

[0035] According to an embodiment, the learnable model is a neural speech codec model, including a NESC codec based on a TADE style layer and / or a learnable codec, and / or a NESC codec with a switch function for switching between noise reduction and no noise reduction (also known as a switch NESC codec with / without noise reduction).

[0036] Note that NESC (Neural End-to-End Speech Codec) is an example of neural coding. Some neural codecs use style layers, such as TADE style layers. NESC uses at least one learnable layer that is configured to process a (multidimensional) audio signal representation of the input audio signal or a processed version thereof to generate an output audio signal representation of the input audio signal. A TADE style layer can be interpreted as a style element that applies conditional feature parameters on the encoder and / or decoder side. The concepts of NESC and TADE style layers are both described in reference [1].

[0037] According to one implementation, mode 12a outputs multiple audio signals, i.e., a mixture of multiple audio signals (e.g., a clean speech signal surrounded by background noise or music). This is indicated by multiple arrows on the output side of decoder 12. If the input signal is a noisy input signal, the output signal will also be a noisy signal when using mode 12a. The second mode 12b can be a noise reduction mode. Even if the input is noisy speech, it can output clean speech in an end-to-end manner. Therefore, on the output side of 12, only one of the multiple audio signals is output as OAS (see one arrow).

[0038] Switch 16 can switch between two modes 12a and 12b. Therefore, the decoder is configured to decode the encoded audio signal AS using modes 12a and 12b, wherein the control signal allows control of the decoder 10 such that multiple audio signals are output as OAS or only one audio signal is output as OAS. Therefore, as... Figure 1The main advantage of the concept discussed in the context is that the present invention provides the user with an optimal trade-off and flexibility by allowing the receiver 1 to freely choose whether to listen to the original background noise or to effectively enhance it. This is particularly advantageous under challenging conditions (e.g., when the receiver has significant background noise), where a temporary improvement in intelligibility is required. It does not require additional models or modules for speech denoising, thus offering potential advantages in terms of algorithmic and architectural complexity. Furthermore, it does not require additional bitstreams for different signals, nor does it require setting up dedicated structures within the bitstreams for different signals.

[0039] As discussed above, a trainable / learnable model can be a neural codec model capable of using different modes to output only one audio signal from multiple input audio signals (extracting one audio signal from multiple audio signals during decoding) or to output multiple (all or part) of multiple input audio signals. Besides neural codec models like NESC, other neural codecs that operate in different modes based on switch states can also be used. Note that the codec, or general speech enhancement module (codec), can be learning-based or not.

[0040] As mentioned above, switch 16 can be set on the receiver 1 or the decoder side, which means that different modes can be applied and selected at the decoder. This principle will be discussed in conjunction with the framework below.

[0041] Figure 2 The overall framework for the embodiments is shown. Figure 2 The encoder, identified by reference numeral 20, is shown on the left, and the decoder side 10 is shown on the right. Encoder 20 receives a mixed signal M of different audio signals S1, S2, ..., SN. This mixed signal M is encoded by encoder 20 to output a bitstream AS. This bitstream AS is referred to as the encoded audio signal AS. The bitstream or encoded audio signal AS is fed to decoder 10 via switch 16. Decoder 10 outputs the encoded audio signal AS in a decoding manner, including the mixed signal OAS. M The output audio signal can be either a single signal, OAS1, or an output audio signal that consists of only one signal. This depends on the control signal used for switch 16.

[0042] This method has a major advantage when one of the signals S1 to SN needs to be separated from the other signals. For example, signal S1 can be clean speech, while signal S2 can be background noise. Therefore, the mixed signal M of the two signals S1 and S2 can be called noisy speech. This mixed signal M can be generated by an encoder (see...) Figure 3The encoder receives multiple (two, three, or more) audio signals and outputs a bitstream (also known as the encoded audio signal AS) to the decoder. Based on the control signal CS provided to switch 16, decoder 10 can output noisy speech (see OAS). M ) or Pure Voice (OAS) S ).

[0043] This framework illustrates the concept that pure speech can be completely separated from background noise or can be partially separated from background noise.

[0044] about Figure 4 Another concept was discussed. According to... Figure 4 For example, the first signal S1 includes the main (loudest) speaker or speech signal or mixed signal component, while signals S2 or SN belong to different speakers (such as speaker 2 or speaker N). All these channels are mixed together to form a speech mixed signal M. The speech mixed signal M is encoded in encoder 20 to output a bitstream or encoded audio signal AS. By using decoder 10 and switch 16 arranged on the decoder side, the use of decoder or decoder can be adjusted to output speech mixed signal OAS when only one channel is selected using control signal CS. M Or the main (loudest) speaker OAS S Switch between them.

[0045] Figure 5 A receiver 1' according to another embodiment is shown, which is based on Figure 1This embodiment is an example, but enhanced to allow the decoder to operate in three or M different modes 12a, 12b, 12c. A first mode can be used to reproduce a mixture of multiple audio signals (e.g., a clean speech signal surrounded by background noise or music), where modes 12b and 12c can extract different target signals. For example, by using mode 12b, a first target signal (such as speech 1) can be reproduced, while mode 12c can reproduce a second target signal (another target signal, such as speech 2). Alternatively, if any subset of the signals from the mixed signal (which may be the most general version) is applicable to 2, 3, or M modes, that subset can be extracted. For each mode or for each possible subset of signals to be extracted, a dedicated switch position can exist. Switching between different modes is performed by controlling the decoder using switch 16 in response to the control signal CS. CS triggers mode m as the decoder's operating mode, and OAS is the result of decoding in mode m. It should be noted that the number of applicable modes can vary, i.e., it can exceed three (4, 5, 6, M). This means that, according to further embodiments, the learnable model is also configured to decode the encoded audio signal (AS) according to one or more other modes to extract additional target signals or additional subsets of signals from the mixed signal.

[0046] Although some methods have been described in the context of the apparatus, it is clear that these aspects also represent descriptions of the corresponding methods, where modules or devices correspond to method steps or features of method steps. Similarly, aspects described in the context of method steps also represent descriptions of features of corresponding modules or devices. Some or all of the method steps may be performed by (or using) hardware devices, such as microprocessors, programmable computers, or electronic circuits. In some embodiments, one or more of the most important method steps may be performed by such devices.

[0047] The encoded audio signals of this invention can be stored on a digital storage medium or transmitted on a transmission medium (such as a wireless transmission medium or a wired transmission medium (such as the Internet)).

[0048] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software. This implementation can be performed using a digital storage medium (e.g., floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or FLASH memory) that stores electrically readable control signals that cooperate (or are capable of cooperating with) a programmable computer system to execute the corresponding methods. Therefore, the digital storage medium can be computer-readable.

[0049] Some embodiments of the invention include a data carrier having electrically readable control signals that can cooperate with a programmable computer system to perform one of the methods described herein.

[0050] Typically, embodiments of the present invention can be implemented as a computer program product having program code that, when run on a computer, enables the execution of one of the methods. For example, the program code can be stored on a machine-readable medium.

[0051] Other embodiments include a computer program stored on a machine-readable medium for performing one of the methods described herein.

[0052] In other words, therefore, an embodiment of the method of the present invention is a computer program having program code that, when run on a computer, performs one of the methods described herein.

[0053] Therefore, another embodiment of the method of the present invention is a data carrier (or digital storage medium, or computer-readable medium) including a computer program recorded thereon for performing one of the methods described herein. The data carrier, digital storage medium, or recording medium is typically tangible and / or non-transient.

[0054] Therefore, another embodiment of the method of the present invention is a data stream or signal sequence, representing a computer program for performing one of the methods described herein. For example, the data stream or signal sequence may be configured to be transmitted via a data communication connection (e.g., via the Internet).

[0055] Other embodiments include processing means, such as a computer or programmable logic device, configured or adapted to perform one of the methods described herein.

[0056] Another embodiment includes a computer having a computer program installed thereon for performing one of the methods described herein.

[0057] Further embodiments of the invention include an apparatus or system configured to transmit (e.g., electrically or optically) a computer program for performing one of the methods described herein to a receiver. For example, the receiver may be a computer, mobile device, storage device, etc. For example, the apparatus or system may include a file server for transmitting the computer program to the receiver.

[0058] In some embodiments, programmable logic devices (such as field-programmable gate arrays) may be used to perform some or all of the functions of the methods described herein. In some embodiments, the field-programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, these methods are preferably performed by any hardware device.

[0059] The above embodiments are merely examples illustrating the principles of the present invention. It should be understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. Therefore, the intent is limited only to the scope of the forthcoming patent claims and not to the specific details presented herein by way of description and explanation of the embodiments.

[0060] References

[0061] [1] Pia, Nicola et al., “NESC: Robust Neural End-2-End Speech Coding with GANs”, https: / / arxiv.org / abs / 2207.03282

[0062] [2] Zeghidour, Neil et al., "SoundStream: An End-to-End Neural AudioCodec", https: / / arxiv.org / abs / 2107.03312

[0063] [3] Défossez Alexandre et al., “High Fidelity Neural Audio Compression”, https: / / arxiv.org / abs / 2210.13438

[0064] [4] Sebastian Braun et al., “Data augmentation and loss normalization for deep noise suppression”, https: / / arxiv.org / abs / 2008.06412.

Claims

1. A receiver for decoding encoded audio signals (AS), the receiver comprising: The decoder (10) is configured as follows: The encoded audio signal (AS) is received, the encoded audio signal (AS) comprising a mixture (M) of multiple audio signals (S1, S2, ..., SN), and The encoded audio signal (AS) is decoded using a learnable model to obtain the audio output signal (OAS). The learnable model is configured to decode the encoded audio signal (AS) to obtain the audio output signal (OAS) comprising the plurality of audio signals (S1, S2, ..., SN) and to obtain the audio output signal (OAS) comprising some or all of the plurality of audio signals (S1, S2, ..., SN). The learnable model is configured to decode in response to a control signal (CS) that is part of the receiver (10) to obtain an audio output signal (OAS) that includes only one of the plurality of audio signals (S1, S2, ..., SN) or to obtain an audio output signal (OAS) that includes some or all of the plurality of audio signals (S1, S2, ..., SN).

2. The receiver according to claim 1, wherein, The control signal (CS) is provided on the decoder side or generated on the decoder side.

3. The receiver according to claim 1 or 2, wherein, The learnable model includes a neural encoder-decoder model or a neural decoder model.

4. The receiver according to any one of the preceding claims, wherein, The learnable model is configured to switch between a first decoding mode and a second decoding mode in response to the control signal, and When operating in the first decoding mode, the audio output signal (OAS) includes only one of the plurality of audio signals (S1, S2, ..., SN), and When operating in the second decoding mode, the audio output signal (OAS) includes some or all of the plurality of audio signals (S1, S2, ..., SN).

5. The receiver according to any one of the preceding claims, wherein, The learnable model is a neural speech codec model that includes a NESC codec based on a TADE style layer and / or a learnable codec.

6. The receiver according to any one of the preceding claims, wherein, The learnable model is configured to decode the audio signal in response to the control signal (CS) in a first mode or a second mode, wherein the first mode is capable of decoding the audio signal to obtain the audio output signal (OAS) comprising only one of the plurality of audio signals (S1, S2, ..., SN), and wherein the second mode is capable of decoding the audio signal to obtain the audio output signal (OAS) comprising some or all of the plurality of audio signals (S1, S2, ..., SN).

7. The receiver according to any one of the preceding claims, wherein, The plurality of audio signals (S1, S2, ..., SN) include audio signals from a plurality of different audio sources (e.g., from different speakers, from different locations), or include background noise or music, voice, or a mixture (M) including at least one of the foregoing elements.

8. The receiver according to claim 7, wherein, The learnable model is configured to decode the encoded audio signal (AS) to obtain the audio output signal (OAS) including the following: Audio signals from only one of multiple audio sources, or A mixture (M) of audio signals from some or all of multiple audio sources.

9. The receiver according to any one of the preceding claims, wherein, The plurality of audio signals (S1, S2, ..., SN) include speech signals, and the learnable model is a neural speech codec model.

10. The receiver according to claim 9, wherein, The mixing of audio signals (M) includes a mixture of speech signals from multiple different audio sources (e.g., different speakers), or a mixture of background noise or music, voice, or at least one of the foregoing elements; and The neural speech codec model is configured to decode the encoded audio signal (AS) to obtain the audio output signal (OAS) including the following: Speech signal from only one of multiple audio sources, or A mixture (M) of speech signals from some or all of multiple audio sources.

11. The receiver according to claim 9, wherein, The mixing of audio signals (M) includes the mixing of clean speech signals and noise signals (such as background noise) (M), and The neural speech codec model is configured to decode the encoded audio signal (AS) to obtain the audio output signal (OAS) including the following: Pure audio signal only, or Noisy speech signals include a mixture of clean speech signals and noise signals (M).

12. The receiver according to any one of the preceding claims, wherein, The decoder (10) is configured to receive the control signal (CS) from an application (such as a speech recognition application) or from a user (such as a person listening to the audio signal), or to automatically generate the control signal (CS); and / or It also includes a switch configured to receive the control signal and provide the decoder with different representations as a response to the control signal.

13. The receiver according to any one of the preceding claims, wherein, The learnable model is also configured to decode the encoded audio signal (AS) according to one or more other modes to extract additional target signals or additional subsets of signals from the mixed signal. The learnable model is configured to decode the audio signal in response to the control signal (CS) to obtain the audio output signal (OAS) that includes only one of the plurality of audio signals (S1, S2, ..., SN), or to obtain the audio output signal (OAS) that includes some or all of the plurality of audio signals (S1, S2, ..., SN), or to obtain the audio output signal (OAS) that includes only the other target signal.

14. A method for decoding an encoded audio signal (AS), the method comprising: The encoded audio signal (AS) is received, the encoded audio signal (AS) comprising a mixture (M) of multiple audio signals (S1, S2, ..., SN), and The encoded audio signal (AS) is decoded using a learnable model to obtain the audio output signal (OAS). The learnable model is configured to decode the encoded audio signal (AS) to obtain the audio output signal (OAS) that includes only one of the plurality of audio signals (S1, S2, ..., SN), and to obtain the audio output signal (OAS) that includes some or all of the plurality of audio signals (S1, S2, ..., SN). The learnable model is configured to decode the audio signal in response to a control signal (CS) that is part of the decoder (10) to obtain the audio output signal (OAS) that includes only one of the plurality of audio signals (S1, S2, ..., SN), or to obtain the audio output signal (OAS) that includes some or all of the plurality of audio signals (S1, S2, ..., SN).

15. A computer program code for executing the method of claim 14 when run on a processor.