Apparatus, method and computer program for selecting modes of input format for audio stream

By selecting the common mode of the audio signal by the control unit, the problem of TF resolution mismatch caused by the difference in the format of different source signals is solved, ensuring that the audio stream maintains high quality during transmission and realizing efficient mixing and transmission of audio signals.

CN121241392APending Publication Date: 2025-12-30NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480035819.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-06-01
Filing Date
2024-05-06
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

In audio streams, the quality of signal components is adversely affected by the different input formats used by audio signals from different sources, especially during mixing and encoding, which may lead to TF resolution mismatch issues.

Method used

The control unit selects a common mode for the audio signals, generates a mixed audio stream based on the available input formats and computational load of each signal, and sends a mode indication before transmission to ensure that the signal formats of each source are consistent, avoiding TF resolution mismatch.

Benefits of technology

It effectively solves the problem of reduced audio quality, ensures that the mixed audio stream maintains high quality during transmission, adapts to differences in signal format from different sources, and improves the overall transmission effect of audio signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121241392A_ABST
    Figure CN121241392A_ABST
Patent Text Reader

Abstract

Examples of the present disclosure relate to selecting a mode of an input format for an audio stream comprising signals mixed from different sources. An indication of one or more modes available for the selected input format of the first audio signal is obtained. An indication of one or more modes available for the selected input format of the second audio signal is also obtained. The second audio signals are to be combined to form an audio stream. In an example, a mode is selected for an audio stream comprising both a first audio signal and a second audio signal, where the selection is based at least in part on one or more common modes available for a selected input format of the first audio signal and a selected input format of the second audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Examples of this disclosure relate to apparatus, methods, and computer programs for selecting a mode for an input format of an audio stream. Some examples relate to apparatus, methods, and computer programs for selecting a mode for an input format of an audio stream that includes signals mixed from different sources. Background Technology

[0002] Audio applications such as teleconferences can obtain audio signals from different sources or capture devices. These different signals can be mixed together to generate an audio stream that can be sent to participants in teleconferences or other audio applications. If the signals from different sources or capture devices use different modes for the input, this can adversely affect the quality of the signal components in the audio stream. Summary of the Invention

[0003] According to various, but not necessarily all, examples of this disclosure, an apparatus is provided including components for performing the following operations: Obtain an indication of one or more modes of a selected input format that can be used for the first audio signal, wherein the first audio signal is received from a first source; Indicates one or more modes of a selected input format that can be used for a second audio signal, wherein the second audio signal is received from a second source, and wherein the first audio signal and the second audio signal will be combined to form an audio stream; Selecting a mode for an audio stream comprising both a first audio signal and a second audio signal, wherein the selection is based at least in part on one or more common modes of a selected input format available for the first audio signal and a selected input format for the second audio signal; and Send the selected mode indication to the corresponding source.

[0004] The indication of available modes can be for receiving more than two audio signals from different sources, and the different audio signals can have different modes available for the selected input format, wherein the audio signals will be combined to form an audio stream.

[0005] The selected mode can include a mode common to two or more audio signals, where the audio signals come from two or more sources.

[0006] The mode can be selected based at least in part on at least one of the following: The most common available patterns for the selected input format used for the source; The computational load for mixing the audio signals from the source into a combined stream in the corresponding available modes for the selected input format; and The computational load for decoding the audio signal from the source and / or encoding the mixed audio stream in the corresponding available modes for the selected input format.

[0007] This component can be used to mix at least a first audio signal and at least a second audio signal using a common mode to generate a mixed audio stream in a selected mode, and to enable the transmission of the mixed audio stream.

[0008] The selected mode can include a first mode for the main audio stream and a second mode for the sub-audio streams.

[0009] Mainstream can include a mixture of two or more audio signals that have a common pattern under the selected input format.

[0010] This component can be used to enable the main stream and substream to be sent without mixing the main stream and substream.

[0011] This component can be used to receive an indication of a preferred mode for an end-user device, and an indication of using the preferred mode to assist in the selection of a mode for an audio stream.

[0012] Audio signals may include metadata-assisted spatial audio signals.

[0013] The selected input format can include metadata-assisted spatial audio formats.

[0014] One or more modes that can be used for the input format may include time-frequency resolution.

[0015] At least one of the available modes may include higher frequency resolution, and at least one of the available modes may include higher time resolution.

[0016] This component can be used to generate spatial metadata using the selected pattern.

[0017] Based on various, but not necessarily all, examples of this disclosure, a method is provided, including: Obtain an indication of one or more modes of a selected input format that can be used for the first audio signal, wherein the first audio signal is received from a first source; Indicates one or more modes of a selected input format that can be used for a second audio signal, wherein the second audio signal is received from a second source, and wherein the first audio signal and the second audio signal will be combined to form an audio stream; Selecting a mode for an audio stream comprising both a first audio signal and a second audio signal, wherein the selection is based at least in part on one or more common modes of a selected input format available for the first audio signal and a selected input format for the second audio signal; and Send the selected mode indication to the corresponding source.

[0018] According to various, but not necessarily all, examples of this disclosure, a computer program is provided that includes program instructions that, when executed by a device, cause the device to perform at least the following: Obtain an indication of one or more modes of a selected input format that can be used for the first audio signal, wherein the first audio signal is received from a first source; Indicates one or more modes of a selected input format that can be used for a second audio signal, wherein the second audio signal is received from a second source, and wherein the first audio signal and the second audio signal will be combined to form an audio stream; Selecting a mode for an audio stream comprising both a first audio signal and a second audio signal, wherein the selection is based at least in part on one or more common modes of a selected input format available for the first audio signal and a selected input format for the second audio signal; and Send the selected mode indication to the corresponding source.

[0019] Although the examples and optional features of this disclosure have been described separately, it will be understood that their provision in all possible combinations and permutations is included within this disclosure. It will be understood that the various examples of this disclosure may include any or all of the features described with respect to other examples of this disclosure, and vice versa. Furthermore, it will be understood that any one or more features (in any combination) may be implemented / included / performed by means of an apparatus, method, and / or computer program instructions as needed and appropriate. Attached Figure Description

[0020] Some examples will now be described with reference to the accompanying drawings, in which:

[0021] Figure 1A and 1B Example system is shown;

[0022] Figure 2 Example metadata frame is shown;

[0023] Figure 3A and 3B Example metadata frame is shown;

[0024] Figure 4 Shows the mixing of audio streams;

[0025] Figure 5 Example methods are shown;

[0026] Figure 6 Shows the mixing of audio streams;

[0027] Figure 7 Shows the mixing of audio streams;

[0028] Figure 8 Example methods are shown;

[0029] Figure 9 An example system is shown; and

[0030] Figure 10 An example device is shown.

[0031] The accompanying drawings are not necessarily to scale. For clarity and brevity, some features and views in the drawings may be shown schematically or enlarged to scale. For example, the dimensions of some elements in the drawings may be enlarged relative to other elements to aid illustration. Corresponding reference numerals are used in the drawings to label corresponding features. For clarity, not all reference numerals may be shown in all drawings. Detailed Implementation

[0032] Examples of this disclosure can be implemented in audio systems that use Immersive Speech and Audio Services (IVAS) codecs.

[0033] Figure 1A and 1B An example system 101 that can be used to implement the examples of this disclosure is illustrated schematically. The example system can use the IVAS codec.

[0034] System 101 is configured to enable audio to be captured by participant devices 105 at one or more locations and transmitted to other participant devices 105 within system 101. This allows audio to be shared between different participant devices 105 within system 101. System 101 can be used for teleconferencing or any other suitable audio application.

[0035] exist Figure 1A In one example, system 101 includes a control unit 103 and three participant devices 105A, 105B, and 105C. In other examples, system 101 may include other types and numbers of devices.

[0036] Control unit 103 is configured to receive audio signals from sources within system 101. In this example, the source may be participant devices 105A, 105B, or 105C. The audio signals from participant devices 105A, 105B, or 105C may include audio generated by user 111 of participant device 105, such as voice signals or any other suitable type of audio.

[0037] Participant devices 105A, 105B, and 105C may include any suitable type of device. Participant devices 105A, 105B, and 105C may include teleconferencing equipment, mobile phones, personal computers, or any other suitable type of device that can be configured to capture audio and provide playback audio signals to user 111.

[0038] System 101 is configured to cause participant devices 105A, 105B, and 105C to send upstream signals 107 to control unit 103. Upstream signals 107 may include audio captured by the respective participant devices 105A, 105B, and 105C. The audio may originate from any sound source at the location of the respective participant devices 105A, 105B, and 105C. Figure 1A In the example, the sound source includes user 111 using participant devices 105A, 105B, and 105C. User 111 may speak during a conference call or communicate in other ways. Therefore, different participant devices 105A, 105B, and 105C provide different audio signal sources.

[0039] The control unit 103 can be configured to receive audio signals from multiple sources. In this case, the audio signals include upstream signals 107A, 107B, and 107C from participant devices 105A, 105B, and 105C.

[0040] The control unit is configured to mix audio signals from multiple sources to generate an audio stream. The control unit 103 can then provide the audio stream to the corresponding participant devices 105A, 105B, and 105C in downstream signals 109A, 109B, and 109C. This allows audio content to be shared among multiple different participant devices 105A, 105B, and 105C at different locations.

[0041] System 101 is configured such that a first participant device 105A receives a downstream signal 109A from control unit 103. The downstream signal 109A comprises packets based on an input format. The input format may be Metadata-Assisted Spatial Audio (MASA) or any other suitable format. The downstream signal 109A, sent from control unit 103 to the first participant device 105A, is formed based on input audio signals from other participant devices 105B and 105C in system 101. Control unit 103 is configured to decode upstream signals received from other participant devices 105B and 105C and mix the decoded signals into a combined format. The mixed signal can then be encoded for transmission to the first participant device 105A. The first participant device 105A receives packets from the downstream signal 109A, decodes the IVAS bitstream within the signal, and renders the signal using a suitable format for playback to a user 111 of the first participant device 105A. For example, the signal can be binaurally rendered for playback via headphones. The control unit can similarly mix and transmit signals for other participant devices 105B and 105C.

[0042] Figure 1B Different systems 101 are shown, in which there is no control unit 103. In system 101, audio signals can be transmitted directly between participant devices 105A, 105B, and 105C.

[0043] In the system of Figure 1, a first audio signal 113A is sent from the first participant device 105A to the second participant device 105B, a second audio signal 113B is sent from the second participant device 105B to the first participant device 105A, a third audio signal 113C is sent from the first participant device 105A to the third participant device 105C, a fourth audio signal 113D is sent from the third participant device 105C to the first participant device 105A, a fifth audio signal 113E is sent from the second participant device 105B to the third participant device 105C, and a sixth audio signal 113E is sent from the third participant device 105C to the second participant device 105B.

[0044] The audio signals transmitted by participant devices 105A, 105B, and 105C may include audio that has been captured by the respective participant devices 105A, 105B, and 105C. The audio signals captured by the participant devices may also be mixed with audio signals received from another source, such as another participant device 105A, 105B, or 105C.

[0045] like Figure 1A and Figure 1B As shown in both examples, more than one user 111 can use participant devices 105A, 105B, and 105C. For example, in Figure 1A and 1B In one example, three users 111 are using the first participant device 105A, only one user 111 is using the second participant device 105B, and only one user 111 is using the third participant device 105C. In other examples, system 101 may include different numbers of users 111 and participant devices 105.

[0046] IVAS codecs can be used in telecommunications systems 101 (such as...). Figure 1A and 1B In System 101), the IVAS codec supports various input formats. These formats include stereo, multichannel (MC), object-based audio (ISM), scene-based audio (SBA), and MASA. Combinations can be supported through various means, such as using objects with MASA (OMASA) or objects with SBA (OSBA). A separate input format exists for binaural audio, which operates in the same way as stereo input.

[0047] The MASA input format uses one or more audio signals and corresponding spatial metadata. MASA spatial metadata parameters describe the spatial characteristics of the captured spatial sound scene. Spatial metadata can include information such as direction in the frequency band and direct-to-total energy ratio, or any other relevant spatial information.

[0048] It can be made by system 101 (such as Figure 1A and 1B The participant device 105 in the system 101 shown obtains a MASA audio stream. For example, the MASA audio stream can be obtained by capturing spatial audio using the microphone of the participant device 105, and the corresponding spatial metadata can be estimated based on the microphone signal.

[0049] MASA streams can also be obtained from other sources by implementing appropriate format conversions, such as spatial audio microphones (e.g., Ambisonics), studio mixes (e.g., 5.1 mixes), or other content.

[0050] The MASA tool can also be used within the codec to encode multichannel signals by converting the multichannel signal into a MASA stream and then encoding that stream.

[0051] This parametric representation of MASA metadata is based on frequency bands. A spatial characteristic is associated with a frequency band, and adjacent frequency bands can exhibit different characteristics. For the MASA format, 24 frequency bands are used. The metadata frame corresponding to a 20 ms audio frame is divided into four subframes, each 5 ms long. Therefore, the parametric representation in each frame includes 24 frequency bands across four time slots, resulting in a total of 96 time-frequency patches.

[0052] The frame size in IVAS is 20 ms, and therefore the time subframe is 5 ms. In addition, MASA supports one or two directions for each time-frequency patch (that is, one or two 2-direction indices for each time-frequency patch, directly relative to the total energy ratio and diffusion coherence parameters).

[0053] Figure 2 The time-frequency resolution of the IVAS MASA metadata frame is illustrated schematically. The metadata frame configured to be processed by the IVAS encoder comprises 96 time-frequency (TF) tiles.

[0054] Various coding methods (especially at lower bit rates) can reduce the effective TF resolution. This reduction may occur by decreasing only the temporal resolution, only the frequency resolution, or both. The effective input TF resolution may also differ from the resolution supported by the MASA format. For example, the same parameter values ​​may be repeated for time subframes 0, 1, 2, 3, thus providing reduced temporal resolution. Similarly, some frequency bands may have the same parameter values ​​as some adjacent frequency bands in the same time subframe, thus providing reduced frequency resolution.

[0055] Figure 3A and 3B This illustration schematically demonstrates a reduction in the TF resolution of a sample IVAS MASA metadata frame. Different TF resolutions can be used in different modes used for the MASA input format.

[0056] Figure 3A The metadata frames in the document show reduced time and frequency resolution. Figure 3A The metadata frame in the document has twelve valid frequency bands and one valid time subframe. This provides twelve TF tiles.

[0057] Figure 3B The metadata frame in the image shows another frame with reduced time and frequency resolution. Figure 3B The metadata frame in the document has six valid frequency bands and four valid time subframes. This provides twenty-four TF tiles.

[0058] Figure 3A and 3BThe metadata frames shown are for illustrative purposes only. Further reductions in TF resolution can be used in different modes of the MASA input format. For example, one or more input modes can use eighteen frequency bands and one time subframe. This will provide eighteen TF tiles. Another one or more input modes can use five frequency bands and four time subframes. This will provide twenty TF tiles.

[0059] Figure 3A and 3B The metadata frame shown is a schematic representation of the frame. An input IVAS MASA metadata frame must have ninety-six TF tiles. The (effective) TF resolution can only be reduced by repeating parameter values. However, in other contexts, the effective TF resolution can correspond to the actual representation. For example, within the codec, only repeated values ​​are "stored".

[0060] As the bit rate decreases, it may also be necessary to reduce the TF resolution as part of the encoding. In scenarios where fewer metadata parameter values ​​are to be encoded, it is easier to encode them all with reasonable precision for reproduction at the decoder / renderer. This reduction can typically be done at least based on the content (values) of the input space metadata, but usually also takes into account the characteristics of the transmitted signal (such as its energy). The characteristics of the transmitted signal can also be analyzed per TF tile.

[0061] Problems may arise in system 101 where audio from different sources needs to be mixed. If the audio sources use different modes of the input format, they may have different TF resolutions. As an illustrative example, in Figure 1A In system 101, the second participant device 105B can generate a MASA input format with full TF resolution. The IVAS encoder on the second participant device 105B encodes the input audio signal according to the negotiated bit rate and sends the encoded signal to the control unit 103. The third participant device 105C can generate a MASA input format with reduced TF resolution. The IVAS encoder on the third participant device 105C also encodes the input audio signal according to the negotiated bit rate and sends the encoded signal to the control unit 103. Due to the bit rate reduction, a significant reduction in TF resolution can be applied to both input streams. The reduction can vary for the respective participant device 105. For example, the encoder on the second participant device 105B might be configured to maintain time resolution and thus lose more frequency resolution, while the encoder on the third participant device 105C would be configured to maintain frequency resolution because the input signal already has a lower effective time resolution.

[0062] The individual input signals from the respective participant devices 105B and 105C can have good quality; however, the difference in TF resolution can cause problems when the different input signals are mixed by the control unit 103. If the mixing is configured to attempt to maintain the best possible quality for both input signals, the resulting mixed signal will have high time resolution (from the second participant device 105B) and high frequency resolution (from the third participant device 105C). However, it may be necessary to perform a new TF resolution reduction to encode the mixed signal for transmission to the first participant device 105A at the available bit rate. In many cases, this may degrade the quality of at least one of the mixed respective input signals or component input signals. In this particular example, the input signal from the third participant device 105C may be significantly affected.

[0063] Figure 4 This schematically illustrates how mixing audio streams and mixing input signals with different TF resolutions can affect the quality of the output signal.

[0064] exist Figure 4 In this process, the control unit 103 receives upstream signal 107B from the second participant device 105B and upstream signal 107C from the third participant device 105C. Figure 4 The metadata frame 401 of the corresponding upstream signals 107B and 107C is shown.

[0065] Metadata frame 401B from the upstream signal 107B of the second participant device 105B shows a reduced frequency resolution. Metadata frame 401B has three effective frequency bands and four effective time subframes. This provides twelve TF tiles.

[0066] Metadata frame 401C from the upstream signal 107C of the third participant device 105C shows reduced frequency resolution and reduced time resolution. Metadata frame 401C has six effective frequency bands and one effective time subframe. This provides six TF tiles.

[0067] The control unit 103 includes a decoder 403, a mixer 405, and an encoder 407.

[0068] Input signals 107B and 107C from the corresponding participant devices 105B and 105D are provided as inputs to the decoder 403. The decoder 403 decodes the input signals 107B and 107C and provides the decoded input signals to the mixer 405.

[0069] Mixer 405 mixes the input signals for decoding. Mixer 405 generates an intermediate mix 409, which is provided as input to encoder 407. Mixer 405 can be configured to retain as much information as possible in the intermediate mix 409. Figure 4 In the example, the intermediate mix has six effective frequency bands and four effective time subframes. This provides twenty-four TF tiles.

[0070] Encoder 407 is configured to perform intermediate mixing to provide a downstream signal 109A that can be sent to the first participant device 105A. Encoder 407 may need to reduce the TF resolution due to bit rate or any other reason.

[0071] In this configuration, the output of encoder 409 has three effective frequency bands and four effective time subframes. This provides twelve TF blocks. The TF resolution of signal 107B from the second participant device 105B is maintained, but the TF resolution of signal 107C from the third participant device 105C is not. Signal 107C from the third participant device 105C does not acquire any time resolution during the operation performed by control unit 103 because this information cannot be added.

[0072] For at least some components of the output signal, this reduction in TF resolution may degrade the audio quality for user 111 of participant device 105. Examples of this disclosure are configured to avoid this degradation in audio quality.

[0073] Figure 5 An example method is illustrated. This method can be implemented by the control unit and / or participant device 105 within example system 101 and / or any other suitable device. This method can be performed during session negotiation or at any other suitable time.

[0074] The method includes: at block 501, obtaining an indication of one or more modes of a selected input format that can be used for a first audio signal, wherein the first audio signal is received from a first source.

[0075] At block 503, the method includes: obtaining an indication of one or more modes of a selected input format that can be used for a second audio signal, wherein the second audio signal is received from a second source. The first audio signal and the second audio signal will be combined to form an audio stream.

[0076] exist Figure 5In the example, the indication of available modes is for receiving two different audio signals from two different sources. The indication of available modes can be for receiving more than two audio signals from different sources, and the different audio signals can have different modes available for the selected input format. Multiple different audio signals from multiple different sources can be combined to form an audio stream.

[0077] The input signal can be as follows Figure 1A The upstream audio signal shown can be any other suitable type of signal. The source from which the audio signal is received can be the participant device 105 within system 101, or any other suitable type of audio signal source.

[0078] The input format specifies the format of the signal used to be supplied to the encoder. Various input formats can include stereo, MC, ISM, SBA, and MASA and / or any other type of format.

[0079] The input format can have different available modes. Different modes can include different operating modes of the input format. For example, the available input modes for the MASA input format can include high frequency resolution (HFR, 1 subframe (1sf) mode) and high temporal resolution mode (HTR, 4 subframe (4sf) mode).

[0080] The modes available for the input format include TF resolution. Different modes may have different TF resolutions. In some examples, at least one of the available modes includes a higher frequency resolution, and at least one of the available modes includes a higher time resolution.

[0081] At block 505, the method includes: selecting a mode for an audio stream that includes both a first audio signal and a second audio signal. The selection of the mode is based at least in part on one or more common modes available for a selected input format for the first audio signal and a selected input format for the second audio signal.

[0082] In some examples, the selected mode may include a mode common to two or more audio signals, where the audio signals come from two or more sources. For example, this could be the selected mode if it has a mode that can be used for both the input format of the first audio signal and the input format of the second audio signal.

[0083] In some examples, the selected mode may be chosen at least in part based on the most common available mode for the selected input format used for the corresponding source. The most common available mode may be a mode indicated as usable for the selected input format for the majority of audio signals. The most common mode may be a mode usable for the largest number of participant devices 105.

[0084] In some examples, a mode can be selected based on the audio quality it provides. If a particular mode offers better audio quality than other available modes, then the mode with the better audio quality can be selected.

[0085] In some examples, the mode can be selected based on the capabilities of the control unit 103 or any other suitable device. For example, the mode can be selected based on the computational load required to mix the input audio signal into an audio stream.

[0086] In some examples, the selected mode can be chosen at least in part based on the computational load of mixing the audio signals from the source into a combined stream in the corresponding available modes for the selected input format. In such examples, modes that result in lower computational loads during mixing can be preferred over modes that result in higher computational loads.

[0087] In some examples, the selected mode can be chosen at least in part based on the computational load of decoding the audio signal from the source and / or encoding the mixed audio stream in the corresponding available modes for the selected input format. In such examples, modes that result in lower computational load for decoding and / or encoding may be preferred over modes that result in higher computational load.

[0088] In some examples, selecting the mode for the audio stream may include selecting more than one mode. In this case, the mixed audio stream may include different components. For example, the audio stream may include an audio main stream and one or more audio substreams. The main stream may be different from the substreams because the main stream may take precedence over any substream and / or the main stream may include more data than one or more substreams. In this case, the selected mode may include a first mode for the audio main stream and a second mode for the audio substreams.

[0089] In some examples, the corresponding components of a mixed audio stream may include a mixture of two or more audio signals. For example, a main stream and / or one or more substreams may include a mixture of two or more audio signals that share a common pattern under a selected input format.

[0090] In examples where the audio stream comprises different components, a mixed audio stream can be sent without mixing the respective components. For example, a mixed audio stream can be sent without mixing the main stream and one or more substreams.

[0091] At box 507, the method includes sending an indication of the selected mode to the corresponding source.

[0092] In some examples, the method may also include Figure 5The additional box is not shown in the image. For example, in some examples, the method may include: using a common mode to mix at least a first audio signal and at least a second audio signal to generate a mixed audio stream in a selected mode, and enabling the transmission of the mixed audio stream. The mixed audio stream may be sent to participant device 105 or any other suitable device to enable playback of audio to a user.

[0093] In some examples, the method may include receiving an indication of a preferred mode for participant device 105, and using the indication of the preferred mode to assist in the selection of a mode for the audio stream. Participant device 105 may be associated with an end user. That is, the participant device may be a device to which the mixed audio stream is to be sent.

[0094] The audio signal used can be a metadata-assisted spatial audio signal or any other suitable type of signal. The selected input format used for the audio signal can include MASA format or any other suitable type of format. The method may also include generating spatial metadata using a selected mode.

[0095] The examples disclosed herein can be used to solve problems such as Figure 4 This illustrates the TF resolution issue that arises when mixing audio signals of different modes. In this example, each participant device 105B, 105C will indicate which modes they have available for the selected input format. Different modes can have different TF resolutions.

[0096] If one or more of the participant devices 105B and 105C have only a single mode available to them, the control unit 103 may request all participant devices 105B and 105C to use the same mode to generate audio signals for input to the decoder 403. If the participant devices 105B and 105C have multiple modes available for a selected input format, the control unit 103 may select the mode to be used for mixing the audio streams. Modes can be selected to maintain quality for most senders at a given bit rate.

[0097] For example, if one or more transmitting devices cannot use a mode with high time resolution (such as 4sf mode), the control unit 103 can select a mode with lower time resolution (such as 1sf mode). The control unit 103 can then send an indication of this selected mode to all relevant participant devices 105 at the lower time resolution. The corresponding participant devices can then use this mode to generate audio signals for transmission.

[0098] In some cases, participant device 105 may need to modify the audio signal to adjust it to a selected mode. In this case, participant device 105 will need to reduce the temporal resolution of the audio signal. The temporal resolution of the audio signal can be reduced by selecting one of the subframe values ​​(for each parameter at each frequency) and overwriting other subframes with that value. In other examples, other modifications for reducing temporal resolution may be used. In some examples, spatial analysis with lower temporal (and possibly higher frequency) resolution may be used.

[0099] Figure 6 The mixing of audio streams is illustrated schematically. In this example, control unit 103 receives input signals 107 from three transmitting participant devices 105. Control unit 103 can be configured to select a suitable mode for mixing the audio streams and generate a mixed audio stream so that it can be sent to receiving participant device 105A.

[0100] exist Figure 6 In this process, the control unit 103 receives upstream signal 107B from the second participant device 105B, upstream signal 107C from the third participant device 105C, and upstream signal 107D from the fourth participant device 105D. Figure 6 Metadata frames 401 for the corresponding upstream signals 107B, 107C, and 107D are shown. Upstream signal 107 can be in MASA format or any other suitable format.

[0101] The upstream signal 107B from the second participant device 105B has a first mode. Figure 6 An example metadata frame 401B is shown for an upstream signal 107B from a second participant device 105B using this first mode. This mode has reduced frequency resolution. Metadata frame 401B has three effective frequency bands and four effective time subframes. This provides twelve TF tiles.

[0102] The upstream signal 107C from the third participant device 105C has the same pattern as the upstream signal 107B from the second participant device 105B. Signal 107C from the third participant device 105C and signal 107B from the second participant device 105B have the same TF resolution. The metadata frame 401C of the upstream signal 107C from the third participant device 105C has three effective frequency bands and four effective time subframes, providing twelve TF tiles.

[0103] The upstream signal 107D from the fourth participant device 105D has a different pattern than the upstream signal 107B from the second participant device 105B. The metadata frame 401D of the upstream signal 107D from the fourth participant device 105D shows reduced frequency and time resolution. The metadata frame 401D has six effective frequency bands and one effective time subframe. This provides six TF tiles.

[0104] The control unit 103 receives an input signal 107 from the corresponding participant device 105. The input signal 107 is provided as input to the decoder 403 of the control unit 103. The decoder 403 decodes the input signal 107 and provides the decoded input signal to the mixer 405. The mixer 405 mixes the decoded input signal to generate an audio stream that is provided as input to the encoder 407. The encoder 407 is configured to mix the audio stream to provide a downstream signal 109A that can be sent to the first participant device 105A.

[0105] The control unit 103 can be configured to select a mode for the audio stream 109. This selection can be based on any common mode available for the corresponding input signal 107 and / or any other suitable factor.

[0106] exist Figure 6 In this example, control unit 103 has been selected to provide different components of the audio stream. In this case, the audio stream includes a main stream 601 and a substream 603.

[0107] In this configuration, the mainstream 601 comprises a mixture of two distinct audio signals. Specifically, the mainstream 601 includes a mixture of upstream signal 107B from the second participant device 105B and upstream signal 107C from the third participant device 105C. These signals share a common mode and therefore have the same TF resolution. The mainstream 601 utilizes this common mode. This results in the mainstream 601 having three effective frequency bands and four effective time subframes, providing twelve TF tiles.

[0108] exist Figure 6 In the example, substream 603 includes an audio signal that does not share a common mode with the audio signal used for mainstream 601. In this case, substream 603 includes an upstream signal 107D from the fourth participant device 105D. Substream 603 uses the mode of the upstream signal 107D from the fourth participant device 105D. This results in substream 603 having six effective frequency bands and one effective time subframe, which provides six TF blocks.

[0109] In this example, a main stream 601 is generated from multiple input signals 107, but a substream 603 is generated from only a single input signal. Main stream 601 may take precedence over substream 603. Main stream 601 may include more data than substream 603.

[0110] An audio stream comprising a main stream 601 and a substream 603 is sent from control unit 103 to first participant device 105A. First participant device 105A can be configured to decode the main stream 601 and substream 603 independently. In some examples, first participant device 105A can render the decoded main stream 601 and decoded substream 603 separately. In some examples, first participant device 105A can combine the decoded main stream 601 and decoded substream 603 into a combined stream, and then render the combined stream.

[0111] Rendering multiple components of an audio stream can be computationally more complex than rendering an audio stream comprising a single component. If participant device 105A prefers to operate in a mode with lower complexity, the receiver can indicate this preference to control unit 103 or any other suitable part of system 101. If participant device 105A has indicated a preference for reduced computational complexity, participant device 105A control unit 103 can select a mode that reduces such computational complexity. For example, control unit 103 can negotiate with other participant devices 105 to use a common mode so that the audio stream for participant device 105A can include only one component. For example, in Figure 6 In the example, if all participating devices 105B, 105C, and 105D are capable of using the mode with a 6x1 TF resolution, the control unit 103 will request to use that mode.

[0112] In some examples, a preferred mixing method can be negotiated between the control unit 103 and the participant device 105, i.e., determining whether the audio stream includes a single component or a main stream and one or more substreams. This negotiation can be based on the computing power and preferences of the participant device 105 and the control unit 103.

[0113] Figure 7 Another audio stream mixing is illustrated schematically. In this example, control unit 103 receives input signals 107 from five transmitting participant devices 105. Control unit 103 can be configured to select a suitable mode for mixing the audio streams and generate a mixed audio stream so that it can be sent to receiving participant device 105A.

[0114] exist Figure 7In this process, the control unit 103 receives upstream signals 107B from the second participant device 105B, 107C from the third participant device 105C, 107D from the fourth participant device 105D, 107E from the fifth participant device 105E, and 107F from the sixth participant device 105F. Figure 7 Metadata frames 401 for the corresponding upstream signals 107B, 107C, 107D, 107E, and 107F are shown. Upstream signal 107 can be in MASA format or any other suitable format.

[0115] exist Figure 7 In the example, the upstream signal 107B from the second participant device 105B and the upstream signal 107C from the third participant device 105C share a common mode. Figure 7 Example metadata frames 401B and 401C from upstream signals 107B and 107C of the second participant device 105B and the third participant device 105C using this common mode are shown. This mode has reduced frequency resolution. Metadata frames 401B and 401C have three effective frequency bands and four effective time subframes. This provides twelve TF tiles.

[0116] The upstream signal 107D from the fourth participant device 105D has a different pattern than the upstream signal 107B from the second participant device 105B. The metadata frame 401D of the upstream signal 107D from the fourth participant device 105D shows reduced frequency and time resolution. The metadata frame 401D has six effective frequency bands and one effective time subframe. This provides six TF tiles.

[0117] The upstream signal 107E from the fifth participant device 105E and the upstream signal 107F from the sixth participant device 105F share a common mode. However, this mode is different from the mode used by the other participant devices 105B, 105C, and 105D. Figure 7 Example metadata frames 401E and 401F from upstream signals 107E and 107F of the fifth participant device 105E and the sixth participant device 105F using this common mode are shown. This mode features reduced frequency resolution and reduced time resolution. Metadata frames 401E and 401F have three effective frequency bands and two effective time subframes. This provides six TF tiles.

[0118] The control unit 103 receives an input signal 107 from the corresponding participant device 105. The input signal 107 is provided as input to the decoder 403 of the control unit 103. The decoder 403 decodes the input signal 107 and provides the decoded input signal to the mixer 405. The mixer 405 mixes the decoded input signal to generate an audio stream that is provided as input to the encoder 407. The encoder 407 is configured to mix the audio stream to provide a downstream signal 109A that can be sent to the first participant device 105A.

[0119] The control unit 103 can be configured to select a mode for the audio stream 109. This selection can be based on any common mode available for the corresponding input signal 107 and / or any other suitable factor.

[0120] exist Figure 7 In this example, control unit 103 has been selected to provide different components of the audio stream. In this case, the audio stream includes a main stream 601 and multiple substreams 603A and 603B. In this case, two substreams 603A and 603B are provided. In other examples, other numbers of substreams 603 may be used.

[0121] In this configuration, the mainstream 601 comprises a mixture of two distinct audio signals. Specifically, the mainstream 601 includes a mixture of upstream signal 107B from the second participant device 105B and upstream signal 107C from the third participant device 105C. These signals share a common mode and therefore have the same TF resolution. The mainstream 601 utilizes this common mode. This results in the mainstream 601 having three effective frequency bands and four effective time subframes, providing twelve TF tiles.

[0122] exist Figure 7 In the example, the first substream 603A includes an upstream signal 107D from the fourth participant device 105D. The first substream 603A uses the pattern of the upstream signal 107D from the fourth participant device 105D. This results in substream 603 having six effective frequency bands and one effective time subframe, which provides six TF blocks.

[0123] The second substream 603B comprises a mixture of the upstream signal 107E from the fifth participant device 105E and the upstream signal 107F from the sixth participant device 105F. These signals share a common mode and therefore have the same TF resolution. The second substream 603B uses the common mode. This results in the second substream 603B having three effective frequency bands and two effective time subframes, which provides six TF tiles.

[0124] An audio stream comprising a main stream 601 and multiple substreams 603A, 603B is sent from control unit 103 to first participant device 105A. First participant device 105A can be configured to decode the main stream 601 and substreams 603A, 603B independently of each other. In some examples, first participant device 105A can render the decoded main stream 601 and the decoded substreams 603A, 603B separately. In some examples, first participant device 105A can combine the decoded main stream 601 and one or more decoded substreams 603A, 603B into a combined stream, and then render the combined stream independently of any other substream 603.

[0125] In the example disclosed herein, mainstream 601 may have a higher preference than substream 603 in terms of processing order. Therefore, if the receiving participant device 105A does not have the capacity to process all streams, it will prioritize mainstream 601 over substream 603. In other cases, the receiving participant device 105A may process mainstream 601 and one or more substreams 603 equally. In this case, mainstream 601 will not have any preference in terms of processing order.

[0126] In some examples, control unit 103 or any other suitable part of receiving participant device 105A or system 101 can determine which component will be the mainstream 601 and which component will be the sub-stream 603. In some cases, mainstream 601 may be the component that mixes most of the input signals. In some examples, mainstream 601 may be the component with the most active input signal. In some examples, mainstream 601 may be the component with the most reliable input signal connection.

[0127] In some examples, the control unit 103 may use other factors to select the mode for the audio stream. For example, the control unit 103 may select one or more modes based on maintaining low computational complexity. In this case, the control unit 103 may mix the input audio signal based on the TF format of the corresponding input signal.

[0128] In some examples, control unit 103 may select one or more modes based on optimizing the output bit rate. In this case, control unit 103 will avoid or minimize the use of substream 603, as this would increase the bit rate.

[0129] In some examples, control unit 103 may select one or more modes based on negotiation with receiving participant device 105A. This negotiation may take into account the receiving participant device 105A's preferences regarding bit rate, quality, and the number of decoding instances. Receiving participant device 105A, indicating a preference for a single decoder instance, will preferably receive an audio stream from control unit 103 comprising a single component, or if the audio stream comprises multiple components, receiving participant device 105A may decode only one of the components. For example, receiving participant device 105A may decode only the main stream 601.

[0130] The mode selected for the input format can be negotiated during session establishment. For IVAS sessions, the mode can be negotiated at the Session Description Protocol (SDP) level.

[0131] For example, the participant device 105 may use a specific operating mode parameter (inf-specific-mode) to indicate the mode available to the participant device 105 for the selected input format.

[0132] Table 1 is a list of examples of IVAS input formats and their corresponding modes. Table 1

[0133] This table only shows the operating modes for MASA and OMASA input formats. Operating modes for other input formats are not listed.

[0134] In the examples in Table 1, the MASA and OMASA input formats have two modes. These two modes are High Time Resolution (HTR) mode and High Frequency Resolution (HFR) mode. The HTR mode uses four time subframes. The HFR mode does not use subframes but can have a higher number of frequency bands than the HTR mode. These modes are examples, and different modes can be used to replace or supplement these examples.

[0135] If a specific operating mode parameter (inf-specific-mode) is used to indicate the available mode, a value can be assigned to the specific operating mode parameter to indicate the available mode.

[0136] The `inf-specific-mode` parameter can be used to indicate the available operating modes for the selected input format. If multiple operating modes are supported within a range, they are indicated by the first and last modes in the range, separated by a hyphen (`inf-specific-mode1` – `inf-specific-mode2`). If the multiple operating modes are not a consecutive range but are individual modes, they can be listed as comma-separated values ​​(`inf-specific-mode1`, `inf-specific-mode2`). Comma-separated values ​​can also be used when the specific operating modes are within a range, but the preferred order of the modes is not the default consecutive range. In both cases (hyphen- or comma-separated list), the available operating modes can be listed in preferred order, from the most preferred mode to the least preferred mode. The `inf-specific-mode-send` and `inf-specific-mode-recv` parameters can be used where different operating modes are used in the transmit and receive directions respectively. If `inf-specific-mode` is not present, this indicates that all possible input modes are available. The `inf-specific-mode` parameter is not required if the selected input format may have only a single operating mode.

[0137] Another parameter (the `disable-inf-specific-mode-switch` parameter) can be used to restrict the sending participant device 105 from switching operating modes during a communication session. For this parameter, flags can be defined to restrict mode switching. Permitted values ​​for this parameter can be 0 and 1. If `disable-inf-specific-mode-switch` is 0 or does not exist, the sending participant device 105 is allowed to switch between negotiation modes for the selected input format during the session. If `disable-inf-specific-mode-switch` is 1, the sending participant device 105 is not allowed to switch negotiation modes for the selected input format during the session.

[0138] In some examples, the available modes can be indicated by using a specific value for the input format parameter (inf). Table 2 shows examples of different values ​​that can be used for the input format parameter. In this case, a unique inf parameter value is assigned to each input format and specific operation mode. In this implementation, the inf-specific-mode parameter is not needed because the information is conveyed in the inf parameter itself. Table 2

[0139] Example 1 below describes a sample SDP offer-response negotiation for initiating a session. In this case, the sender provides a MASA input format with HTR and HFR modes. The receiver prefers the HTR MASA mode and includes only the HTR in the SDP response for the inf-specific-mode parameter. Example 1 – Example SDP provides a response scenario

[0140] The media line (m line) describes the port (49152) used for the session. RTP / AVP stands for RTP profile for audio and video, and 96 is an indicator for the dynamic payload type. The payload type number 96 is further described on the rtpmap and fmtp lines.

[0141] The rtpmap line indicates the use of the IVAS codec with a 16 kHz timestamp clock frequency. The clock frequency used for IVAS has not yet been determined and may change before the standard is finalized. The Enhanced Voice Services (EVS) codec can use a 16 kHz clock frequency. The timestamp is one of the fields in the fixed RTP header. It increments throughout the session and reflects the packet flow from sender to receiver. In the case of 20 ms voice frame blocks and a 16 kHz timestamp clock frequency, the timestamp value increases by 320 for each consecutive frame block.

[0142] In the SDP provision, the fmtp line indicates the selected input format for the transmitting participant device 105 in the inf parameter. This value refers to the inf value in Table 3 below (8=MASA). The inf-specific-mode parameter indicates the available operating modes for the provided IVAS input format. Transmitting device 105 provides two different operating modes for MASA: HTR (High Time Resolution) and HFR (High Frequency Resolution). A bit rate of 512 kbps is provided for the session.

[0143] The receiver sends an SDP reply to the sender using the modified FMTP lines. The receiver has selected a preferred input format for the sender's use (8=MASA). The receiver prefers the HTR mode and includes it only in the SDP reply. The sender should then use only the HTR MASA input mode during the session.

[0144] In this example, the parameters br, ptime, and maxptime use the definitions provided in the Enhanced Voice Services (EVS) specification (3GPP TS26.445). In general, these parameters represent: br: Indicates the bit rate (in kilobits per second) used for the session. This parameter can have a single value (br0) or a hyphenated pair of two bit rates (br1–br2), where br1 and br2 are used as the minimum and maximum bit rates, respectively. ptime: Packet time, the length of time (in milliseconds) represented by the media within a packet. In IVAS, ptime is set to 20 ms. `maxptime`: Indicates the maximum amount of media (in milliseconds) that can be encapsulated in each packet. For frame-based codecs (such as IVAS), the time should be an integer multiple of the frame size (20 ms for IVAS).

[0145] Table 3 shows the IVAS input format and its assigned attribute values. These values ​​were used in Example 1 above and Example 2 below. Table 3

[0146] Example 1 describes another example SDP provide-response scenario. In this case, the sender provides IVAS input format 8 (MASA) and a bit rate range of 13.2–160 kbps under two different operating modes (HTR and HFR). The receiver prefers the HTR mode and indicates this by listing it as the first option in the SDP response inf-specific-mode parameter value. Additionally, the receiver wishes to avoid switching between specific operating modes during the session and indicates this using the parameter disable-inf-specific-mode-switch=1. The sender should use the MASA HTR input format during the session. Example 2 – Example SDP provides a response scenario

[0147] Figure 8 Another example method that can be used to implement the examples of this disclosure is shown.

[0148] Can be used as Figure 1A The method can be implemented using the system 101 shown or any other suitable system 101. Figure 8 In the example method, some boxes are implemented by control unit 103 or other similar components, and some boxes are implemented by participant device 105.

[0149] At box 801, control unit 103 receives audio signals from two or more transmitting participant devices 105. The received audio signals may include information that will be mixed into an audio stream.

[0150] At block 803, control unit 103 receives an indication of one or more modes that can be used to transmit a selected input format to participant device 105. For example, if the selected input format is MASA, the available modes could be high frequency resolution (HFR) mode and high time resolution (HTR) mode. Other modes can be used in place of or to supplement these modes.

[0151] At block 805, control unit 103 receives an indication of mode preference from receiving participant device 105. For example, receiving participant device 105 may indicate whether it prefers HFR or HTR. In other examples, receiving participant device 105 may indicate information that control unit 103 can use to select the mode to use. For example, receiving participant device 105 may indicate whether it prefers low computational complexity or has any other criteria that should be considered.

[0152] At box 807, it is determined whether there are any additional input signals to be mixed. Additional input signals may be received from other transmitting participant devices 105. If there are additional signals to be mixed, the method returns to box 801 and repeats boxes 801 through 805 for the next input signal. If there are no additional input signals, the method continues to box 809.

[0153] At block 809, the control unit determines whether there is a common mode available for all input signals. If a common mode is available, the method proceeds to block 811. If no common mode is available for all input signals, the method proceeds to block 821.

[0154] At block 811, the method includes selecting a common mode for use by all transmitting participant devices 105. At block 813, the control unit 103 indicates the selected mode to the participant devices 105. The control unit 103 may indicate the mode to be used to the transmitting participant devices 105. Once the transmitting participant devices 105 have received the indication of the selected mode, they will use that mode to generate an audio signal.

[0155] At block 815, control unit 103 mixes the received input audio signals into an audio stream. Using a common mode for all input signals enables the generation of a mixed audio stream without any quality degradation.

[0156] At block 817, control unit 103 sends the mixed audio stream to receiving participant device 105. At block 819, receiving participant device 105 receives the mixed audio stream and performs decoding and rendering of the mixed audio stream. This enables the playback of audio content to the user of receiving participant device 105.

[0157] If no common pattern is available for all input signals, the method proceeds to block 821. In this case, the audio stream will be mixed into a main stream 601 and one or more substreams 603. At block 821, control unit 103 determines whether there is any common pattern for a subset of the input signals. This could be a pattern common to some, but not all, of the transmitting participant devices 105.

[0158] At block 823, control unit 103 determines the configuration for mainstream 601 and the configuration for one or more sub-streams 603. The configuration for a given stream may include the input signals to be mixed into the stream and the modes to be used. For example, control unit 103 may determine that signals with a common mode can be mixed into mainstream 601, while information without a common mode can be mixed into sub-streams 603. Mainstream 601 and sub-streams 603 may be configured such that signals using different modes are not mixed together.

[0159] At box 825, control unit 103 indicates the selected mode to participant device 105. Control unit 103 may indicate the mode to be used to transmitting participant device 105. The selected mode may be the mode to be used for main stream 601 and one or more substreams 603. Once transmitting participant device 105 has received the indication of the selected mode, transmitting participant device 105 will use the mode to generate an audio signal.

[0160] At block 827, control unit 103 mixes the received input audio signal into a main stream 601 and one or more substreams 603. A subset of the input signal sharing a common mode can be mixed into the main stream 601, and the remaining input signal can be mixed into one or more substreams.

[0161] At block 829, control unit 103 sends a mixed audio stream comprising a main stream 601 and one or more substreams 603 to receiving participant device 105. At block 831, receiving participant device 105 receives the mixed audio stream and performs decoding and rendering of the mixed audio stream. In some examples, receiving participant device 105 may decode the main stream 601 and one or more substreams 603 separately. Receiving participant device 105 can then enable the playback of audio content to a user of receiving participant device 105.

[0162] Figure 9 An example system 101 is shown that can be used to implement the examples of this disclosure. System 101 can be an IVAS system. The system can use MASA format or any other suitable type of format.

[0163] Figure 9The example system 101 includes a control unit 103 and three participant devices 105. In this example, a first participant device 105A is configured to receive participant devices 105A, and a second participant device 105B and a third participant device 105C are transmitting participant devices 105. It will be understood that in the examples of this disclosure, the respective participant devices 105 can both transmit and receive audio signals.

[0164] In this example, control unit 103 includes session controller module 901, encoder / decoder module 903, selection module 905, and transport stream module 907. In other examples, control unit 103 may include different modules and / or combinations of modules.

[0165] The session controller module 901 can be configured to establish a communication session between the control unit 103 and the corresponding participant device 105. The session controller module 901 can be configured to receive session negotiation signaling from the corresponding participant device 105. The session controller module 901 can be configured to negotiate parameters for the multimedia session between the sending participant devices 105B and 105C and the receiving participant device 105A.

[0166] The encoder / decoder module 903 is configured to decode audio signals received from the respective transmitting participant devices 105B, 105C. The encoder / decoder module 903 is also configured to encode mixed audio streams for transmission to the receiving participant device 105A. The encoder / decoder module 903 may include a MASA encoder / decoder or any other suitable type of encoder.

[0167] Selection module 905 can be configured to select a mode to be used for mixing audio streams. Selection module 905 can receive information indicating available modes for the selected input format used by the corresponding transmitting participant device 105. Selection module 905 can also receive information related to the preferences (such as computational requirements) of receiving participant device 105A. Selection module 905 can use any of the methods or processes described herein to select the mode to be used for mixing.

[0168] The transport stream module 907 can be configured to generate a transport audio stream. The transport stream can use any suitable protocol, such as the Real-Time Transport Protocol (RTP). The transport stream module 907 provides an encoded stream as the payload.

[0169] The corresponding participant device 105 also includes a microphone array 909, a session controller module 911, an encoder / decoder module 913, a processing module 915, and a transport stream module 917. In other examples, the control unit 103 may include different modules and / or combinations of modules.

[0170] Microphone array 909 may include multiple microphones. These microphones are configured to capture sound and generate electrical output signals. The microphones within microphone array 909 may be spatially arranged to enable the capture of spatial audio.

[0171] Participant device 105 is configured such that the output signal from microphone array 909 is provided to processing module 915. Processing module 915 is configured to process the microphone signal into a selected format. Processing module 915 is configured to process the microphone signal into a selected mode of the selected format.

[0172] Participant device 105 is configured such that the output signal from processing module 915 is provided as input to encoder module 913. Encoder module 913 is configured to encode the processed microphone signal for transmission to control unit 103. Encoder module 913 may include a MASA encoder or any other suitable type of encoder.

[0173] The encoded signal is provided to the transport stream module 917. The transport stream module 917 can be configured to generate a transport audio stream. The transport stream can use any suitable protocol, such as Real-Time Transport Protocol (RTP). The transport stream module 917 provides the encoded stream as a payload. This encoded stream can be sent to the control unit 103.

[0174] Session controller module 911 can be configured to establish a communication session between control units 103. Session controller module 911 can be configured to send session negotiation signaling to control unit 103. Session controller module 911 can be configured to negotiate parameters for a multimedia session between participant device 105 and control unit 103. The session controller module 911 of participant device 105 is configured to send signals to the session controller module 911 of control unit 103.

[0175] The participant device can also be configured to receive head tracking information. Head tracking information can be received from one or more positioning devices or sensors configured to monitor user 111. The participant device 105 can use this information to determine the user's head orientation and render the received audio stream.

[0176] exist Figure 9 In the example, participant device 105 is the same. In other examples, participant device 105 may be different.

[0177] Figure 10An apparatus 1001 that can be used to implement an example of this disclosure is schematically shown. In this example, apparatus 1001 includes a controller 1003. The controller 1003 may be a chip or a chipset. Apparatus 1001 may be located within a control unit 103 or participant device 105 or any other suitable device.

[0178] exist Figure 10 In the examples, controller 1003 can be implemented as a controller circuit. In some examples, controller 1003 can be implemented solely in hardware, possess certain aspects of software (including firmware) alone, or be a combination of hardware and software (including firmware).

[0179] like Figure 10 As shown, the controller 1003 can be implemented using instructions that implement hardware functions, for example by using executable instructions of a computer program 1009 in a general-purpose or special-purpose processor 1005, which can be stored on a computer-readable storage medium (disk, memory, etc.) for execution by such processor 1005.

[0180] Processor 1005 is configured to read from and write to memory 1007. Processor 1005 may also include an output interface through which processor 1005 outputs data and / or commands; and an input interface through which data and / or commands are input to processor 1005.

[0181] Memory 1007 stores a computer program 1009, including computer program instructions (computer program code 1011), which, when loaded into processor 1005, controls the operation of controller 1003. The computer program instructions of computer program 1009 provide logic and routines that enable controller 1003 to execute the methods shown in the figures. By reading memory 1007, processor 1005 can load and execute computer program 1009.

[0182] Device 1001 includes: At least one processor 1005; and At least one memory 1007 stores instructions that, when executed by at least one processor 1005, cause at least the following to be performed by the device 1001: 501 indicates one or more modes of a selected input format that can be used for the first audio signal, wherein the first audio signal is received from a first source; 503 indicates one or more modes of a selected input format that can be used for the second audio signal, wherein the second audio signal is received from a second source, and wherein the first audio signal and the second audio signal will be combined to form an audio stream; Selection 505 is used for a mode of an audio stream comprising both a first audio signal and a second audio signal, wherein the selection is based at least in part on one or more common modes of a selected input format available for the first audio signal and a selected input format for the second audio signal; and Send the 507 mode selection instruction to the appropriate source.

[0183] like Figure 10 As shown, computer program 1009 can reach controller 1003 via any suitable transmission mechanism 1013. Transmission mechanism 1013 can be, for example, a machine-readable medium, a computer-readable medium, a non-transitory computer-readable storage medium, a computer program product, a storage device, a recording medium (such as an optical disc read-only memory (CD-ROM) or digital versatile optical disc (DVD) or solid-state memory), or an article of manufacture that includes or tangibly embodies computer program 1009. The transmission mechanism can be a signal configured to reliably transmit computer program 1009. Controller 1003 can propagate or transmit computer program 1009 as a computer data signal. In some examples, wireless protocols such as Bluetooth, Bluetooth Low Energy, Bluetooth Smart, and 6LoWpan (IP over Low Energy Personal Area Network) can be used. v 6) The computer program 1009 is sent to the controller 1003 via ZigBee, ANT+, Near Field Communication (NFC), Radio Frequency Identification, Wireless Local Area Network (Wireless LAN) or any other suitable protocol.

[0184] Computer program 1009 includes computer program instructions for causing device 1001 to perform at least the following operations, or the computer program instructions are configured to perform at least the following operations: 501 indicates one or more modes of a selected input format that can be used for the first audio signal, wherein the first audio signal is received from a first source; 503 indicates one or more modes of a selected input format that can be used for the second audio signal, wherein the second audio signal is received from a second source, and wherein the first audio signal and the second audio signal will be combined to form an audio stream; Selection 505 is used for a mode of an audio stream comprising both a first audio signal and a second audio signal, wherein the selection is based at least in part on one or more common modes of a selected input format available for the first audio signal and a selected input format for the second audio signal; and Send the 507 mode selection instruction to the appropriate source.

[0185] Computer program instructions may be included in computer program 1009, non-transitory computer-readable medium, computer program product, or machine-readable medium. In some, but not necessarily all, examples, computer program instructions may be distributed across multiple computer programs 1009.

[0186] Although memory 1007 is shown as a single component / circuit, it can be implemented as one or more separate components / circuits, some or all of which can be integrated / removable, and / or provide permanent / semi-permanent / dynamic / cached storage.

[0187] Although processor 1005 is shown as a single component / circuit, it can be implemented as one or more separate components / circuits, some or all of which can be integrated / removable. Processor 1005 can be a single-core or multi-core processor.

[0188] References to “computer-readable storage medium,” “computer program product,” “computer program tangibly embodied,” or “controller,” “computer,” “processor,” etc., should be understood to include not only computers with different architectures (such as single / multiprocessor architectures and sequential (von Neumann) / parallel architectures), but also special-purpose circuits such as field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), signal processing devices, and other processing circuits. References to computer programs, instructions, code, etc., should be understood to include software for programmable processors or firmware, such as programmable content for hardware devices, whether instructions for processors or configuration settings for fixed-function devices, gate arrays, or programmable logic devices, etc.

[0189] As used in this application, the term "circuit" may refer to one or more or all of the following: (a) Hardware circuit implementation only (such as implementation in analog and / or digital circuits only), and (b) A combination of hardware circuitry and software, such as (if applicable): (i) A combination of analog and / or digital hardware circuitry with software / firmware, and (ii) Any part of a hardware processor (including a digital signal processor), software, and memory that works together to enable a device (such as a mobile phone or server) to perform various functions, and (c) Hardware circuitry and / or processors, such as microprocessors or a portion thereof, which require software (e.g., firmware) to operate, but where operation does not require software, the software may not exist.

[0190] The definition of "circuit" applies to all uses of the term in this application (including in any claim). As a further example, as used in this application, the term "circuit" also covers implementations of hardware circuits or processors and their accompanying software and / or firmware. The term "circuit" also covers (e.g., and if applicable to a particular claim element) baseband integrated circuits for mobile devices or similar integrated circuits in servers, cellular network devices, or other computing or networking devices.

[0191] Figure 4 and 8 The boxes shown may represent steps in the method and / or code portions in the computer program 1009. A specific sequence diagram of the boxes does not necessarily imply a required or preferred order, and the order and arrangement of the boxes may be changed. Furthermore, some boxes may be omitted.

[0192] The term "includes" as used in this document has an inclusive meaning, not an exclusive meaning. That is, any reference to X including Y indicates that X may include only one Y or may include multiple Ys. If the exclusive meaning of "includes" is intended to be used, it will be made clear in the context by referring to "includes only one..." or by using "comprises".

[0193] In this description, the terms “connection,” “coupling,” and “communication,” and their derivatives, refer to operational connection / coupling / communication. It should be understood that any number or combination of intermediate components (including no intermediate components) may exist to provide direct or indirect connection / coupling / communication. Any such intermediate component may include hardware and / or software components.

[0194] As used herein, the term "determine" (and its grammatical variations) can include, but is not limited to: calculation, estimation, processing, derivation, measurement, investigation, identification, lookup (e.g., searching in a table, database, or another data structure), ascertainment, etc. Furthermore, "determine" can include receiving (e.g., receiving information), accessing (e.g., accessing data in memory), obtaining, etc. Additionally, "determine" can include parsing, selecting, picking, building, etc.

[0195] Various examples have been referenced in this description. Descriptions of features or functions in examples indicate that such features or functions exist in that example. The use of the terms "example," "for example," "may," or "possibly" throughout this document (whether explicitly stated or not) implies that such features or functions exist at least in the described example (whether or not described as examples), and that they may, but not necessarily, exist in some or all other examples. Therefore, "example," "for example," "may," or "possibly" refers to a specific instance within a class of examples. The characteristics of an instance can be characteristics of only that instance, or characteristics of the class, or characteristics of a subclass of the class (which includes some, but not all, instances of that class). Thus, it is implicitly disclosed that a feature described with reference to one example but not to another may (if possible) be used as part of a combination of works in that other example, but is not necessarily required to be used in that other example.

[0196] Although examples have been described in the preceding paragraphs with reference to various examples, it should be understood that modifications can be made to the given examples without departing from the scope of the claims.

[0197] The features described above can be used in combinations other than those explicitly described above.

[0198] Although some features have been described with reference to certain features, these features can be performed by other features (whether or not they are described).

[0199] Although features have been described with reference to some examples, these features may also exist in other examples (whether or not they are described).

[0200] The terms “a,” “an,” or “that” as used in this document have an inclusive rather than an exclusive meaning. That is, any reference to X including a / an / that Y indicates that X may include only one Y or may include more than one Y, unless the context clearly indicates the opposite. If “a,” “an,” or “that” is intended to have an exclusive meaning, it will be clarified in the context. In some cases, “at least one” or “one or more” may be used to emphasize inclusiveness, but the absence of these terms should not be taken for granted that any exclusive meaning can be inferred.

[0201] The feature (or combination of features) in the claims refers to the feature (or combination of features) itself, and also to features that achieve substantially the same technical effect (equivalent features). Equivalent features include, for example, features that are variations and achieve substantially the same result in substantially the same manner. Equivalent features include, for example, features that perform substantially the same function in substantially the same manner to achieve substantially the same result.

[0202] In this description, various examples have been referenced that use adjectives or adjective phrases to describe the characteristics of the example. Such a description of the characteristic of the example indicates that the characteristic exists exactly as described in some examples, and substantially as described in others.

[0203] The foregoing description illustrates some examples of this disclosure, but those skilled in the art will recognize possible alternative structures and methodological features that provide functionality equivalent to specific examples of such structures and features described above, and which have been omitted from the foregoing description for the sake of brevity and clarity. However, the foregoing description should be understood to implicitly include references to such alternative structures and methodological features that provide equivalent functionality, unless such alternative structures or methodological features are expressly excluded in the foregoing description of the examples of this disclosure.

[0204] When the foregoing specification focuses on those features deemed important, it should be understood that the applicant may seek protection by means of the claims for any patentable feature or combination of features (whether emphasized or not) mentioned above and / or shown in the figures.

Claims

1. An apparatus comprising means for: obtaining an indication of one or more modes available for a selected input format of the first audio signal, wherein, obtaining an indication of one or more modes available for a selected input format of a first audio signal, wherein the first audio signal is received from a first source; obtaining an indication of one or more modes available for a selected input format of a second audio signal, wherein the second audio signal is received from a second source, and wherein the first audio signal and the second audio will be combined to form an audio stream; selecting a mode for the audio stream comprising both the first audio signal and the second audio signal, wherein the selection is based at least in part on one or more common modes available for the selected input format of the first audio signal and the selected input format of the second audio signal; and sending an indication of the selected mode to the respective sources.

2. The apparatus of claim 1, wherein, The indication of the available modes is received for more than two audio signals from different sources, and different audio signals can have different modes available for the selected input format, and wherein the audio signals will be combined to form an audio stream.

3. The device of any of the preceding claims, wherein, The selected mode comprises a mode common to two or more audio signals, wherein the audio signals are from two or more sources.

4. The device of any of the preceding claims, wherein, The mode is selected based at least in part on at least one of: a most common available mode for the selected input format for the sources; a computational load for mixing the audio signals from the sources into a combined stream in the respective available modes for the selected input format; and a computational load for decoding the audio signals from the sources and / or encoding a mixed audio stream in the respective available modes for the selected input format.

5. The device of any of the preceding claims, wherein, The means are for mixing at least the first audio signal and at least the second audio signal using the common mode to generate a mixed audio stream in the selected mode, and enabling transmission of the mixed audio stream.

6. The apparatus of any one of claims 1 to 4, wherein, The selected mode comprises a first mode for an audio main stream and a second mode for an audio sub stream.

7. The apparatus of claim 6, wherein, The main stream can comprise a mix of two or more audio signals having a common mode in the selected input format.

8. The apparatus of any one of claims 6-7, wherein, The means are for enabling the main stream and the sub stream to be transmitted without mixing the main stream and the sub stream.

9. The device of any of the preceding claims, wherein, The means are for receiving an indication of a preferred mode for an end user device, and using the indication of the preferred mode to assist in the selection of the mode for the audio stream.

10. The device of any of the preceding claims, wherein, The audio signals comprise metadata assisted spatial audio signals.

11. The device of any of the preceding claims, wherein, The selected input format comprises a metadata assisted spatial audio format.

12. The device of any one of the preceding claims, wherein, The one or more modes available for an input format comprise time frequency resolution.

13. The apparatus of claim 12, wherein, At least one of the available modes comprises a higher frequency resolution, and at least one of the available modes comprises a higher time resolution.

14. The device of any of the preceding claims, wherein, The means are for generating spatial metadata using the selected mode.

15. A method comprising: obtaining an indication of one or more modes available for a selected input format of a first audio signal, wherein the first audio signal is received from a first source; obtaining an indication of one or more modes available for a selected input format of a second audio signal, wherein the second audio signal is received from a second source, and wherein the first audio signal and the second audio will be combined to form an audio stream; selecting a mode for the audio stream comprising both the first audio signal and the second audio signal, wherein the selection is based at least in part on one or more common modes available for the selected input format of the first audio signal and the selected input format of the second audio signal; and sending an indication of the selected mode to the respective sources. obtaining an indication of one or more modes available for a selected input format of a second audio signal, wherein the second audio signal is received from a second source, and wherein the first audio signal and the second audio are to be combined to form an audio stream; selecting a mode for the audio stream including both the first audio signal and the second audio signal, wherein the selection is based at least in part on one or more common modes available for the selected input format of the first audio signal and the selected input format of the second audio signal; and sending an indication of the selected mode to the respective sources.

16. A computer program comprising program instructions that, when executed by an apparatus, cause the apparatus to perform at least: obtaining an indication of one or more modes available for a selected input format of the first audio signal, wherein, the first audio signal is received from a first source; obtaining an indication of one or more modes available for a selected input format of a second audio signal, wherein the second audio signal is received from a second source, and wherein the first audio signal and the second audio are to be combined to form an audio stream; selecting a mode for the audio stream including both the first audio signal and the second audio signal, wherein the selection is based at least in part on one or more common modes available for the selected input format of the first audio signal and the selected input format of the second audio signal; and sending an indication of the selected mode to the respective sources.

17. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to perform at least: obtaining an indication of one or more modes available for a selected input format of a first audio signal, wherein the first audio signal is received from a first source; obtaining an indication of one or more modes available for a selected input format of a second audio signal, wherein the second audio signal is received from a second source, and wherein the first audio signal and the second audio are to be combined to form an audio stream; selecting a mode for the audio stream including both the first audio signal and the second audio signal, wherein the selection is based at least in part on one or more common modes available for the selected input format of the first audio signal and the selected input format of the second audio signal; and sending an indication of the selected mode to the respective sources.

18. The apparatus of claim 17, wherein, the selected mode includes a mode common to two or more audio signals, wherein the audio signals are from two or more sources.

19. The apparatus of any one of claims 17 or 18, wherein, the mode is selected based at least in part on at least one of: a most common available mode for a selected input format of the source; a computational load for mixing the audio signals from the sources into a combined stream in a respective available mode for the selected input format; and a computational load for decoding the audio signals from the sources and / or encoding a mixed audio stream in a respective available mode for the selected input format.

20. The apparatus of any one of claims 17-19, wherein, further causing the apparatus to: mixing at least the first audio signal and at least the second audio signal using a public mode to generate a mixed audio stream in the selected mode; and enabling transmission of the mixed audio stream.

21. The apparatus of any one of claims 17-20, wherein, causing the apparatus to generate spatial metadata using the selected mode.

22. The apparatus of any one of claims 17-21, wherein, The audio signals comprise metadata-aided spatial audio signals.

23. The apparatus of any one of claims 17-22, wherein, The selected input format comprises a metadata-aided spatial audio format.