Beamforming control for 6-degrees of freedom audio rendering
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- NOKIA TECHNOLOGIES OY
- Filing Date
- 2024-06-19
- Publication Date
- 2026-05-13
AI Technical Summary
Current MPEG-I Immersive audio renderers face challenges in accurately estimating audio source spectral properties when higher-order Ambisonics (HOA) sources are obscured by intermediate audio sources, leading to lower quality 6-degrees-of-freedom (6DoF) rendering and increased computational complexity.
The method involves identifying unknown audio sources interfering with known sources and switching to alternative microphones for beamforming, allowing the renderer to capture audio from known sources while the unknown source is between the microphone and the known source, thereby improving sound localization and reducing computational overhead.
This approach enhances audio quality and reduces computational resources by accurately accounting for obscured audio sources, providing better 6DoF rendering and efficient processing in dynamic environments.
Smart Images

Figure EP2024067051_16012025_PF_FP_ABST
Abstract
Description
[0001]BEAMFORMING CONTROL FOR 6-DEGREES OF FREEDOM AUDIO RENDERING Field The present application relates to apparatus and methods for beamforming control for 6-degrees of freedom audio rendering, but not exclusively for beamforming control for 6-degrees of freedom audio rendering in augmented reality and / or virtual reality apparatus. Background The current implementations of the MPEG-I Immersive audio standard (ISO / IEC 23090-4 WD3) renderer is configured to support 6 degrees-of-freedom (6DoF) rendering of audio scenes comprising multiple first order or higher-order Ambisonics (FOA, HOA) microphone recordings or synthesized signals. The renderer implementations can be able to provide a binaural signal ‘at’ the listener position (and orientation) based on the recorded FOA / HOA signals and their positions. That is, the renderer is able to provide a binaural signal at a non-sampled position in the scene, thus providing a 6DoF experience for the listener. Figure 1 shows an example scene 100 with microphones m1101, m2103, m3105, m4 107, and m5109 at positions p1, p2, p3, p4 and p5 respectively and a listener 111 l1 at position pl. Current MPEG-I audio renderer implementations are able to utilize information (e.g., position) about audio sources present in the scene (e.g., during the recording or audio sources present in synthetic ambisonics sources) such as described in WO2022136725. The positions of any of the audio sources present in the recorded or synthetic audio scene may be given as input information to the renderer in addition to the audio signals and positions of the FOA or HOA microphones. Using this additional information about the audio sources allows for the renderer to provide improved sound quality through improved localization and distance attenuation behavior for the known sources. To achieve this, the renderer can be configured to perform beamforming (from a FOA / HOA signal) towards a known source to estimate the source properties (energy at different frequency bands). Based on this and the position of the known source, the renderer can be configured to determine from which direction, from the listener’s perspective, and in which frequency bands, how much energy is being contributed by the known source. However during estimation of audio source spectral properties in order to perform 6DoF rendering of scenes comprising multiple HOA sources, a problem can occur if the HOA source and the audio source whose position information is provided to the renderer does not consider the presence of other audio sources between a HOA source and an informed source (whose rendering is being improved selectively). There is current research into rendering and improving the estimate of spectral properties of an audio source (for example with beamforming) and where the audio source is ‘obscured’ or ‘blocked’ or ‘masked’ by an intermediate audio source. Summary There is provided according to a first aspect a method for a rendering of a six degrees-of-freedom audio scene, the method comprising: providing at least two microphones located at respective position for audio capturing; identifying at least one known audio source being captured based on a beamforming by at least one microphone of the at least two microphones, the at least one audio source is located in the six degrees-of-freedom audio scene relative to the at least two microphones; identifying at least one unknown audio source in the six degrees-of- freedom audio scene, wherein the at least one unknown audio source at least partially interferes with the capturing of the at least one known audio source by the at least one microphone; and controlling the beamforming based on identifying at least one alternative microphone of the at least two microphones, the controlling of beamforming is such that the at least one alternative microphone starts capturing audio from the at least one known audio source while the at least one unknown audio source is at least partially located between the at least one microphone and the at least one known audio source. The at least two microphones may be higher-order Ambisonics microphones. Controlling the beamforming based on identifying at least one alternative microphone of the at least two microphones may comprise determining the at least one unknown audio source is masking the at least one known audio source with respect to the beamforming by at least one microphone of the at least two microphones. Controlling the beamforming based on identifying at least one alternative microphone of the at least two microphones may further comprise generating a signal, the signal indicating that the at least one alternative microphone starts capturing audio from the at least one known audio source while the at least one unknown audio source is at least partially located between the at least one microphone and the at least one known audio source. Controlling the beamforming based on identifying at least one alternative microphone of the at least two microphones may further comprise generating a signal, the signal associated with the at least one known audio source and indicating which the at least one microphone of the two microphones and the at least one alternative microphone is to be selected. Controlling the beamforming based on identifying at least one alternative microphone of the at least two microphones may further comprise switching from the one of the at least two microphones to the at least one alternative microphone. Controlling the beamforming based on identifying at least one alternative microphone of the at least two microphones may comprise obtaining a signal associated with the at least one known audio source, the signal indicating where the at least one microphone of the two microphones and the at least one alternative microphone is to be selected. The method may further comprising receiving a six degrees-of-freedom audio scene description wherein the six degrees-of-freedom audio scene description comprises the signal associated with the at least one known audio source. Controlling the beamforming based on identifying at least one alternative microphone of the at least two microphones may comprise activating one of: the at least one alternative microphone; or the at least one microphone, for capturing audio from the at least one known audio source and deactivating the other of the at least one alternative microphone or the at least one microphone. Controlling the beamforming based on identifying at least one alternative microphone of the at least two microphones may comprise controlling disabling or enabling of the at least one known source. The method may further comprise determining the at least one unknown audio source at least partially interferes with the capturing of the at least one known audio source by the at least one microphone. Determining the at least one unknown audio source at least partially interferes with the capturing of the at least one known audio source by the at least one microphone may comprise determining the position of the at least one unknown audio source is between or substantially between the at least one known audio source and the at least one of the at least two microphones. Determining the at least one unknown audio source at least partially interferes with the capturing of the at least one known audio source by the at least one microphone may comprise determining a distance between the at least one unknown audio source and the at least one known audio source is less than a threshold distance. Determining the at least one unknown audio source at least partially interferes with the capturing of the at least one known audio source by the at least one microphone may comprise determining an energy estimate of an energy of the at least one known audio source is lower than a threshold energy. The method may furthermore comprise determining at least one spatial metadata parameter associated with the at least one known audio source based on an analysis of the beamformed at least one known audio source. According to a second aspect there is provided an apparatus for a rendering of a six degrees-of-freedom audio scene, the apparatus comprising means configured to: provide at least two microphones located at respective position for audio capturing; identify at least one known audio source being captured based on a beamforming by at least one microphone of the at least two microphones, the at least one audio source is located in the six degrees-of-freedom audio scene relative to the at least two microphones; identifying at least one unknown audio source in the six degrees-of-freedom audio scene, wherein the at least one unknown audio source at least partially interferes with the capturing of the at least one known audio source by the at least one microphone; and control the beamforming based on identifying at least one alternative microphone of the at least two microphones, the control of beamforming is such that the at least one alternative microphone starts capturing audio from the at least one known audio source while the at least one unknown audio source is at least partially located between the at least one microphone and the at least one known audio source. The at least two microphones may be higher-order Ambisonics microphones. The means configured to control the beamforming based on identifying at least one alternative microphone of the at least two microphones may be configured to determine the at least one unknown audio source is masking the at least one known audio source with respect to the beamforming by at least one microphone of the at least two microphones. The means configured to control the beamforming based on identifying at least one alternative microphone of the at least two microphones may be configured to generate a signal, the signal indicating that the at least one alternative microphone starts capturing audio from the at least one known audio source while the at least one unknown audio source is at least partially located between the at least one microphone and the at least one known audio source. The means configured to control the beamforming based on identifying at least one alternative microphone of the at least two microphones may be configured to generate a signal, the signal associated with the at least one known audio source and indicating which the at least one microphone of the two microphones and the at least one alternative microphone is to be selected. The means configured to control the beamforming based on identifying at least one alternative microphone of the at least two microphones may be configured to switch from the one of the at least two microphones to the at least one alternative microphone. The means configured to control the beamforming based on identifying at least one alternative microphone of the at least two microphones may be configured to obtain a signal associated with the at least one known audio source, the signal indicating where the at least one microphone of the two microphones and the at least one alternative microphone is to be selected. The means may be further configured to receive a six degrees-of-freedom audio scene description wherein the six degrees-of-freedom audio scene description comprises the signal associated with the at least one known audio source. The means configured to control the beamforming based on identifying at least one alternative microphone of the at least two microphones may be configured to activate one of: the at least one alternative microphone; or the at least one microphone, for capturing audio from the at least one known audio source and deactivating the other of the at least one alternative microphone or the at least one microphone. The means configured to control the beamforming based on identifying at least one alternative microphone of the at least two microphones may be configured to control disabling or enabling of the at least one known source. The means may be further configured to determine the at least one unknown audio source at least partially interferes with the capturing of the at least one known audio source by the at least one microphone. The means configured to determine the at least one unknown audio source at least partially interferes with the capturing of the at least one known audio source by the at least one microphone may be configured to determine the position of the at least one unknown audio source is between or substantially between the at least one known audio source and the at least one of the at least two microphones. The means configured to determine the at least one unknown audio source at least partially interferes with the capturing of the at least one known audio source by the at least one microphone may be configured to determine a distance between the at least one unknown audio source and the at least one known audio source is less than a threshold distance. The means configured to determine the at least one unknown audio source at least partially interferes with the capturing of the at least one known audio source by the at least one microphone may be configured to determine an energy estimate of an energy of the at least one known audio source is lower than a threshold energy. The means may furthermore be configured to determine at least one spatial metadata parameter associated with the at least one known audio source based on an analysis of the beamformed at least one known audio source. According to a third aspect there is provided an apparatus for a rendering of a six degrees-of-freedom audio scene, the apparatus comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the system at least to perform: providing at least two microphones located at respective position for audio capturing; identifying at least one known audio source being captured based on a beamforming by at least one microphone of the at least two microphones, the at least one audio source is located in the six degrees-of-freedom audio scene relative to the at least two microphones; identifying at least one unknown audio source in the six degrees-of- freedom audio scene, wherein the at least one unknown audio source at least partially interferes with the capturing of the at least one known audio source by the at least one microphone; and controlling the beamforming based on identifying at least one alternative microphone of the at least two microphones, the controlling of beamforming is such that the at least one alternative microphone starts capturing audio from the at least one known audio source while the at least one unknown audio source is at least partially located between the at least one microphone and the at least one known audio source. The at least two microphones may be higher-order Ambisonics microphones. The apparatus caused to perform controlling the beamforming based on identifying at least one alternative microphone of the at least two microphones may be caused to perform determining the at least one unknown audio source is masking the at least one known audio source with respect to the beamforming by at least one microphone of the at least two microphones. The apparatus caused to perform controlling the beamforming based on identifying at least one alternative microphone of the at least two microphones may further be caused to perform generating a signal, the signal indicating that the at least one alternative microphone starts capturing audio from the at least one known audio source while the at least one unknown audio source is at least partially located between the at least one microphone and the at least one known audio source. The apparatus caused to perform controlling the beamforming based on identifying at least one alternative microphone of the at least two microphones may further be caused to perform generating a signal, the signal associated with the at least one known audio source and indicating which the at least one microphone of the two microphones and the at least one alternative microphone is to be selected. The apparatus caused to perform controlling the beamforming based on identifying at least one alternative microphone of the at least two microphones may further be caused to perform switching from the one of the at least two microphones to the at least one alternative microphone. The apparatus caused to perform controlling the beamforming based on identifying at least one alternative microphone of the at least two microphones may be caused to perform obtaining a signal associated with the at least one known audio source, the signal indicating where the at least one microphone of the two microphones and the at least one alternative microphone is to be selected. The apparatus may be further caused to perform receiving a six degrees-of- freedom audio scene description wherein the six degrees-of-freedom audio scene description comprises the signal associated with the at least one known audio source. The apparatus caused to perform controlling the beamforming based on identifying at least one alternative microphone of the at least two microphones may be caused to perform activating one of: the at least one alternative microphone; or the at least one microphone, for capturing audio from the at least one known audio source and deactivating the other of the at least one alternative microphone or the at least one microphone. The apparatus caused to perform controlling the beamforming based on identifying at least one alternative microphone of the at least two microphones may be caused to perform controlling disabling or enabling of the at least one known source. The apparatus may be further caused to perform determining the at least one unknown audio source at least partially interferes with the capturing of the at least one known audio source by the at least one microphone. The apparatus caused to perform determining the at least one unknown audio source at least partially interferes with the capturing of the at least one known audio source by the at least one microphone may be caused to perform determining the position of the at least one unknown audio source is between or substantially between the at least one known audio source and the at least one of the at least two microphones. The apparatus caused to perform determining the at least one unknown audio source at least partially interferes with the capturing of the at least one known audio source by the at least one microphone may be caused to perform determining a distance between the at least one unknown audio source and the at least one known audio source is less than a threshold distance. The apparatus caused to perform determining the at least one unknown audio source at least partially interferes with the capturing of the at least one known audio source by the at least one microphone may be caused to perform determining an energy estimate of an energy of the at least one known audio source is lower than a threshold energy. The apparatus may be further caused to perform determining at least one spatial metadata parameter associated with the at least one known audio source based on an analysis of the beamformed at least one known audio source. According to a fourth aspect there is provided an apparatus for a rendering of a six degrees-of-freedom audio scene, the apparatus comprising: circuitry configured to determine respective positions for audio capturing for at least two microphones; identifying circuitry configured to identify at least one known audio source being captured based on a beamforming by at least one microphone of the at least two microphones, the at least one audio source is located in the six degrees-of-freedom audio scene relative to the at least two microphones; identifying circuitry configured to identify at least one unknown audio source in the six degrees-of-freedom audio scene, wherein the at least one unknown audio source at least partially interferes with the capturing of the at least one known audio source by the at least one microphone; and controlling circuitry configured to control the beamforming based on identifying at least one alternative microphone of the at least two microphones, the controlling of beamforming is such that the at least one alternative microphone starts capturing audio from the at least one known audio source while the at least one unknown audio source is at least partially located between the at least one microphone and the at least one known audio source. The at least two microphones may be higher-order Ambisonics microphones. According to a fifth aspect there is provided a computer program comprising instructions [or a computer readable medium comprising instructions] for causing an apparatus, for a rendering of a six degrees-of-freedom audio scene , the apparatus caused to perform at least the following: providing at least two microphones located at respective position for audio capturing; identifying at least one known audio source being captured based on a beamforming by at least one microphone of the at least two microphones, the at least one audio source is located in the six degrees-of-freedom audio scene relative to the at least two microphones; identifying at least one unknown audio source in the six degrees-of- freedom audio scene, wherein the at least one unknown audio source at least partially interferes with the capturing of the at least one known audio source by the at least one microphone; and controlling the beamforming based on identifying at least one alternative microphone of the at least two microphones, the controlling of beamforming is such that the at least one alternative microphone starts capturing audio from the at least one known audio source while the at least one unknown audio source is at least partially located between the at least one microphone and the at least one known audio source. The at least two microphones may be higher-order Ambisonics microphones. According to a sixth aspect there is provided a non-transitory computer readable medium comprising program instructions for causing an apparatus, for a rendering of a six degrees-of-freedom audio, to perform at least the following: providing at least two microphones located at respective position for audio capturing; identifying at least one known audio source being captured based on a beamforming by at least one microphone of the at least two microphones, the at least one audio source is located in the six degrees-of-freedom audio scene relative to the at least two microphones; identifying at least one unknown audio source in the six degrees-of-freedom audio scene, wherein the at least one unknown audio source at least partially interferes with the capturing of the at least one known audio source by the at least one microphone; and controlling the beamforming based on identifying at least one alternative microphone of the at least two microphones, the controlling of beamforming is such that the at least one alternative microphone starts capturing audio from the at least one known audio source while the at least one unknown audio source is at least partially located between the at least one microphone and the at least one known audio source. The at least two microphones may be higher-order Ambisonics microphones. According to a seventh aspect there is provided an apparatus, for a rendering of a six degrees-of-freedom audio scene, comprising: means for providing at least two microphones located at respective position for audio capturing; means for identifying at least one known audio source being captured based on a beamforming by at least one microphone of the at least two microphones, the at least one audio source is located in the six degrees-of-freedom audio scene relative to the at least two microphones; means for identifying at least one unknown audio source in the six degrees-of-freedom audio scene, wherein the at least one unknown audio source at least partially interferes with the capturing of the at least one known audio source by the at least one microphone; and means for controlling the beamforming based on identifying at least one alternative microphone of the at least two microphones, the controlling of beamforming is such that the at least one alternative microphone starts capturing audio from the at least one known audio source while the at least one unknown audio source is at least partially located between the at least one microphone and the at least one known audio source. The at least two microphones may be higher-order Ambisonics microphones. According to an eighth aspect there is provided a computer readable medium comprising instructions for causing an apparatus, for a rendering of a six degrees-of-freedom audio scene, to perform at least the following: providing at least two microphones located at respective position for audio capturing; identifying at least one known audio source being captured based on a beamforming by at least one microphone of the at least two microphones, the at least one audio source is located in the six degrees-of-freedom audio scene relative to the at least two microphones; identifying at least one unknown audio source in the six degrees-of- freedom audio scene, wherein the at least one unknown audio source at least partially interferes with the capturing of the at least one known audio source by the at least one microphone; and controlling the beamforming based on identifying at least one alternative microphone of the at least two microphones, the controlling of beamforming is such that the at least one alternative microphone starts capturing audio from the at least one known audio source while the at least one unknown audio source is at least partially located between the at least one microphone and the at least one known audio source. The at least two microphones may be higher-order Ambisonics microphones. An apparatus comprising means for performing the actions of the method as described above. An apparatus configured to perform the actions of the method as described above. A computer program comprising instructions for causing a computer to perform the method as described above. A computer program product stored on a medium may cause an apparatus to perform the method as described herein. An electronic device may comprise apparatus as described herein. A chipset may comprise apparatus as described herein. Embodiments of the present application aim to address problems associated with the state of the art. Summary of the Figures For a better understanding of the present application, reference will now be made by way of example to the accompanying drawings in which: Figure 1 shows an example audio scene; Figure 2 shows the example audio scene with example audio sources and beamforming used to provide the audio signals to implement source estimations; Figure 3 shows the example audio scene with example audio sources and a ‘obscured’ audio source where the ‘obscuring’ or ‘masking’ audio source prevents beamforming as shown in Figure 2 to be used to provide the audio signals to implement source estimations; Figure 4 shows the example audio scene as shown in Figure 3 where an alternative microphone beamforming is selected to overcome the ‘obscuring’ audio source from interfering with the beamforming as shown in Figure 2; Figure 5 shows a schematic apparatus suitable for implementing some embodiments; Figures 6a and 6b show the example audio scenes shown in Figures 2 and 3 having applied some embodiments; Figure 7 shows the example audio scene shown in Figure 4 having applied some embodiments; Figure 8 shows a system suitable for implementing some embodiments; Figure 9 shows an example beamforming HOA source chooser and pseudocode for implementing the choosing according to some embodiments; Figure 10 shows an example closest HOA source finder and obscuring source checker according to some embodiments; Figure 11 shows an example microphone switching for beamforming and EIF example switch information according to some embodiments; Figure 12 shows an example MPEG-I Audio bitstream formatting change according to some embodimdents; Figure 13 shows a further example microphone switching for beamforming and EIF example switch information according to some embodiments; and Figure 14 shows an example device suitable for implementing the apparatus shown in previous figures. Embodiments of the The following describes in further detail suitable apparatus and possible mechanisms for implementing rendering. As discussed above to achieve this, the renderer can be configured to perform beamforming (from a FOA / HOA signal) towards a known source to estimate the source properties (energy at different frequency bands). Based on this and the position of the known source, the renderer can be configured to determine from which direction, from the listener’s perspective, and in which frequency bands, how much energy is being contributed by the known source. For example Figure 2 shows an example scene such as shown in Figure 1 with two known audio sources s1201 and s2203 at positions ps1 and ps2. Further is shown in Figure 2 that there is a beamforming 204 from the microphone 101 m1to a source s2203 at position ps2 as well as beamforming 202 from the microphone m2 103 to source s1201 at position ps1 in order to estimate the source properties. An example scene is a recording of a couple of musicians playing on a busy street. There are people around the musicians as well as people walking on the street. Some people on the street are listening and contributing to the audio scene by applauding or talking. The song is recorded using multiple Ambisonics microphones, placed some distance apart. The positions of at least some of the musicians are known. The scene is captured and then subsequently turned into a VR or more generally a spatial audio scene by a content creator. The content creator creates a scene description with the positions of the recording microphones and any of the known sources. The scene description is then encoded into a bitstream along with the audio and is provided to the listener for consumption, for example as part of a VR or spatial audio application. There can be a problem, in the estimation of audio source spectral properties in order to perform 6DoF rendering of scenes comprising multiple HOA sources, if the HOA source and the audio source whose position information is provided to the renderer without considering the presence of other audio sources between a HOA source and an informed source (whose rendering is being improved selectively). In other words an estimation of the audio source spectral properties for an audio source can be inaccurate where there is an occluding (obscuring or intermediate or masking) audio source between the microphone and the audio source. Figure 3 shows the audio scene shown previously in Figures 1 and 2 when an example situation where an unknown source sx303 at position psx is situated between the beamforming FOA / HOA microphone 103 and the known source s1201. This unknown source sx 303 will cause the estimated properties of the known source s1201 be erroneous. For the listener this will likely cause a lower quality listening experience as localization is adversely affected. Performing informed mode rendering (in other words rendering when there are active informed sources present in the scene) requires higher computational complexity. Thus, obscured informed sources not only likely produce a lower quality rendering but will likely also consume additional computational resources. Such a situation should be avoided. Furthermore, some of the sources in the scene may be intermittent, or in other words produce audio intermittently. For example, a person talking in the scene may not be talking all of the time (or in other words the speech may comprise pauses, some of which may be long pauses). In these situations where the sources are intermittent, the renderer is expending computation (for estimating audio source information) even when the source is not generating any sound. As such a further aspect as discussed in the embodiments herein is the apparatus and methods for providing a mechanism for switching off ‘informed’ source rendering for informed sources for a determined period is needed. In other words, a method is required to perform ‘informed’ source rendering only when necessary to avoid any penalty in unnecessary computations. Furthermore, such a method also prevents any undesired effects from the beam forming performed in the direction of ‘informed’ source when it is not emitting audio. In the example scene therefore the embodiments attempt to produce better performance rendering, for example based on audio quality and computational requirements, in situations such as described above when the musicians move or when people on the busy street walk (while talking, for example) between the capture microphones and the musicians. In this example the musicians’ (musician audio sources) positions are known and beamforming is employed directing a ‘beam’ towards them from the microphones and aim to avoid or mitigate the issues mentioned above. Additionally in some embodiments the content creator would like to be able to control the beamforming when creating VR / AR scenes out of the content. Thus the embodiments can relate to rendering 6DoF audio scenes captured with multiple higher-order Ambisonics microphones (HOA). In these embodiments there can be provided a method for controlling beamforming from the captured microphone array data towards audio sources in the scene to achieve improved rendering quality even when the target of beamforming for a microphone array is obscured. This can be achieved in some embodiments by switching the beamforming microphone array based on content creator instructions or a determination based on the known audio source positions. In summary, in some embodiments, the content creator can be configured to control the beamforming towards informed source positions (or positions of known audio sources in the scene). This in some embodiments be implemented by generating and adding control information to the scene description. The control information can be used for controlling which HOA sources are used for beamforming towards which informed source as well as enabling / disabling informed source rendering for any of the defined informed sources. The implementation of selective source rendering according to the embodiments disclosed herein can increase (audio) quality as the situations described above are avoided. That is the use of the control information enables the beamforming operations to be controlled when a beamform target audio source is obscured by another audio source. In some embodiments the implementation alternatives / use cases / contexts can be as follows: A content creator listens to the scene in the content creation phase and indicates in the scene description when to switch beamforming microphones or disable beamforming towards a known audio source. (This can be understood to be a ‘fully manual’ mode of operation); the encoder, during the bitstream creation determines if beamforming is obscured. This can be implemented based on the positions of the known sources in the scene. When a known source is determined to be between a beamforming microphone and the beamforming target, the encoder adds into the bitstream information about switching beamforming microphones or disabling / enabling beamforming for a known source. (An ‘automatic’ encoder mode); the renderer determines that a known source is between a beamforming microphone and the beamforming target, and switches beamforming microphones or disables / enables beamforming for a known source. (An ‘automatic’ renderer mode); a combination of the manual, automatic encoder and automatic render mode. Thus, for example, Figure 4 shows an example of the scene as shown in Figures 1 to 3 where the beamforming 401 is implemented from the HOA microphone ^^105 located at ^^rather than beamforming from the HOA microphone ^^103 located at ^^because of the audio source sx303 located between the microphone 103 and the known source s1201. With respect to Figure 5 is shown a schematic view of an example renderer capable of implementing MPEG-I Immersive Audio system for 6DoF Multi-Point HOA rendering of audio scenes comprising HOA / FOA content (recorded or synthesized). This is an example implementation of rendering such a scene, and in some embodiments it would be possible to make some modifications to this example implementation. In some embodiments the renderer comprises a pre-processor 501. The pre- processor 501 is configured to receive information on the information of the scene (such as the HOA / FOA sources at positions ^^..^^504) and based on these performs Delaunay triangulation to provide a set of triangles ^^..^^512 that partition the scene into triangle sections. These triangulations are shown as the example triangulations in Figure 1 with references 151 (with vertices defined by microphones 101, 105, 109), 153 (with vertices defined by microphones 107, 105, 109), 155 (with vertices defined by microphones 101, 105, 103), 157 (with vertices defined by microphones 107, 105, 103). The pre-processor 501 furthermore is configured to receive head-related impulse responses (HRIRs) 500 which are converted to frequency domain head- related transfer functions HRTF 528 using a short-term Fourier transform. In some embodiments (for example the MPEG-I Immersive Audio case), an alias-free STFT algorithm is employed. The HRTFs 528 in some embodiments can also be used to calculate an Ambisonics-to-binaural transform matrix ^^^^^^^^(^) , for each frequency band ^. In some embodiments the renderer comprises a position pre-processor 505 is configured to receive the listener or renderer position and the source positions to determine which HOA / FOA sources that are close to the listener for later processing purposes. For example during rendering, for each input frame ^ , the position pre- processor 505 is configured to take as an input the listener position 506 ^^, the HOA / FOA source positions 504 ^^..^^and the set of triangles 512 ^^..^^created in the pre-processor 501. Then based on the input, the position pre-processor 505 determines interpolation weights 518 ^^(^, ^) for the HOA / FOA source positions 504 ^^. In some embodiments the pre-processor is configured to determine in which triangle ^^, the listener is in and then calculating the barycentric coordinates for the triangle ^^and the listener position. The barycentric coordinates can then be used as the weights. The weights can sum to one and the closer the listener is to an HOA / FOA source, the higher the weight will be. At an edge of a triangle, the weight for the HOA / FOA source that is NOT part of that edge is 0. The weights 518 ^^(^, ^) and the triangle that the listener is in (active triangle) 514 ^^(^) are provided as output. In some embodiments when the listener moves from a triangle to another, the switching of the active triangle may be delayed, and only switched after a few frames of audio. This is because processing of the audio for HOA / FOA sources for the new triangle takes a few frames of audio to provide meaningful output. Thus, the active triangle is not always the triangle that the listener is in. The position pre- processor in some embodiments is also configured to determine the HOA / FOA source that is closest to the listener. In some embodiments the renderer comprises a spatial analyser 503. The spatial analyser 503 can be configured to provide spatial metadata parameters 520 for frames of HOA / FOA signals that positioned near the listener. These are later used, in the metadata interpolation block, to estimate spatial metadata parameters at the listener position. The spatial analyzer 503 this is configured to receive at least one frame (for example, 256 samples) of HOA / FOA source signals 502 ^^^^(^, ^) for each source ^ and frame ^ and the current active triangle 514 ^^(^). For each HOA / FOA source ^, belonging to the triangle ^^(^), the spatial analyzer 503 can be configured to perform STFT processing to obtain time-frequency domain signals ^(^, ^, ^) 516, where ^ refers to a sub-frame (for example, 128 samples). The spatial analyser 503 can furthermore be configured to calculate spatial metadata 520 from the time-frequency domain signals 502. The spatial metadata 520 comprises the energy e(i,j,k,b), direction information (azimuth ^(i,j,b) and elevation ^(i,j,b)) and diffuseness information (direct-to-total energy ratio r(i,j,b)) and is obtained as follows: First, a signal analysis vector ^, ^)is calculated: where ^^,^(^, ^, ^)is the value in matrix corresponding to channel ^ and frequency bin ^. From the signal analysis vector, a signal intensity vector is calculated: where^^^^^(^^^,^^^^,^^^^,^^^^)denotes the complex conjugate of ^^(^, ^, ^, ^).Signal energy is then calculated as follows: An average intensity vector and average energy is calculated as follows: The direction data (azimuth and elevation) are calculated as follows: where ^̃^(^, ^, ^)is the nth element of the average intensity vector ^), ^^(^, ^, ^)is the azimuth and ^^(^, ^, ^)is the elevation. The direct-to-total energy ration is calculated as follows: The energy for the subframes ^ are then obtained as follows: ^(^, ^, ^, ^)= ^̃(^, ^, ^), ^ ∈ 1.. ^^^In some embodiments the renderer comprises a spatial metadata interpolation 507. The spatial metadata interpolator 507 can be configured to provide an estimate of the spatial metadata at the listener position based on the spatial metadata for the HOA / FOA sources belonging to the triangle that the listener is in, and the interpolation weights calculated in the position pre-processing block. First, the spatial metadata 522 is converted into vector form: The vectors are rotated according to the listener’s head oi 508 and source ol 510 orientations: An interpolated spatial metadata vector is then calculated by a weighted average of spatial metadata vectors: And finally, the interpolated spatial metadata vector is converted to spatial metadata parameters as follows: Azimuth: ^ ^, ^))^ Direct-to-total energy ratio: In some embodiments the renderer comprises a signal Interpolator 509. The signal interpolator is configured to provide an interpolated signal 528 at the listener position, that is, an estimate of the signal at the listener position. This is used later (by the mixer 511) in conjunction with the interpolated metadata at the listener position to provide the final binaural output. The interpolated signal ^^^,^(^, ^)528 can in some embodiments be obtained by taking the HOA / FOA signal 516 that is closest to the listener and applying an EQ on the signal: ^^^,^(^, ^) = ^^^(^, ^, ^)^^,^(^^(^), ^, ^) where ^^(^) is index of the HOA source chosen for interpolation (closest to listener in most cases) and: The renderer in some embodiments comprises a mixer 511 The mixer 511 in some embodiments takes as input the interpolated spatial metadata 522 at the listener position and the interpolated signal at the listener position and provides the binaural time-frequency domain output signal. The mixer in some embodiments is configured to create a signal covariance matrix from the interpolated spatial metadata 522 which describes the desired (or target) spatial characteristics of the signal at the listener position. An optimal mixing algorithm is then used to obtain a mixing matrix that when multiplied with the interpolated signal, the resulting signal is inline with the desired spatial characteristics. First, a binaural prototype signal is created from the interpolated signal: Where ^^^(^) is a rotation matrix taking into account the listener and HOA / FOA source rotation and ^^^^^^^^(^)is the Ambisonics to binaural matrix. Then, a signal covariance matrix ^^is calculated for the prototype signal: Recursive averaging is applied to get the signal covariance matrix for frame ^: ^^(^, ^) = (1 − ^)^^^^^(^, ^) + ^^^(^ − 1, ^) Where ^ = 0.9. Next, a signal covariance matrix ^^is calculated from the interpolated spatial metadata at the listener position. First the direct portion of ^^is calculated: where ^(^, ^)refers to the HRTF value at frequency bin ^, in direction. Second, the diffuse portion of ^^is calculated. ^^is then obtained as follows: ^^^^(^, ^) =^^^^^^( ) ^^^^^^^ ^ ^^^, ^ + ^^(^, ^) And recursive averaging: The signal covariance matrices are then used to obtain mixing matrices 530 which are used to obtain the binaural output like so: Where ^(^, ^)is a decorrelated time-frequency domain signal obtained from a buffer of previous binaural signals ^. The matrices ^^(^, ^, ^)and ^^^(^, ^, ^)are obtained from the optimal mixing procedure outlined in Vilkamo, J., Bäckström, T., & Kuntz, A. (2013). Optimized Covariance Domain Framework for TimeFrequency Processing of Spatial Audio. Journal of the Audio Engineering Society, 61(6), 403-411. After applying the mixing matrices on the binaural prototype signal ^, the result output 530 ^ has the spatial characteristics of the spatial metadata at the listener position. In some embodiments the renderer comprises an output 513. The output 513 takes the obtained binaural output signal 530 ^(^, ^, ^)and performs inverse STFT on it to produce the final time domain binaural output signal 532. As mentioned earlier, information about known audio sources in the scene may be used to improve rendering quality. The improvement comes via being able to calculate a more accurate target signal covariance matrix ^^via the use of beamforming to estimate the characteristics of the known audio sources. Note that no audio signal is required from the known audio sources, only their positions. This processing as described herein is further described in GB patent publication 2020239.6. In some embodiments the informed source information is provided by the content creator in the EIF as follows. <InformedSource> An informed source description for MPHOA rendering. The informed sources represent the known audio sources within the scene. Attribute Type Flags Default Description id ID R Identifier position Position R, M Position listenerThreshold Value O 3.0 Distance from the listener to the informed source in meters where the informed source is active.0.0 distance disables the threshold. If not defined in EIF, 3.0 will be used. hoaSourceThresholdValue O 3.0 Distance from the HOASources to the informed source in meters where the informed source is active. 0.0 distance disables the threshold. If not defined in EIF, 3.0 will be used. priorityValue Integer O 0 Determines the processing priority between informed sources that are close to each other. In some examples the EIF information is encoded into the bitstream using the following syntax: Syntax No. of Mnemonic bits InformedSourceInfoStruct() { informedSourceId = GetID(); informedSourcePositionX; 31 uimsbf informedSourcePositionY; 31 uimsbf informedSourcePositionZ; 31 uimsbf listenerThresholdPresent; 1 bslbf hoaSourceThresholdPresent; 1 bslbf if (listenerThresholdPresent == True) { lpdInformedSourceEnableThreshold; 16 uimsbf } if (hoaSourceThresholdPresent == True) { hpdInformedSourceEnableThreshold; 16 uimsbf } priorityValue; 6 uimsbf } In some embodiments the pre-processor 501 is further configured to implement the following pre-processing In addition to the pre-processing described in the previous section the following is performed when informed sources are present: First vectors from each informed source position ^^^to HOA / FOA source position ^^is determined, ^(^, ^)= ^^^− ^^which are then rotated according to the orientation of the HOA / FOA source: ^^^^(^, ^)= ^^^^(^)^(^, ^) where ^^^^(^) is the rotation matrix for according to the orientation of HOA / FOA source ^. The time-delay (in samples) between the HOA / FOA source / informed source pairs is calculated as follows: where ^ is the speed of sound and ^^is the sampling frequency. The time-delay between two HOA sources for an informed source is obtained with: The delay in short-time Fourier transform (STFT) hop sizes shall be obtained with: ^^^^(^^, ^^, ^)= round where ^^^^is the set hop length. In addition, for each informed source, the closest HOA / FOA source ^^^^^^^^(^) is determined based on the lengths of ^(^, ^). The attenuation gain between ^^^^^^^^(^)and some other HOA source ^ for informed source ^ is obtained with: In addition, beamforming weights ^(^^^^^^^^(^), ^) that focus a beam pattern from the closest HOA / FOA source towards the informed source ^ are calculated: where ^(^^^^^^^^(^), ^) is the set of real spherical harmonic functions for the pair (^^^^^^^^(^), ^) for all the orders ^ ∈ and degrees ^ ∈ [−^, ^]. Real spherical harmonics defined as: 0 for each HOA source ^ and informed source ^ pairs, the outer product of ^(^, ^)and ^(^, ^)^is also calculated up to order 1: ^^(^, ^)= ^(^, ^)^(^, ^)^, ^^^ ^ ≤ 1 Additionally the position pre-processor 503 in some embodiments is further configured to determine from the informed sources which are the active informed sources. An Informed source is active when the certain conditions hold: 1. The listener is closer than a set threshold (lpdInformedSourceEnableThreshold in the bitstream) to an informed source 2. The closest HOA source to an informed source is closer than a set threshold (hpdInformedSourceEnableThreshold in the bitstream) In the spatial analyser 503 the spatial analysis is performed for HOA / FOA sources that belong to the triangle that the listener is in. In addition, the spatial analyser 503 is also performed for HOA / FOA sources which are the closest sources to active informed sources. In some embodiments for each informed source, a signal is obtained via beamforming. The beamformed signals are post-filtered and time-aligned at the HOA source positions and at the listener position. For each active informed source for the current frame, beamforming is performed as follows: The obtained signal is post-filtered: where a cross-pattern coherence (CroPaC) based post-filter is used. Time-alignment at HOA source positions is done as follows: ^^^^^(^, ^, ^, ^, ^)= ^^^^^(^(^, ^, ^), ^, ^, ^)where Time-alignment at listener position: ^^^.^^^(^, ^, ^, ^) = ^^^^^^^(^, ^^,^^^^^^^, ^), ^, ^, ^^ Furthermore the spatial analysis is configured to implement analysis which is configured to generate spatial metadata parameters which have contribution from the informed sources only for the active HOA / FOA sources. The signal covariance matrix from the time-aligned beamformed signals ^^^^^(^, ^, ^, ^, ^)for the active HOA / FOA sources can be determined as: The signal covariance matrix from the listener towards a source is obtained as follows: where ^^^^^^^,^^^(^, ^), ^^ is the HRTF from the listener towards the informed source ^ and is an attenuation gain and ^^^,^^^(^, ^, ^) is the energy of the time-aligned signal ^^^,^^^(^, ^, ^). Now, ^^is the signal covariance matrix at the listener position with contribution only from the informed sources. For the output the contribution for all other sources present in the scene is to be determined. Note that the signal covariance matrix ^^as calculated in the previous section as that has the contribution from all sources in the scene (informed and non-informed). This was calculated based on interpolated spatial metadata at the listener position. Thus, it would be beneficial to be able to calculate interpolated spatial metadata at the listener position without the contribution of the informed sources. This is done as follows: First, calculate intensity vectors and energy at the HOA / FOA sources from the signal covariance matrices ^^^^(^, ^, ^)(informed source contribution at the Similarly, intensity ^^^^(^, ^, ^, ^)and energy ^^^^(^, ^, ^, ^)are calculated for the HOA / FOA signals. These contain the contribution from all sources in the scene. Next, we obtain intensity and energy which has only the contribution from non- informed sources by subtraction: ^^^^(^, ^, ^, ^)= ^^^^(^, ^, ^, ^)− ^^^^(^, ^, ^, ^)^^^^(^, ^, ^, ^)= ^^^^(^, ^, ^, ^)− ^^^^(^, ^, ^, ^)From these, spatial metadata at HOA / FOA source positions are obtained that have contribution only from the non-informed sources: Azimuth: Direct-to-total energy ratio: These are then interpolated in the same manner as in the metadata interpolation block described in the previous section resulting in ^^^^,^^^^^(^, ^, ^), ^^^^,^^^^^(^, ^, ^), ^^^^,^^^^^(^, ^, ^) and ^^^^,^^^^^(^, ^, ^). A signal covariance matrix can then be calculated from the interpolated metadata in the same manner as ^^was calculated previously. This generates a signal covariance matrix ^^^^with contribution from only the non-informed sources. The final target covariance matrix is then obtained as follows: ^^(^, ^)= ^^(^, ^)+ ^^^^(^, ^)The rest of the processing follows the methods as described above. As was described above, in the current system, beamforming towards an informed source m is done from the HOA / FOA source ^^^^^^^^(^), that is, the closest HOA / FOA source to the informed source. Furthermore, whether or not an informed source is active is controlled by the listener and HOA / FOA source distance thresholds: If the informed source is too far away from either a lister or the closest HOA / FOA source, the informed source is set not active. In addition to these controls for informed mode rendering, there is implemented two new controllable features for informed mode rendering. These can be implemented separately in some embodiments and can be summarized as control from which HOA / FOA source beamforming is done towards an informed source and control of the active state of informed sources. The control from which HOA / FOA source beamforming is done towards an informed source can in some embodiments be determined by a content creator implementation which defines this manually based on a listening of the scene in the content creation phase, and / or a selection of HOA / FOA source for beamforming may be time-dependent and / or the detection of when to switch beamforming HOA / FOA source may be done automatically. The control from which the active state of informed sources can be implemented in the following manner: informed sources may be switched off by the content creator in a situation where he thinks they are not needed to be rendered by the informed mode rendering (when quiet, for example); and / or can be a time-dependent modification (sometimes active, sometimes not active); and / or can be an automatic detection of when to activate / deactivate informed sources is possible. Thus, as shown in Figures 6a and 6b which shows that the information 601 provided enables the selection of the microphone m2103 where there is no obscuring source sx 303 and selection of microphone m3105 where there is an obscuring source sx 303. With respect to Figure 7 the automatic switching is shown where there is a beamforming 401 from the microphone 105 to the source s1201 and a beamforming 202 from the microphone m2103 to the source s1201. In this example instead of adding the information about changing the beamforming HOA source manually into the bitstream signaling, the determination that the beamforming HOA source needs to be changed is implemented in the renderer at runtime. For example, the case where a first informed source sx 303 moves between the HOA microphone m2103 source used for beamforming 202 and a second informed source s1201, the microphone and beamforming for the second informed source 201 is changed to the second closest one, the microphone m3105. tIn normal operation, the renderer selects the closest HOA source for each informed source for beamforming. In the proposed system, the automatically (default selection by the renderer without this invention) selected HOA source can be overridden based on the new proposed information in the bitstream (as part of this invention). Furthermore, the HOA source may be changed over time during rendering via the MPEG-I Immersive Audio update mechanism (timed updates, user position dependent updates or dynamic updates). In some embodiments the selection of the HOA / FOA source from which beamforming is implemented is controlled via the bitstream and determining which informed sources are active may be controlled by the content creator via the bitstream. This is shown in Figure 11. The MPEG-I Audio Encoder Input Format defines an <InformedSource> element, which controls the informed mode rendering. In this embodiment, two new attributes are added to the <InformedSource> element, beamformHOASource and active. <InformedSource> An informed source description for MPHOA rendering. The informed sources represent the known audio sources within the scene. Attribute Type Flags Default Description id ID R Identifier position Position R, M Position listenerThreshold Value O 3.0 Distance from the listener to the informed source in meters where the informed source is active. 0.0 distance disables the threshold. If not defined in EIF, 3.0 will be used. hoaSourceThreshold Value O 3.0 Distance from the HOASources to the informed source in meters where the informed source is active. 0.0 distance disables the threshold. If not defined in EIF, 3.0 will be used. priorityValue Integer O 0 Determines the processing priority between informed sources that are close to each other. beamformHOASourceID O, M Determines which HOA source is used for beamforming for this informed source. active bool O, M true Determines if this informed source is active. The updated bitstream syntax looks like the syntax uses in Figure 12 and represented below: Syntax No. of Mnemonic bits InformedSourceInfoStruct() { informedSourceId = GetID(); informedSourcePositionX; 31 uimsbf informedSourcePositionY; 31 uimsbf informedSourcePositionZ; 31 uimsbf listenerThresholdPresent; 1 bslbf hoaSourceThresholdPresent; 1 bslbf if (listenerThresholdPresent == True) { lpdInformedSourceEnableThreshold; 16 uimsbf } if (hoaSourceThresholdPresent == True) { hpdInformedSourceEnableThreshold; 16 uimsbf } priorityValue; 6 uimsbf beamformHOASourcepresent; 1 bslbf if (beamformHOASourcepresent == True) { beamformHOASource = GetID(); } informedSourceActive; 1 bslbf } When beamformHOASource is set, the defined HOA source is used for beamforming instead of the closest HOA source. In the description above, ^^^^^^^^(^)is not determined by distance between the informed source and the HOA sources, but by the value indicated by the beamformHOASource, i.e. the HOA source identifier, in the bitstream. When informedSourceActive is set to true, informed source processing may be performed for this informed source, if the other conditions for informed source rendering are met (defined by lpdInformedSourceEnableThreshold, hpdInformedSourceEnableThreshold and priorityValue). When informedSourceActive is set to false, informed source processing is not performed for this informed source. Once the information is read from the bitstream, the renderer then performs informed mode rendering based on this information. The selection of the beamforming microphone is implemented (according to bitstream information), depending on implementation, in the pre-processor 501 or the position pre- processor 505. If the beamforming microphone is allowed to be changed during runtime, the process is implemented in the position pre-processing block as this block is processed for every audio frame. If the beamforming HOA sources are set only at initialization, the processing may be done in the pre-processor 501. In some embodiments the detection of when to switch the HOA / FOA source for beamforming and the active state of the informed source may be implemented automatically, during rendering. For example in some embodiments there can be a switch in beamforming HOA / FOA source when a second informed source moves between the beamforming source and the informed source. In some embodiments, the second informed source causes a switch in beamforming HOA / FOA source only if it is an active informed source. In other embodiments the listener and HOA thresholds are not exceeded and informedSourceActive is set to true. The positions of the informed sources (^^^) are known at rendering time. When an informed source is active, the system monitors all other informed sources for their position, if the position of any informed source moves within a threshold distance from the line connecting the informed source and the HOA / FOA source used for beamforming for that source (^(^^^^^^^^(^), ^)), the system switches the HOA / FOA source used for beamforming. This can in some embodiments be implemented by finding the HOA / FOA source that is the next closest to the informed source. In some situations the second closest HOA / FOA source may also be obscured by another informed source or there is no suitable HOA / FOA source found. In such embodiments, the informed source is set not active. In some embodiments there is an immediate switch of an informed source to not active when it is obscured (in other words no switching of beamforming HOA source). In some embodiments the selection of beamforming microphone is implemented in a Beamforming HOA source chooser, which can be located in the position pre-processor 503. An example beamforming HOA source chooser can be shown in Figure 9. The beamforming HOA source chooser 901 can be configured to receive the positions of the informed sources 900 as well as the HOA / FOA source positions 902. As output it provides a list of informed source / HOA source pairs 904 that describe which HOA / FOA source is used for which informed source. The following pseudo-code describes the processing for finding a HOA / FOA source (chosen_mic) to be used for a given informed source (p_s): choose_beamform_array(p_s, p_m_1..m_N, p_s_1..s_N) min_dist = LARGE_NUMBER; chosen_mic = None; while j < m_N: if distance(p_s, p_m_j) < min_dist: if !obscured(p_s, p_m_j, p_s_1..s_N): min_dist = distance(p_s, p_m_j); chosen_mic = p_m_j; break; return chosen_mic; The distance function returns the Euclidian distance between the two positions. The obscured function describes whether the informed source (p_s) is obscured by any of the other informed sources (p_s_1..s_N) if being beamformed to from HOA(FOA source p_m_j). The obscuring check is done by comparing the angle of vectors drawn from the HOA / FOA source towards the informed source (p_s) and all other informed source and checking for the angle difference between them. If the angle difference for any of the other informed sources is below a threshold, the obscured function returns a true value. With respect to Figure 10 an alternative source selector configuration is shown where the informed sources are deemed inactive if they are obscured. In these embodiments there is no alternate HOA / FOA source searched for in the case of an informed source being obscured. The source selector in some embodiments comprises closest HOA source finder 1001, which can be located in the position pre-processor 503. The closest HOA source finder 1001 is configured to receive the positions of the informed sources 900 as well as the HOA / FOA source positions 902. As an output it provides a closest informed source / HOA source 1004 that describes which HOA / FOA source is used for which informed source. Furthermore is shown an obscuring checker 1003 which receives a closest HOA 1004 and further the informed sources 900 as well as the HOA / FOA source positions 902 and from this determine whether the informed source is active or deactivated once a threshold distance is used to check when informed sources are close to each other. In the case they are too close, both informed sources are set not active. In these embodiments the obscure function checks for distance differences between informed sources instead of angle differences as described above. With respect to Figure 13 is shown the further example encoder input format <InformedSource> An informed source description for MPHOA rendering. The informed sources represent the known audio sources within the scene. Attribute Type Flags Default Description id ID R Identifier position Position R, M Position listenerThreshold Value O 3.0 Distance from the listener to the informed source in meters where the informed source is active. 0.0 distance disables the threshold. If not defined in EIF, 3.0 will be used. hoaSourceThresholdValue O, M 3.0 Distance from the HOASources to the informed source in meters where the informed source is active. 0.0 distance disables the threshold. If not defined in EIF, 3.0 will be used. priorityValue Integer O 0 Determines the processing priority between informed sources that are close to each other. In some embodiments after beamforming, when estimating, the informed source energy (^^^^^(^, ^, ^, ^)), if energy in all frequency bands is low, abort further informed source processing for this informed source to lower computational complexity. Informed source processing requires more computations than regular processing without informed sources. Each separate active informed source increases the computational complexity somewhat, thus it is beneficial to save on the computation, when an informed source is quiet, and no quality improvements are to be gained from informed source processing. Thus as shown in Figure 8 there shows an example implementation of some embodiments in a MPEG-I Immersive audio context. The content creator 801 in some embodiments is configured to receive a recording of the scene in the form of audio data 802, the creation of the scene description (EIF) shown by the audio scene description 800, the encoding of the audio data and the scene description within the MPEG-I encoder 803 generating 6DoF audio bitstream 804 and a MPEG- H encoder 805 configured to generate MPEG-H 3DA (ISO / IEC 23008-3) audio bitstream 806. Once the scene has been encoded, it is stored in a server 811 for retrieval by users for listening. The server 811 can comprise bitstream storage 813 which receives the 6DoF audio bitstream 804 and a MPEG-H 3DA encoder 805 configured to generate MPEG-H 3DA audio bitstream 806 and output a 6DoF audio bitstream 804 and a MPEG-H 3DA encoder 805 configured to generate 6DoF audio bitstream 814. The player 821 functionality includes fetching of the scene data, decoding it and then rendering to the listener. In some implementation embodiments, the beamformHOAsource may be selected such that beamforming needs to be performed for as few HOA sources as possible without any significant adverse impact on the rendering quality. The player 821 in some embodiments comprise a playback device 823. The playback device 823 comprises a bitstream parser 827 which generates an audio and metadata from the 6DoF audio bitstream 814. The playback device 823 comprises a MPEG-I audio renderer 829 which receives 6DoF tracking 824 from a head mounted device 825 and further generate audio 826 to be passed to the head mounted device 825. With respect to Figure 14 an example electronic device which may be used as any of the apparatus parts of the system as described above. The device may be any suitable electronics device or apparatus. For example in some embodiments the device 2000 is a mobile device, user equipment, tablet computer, computer, audio playback apparatus, etc. The device may for example be configured to implement the encoder or the renderer or any functional block as described above. In some embodiments the device 2000 comprises at least one processor or central processing unit 2007. The processor 2007 can be configured to execute various program codes such as the methods such as described herein. In some embodiments the device 2000 comprises a memory 2011. In some embodiments the at least one processor 2007 is coupled to the memory 2011. The memory 2011 can be any suitable storage means. In some embodiments the memory 2011 comprises a program code section for storing program codes implementable upon the processor 2007. Furthermore in some embodiments the memory 2011 can further comprise a stored data section for storing data, for example data that has been processed or to be processed in accordance with the embodiments as described herein. The implemented program code stored within the program code section and the data stored within the stored data section can be retrieved by the processor 2007 whenever needed via the memory-processor coupling. In some embodiments the device 2000 comprises a user interface 2005. The user interface 2005 can be coupled in some embodiments to the processor 2007. In some embodiments the processor 2007 can control the operation of the user interface 2005 and receive inputs from the user interface 2005. In some embodiments the user interface 2005 can enable a user to input commands to the device 2000, for example via a keypad. In some embodiments the user interface 2005 can enable the user to obtain information from the device 2000. For example the user interface 2005 may comprise a display configured to display information from the device 2000 to the user. The user interface 2005 can in some embodiments comprise a touch screen or touch interface capable of both enabling information to be entered to the device 2000 and further displaying information to the user of the device 2000. In some embodiments the user interface 2005 may be the user interface for communicating. In some embodiments the device 2000 comprises an input / output port 2009. The input / output port 2009 in some embodiments comprises a transceiver. The transceiver in such embodiments can be coupled to the processor 2007 and configured to enable a communication with other apparatus or electronic devices, for example via a wireless communications network. The transceiver or any suitable transceiver or transmitter and / or receiver means can in some embodiments be configured to communicate with other electronic devices or apparatus via a wire or wired coupling. The transceiver can communicate with further apparatus by any suitable known communications protocol. For example in some embodiments the transceiver can use a suitable universal mobile telecommunications system (UMTS) protocol, a wireless local area network (WLAN) protocol such as for example IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth, or infrared data communication pathway (IRDA). The input / output port 2009 may be configured to receive the signals. In some embodiments the device 2000 may be employed as at least part of the renderer. The input / output port 2009 may be coupled to headphones (which may be a headtracked or a non-tracked headphones) or similar. Thus in summary the embodiments as described above show: A normative bitstream comprising: Information specifying the triggers and the guidance parameters for dynamically modifying the reverb predelay parameter; a bitstream description the parameters for which the renderer is expected to react (e.g., lower order early reflections, complexity or network bottleneck, etc.) and modifying the reverberation rendering dynamically based on the triggers. Additionally in some embodiments the normative bitstream comprises trigger and predelay modification parameters described using the syntax described herein. The bitstream in some embodiments is streamed to end-user devices or made available for download or stored. In some embodiments the normative renderer is configured to decode the bitstream to obtain the scene, reverberation parameters and dynamic reverb adjustment parameters and perform the modification to reverberator parameters as described herein. Moreover in some embodiments the renderer is configured to implement reverberation and early reflections rendering. In some embodiments the complete normative renderer can also obtain other parameters from the bitstream related to room acoustics and sound source properties, and use them to render the direct sound, diffraction, sound source spatial extent or width, and other acoustic effects in addition to diffuse late reverberation and early reflections. Thus in summary the concept is on in which there is the capacity for dynamic modification of rendering of reverberation based on the various triggers specified in the bitstream to enable bitrate and computational scalability based on suboptimal early reflections or other missing acoustic effects. In general, the various embodiments of the invention may be implemented in hardware or special purpose circuits, software, logic or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the invention is not limited thereto. While various aspects of the invention may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof. The embodiments of this invention may be implemented by computer software executable by a data processor of the mobile device, such as in the processor entity, or by hardware, or by a combination of software and hardware. Further in this regard it should be noted that any blocks of the logic flow as in the Figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media such as hard disk or floppy disks, and optical media such as for example DVD and the data variants thereof, CD. The memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The data processors may be of any type suitable to the local technical environment, and may include one or more of general-purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASIC), gate level circuits and processors based on multi-core processor architecture, as non-limiting examples. Embodiments of the inventions may be practiced in various components such as integrated circuit modules. The design of integrated circuits is by and large a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate. Programs, such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design, of San Jose, California automatically route conductors and locate components on a semiconductor chip using well established rules of design as well as libraries of pre-stored design modules. Once the design for a semiconductor circuit has been completed, the resultant design, in a standardized electronic format (e.g., Opus, GDSII, or the like) may be transmitted to a semiconductor fabrication facility or "fab" for fabrication. As used in this application, the term “circuitry” may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry) and (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions) and I hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation. This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device. The term “non-transitory,” as used herein, is a limitation of the medium itself (i.e., tangible, not a signal ) as opposed to a limitation on data storage persistency (e.g., RAM vs. ROM). As used herein, “at least one of the following: ” and “at least one of ” and similar wording, where the list of two or more elements are joined by “and” or “or”, mean at least any one of the elements, or at least any two or more of the elements, or at least all the elements The foregoing description has provided by way of exemplary and non- limiting examples a full and informative description of the exemplary embodiment of this invention. However, various modifications and adaptations may become apparent to those skilled in the relevant arts in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this invention will still fall within the scope of this invention as defined in the appended claims.
Claims
CLAIMS:
1. A method for a rendering of a six degrees-of-freedom audio scene, the method comprising: providing at least two microphones located at respective position for audio capturing; identifying at least one known audio source being captured based on a beamforming by at least one microphone of the at least two microphones, the at least one audio source is located in the six degrees-of-freedom audio scene relative to the at least two microphones; identifying at least one unknown audio source in the six degrees-of-freedom audio scene, wherein the at least one unknown audio source at least partially interferes with the capturing of the at least one known audio source by the at least one microphone; and controlling the beamforming based on identifying at least one alternative microphone of the at least two microphones, the controlling of beamforming is such that the at least one alternative microphone starts capturing audio from the at least one known audio source while the at least one unknown audio source is at least partially located between the at least one microphone and the at least one known audio source.
2. The method as claimed in claim 1, wherein the at least two microphones are higher-order Ambisonics microphones.
3. The method as claimed in any of claims 1 or 2, wherein controlling the beamforming based on identifying at least one alternative microphone of the at least two microphones comprises determining the at least one unknown audio source is masking the at least one known audio source with respect to the beamforming by at least one microphone of the at least two microphones.
4. The method as claimed in any of claims 1 to 3, wherein controlling the beamforming based on identifying at least one alternative microphone of the at least two microphones further comprises generating a signal, the signal indicatingthat the at least one alternative microphone starts capturing audio from the at least one known audio source while the at least one unknown audio source is at least partially located between the at least one microphone and the at least one known audio source.
5. The method as claimed in any of claims 1 to 3, wherein controlling the beamforming based on identifying at least one alternative microphone of the at least two microphones further comprises generating a signal, the signal associated with the at least one known audio source and indicating which the at least one microphone of the two microphones and the at least one alternative microphone is to be selected.
6. The method as claimed in any of claims 1 to 3, wherein controlling the beamforming based on identifying at least one alternative microphone of the at least two microphones further comprises switching from the one of the at least two microphones to the at least one alternative microphone.
7. The method as claimed in any of claims 1 to 6, wherein controlling the beamforming based on identifying at least one alternative microphone of the at least two microphones comprises obtaining a signal associated with the at least one known audio source, the signal indicating where the at least one microphone of the two microphones and the at least one alternative microphone is to be selected.
8. The method as claimed in claim 7, further comprising receiving a six degrees-of-freedom audio scene description wherein the six degrees-of-freedom audio scene description comprises the signal associated with the at least one known audio source.
9. The method as claimed in any of claims 1 to 8, wherein controlling the beamforming based on identifying at least one alternative microphone of the at least two microphones comprises activating one of: the at least one alternative microphone; orthe at least one microphone, for capturing audio from the at least one known audio source and deactivating the other of the at least one alternative microphone or the at least one microphone.
10. The method as claimed in any of claims 1 to 9, wherein controlling the beamforming based on identifying at least one alternative microphone of the at least two microphones comprises controlling a disabling or enabling of the at least one known source.
11. The method as claimed in any of claims 1 to 10, further comprising determining the at least one unknown audio source at least partially interferes with the capturing of the at least one known audio source by the at least one microphone.
12. The method as claimed in claim 11, wherein determining the at least one unknown audio source at least partially interferes with the capturing of the at least one known audio source by the at least one microphone comprises determining the position of the at least one unknown audio source is between or substantially between the at least one known audio source and the at least one of the at least two microphones.
13. The method as claimed in claim 11, wherein determining the at least one unknown audio source at least partially interferes with the capturing of the at least one known audio source by the at least one microphone comprises determining a distance between the at least one unknown audio source and the at least one known audio source is less than a threshold distance.
14. The method as claimed in claim 11, wherein determining the at least one unknown audio source at least partially interferes with the capturing of the at least one known audio source by the at least one microphone comprises determining an energy estimate of an energy of the at least one known audio source is lower than a threshold energy.
15. The method as claimed in any of claims 1 to 14, furthermore comprising determining at least one spatial metadata parameter associated with the at least one known audio source based on an analysis of the beamformed at least one known audio source.
16. An apparatus comprising means for performing the method of any of claims 1 to 15.
17. A computer program comprising instructions, which, when executed by an apparatus, cause the apparatus to perform the method of any of claims 1 to 15.
18. An apparatus for a rendering of a six degrees-of-freedom audio scene, the apparatus comprising means configured to: provide at least two microphones located at respective position for audio capturing; identify at least one known audio source being captured based on a beamforming by at least one microphone of the at least two microphones, the at least one audio source is located in the six degrees-of-freedom audio scene relative to the at least two microphones; identifying at least one unknown audio source in the six degrees-of-freedom audio scene, wherein the at least one unknown audio source at least partially interferes with the capturing of the at least one known audio source by the at least one microphone; and control the beamforming based on identifying at least one alternative microphone of the at least two microphones, the control of beamforming is such that the at least one alternative microphone starts capturing audio from the at least one known audio source while the at least one unknown audio source is at least partially located between the at least one microphone and the at least one known audio source.
19. An apparatus for a rendering of a six degrees-of-freedom audio scene, the apparatus comprising at least one processor and at least one memory storinginstructions that, when executed by the at least one processor, cause the apparatus at least to: provide at least two microphones located at respective position for audio capturing; identify at least one known audio source being captured based on a beamforming by at least one microphone of the at least two microphones, the at least one audio source is located in the six degrees-of-freedom audio scene relative to the at least two microphones; identifying at least one unknown audio source in the six degrees-of-freedom audio scene, wherein the at least one unknown audio source at least partially interferes with the capturing of the at least one known audio source by the at least one microphone; and control the beamforming based on identifying at least one alternative microphone of the at least two microphones, the control of beamforming is such that the at least one alternative microphone starts capturing audio from the at least one known audio source while the at least one unknown audio source is at least partially located between the at least one microphone and the at least one known audio source.
20. A non-transitory computer readable medium comprising program instructions for causing an apparatus, for a rendering of a six degrees-of-freedom audio, at least to: provide at least two microphones located at respective position for audio capturing; identify at least one known audio source being captured based on a beamforming by at least one microphone of the at least two microphones, the at least one audio source is located in the six degrees-of-freedom audio scene relative to the at least two microphones; identifying at least one unknown audio source in the six degrees-of-freedom audio scene, wherein the at least one unknown audio source at least partially interferes with the capturing of the at least one known audio source by the at least one microphone; andcontrol the beamforming based on identifying at least one alternative microphone of the at least two microphones, the control of beamforming is such that the at least one alternative microphone starts capturing audio from the at least one known audio source while the at least one unknown audio source is at least partially located between the at least one microphone and the at least one known audio source.