Spatial-surround sound system and method for hybrid audio rendering
The spatial-surround sound system combines conventional and object-based rendering techniques to achieve a symmetric soundstage and accurate location-based rendering in vehicles, addressing the limitations of existing systems by using a hierarchical approach with separate rendering zones and gain adjustments.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-19
- Publication Date
- 2026-03-26
AI Technical Summary
Conventional audio rendering systems for vehicles struggle to achieve a symmetric soundstage while maintaining seat-position independent location-based rendering, particularly when dealing with stereo and object-based audio formats.
A spatial-surround sound system that combines conventional and object-based rendering techniques by generating output audio signals for multiple speakers based on virtual positions, using a hierarchical approach with separate rendering zones and gain adjustments for each audio object, ensuring a symmetric front soundstage and location-based rendering in other areas.
The system provides a symmetric soundstage for front seat occupants while accurately rendering audio objects from different positions, maintaining stereo width and improving the perception of the listening environment, compatible with both stereo and object-based audio formats.
Smart Images

Figure EP2024076202_26032026_PF_FP_ABST
Abstract
Description
[0001] SPATIAL-SURROUND SOUND SYSTEM AND METHOD FOR HYBRID AUDIO RENDERING
[0002] TECHNICAL FIELD
[0003] The present disclosure relates to audio processing in general. For example, the disclosure relates to a spatial-surround sound system and method for hybrid audio rendering, in particular inside of a vehicle.
[0004] BACKGROUND
[0005] Audio rendering for cars is conventionally done in one of two ways. One approach was to upmix and render the input signals so that a symmetric soundstage was achieved. Using stereo audio input signals this could be achieved relatively easy by simply adding a center channel which reproduced a filtered sum of the front left and front right signals. With the new object-based audio formats different rendering methods were developed where the target is that objects are perceived from the same point for all listeners inside of a car.
[0006] Conventional solutions either only focused on rendering sounds towards a given speaker setup trying to achieve a symmetrical stage or they focused on rendering sounds based on object positions. Rendering sounds towards given speaker setups was simply done by using some kind of remixing or upmixing. Object position-based rendering usually was done using methods like VBAP, e.g. vector-based amplitude panning. One reason for this hard split was also caused by the format of the available input streams. These were either encoded towards a predefined speaker setup or really object based, thus either one or the other format was available.
[0007] SUMMARY
[0008] It is an objective to provide an improved spatial-surround sound system and method for hybrid audio rendering, in particular inside of a vehicle.
[0009] The foregoing and other objectives are achieved by the subject matter of the independent claims. Further implementation forms are apparent from the dependent claims, the description and the figures. In this disclosure some or more of the following abbreviations, acronyms and / or definitions will be used:
[0010] Vector Based Amplitude Panning VBAP
[0011] As used herein a speaker or loudspeaker may comprise a set of woofer, midrange and tweeter drivers which can be arranged in one or more enclosures at different locations. In the automotive industry it is often the case that what is called a front left speaker is a combination of a front left woofer driver which is installed in the front left door panel in the bottom, a front left midrange driver installed in the upper front part of the left door and a front left tweeter driver installed in the A-pillar of the car.
[0012] As used herein a loudspeaker driver may be a single loudspeaker chassis, for instance, a subwoofer driver, a woofer driver, a midrange driver or a tweeter driver. To avoid the need to specify drivers by writing all their properties like “Front Left Woofer”, instead abbreviations are used consisting of the first letter of the words. Thus, a front left woofer is abbreviated as FLW. Thus, the following abbreviations will be used herein for loudspeaker drivers:
[0013] F Front B Back (this usually refers to speakers installed in the back doors of a vehicle)
[0014] S Surround (this is usually the location behind the rear passengers like on the parcel shelf) L Left
[0015] R Right
[0016] C Center
[0017] W Woofer
[0018] M Midrange
[0019] T Tweeter
[0020] A subwoofer driver is herein referred to as “SUB”. Filters that are applied will use the common technical letter indicating a transfer function which is “H”:
[0021] H Transfer function
[0022] Thus, a filter for the front left woofer will be abbreviated as “HFLW”. A filter like “HLR” indicates a filter that is identical for the left and right channels.
[0023] According to a first aspect a spatial-surround sound system, in particular an automotive spatial-surround sound system, e.g. a spatial-surround system for the interior of a vehicle, with a plurality of speakers, including a front center speaker located at a front center position and a front left speaker located at a front left position and front right speaker located at a front right position, is provided for generating spatial-surround sound, e.g. a soundfield. The spatial-surround sound system comprises processing circuitry configured to generate a plurality of output audio signals for the plurality of speakers based on a plurality of input audio signals for rendering one or more audio objects depending on a respective virtual position associated with each audio object in a first rendering zone with a first audio rendering scheme and in a second rendering zone with a second audio rendering scheme. For generating the plurality of output audio signals for the one or more audio objects in the first rendering zone with the first audio rendering scheme the processing circuitry is configured to render the one or more audio objects to a virtual speaker setup comprising the plurality of speakers except, e.g. without the front center speaker, and to upmix the virtual front left and front right signals for generating the respective output audio signals for the front left speaker, the front right speaker and the front center speaker. As will be appreciated, the spatial-surround sound system according to the first aspect allows combining the two conventional audio rendering approaches, in particular in the interior of a vehicle, by transcoding the audio input signal(s), regardless if it is a conventional speaker setup encoding like Stereo or Dolby Digital or an object based model into a virtual object layer. This allows recreating the speaker signals or, if desired, directly applying object based rendering. Combining the two rendering approaches in the proposed manner furthermore allows combining their advantages, one being that objects can be rendered according to their object position, the second one that this is achieved for both front seat positions, so a symmetric stage.
[0024] In a further possible implementation form, for generating the plurality of output audio signals for the one or more audio objects in the second rendering zone with the second audio rendering scheme the processing circuitry of the spatial-surround sound system according to the first aspect is configured to combine for each audio object the virtual front left and virtual front right signal using a weighted sum with a gain depending on the position of the respective audio object for generating the output audio signal for the front center speaker of the respective audio object, and for each of the plurality of speakers sum the corresponding audio signals associated with the one or more audio objects for generating the output audio signal for the respective speaker. In a further possible implementation form, for generating the plurality of output audio signals for the one or more audio objects in the second rendering zone with the second audio rendering scheme the processing circuitry of the spatial-surround sound system according to the first aspect is configured to sum up the virtual left signals generated by the first audio rendering schemes for the one or more audio objects to generate the output audio signal for the front left speaker and to sum up the virtual right signals generated by the first audio rendering scheme for the one or more audio objects to generate the output audio signal for the front right speaker. Moreover, the processing circuitry of the spatial-surround sound system according to the first aspect is configured to combine the virtual front left and front right signals of the one or more audio objects using a weighted sum with the same gains for the one or more audio objects for generating the output audio signal for the front center speaker.
[0025] In a further possible implementation form, the processing circuitry of the spatial-surround sound system according to the first aspect is further configured to filter the respective output audio signal for the front left, front right and front center speaker such that the front center speaker, the front left speaker and the front right speaker implement a symmetric front stage.
[0026] In a further possible implementation form, the plurality of input audio signals comprise one or more non-object based form audio signals and for generating one or more of the plurality of output audio signals the processing circuitry of the spatial- surround sound system according to the first aspect is configured to convert the one or more non-object based form audio signals into the one or more audio objects with a respective virtual position.
[0027] In a further possible implementation form, for converting the one or more non-object based form audio signals into the one or more audio objects with a respective virtual position the processing circuitry of the spatial-surround sound system according to the first aspect is configured to upmix a stereo input audio signal into a virtual speaker setup of a plurality of speakers, wherein the virtual positions of the speakers are defined.
[0028] In a further possible implementation form, the respective virtual position associated with each audio object is defined by an azimuth and / or an elevation angle relative to a reference direction and wherein the first rendering zone is defined by a first range of azimuth and / or elevation angles and the second rendering zones is defined by a second range of azimuth and / or elevation angles.
[0029] In a further possible implementation form, the respective virtual position associated with each audio object is further defined by a distance from a reference position, wherein the processing circuitry of the spatial-surround sound system according to the first aspect is configured to generate the plurality of output audio signals for the plurality of speakers based on the plurality of input audio signals for rendering the one or more audio objects such that an audio object at larger distance has a smaller amplitude than an audio object at a smaller distance from the reference point.
[0030] In a further possible implementation form, each audio object is associated with a direct sound or an indirect sound, wherein the processing circuitry of the spatial-surround sound system according to the first aspect is configured to generate the plurality of output audio signals for the plurality of speakers based on the plurality of input audio signals for rendering the one or more audio objects associated with a direct sound with a different audio rendering scheme than the one or more audio objects associated with an indirect sound.
[0031] According to a second aspect a method for operating a sound system with a plurality of speakers, including a front center speaker located at a front center position and a front left speaker located at a front left position and front right speaker located at a front right position, for generating spatial-surround sound, e.g. a soundfield, wherein the method according to the second aspect comprises: generating a plurality of output audio signals for the plurality of speakers based on a plurality of input audio signals for rendering one or more audio objects depending on a respective virtual position associated with each audio object in a first rendering zone with a first audio rendering scheme and in a second rendering zone with a second audio rendering scheme, wherein generating the plurality of output audio signals for the one or more audio objects in the first rendering zone with the first audio rendering scheme comprises: rendering the one or more audio objects to a virtual speaker setup comprising the plurality of speakers except the front center speaker, and upmixing a virtual front left and front right signal for generating the respective output audio signals for the front left, front right and front center speaker.
[0032] The method according to the second aspect can be performed by the spatial surround sound system according to the first aspect. Thus, further features of the method according to the second aspect result directly from the functionality of the spatial surround sound system according to the first aspect as well as its different implementation forms and embodiments described above and below.
[0033] According to a third aspect a computer program product is provided, comprising a computer-readable storage medium for storing program code which causes a computer or a processor to perform the method according to the second aspect, when the program code is executed by the computer or the processor.
[0034] Details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims.
[0035] BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In the following, embodiments of the present disclosure are described in more detail with reference to the attached figures and drawings, in which:
[0037] Fig. 1 is a schematic diagram illustrating a spatial surround sound system according to an embodiment;
[0038] Fig. 2 is a schematic diagram illustrating a spatial surround sound system according to a further embodiment for object specific rendering;
[0039] Fig. 3 is a schematic diagram illustrating in more detail a position-dependent object rendering stage of the spatial surround sound system of figure 2;
[0040] Figs. 4a and 4b show schematic diagram illustrating two and three audio rendering zones used by a spatial surround sound system according to embodiments; and
[0041] Fig. 5 is a flow diagram illustrating steps of a method according to an embodiment for operating a spatial surround sound system.
[0042] In the following, identical reference signs refer to identical or at least functionally equivalent features. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] In the following description, reference is made to the accompanying figures, which form part of the disclosure, and which show, by way of illustration, specific aspects of embodiments of the present disclosure or specific aspects in which embodiments of the present disclosure may be used. It is understood that embodiments of the present disclosure may be used in other aspects and comprise structural or logical changes not depicted in the figures. The following detailed description, therefore, is not to be taken in a limiting sense, and the scope of the present disclosure is defined by the appended claims.
[0044] For instance, it is to be understood that a disclosure in connection with a described method may also hold true for a corresponding device or system configured to perform the method and vice versa. For example, if one or a plurality of specific method steps are described, a corresponding device may include one or a plurality of units, e.g. functional units, to perform the described one or plurality of method steps (e.g. one unit performing the one or plurality of steps, or a plurality of units each performing one or more of the plurality of steps), even if such one or more units are not explicitly described or illustrated in the figures. On the other hand, for example, if a specific apparatus is described based on one or a plurality of units, e.g. functional units, a corresponding method may include one step to perform the functionality of the one or plurality of units (e.g. one step performing the functionality of the one or plurality of units, or a plurality of steps each performing the functionality of one or more of the plurality of units), even if such one or plurality of steps are not explicitly described or illustrated in the figures. Further, it is understood that the features of the various exemplary embodiments and / or aspects described herein may be combined with each other, unless specifically noted otherwise.
[0045] Figure 1 is a schematic diagram illustrating a spatial surround sound system 100 according to an embodiment as well as processing steps implemented by a processing circuitry of the spatial surround sound system 100. As illustrated in figure 1, the spatial surround sound system 100 comprises a plurality of speakers 1 lOa-m and processing circuitry configured to process one or more audio input signals into respective output signals for driving the plurality of speakers 1 lOa-m. In an embodiment, the spatial surround sound system 100 is an automotive spatial surround sound system 100 (e.g. configured to be implemented and operated for generating surround sound), e.g. a soundfield in the interior of a vehicle, e.g. car. The processing circuitry of the spatial surround sound system 100 may be implemented in hardware and / or software and may comprise digital circuitry, or both analog and digital circuitry. Digital circuitry may comprise components such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), or general-purpose processors. The spatial surround sound system 100 may further comprise a memory configured to store executable program code which, when executed by the processing circuitry, causes the spatial surround sound system 100 to perform the functions and methods described herein.
[0046] As alreadv described above, the conventional approaches for audio rendering in the interior of a vehicle do not allow to achieve a symmetric tuning in a front sound zone, while at the same time achieving a seat-position independent location-based rendering for a surround / back sound zone. To solve this problem the processing circuitry of the spatial-surround sound system 100 is configured to generate the plurality of output audio signals for the plurality of speakers 1 lOa-m based on the plurality of input audio signals for rendering one or more audio objects depending on a respective position associated with each audio object in a first rendering zone, such as front sound zone 410 illustrated in figures 4a and 4b, with a first audio rendering scheme and in a second rendering zone, such as a surround / back sound zone 420 illustrated in figures 4a and 4b, with a second audio rendering scheme that is different from the first audio rendering scheme, as will be described in more detail below. For generating the plurality of output audio signals for the one or more audio objects in the first rendering zone, e.g. the front sound zone 410 illustrated in figures 4a and 4b, with the first audio rendering scheme the processing circuitry of the spatial-surround sound system 100 is configured to render the one or more audio objects to a virtual speaker setup comprising the plurality of speakers 1 lOa-m, but not the front center speaker, and to upmix the virtual front left and front right signals for generating the respective output audio signals for the front left speaker, the front right speaker and the front center speaker.
[0047] Thus, as will be appreciated, embodiments disclosed herein are based on a hierarchical approach. More specifically, in a processing stage 101, 102 illustrated in figure 1, if the audio input stream is not already object based, the processing circuitry of the spatial-surround sound system 100 is configured to convert the audio input stream into one or more audio objects. In an embodiment, this may be achieved using conventional passive upmixing, active upmixing, neural network based upmixing or the like. Passive and active upmixing according to stage 112, for instance, may convert a stereo signal into a 5.1 or 7.1 or any other channel configuration object-based representation. As will be appreciated, in such a representation, the generated signals for the one or more audio objects comprise position information, e.g. are associated with one or more positions, which, however, do not need to match the real speaker layout. Neural network based upmixing is more focusing on decomposing different components of the recording like the voice, drums, guitar and the rest. Also, in this case the resulting objects may be given some new position information. Combinations of both conventional and neural network based upmixing are possible. This approach allows, for instance, extracting an ambience signal using active upmixing and then decomposing the direct part into the different instruments, voices or other sound types. After converting the input signals into an object-based format in processing stage 112 of figure 1, a second transformation stage 103 may be applied. For instance, in processing stage 103 a voice may be transformed into a direct part and several indirect wall reflections, each representing different audio objects, wherein each of those objects may have a position that can be real or virtual. An indirect sound reflected from a wall will be perceived as coming from behind the wall, thus it is a virtual position. Sounds that directly arrive without any reflection are called direct sounds. In an embodiment, this information (direct sound or indirect sound) may be stored together with the object. In a further processing stage 105 the processing circuitry of the spatial-surround sound system 100 is configured to render all the generated audio objects using, for instance, VBAP rendering to a virtual speaker setup. The virtual speaker setup is characterized by the fact that it consists of the plurality of physical speakers 1 lOa-m except the center speaker(s) (e.g. the center speaker C illustrated in figure 4a). In most cars there is only one center speaker (usually a combination of a center midrange driver and a center tweeter driver), but it is possible to also have a center height speaker or center speakers in the back where the same principle may be applied. As will be appreciated, by rendering the sound to this virtual speaker setup, in a car speaker setup, sounds that are in the front will be rendered to the left and right front virtual speakers. This then allows using a symmetric stereo rendering approach only on the left and right virtual speaker signals which then generates the real left, right and center signals. The symmetric stereo rending is done by first summing the virtual left and the virtual right using the adder 106 to create a new center signal and then filter the virtual left, the new center and the virtual right in stages 1071-107n. The signals that were not in the front rendering zone are also sent to a stage of filters 107a-107e. All signals rendered in a conventional way as well as the new left, right and center signals are then sent to a post processing stage 109. This converts the speaker signals into appropriate signals for the final drivers. As an example, the front left speaker signal is split up into a front left woofer, a front left midrange and a front left tweeter driver signal. Further processing like equalizing, time adjustments, compression, limiting and amplification may be applied to these driver signals. The complete processing will then create a symmetric front stage, while at the same time objects not in the front stage are rendered location based. This means that a centered signal in the front will be perceived directly in front of both the driver and the front passenger. All other sounds, whose angle location information is not in between the left and right front or back speakers, do not need to be remixed further, even if this is also possible.
[0048] As will be appreciated, in comparison with conventional approaches, the spatial-surround sound system 100 according to the embodiment illustrated in figure 1 allows implementing a symmetric front stage in a front of a car (such as the front sound zone 410 illustrated in figure 4a), while location-based rendering is performed in the other areas of the car interior (such as the surround / back sound zone 420 illustrated in figure 4a). Thus, the spatial-surround sound system 100 according to the embodiment illustrated in figure 1 combines the advantages of classical channel-based rendering approaches for symmetric sound stage creation with modem object-based rendering. Moreover, the embodiment illustrated in figure 1 is computationally very efficient, because in a first stage all objects are rendered to one virtual output and the conventional center algorithm is applied to the virtual output.
[0049] For the embodiment of the spatial-surround sound system 100 illustrated in figure 1 the same rendering approach may be used for all objects. This might cause unwanted effects like incorrect virtually perceived position or distance from objects. Furthermore, objects which are already outside of the front sound zone, but still close to the border, may cause some signal to be present in the virtual front left or virtual front right signals. This may cause those signals to be rendered also in the front sound zone 410 via the center channel causing the stereo width to be reduced, which might not be desired. To overcome this limitation, a position dependent rendering is implemented by the embodiment of the spatial-surround sound system 100 illustrated in figure 2. As will be appreciated, instead of the generic rendering approach for all objects (as implemented by processing stages 105 and 106 of figure 1), in the embodiment of figure 2 an object specific rendering approach is used, as implemented by the processing stage 105 of figure 2, which is illustrated in more detail in figure 3. Besides this change, the rendering approach of figure 2 is identical to the rendering approach of figure 1. So, similar to the one of figure 1 , first in stage 101 it is decided if the audio input is object based or not, if it is then it is directly forwarded to stage 103, otherwise it is converted into an object format in stage 112 and then sent to stage 103. After the object specific rendering stage 105’, the outputs signals of this stage are then first sent to channel-specific filters 107a-107n, and then post-processed typically using processing like compression, limiting, delaying and amplification, before being sent to the speakers l lOa-HOm.The object specific rendering approach implemented by the processing stage 105 of figure 2 and illustrated in more detail in figure 3 is very similar to the rendering approach of the embodiment of the spatial-surround sound system 100 illustrated in figure 1. The main difference is, that in the embodiment of figures 2 and 3 for each object the processing circuitry of spatial-surround sound system 100 is configured to first render each object individually to the given virtual speaker setup lacking the center speaker indicated in stages 105a-105n, then to compute a specific gain for the summed center channel of each of those stages 105a- 105n by mapping the angle of the audio object to a gain (as illustrated by the processing stages 11 la-n of figure 3). For objects in the middle of the front sound zone 410 this gain may be 1. For objects within a transition zone 415, e.g. for objects whose angle is close to the border between the front sound zone 410 and the surround / rear sound zone 420 illustrated in figures 4a and 4b the gain may be lower so that in this case less signal will be sent to the center speaker. The curve mapping the angle to a gain may be using a 4-point linear interpolation approach. In further embodiments a smoother curve with more points or other functions like the classical windowing functions may be used for the mapping. Only after applying the object specific gains for each object individually the results are then summed up in stage 108. This is the main difference to the embodiment of figure 1.
[0050] As will be appreciated, without the position-dependent gain implemented by the embodiment of figures 2 and 3, the mixing to the center speaker would only stop if the objects angle points to a direction that is behind the first speaker outside of the front sound zone. In a car, these speakers are usually the surround speakers installed in the back doors. This would mean that the front zone rendering, where some signal from the virtual front signals is copied to the center, would nearly be extended to the back doors which is usually undesired. With the object specific rendering, the area where signals are copied from the virtual front signals to the center can be limited. The area, where the gain will decrease but did not yet reach 0, may be implemented as a transition zone 415, as illustrated in figure 4b. The transition zone 415 might be defined to be completely inside of the front left and right speaker positions. This will cause signals, that are in the direction of the left or right front speakers, to not be copied to the center, this way the stereo width for such signals will be higher. The transition zone 415 may also overlap with the position of the front left and right speakers or be outside, this will allow defining the width of the front stage 410 and the stereo width. The wider the front stage, the less the stereo width will be present.
[0051] As will be appreciated, in comparison with the embodiment of figure 1, the embodiment of the spatial-surround sound system 100 illustrated in figures 2 and 3 allows a more precise definition of the front zone 410 where symmetry is desired. In the embodiment of figure 1, the front zone borders are defined by the speaker layout. In a usual car it will start from the back-door speakers, continue over the front left, center, right and reach until the back-right door speaker. Even if, due to the rendering algorithm, the effect at the back-door speakers is still zero and only starts to fade in there, a signal at an angle in the middle of the back-left door speaker and the front left speaker would already be applied with the same gain to the back left and front left speakers. As the front stage rendering may only be working optimal for signals that are in between the front left and front right speaker angles, this may cause some quality degradation. The embodiment of the spatial-surround sound system 100 illustrated in figures 2 and 3 addresses this issue by applying an object angle dependent gain to the center signal. In this way, the width of the front zone 410 may be controlled very accurately . By applying an angle-dependent gain, a smooth transition 415 between the zones 410 and 420 is achieved which avoids any discontinuities that might cause audible clicks when objects are moving from one zone to the other. Having full control over the front zone area 410 also avoids reducing the stereo width for signals that are far out of the center.
[0052] As will be appreciated, the embodiments of the spatial-surround sound system 100 described above allow combining a front stage rendering with an object-based rendering. In this way, playing back music, for both the driver and the front passenger the different instruments of music in the front will have a clear location, both the driver and passenger can perceive a clear distance and angle information for all instruments of the music recording. Furthermore, both will perceive the stage nearly symmetrically in front of them, both will perceive the center of the stage directly in front of them. At the same time, objects that are not encoded to be in the front will also be perceived from both the driver and the front passenger from one position outside of the front. Even if the position will not be exactly the same for both the driver and front passenger, it is a good approximation. So, sounds in the front will be very accurately rendered with the most audiophile rendering approach. Sounds from the sides, where this symmetrical rendering approach is not possible, will also be rendered with the best approach for those locations. Moreover, the embodiments of the spatial-surround sound system 100 described above work for both conventionally encoded formats (stereo, 5.1) as well as for object-based formats. So, the passengers can enjoy such benefits with all kinds of audio source formats. Furthermore, the audio processing stage, where a room model is applied, may furthermore improve the perception of the listening room, the sound in a car will sound less like a car but more like the original recording room.
[0053] Figure 5 is a flow diagram illustrating a method 500 for operating the spatial-surround sound system 100. The method 500 comprises a step 501 of generating a plurality of output audio signals for the plurality of speakers 11 Oa-m of the spatial-surround sound system 100 based on a plurality of input audio signals for rendering one or more audio objects depending on a respective position associated with each audio object in a first rendering zone, such as the first audio rendering zone 410 illustrated in figures 4a and 4b, with a first audio rendering scheme and in a second rendering zone, such as the second audio rendering zone 420 illustrated in figures 4a and 4b with a second audio rendering scheme. The step 501 of generating the plurality of output audio signals for the one or more audio objects in the first rendering zone with the first audio rendering scheme comprises a step 503 of rendering the one or more audio objects to a virtual speaker setup comprising the plurality of speakers HOa-m except the front center speaker, and a step 505 of upmixing a virtual front left and front right signal for generating the respective output audio signals for the front left speaker, the front right speaker and the front center speaker.
[0054] The method 500 can be performed by the spatial-surround sound system 100 according to an embodiment. Thus, further features of the method 500 result directly from the functionality of the spatial-surround sound system 100 as well as its different embodiments described above and below.
[0055] The person skilled in the art will understand that the "blocks" ("units") of the various figures (method and apparatus) represent or describe functionalities of embodiments of the present disclosure (rather than necessarily individual "units" in hardware or software) and thus describe equally functions or features of apparatus embodiments as well as method embodiments (unit = step). In the several embodiments provided in the present application, it should be understood that the disclosed system, apparatus, and method may be implemented in other manners. For example, the described embodiment of an apparatus is merely exemplary. For example, the unit division is merely logical function division and may be another division in an actual implementation. For example, a plurality of units or components may be combined or integrated into another system, or some features may be ignored or not performed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections may be implemented by using some interfaces. The indirect couplings or communication connections between the apparatuses or units may be implemented in electronic, mechanical, or other forms.
[0056] The units described as separate parts may or may not be physically separate, and parts displayed as units may or may not be physical units, may be located in one position, or may be distributed on a plurality of network units. Some or all of the units may be selected according to actual needs to achieve the objectives of the solutions of the embodiments.
[0057] In addition, functional units in the embodiments of the invention may be integrated into one processing unit, or each of the units may exist alone physically, or two or more units are integrated into one unit.
Claims
CLAIMS1. A spatial-surround sound system (100) with a plurality of speakers (1 lOa-m), including a front center speaker, a front left speaker and a front right speaker, for generating spatial-surround sound, wherein the spatial-surround sound system (100) comprises processing circuitry configured to: generate a plurality of output audio signals for the plurality of speakers (1 lOa-m) based on a plurality of input audio signals for rendering one or more audio objects depending on a respective position associated with each audio object in a first rendering zone (410) with a first audio rendering scheme and in a second rendering zone (420) with a second audio rendering scheme, wherein for generating the plurality of output audio signals for the one or more audio objects in the first rendering zone (410) with the first audio rendering scheme the processing circuitry is configured to: render the one or more audio objects to a virtual speaker setup comprising the plurality of speakers (1 lOa-m) except the front center speaker, and upmix the virtual front left and front right signals for generating the respective output audio signals for the front left speaker, the front right speaker and the front center speaker.
2. The sound system (100) of claim 1, wherein for generating the plurality of output audio signals for the one or more audio objects in the second rendering zone (420) with the second audio rendering scheme the processing circuitry is configured to: combine for each audio object the virtual front left signal and the virtual front right signal using a weighted sum with a gain depending on the position of the respective audio object for generating the output audio signal for the front center speaker of the respective audio object, and for each of the plurality of speakers (HOa-m) sum the corresponding audio signals associated with the one or more audio objects for generating the output audio signal for the respective speaker.
3. The sound system (100) of claim 1 or 2, wherein for generating the plurality of output audio signals for the one or more audio objects in the second rendering zone (420) with the second audio rendering scheme the processing circuitry is configured to: sum up the virtual left signals generated by the first audio rendering scheme for the one or more audio objects to generate the output audio signal for the front left speaker; sum up the virtual right signals generated by the first audio rendering scheme for the one or more audio objects to generate the output audio signal for the front right speaker; and combine the virtual front left and front right signals of the one or more audio objects using a weighted sum with the same gains for the one or more audio objects for generating the output audio signal for the front center speaker.
4. The sound system (100) of any one of the preceding claims, wherein the processing circuitry is further configured to filter the respective output audio signal for the front left speaker, the front right speaker and the front center speaker such that the front center speaker, the front left speaker and the front right speaker implement a symmetric front stage.
5. The sound system (100) of any one of the preceding claims, wherein the plurality of input audio signals comprises one or more non-object based form audio signals and wherein for generating one or more of the plurality of output audio signals the processing circuitry is configured to convert the one or more non-object based form audio signals into the one or more audio objects with a respective position.
6. The sound system (100) of claim 5, wherein for converting the one or more non-object based form audio signals into the one or more audio objects with a respective position the processing circuitry is configured to upmix a stereo input audio signal into a virtual speaker setup of a plurality of speakers, wherein the virtual positions of the speakers are defined.
7. The sound system ( 100) of any one of the preceding claims, wherein the respective position associated with each audio object is defined by an azimuth and / or an elevation angle and wherein the first rendering zone (410) is defined by a first range of azimuth and / or elevation angles and the second rendering zone (420) is defined by a second range of azimuth and / or elevation angles.
8. The sound system (100) of claim 7, wherein the respective position associated with each audio object is further defined by a distance and wherein the processing circuitry is configured to generate the plurality of output audio signals for the plurality of speakers based on the plurality of input audio signals for rendering the one or more audio objects such that an audio object at larger distance has a smaller amplitude than an audio object at a smaller distance.
9. The sound system (100) of claim 7 or 8, wherein each audio object is associated with a direct sound or an indirect sound and wherein the processing circuitry is configured to generate the plurality of output audio signals for the plurality of speakers based on the plurality of input audio signals for rendering the one or more audio objects associated with a direct sound with a different audio rendering scheme than the one or more audio objects associated with an indirect sound.
10. A method (500) for operating a spatial-surround sound system with a plurality of speakers (1 lOa-m), including a front center speaker, a front left speaker, and a front right speaker, for generating spatial-surround sound, wherein the method (500) comprises: generating (501) a plurality of output audio signals for the plurality of speakers (1 lOa-m) based on a plurality of input audio signals for rendering one or more audio objects depending on a respective position associated with each audio object in a first rendering zone (410) with a first audio rendering scheme and in a second rendering zone (420) with a second audio rendering scheme, wherein generating (501) the plurality of output audio signals for the one or more audio objects in the first rendering zone (410) with the first audio rendering scheme comprises: rendering (503) the one or more audio objects to a virtual speaker setup comprising the plurality of speakers (1 lOa-m) except the front center speaker, and upmixing (505) a virtual front left and front right signal for generating the respective output audio signals for the front left speaker, the front right speaker and the front center speaker.
11. A computer program product comprising a computer-readable storage medium for storing program code which causes a computer or a processor to perform the method (500) of claim 10 when the program code is executed by the computer or the processor.
Citation Information
Patent Citations
Apparatus and Method for Multi-Channel Parameter Transformation
US20110013790A1
System and method for processing audio signal
US20170086005A1
Audio providing apparatus and audio providing method
US20180359586A1