Sound processing device, method, and program
By performing rendering processing for each speaker layout with matching playback bands and using filters to direct audio signals to appropriate speakers, the technology ensures high-quality audio reproduction despite varying speaker capabilities.
Patent Information
- Application Number
- JP2022547497
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-09-09
- Filing Date
- 2021-08-27
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2041-08-27
AI Technical Summary
When rendering and playing back object-based audio using a speaker layout with multiple speakers of different playback bands, the sound quality deteriorates due to mismatched frequency bands and localization positions, leading to issues like sound disappearance.
Perform rendering processing for each speaker layout with the same playback band, generating separate output audio signals for speakers with different playback bands, and using high-pass, band-pass, and low-pass filters to ensure each frequency component is reproduced by the appropriate speakers.
This approach suppresses sound quality deterioration, enabling high-quality audio playback by ensuring each frequency component is reproduced by the correct speakers, thus maintaining sound integrity across different localization positions.
Smart Images

Figure 0007771963000001 
Figure 0007771963000002 
Figure 0007771963000003
Abstract
Description
[Technical Field]
[0001] The present technology relates to an audio processing device, method, and program, and more particularly to an audio processing device, method, and program that enable audio reproduction with higher sound quality. [Background technology]
[0002] In recent years, object-based audio technology has been attracting attention.
[0003] In object-based audio, audio data is composed of a waveform signal (audio signal) for an object and metadata indicating localization information that indicates the relative position of the object as viewed from a predetermined reference listening point (listening position). Based on the metadata, the waveform signal is rendered into a desired number of channels using, for example, VBAP (Vector Based Amplitude Panning) and played back (see, for example, Non-Patent Document 1 and Non-Patent Document 2). [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] ISO / IEC 23008-3 Information technology - High efficiency coding and media delivery in heterogeneous environments - Part 3: 3D audio [Non-patent document 2] Ville Pulkki, “Virtual Sound Source Positioning Using Vector Base Amplitude Panning”, Journal of AES, vol.45, no.6, pp.456-466, 1997 Summary of the Invention [Problem to be solved by the invention]
[0005] Incidentally, when rendering and playing back an object using a speaker layout in which multiple speakers are arranged in a three-dimensional space, many speakers will be used, but there may be cases in which all speakers do not have the same playback band.
[0006] For example, in-car audio is a use case where many speakers can be arranged, and is generally configured with a speaker layout that combines a speaker called a woofer that reproduces low frequencies, a speaker called a squawker that reproduces mid-frequency frequencies, and a speaker called a tweeter that reproduces high frequencies.
[0007] However, when rendering VBAP or the like of object audio is performed with such a speaker layout, the reproduction band of the speakers used for reproduction differs depending on the localized position of the object.
[0008] Therefore, for example, when the sound of an object containing only high-frequency components is reproduced by a woofer located near the object's localization position, depending on the frequency band and localization position of the object's sound, the sound quality may deteriorate, such as the sound disappearing.
[0009] The present technology has been made in view of such circumstances, and makes it possible to play back audio with higher sound quality. [Means for solving the problem]
[0010] An audio processing device according to one aspect of the present technology includes a first rendering processing unit that performs rendering processing based on an audio signal and generates a first output audio signal for outputting sound from a plurality of first speakers, and a second rendering processing unit that performs rendering processing based on the audio signal and generates a second output audio signal for outputting sound from a plurality of second speakers having a playback band different from that of the first speakers.
[0011] An acoustic processing method or program according to one aspect of the present technology includes the steps of performing a rendering process based on an audio signal, generating a first output audio signal for outputting sound from a plurality of first speakers, and performing a rendering process based on the audio signal, generating a second output audio signal for outputting sound from a plurality of second speakers having a playback band different from that of the first speakers.
[0012] In one aspect of the present technology, a rendering process is performed based on an audio signal to generate a first output audio signal for outputting sound from a plurality of first speakers, and a rendering process is performed based on the audio signal to generate a second output audio signal for outputting sound from a plurality of second speakers having a playback band different from that of the first speakers. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 1 is a diagram illustrating the present technology. [Figure 2] FIG. 1 is a diagram illustrating an example of the configuration of an audio playback system. [Figure 3] 10A and 10B are diagrams illustrating examples of frequency characteristics of an HPF, a BPF, and an LPF. [Figure 4] 10 is a flowchart illustrating a playback process. [Figure 5] FIG. 1 is a diagram illustrating an example of the configuration of an audio playback system. [Figure 6] 10 is a flowchart illustrating a playback process. [Figure 7] FIG. 1 is a diagram illustrating an example of the configuration of an audio playback system. [Figure 8] 10 is a flowchart illustrating a playback process. [Figure 9] FIG. 1 is a diagram illustrating an example of the configuration of an audio playback system. [Figure 10] 10 is a flowchart illustrating a playback process. [Figure 11]FIG. 1 is a diagram illustrating an example of the configuration of an audio playback system. [Figure 12] 10A and 10B are diagrams illustrating examples of frequency characteristics of an HPF and an LPF. [Figure 13] 10 is a flowchart illustrating a playback process. [Figure 14] FIG. 1 illustrates an example of the configuration of a computer. DETAILED DESCRIPTION OF THE INVENTION
[0014] Hereinafter, embodiments to which the present technology is applied will be described with reference to the drawings.
[0015] First Embodiment About this technology When object-based audio is played back using a speaker system consisting of multiple speakers with different playback bands, this technology performs rendering processing for each speaker layout consisting of speakers with the same playback band, thereby achieving higher quality audio playback.
[0016] For example, in this technology, as shown in FIG. 1, a plurality of speakers SP11-1 to SP11-18 are arranged on the surface of a sphere P11 centered around a user U11 who is a listener of object-based audio, surrounding the user U11.
[0017] The object-based audio is reproduced using a speaker system consisting of these speakers SP11-1 to SP11-18.
[0018] In the following description, when there is no need to particularly distinguish between the speakers SP11-1 to SP11-18, they will also be simply referred to as the speaker SP11.
[0019] In this example, the plurality of speakers SP11 include speakers with different reproduction bands, so rendering processing is performed for each reproduction band.
[0020] For example, a speaker group consisting of speakers SP11 having the same reproduction band, or more specifically, a three-dimensional arrangement of the speakers SP11 that make up the speaker group, will be called one speaker layout.
[0021] At this time, rendering processing is performed for each speaker layout that constitutes the speaker system, and speaker playback signals for playing back the sounds of objects (audio objects) in the speaker layout are generated.
[0022] The rendering process may be any process such as VBAP or panning.
[0023] When rendering processing is performed on one speaker layout, speaker playback signals are generated for each speaker SP11 of that speaker layout.
[0024] When VBAP is performed as the rendering process, one or more meshes are formed on the surface of the sphere P11 by all the speakers SP11 that make up the speaker layout.
[0025] A triangular area on the surface of the sphere P11 surrounded by the three speakers SP11 that make up the speaker layout is one mesh.
[0026] Now, suppose that VBAP is performed for one object with a predetermined speaker layout.
[0027] Also, object data of the object is supplied, and the object data consists of an object signal, which is an audio signal for reproducing the sound of the object, and metadata, which is information about the object.
[0028] The metadata includes at least the position of the object, that is, position information indicating the localized position of the sound image of the object.
[0029] The position information of the object may be, for example, coordinate information indicating the relative position of the object as viewed from the head position of the user U11, which is a predetermined reference listening position. In other words, the position information is information indicating the relative position of the object with respect to the head position of the user U11.
[0030] In VBAP, one mesh that includes the position indicated by the object's position information (hereinafter also referred to as the object position) is selected from the meshes formed by the speaker SP11 of the speaker layout. Here, the selected mesh is called the selected mesh.
[0031] Next, based on the positional relationship between the placement position of each speaker SP11 that constitutes the selected mesh and the object position, a VBAP gain is calculated for each speaker SP11, and the gain of the object signal is adjusted using the VBAP gain to produce a speaker playback signal.
[0032] That is, the signal obtained by adjusting the gain of the object signal based on the VBAP gain calculated for the speaker SP11 is the speaker playback signal of that speaker SP11. Note that, of all the speakers SP11 in the speaker layout, the speaker playback signals of the speakers SP11 other than the speaker SP11 that constitutes the selected mesh are set to zero signals. In other words, the VBAP gains of the speakers SP11 other than the speaker SP11 that constitutes the selected mesh are set to 0.
[0033] When sound is output from each speaker SP11 in the speaker layout obtained in this way based on the speaker playback signals of those speakers SP11, the sound of the object is reproduced so that the sound image is localized at the object position indicated by the position information.
[0034] In addition, for example, panning can be used to generate speaker playback signals for each speaker SP11 in the speaker layout.
[0035] In such a case, for example, the gain for each speaker SP11 is calculated based on the positional relationship between each speaker SP11 in the speaker layout and the object in each direction in the drawing, such as the front-to-back direction, the left-to-right direction, and the up-to-down direction. Then, the gain of the object signal is adjusted using the calculated gain for each speaker SP11, and a speaker playback signal for each speaker SP11 is generated.
[0036] In this way, the rendering process for each speaker layout may be any process such as VBAP or panning, but the following description will be given of a case where VBAP is performed as the rendering process.
[0037] In a speaker system, rendering processing is performed for each of the multiple speaker layouts that make up the speaker system, each with a different playback band, and speaker playback signals are generated for all speakers SP11 that make up the speaker system. In other words, multiple speaker layout configurations are prepared for each playback band, and rendering processing is performed for each of these playback bands.
[0038] By doing this, even when speakers SP11 with different playback bands are mixed, the present technology can suppress deterioration in sound quality due to the playback bands of the speakers SP11, thereby enabling audio playback with higher sound quality.
[0039] For example, suppose that a mesh is formed from all the speakers SP11 that make up the speaker system, and VBAP is performed as the rendering process.
[0040] In this case, if the object position is within a mesh formed by speakers SP11-1, SP11-2, and SP11-5, for example, the sound of the object is reproduced by speakers SP11-1, SP11-2, and SP11-5.
[0041] In this case, if the sound of the object consists only of high-frequency components and the speakers SP11-1, SP11-2, and SP11-5 have a low-frequency reproduction band, the speaker SP11-1 will not be able to reproduce the sound of the object with sufficient sound pressure, which will result in a deterioration in sound quality, such as the sound of the object becoming too quiet to be heard.
[0042] In contrast, with this technology, rendering processing is performed for each of multiple playback bands, so that the components of each frequency band are always reproduced by the speaker SP11 with a playback band that includes that frequency band. This suppresses deterioration in sound quality due to the playback band of the speaker SP11, enabling audio reproduction with higher sound quality.
[0043] In the present technology, the number of speakers SP11 constituting the speaker system, the reproduction bands of each speaker SP11, and the placement positions of the speakers SP11 for each reproduction band can be any number, reproduction band, and placement positions.
[0044] <Example of audio playback system configuration> FIG. 2 is a diagram showing an example of the configuration of an embodiment of an audio playback system to which the present technology is applied.
[0045] The audio reproduction system 11 shown in FIG. 2 includes an audio processing device 21 and a speaker system 22, and reproduces object-based audio content based on supplied object data.
[0046] In this example, the content consists of N objects, and object data for these N objects is supplied, but the number of objects may be any number. Also, as described above, the object data for one object includes an object signal for playing the sound of that object, and object metadata.
[0047] The sound processing device 21 includes a playback signal generating unit 31, D / A (Digital / Analog) converting units 32-1-1 through 32-3-Nw, and amplifying units 33-1-1 through 33-3-Nw.
[0048] The reproduction signal generation unit 31 performs rendering processing for each reproduction band, and generates a speaker reproduction signal, which is an output audio signal to be output.
[0049] The playback signal generation unit 31 has rendering processing units 41-1 to 41-3, HPFs (High Pass Filters) 42-1 to 42-Nt, BPFs (Band Pass Filters) 43-1 to 43-Ns, and LPFs (Low Pass Filters) 44-1 to 44-Nw.
[0050] The speaker system 22 includes speakers 51-1-1 to 51-1-Nt, speakers 51-2-1 to 51-2-Ns, and speakers 51-3-1 to 51-3-Nw, each having a different reproduction band.
[0051] In the following description, when there is no need to particularly distinguish between the speakers 51-1-1 to 51-1-Nt, they will also be simply referred to as the speaker 51-1.
[0052] Similarly, hereinafter, when there is no need to particularly distinguish between the speakers 51-2-1 to 51-2-Ns, they will also be simply referred to as speakers 51-2, and when there is no need to particularly distinguish between the speakers 51-3-1 to 51-3-Nw, they will also be simply referred to as speakers 51-3.
[0053] In the following description, when there is no need to particularly distinguish between the speakers 51-1 to 51-3, they will also be simply referred to as speakers 51. The speaker 51 constituting the speaker system 22 corresponds to the speaker SP11 shown in FIG.
[0054] The rendering processing units 41-1 to 41-3 perform rendering processing such as VBAP based on the object signals and metadata that make up the supplied object data, and generate speaker playback signals for the speakers 51.
[0055] For example, the rendering processing unit 41-1 performs rendering processing for each of N objects, and generates, for each object, a speaker playback signal whose output destination is one of the speakers 51-1-1 to 51-1-Nt.
[0056] The rendering processing unit 41-1 also adds the speaker playback signals generated for each object for the same speaker 51-1 to generate a final speaker playback signal for that speaker 51-1. The sound based on the speaker playback signal thus obtained contains the sounds of each of the N objects.
[0057] The rendering processing unit 41-1 supplies the final speaker playback signals generated for the speakers 51-1-1 to 51-1-Nt to the HPFs 42-1 to 42-Nt.
[0058] Similar to the rendering processing unit 41-1, the rendering processing unit 41-2 generates speaker playback signals for each speaker 51-2 for reproducing the sounds of N objects, with each speaker 51-2-1 to speaker 51-2-Ns being the final output destination, and supplies these to BPFs 43-1 to 43-Ns.
[0059] Similar to the rendering processing unit 41-1, the rendering processing unit 41-3 generates speaker playback signals for each speaker 51-3 for reproducing the sounds of N objects, with the speakers 51-3-1 to 51-3-Nw as the final output destinations, and supplies these signals to LPFs 44-1 to 44-Nw.
[0060] Hereinafter, when there is no need to particularly distinguish between the rendering processing units 41-1 to 41-3, they will also be simply referred to as the rendering processing unit 41.
[0061] The HPFs 42-1 to 42-Nt are HPFs that pass a frequency band that includes at least the reproduction band of the speaker 51-1, that is, high-frequency components, and cut off mid- and low-frequency components.
[0062] The HPFs 42-1 through 42-Nt perform filtering processing on the speaker playback signals supplied from the rendering processing unit 41-1, and supply the resulting speaker playback signals containing only high-frequency components to the D / A conversion units 32-1-1 through 32-1-Nt.
[0063] In the following, when there is no need to particularly distinguish between the HPFs 42-1 to 42-Nt, they will also be simply referred to as HPFs 42. The HPF 42 can be said to function as a band-limiting processing unit that performs HPF filtering processing, which is band-limiting processing according to the playback band of the speaker 51-1, on the input speaker playback signal to generate a band-limited speaker playback signal (band-limited signal).
[0064] The BPFs 43-1 to 43-Ns are BPFs that pass a frequency band that includes at least the reproduction band of the speaker 51-2, that is, mid-range components, and cut off other components.
[0065] The BPFs 43-1 to 43-Ns perform filtering processing on the speaker playback signals supplied from the rendering processing unit 41-2, and supply the resulting speaker playback signals containing only mid-frequency components to the D / A conversion units 32-2-1 to 32-2-Ns.
[0066] Hereinafter, when there is no need to particularly distinguish between BPF43-1 to BPF43-Ns, they will also be simply referred to as BPF 43. BPF 43 can be said to function as a band-limiting processing unit that performs BPF filtering processing, which is band-limiting processing according to the playback band of speaker 51-2, on the input speaker playback signal to generate a band-limited speaker playback signal (band-limited signal).
[0067] The LPFs 44-1 to 44-Nw are LPFs that pass a frequency band that includes at least the reproduction band of the speaker 51-3, that is, low-frequency components, and cut off mid- and high-frequency components.
[0068] The LPFs 44-1 through 44-Nw perform filtering processing on the speaker playback signals supplied from the rendering processing unit 41-3, and supply the resulting speaker playback signals containing only low-frequency components to the D / A conversion units 32-3-1 through 32-3-Nw.
[0069] Hereinafter, when there is no need to particularly distinguish between LPF 44-1 to LPF 44-Nw, they will be simply referred to as LPF 44. LPF 44 can be said to function as a band-limiting processing unit that performs LPF filtering processing, which is a band-limiting processing according to the reproduction band of speaker 51-3, on the input speaker reproduction signal to generate a band-limited speaker reproduction signal (band-limited signal).
[0070] The D / A conversion units 32-1-1 through 32-1-Nt D / A convert the speaker playback signals supplied from the HPFs 42-1 through 42-Nt, and supply the resulting analog speaker playback signals to the amplification units 33-1-1 through 33-1-Nt.
[0071] Hereinafter, when there is no need to particularly distinguish between the D / A conversion units 32-1-1 to 32-1-Nt, they will also be simply referred to as D / A conversion unit 32-1.
[0072] The D / A conversion units 32-2-1 through 32-2-Ns D / A convert the speaker playback signals supplied from the BPFs 43-1 through 43-Ns, and supply the resulting analog speaker playback signals to the amplification units 33-2-1 through 33-2-Ns.
[0073] Hereinafter, when there is no need to particularly distinguish between the D / A conversion units 32-2-1 to 32-2-Ns, they will also be simply referred to as D / A conversion units 32-2.
[0074] The D / A conversion units 32-3-1 through 32-3-Nw D / A convert the speaker playback signals supplied from the LPFs 44-1 through 44-Nw, and supply the resulting analog speaker playback signals to the amplification units 33-3-1 through 33-3-Nw.
[0075] Hereinafter, when there is no need to particularly distinguish between the D / A conversion units 32-3-1 to 32-3-Nw, they will also be simply referred to as D / A conversion units 32-3. Also, hereinafter, when there is no need to particularly distinguish between the D / A conversion units 32-1 to 32-3, they will also be simply referred to as D / A conversion units 32.
[0076] The amplifiers 33-1-1 through 33-1-Nt amplify the speaker playback signals supplied from the D / A converters 32-1-1 through 32-1-Nt, and supply the amplified signals to the speakers 51-1-1 through 51-1-Nt.
[0077] The amplifiers 33-2-1 through 33-2-Ns amplify the speaker playback signals supplied from the D / A converters 32-2-1 through 32-2-Ns, and supply the amplified signals to the speakers 51-2-1 through 51-2-Ns.
[0078] The amplifiers 33-3-1 through 33-3-Nw amplify the speaker playback signals supplied from the D / A converters 32-3-1 through 32-3-Nw, and supply the amplified signals to the speakers 51-3-1 through 51-3-Nw.
[0079] Hereinafter, when there is no need to particularly distinguish between the amplifiers 33-1-1 to 33-1-Nt, they will also be simply referred to as amplifiers 33-1, and when there is no need to particularly distinguish between the amplifiers 33-2-1 to 33-2-Ns, they will also be simply referred to as amplifiers 33-2.
[0080] Hereinafter, when there is no need to particularly distinguish between the amplifiers 33-3-1 to 33-3-Nw, they will also be simply referred to as amplifiers 33-3, and when there is no need to particularly distinguish between the amplifiers 33-1 to 33-3, they will also be simply referred to as amplifiers 33.
[0081] The D / A conversion unit 32 and the amplification unit 33 may be provided outside the sound processing device 21.
[0082] The speakers 51-1-1 through 51-1-Nt output sounds based on the speaker playback signals supplied from the amplifiers 33-1-1 through 33-1-Nt.
[0083] Each of the Nt speakers 51-1 constituting the speaker system 22 is a speaker called a tweeter that has a reproduction band mainly for high frequencies (high frequencies). In the speaker system 22, these Nt speakers 51-1 form one speaker layout for high frequencies.
[0084] The speakers 51-2-1 through 51-2-Ns output sounds based on the speaker playback signals supplied from the amplifiers 33-2-1 through 33-2-Ns.
[0085] Each of the Ns speakers 51-2 constituting the speaker system 22 is a speaker called a squawker, which has a reproduction band mainly in the mid-band (mid-range). In the speaker system 22, these Ns speakers 51-2 form one speaker layout for the mid-range.
[0086] The speakers 51-3-1 through 51-3-Nw output sounds based on the speaker playback signals supplied from the amplifiers 33-3-1 through 33-3-Nw.
[0087] Each of the Nw speakers 51-3 constituting the speaker system 22 is a speaker called a woofer that has a playback band mainly for the low band (low frequency range). In the speaker system 22, these Nw speakers 51-3 form one speaker layout for the low frequency range.
[0088] In this way, the speaker system 22 is composed of a plurality of speakers 51 having different playback bands, i.e., high, mid, and low bands. In other words, a plurality of speakers 51 having different playback bands are arranged around a listener who listens to content.
[0089] Although an example will be described here in which the speaker system 22 consisting of the speakers 51-1 to 51-3 is provided separately from the sound processing device 21, the speaker system 22 may be provided in the sound processing device 21. In other words, the speaker system 22 may be included in the sound processing device 21.
[0090] As described above, in the audio reproduction system 11, rendering processing is performed for each reproduction band of the speaker 51, that is, for each speaker layout of each reproduction band.
[0091] Therefore, for example, when VBAP is performed as the rendering process in the rendering process unit 41-1, the rendering process unit 41-1 selects the above-mentioned selected mesh from the meshes formed by the Nt speakers 51-1.
[0092] Similarly, the rendering processing unit 41-2 selects the above-mentioned selected mesh from the mesh formed by the Ns speakers 51-2, and the rendering processing unit 41-3 selects the above-mentioned selected mesh from the mesh formed by the Nw speakers 51-3.
[0093] The frequency characteristics of the HPF 42, BPF 43, and LPF 44, which function as band limiting processing units, i.e., the limiting band (pass band), are as shown in Fig. 3. In Fig. 3, the horizontal axis represents frequency (Hz), and the vertical axis represents sound pressure level (dB).
[0094] In FIG. 3, a broken line L11 indicates the frequency characteristics of the HPF 42, a broken line L12 indicates the frequency characteristics of the BPF 43, and a broken line L13 indicates the frequency characteristics of the LPF 44.
[0095] As can be seen from the broken line L11, the HPF 42 performs high-pass filtering, passing components in a higher frequency band than the other BPFs 43 and LPFs 44, that is, high-frequency components.
[0096] It can also be seen that the BPF 43 performs mid-pass filtering, which passes components in a frequency band higher than the LPF 44 and lower than the HPF 42, i.e., mid-frequency components. It can also be seen that the LPF 44 performs low-pass filtering, which passes components in a frequency band lower than the other BPFs 43 and HPF 42, i.e., low-frequency components.
[0097] Furthermore, here, the passbands of the HPF 42 and the BPF 43 cross over, and the passbands of the BPF 43 and the LPF 44 also cross over. Here, an example has been given in which the passbands of the HPF 42 and the BPF 43 and the passbands of the BPF 43 and the LPF 44 cross over, but this is not limiting. For example, neither the passbands of the HPF 42 and the BPF 43 nor the passbands of the BPF 43 and the LPF 44 need cross over, and one of them may have a crossover characteristic.
[0098] In the audio reproduction system 11, the Nt HPFs 42 have the same characteristics (frequency characteristics), but these Nt HPFs 42 may be filters (HPFs) having different characteristics from each other.
[0099] Alternatively, the HPF 42 may not be provided between the rendering processing unit 41-1 and the speaker 51-1, and the speaker playback signal obtained by the rendering processing unit 41-1 may be supplied to the speaker 51-1 via the D / A conversion unit 32-1 and the amplification unit 33-1. In other words, the filtering process (bandwidth limiting process) by the HPF 42 may not be performed, and sound based on the speaker playback signal may be played back by the speaker 51-1.
[0100] Similarly, the Ns BPFs 43 are assumed to have the same characteristics (frequency characteristics), but these BPFs 43 may have different characteristics from each other, or no BPF 43 may be provided between the rendering processing unit 41-2 and the speaker 51-2.
[0101] Furthermore, although the Nw LPFs 44 are assumed to have the same characteristics (frequency characteristics), these LPFs 44 may have different characteristics from each other, or no LPF 44 may be provided between the rendering processing unit 41-3 and the speaker 51-3.
[0102] <Explanation of the regeneration process> Next, the operation of the audio reproduction system 11 will be described. That is, the reproduction process by the audio reproduction system 11 will be described below with reference to the flowchart in Fig. 4. This reproduction process starts when object data of N objects constituting the content is supplied to each rendering processing unit 41.
[0103] In step S11, the rendering processing unit 41-1 performs rendering processing for the high-frequency speaker 51-1 based on the supplied N pieces of object data, and supplies the resulting speaker playback signal to the HPF .
[0104] That is, rendering is performed on a speaker layout consisting of Nt speakers 51-1, and speaker playback signals are generated as output audio signals. For example, in step S11, VBAP is performed as rendering processing using a mesh formed by the Nt speakers 51-1.
[0105] In step S12, the HPF 42 performs filtering processing (band-limiting processing) using the HPF on the speaker reproduction signal supplied from the rendering processing unit 41-1, and supplies the resulting band-limited speaker reproduction signal to the D / A conversion unit 32-1.
[0106] The D / A conversion unit 32-1 D / A converts the speaker playback signal supplied from the HPF 42 and supplies it to the amplification unit 33-1, and the amplification unit 33-1 amplifies the speaker playback signal supplied from the D / A conversion unit 32-1 and supplies it to the speaker 51-1.
[0107] In step S13, the rendering processing unit 41-2 performs rendering processing for the mid-band speaker 51-2 based on the supplied N pieces of object data, and supplies the resulting speaker playback signal to the BPF 43.
[0108] For example, in step S13, a mesh formed by Ns speakers 51-2 is used to perform VBAP as the rendering process.
[0109] In step S14, the BPF 43 performs filtering processing (band limiting processing) using the BPF on the speaker reproduction signal supplied from the rendering processing unit 41-2, and supplies the resulting band-limited speaker reproduction signal to the D / A conversion unit 32-2.
[0110] The D / A conversion unit 32-2 D / A converts the speaker playback signal supplied from the BPF 43 and supplies it to the amplification unit 33-2, and the amplification unit 33-2 amplifies the speaker playback signal supplied from the D / A conversion unit 32-2 and supplies it to the speaker 51-2.
[0111] In step S15, the rendering processing unit 41-3 performs rendering processing for the low-band speaker 51-3 based on the supplied N pieces of object data, and supplies the resulting speaker playback signal to the LPF 44.
[0112] For example, in step S15, a mesh formed by Nw speakers 51-3 is used, and VBAP is performed as the rendering process.
[0113] In step S16, the LPF 44 performs filtering (band-limiting) using the LPF on the speaker reproduction signal supplied from the rendering processing unit 41-3, and supplies the resulting band-limited speaker reproduction signal to the D / A conversion unit 32-3.
[0114] The D / A conversion unit 32-3 D / A converts the speaker playback signal supplied from the LPF 44 and supplies it to the amplification unit 33-3, and the amplification unit 33-3 amplifies the speaker playback signal supplied from the D / A conversion unit 32-3 and supplies it to the speaker 51-3.
[0115] In step S17, all the speakers 51 that make up the speaker system 22 output sounds based on the speaker playback signals supplied from the amplifier 33, and the playback process ends.
[0116] When sounds based on the speaker playback signals are output from all speakers 51, the sounds of N objects are reproduced for each playback band according to the speaker layout for each playback band. The sound images of each of the N objects are then localized at the object positions indicated by the position information included in the metadata of each object.
[0117] In this way, the audio playback system 11 plays back content by performing rendering processing for each playback band of the speaker 51, i.e., for each speaker layout of each of the multiple playback bands. In this way, it is possible to suppress deterioration in sound quality due to the playback band of the speaker 51, and to play back audio with higher sound quality.
[0118] Specifically, for example, the audio reproduction system 11 includes a mixture of speakers 51 with different reproduction bands.
[0119] However, in the audio reproduction system 11, a speaker layout configuration is prepared for each of a plurality of reproduction bands, and each object is rendered and reproduced for each reproduction band.
[0120] Therefore, objects are played back at appropriate positions for each speaker layout of each playback band, realizing more appropriate object-based audio rendering playback. This makes it possible to avoid degradation of sound quality, such as sound disappearing due to the frequency band and localization position of the object. In other words, audio playback with higher sound quality can be achieved.
[0121] Second Embodiment <Example of audio playback system configuration> In the above, an example has been described in which the output of the rendering processing unit 41 is subjected to band-limiting filtering processing in accordance with the target speaker layout.
[0122] However, the present invention is not limited to this. For example, the object signal input to the rendering processing unit 41 may be subjected to band-limiting filtering processing in accordance with the target speaker layout.
[0123] In such a case, the audio reproduction system may have the configuration shown in Fig. 5. In Fig. 5, the same reference numerals are used to designate parts that correspond to those in Fig. 2, and the description thereof will be omitted where appropriate.
[0124] The audio reproduction system 81 shown in FIG. 5 includes a sound processing device 91 and a speaker system 22.
[0125] The sound processing device 91 also includes a playback signal generating unit 101, D / A conversion units 32-1-1 through 32-3-Nw, and amplification units 33-1-1 through 33-3-Nw.
[0126] The reproduction signal generation unit 101 has HPFs 42-1 to 42-N, BPFs 43-1 to 43-N, LPFs 44-1 to 44-N, and rendering processing units 41-1 to 41-3.
[0127] The configuration of the audio reproduction system 81 differs from that of the audio reproduction system 11 shown in FIG. 2 in that an audio processing device 91 is provided instead of the audio processing device 21, but in other respects it has the same configuration as the audio reproduction system 11.
[0128] In particular, the sound processing device 91 has a configuration in which the reproduction signal generating unit 31 of the sound processing device 21 is replaced with a reproduction signal generating unit 101 .
[0129] As described above, in the reproduction signal generating unit 31, the HPF 42, BPF 43, and LPF 44 are provided after the rendering processing unit 41.
[0130] On the other hand, in the reproduction signal generating unit 101, an HPF 42, a BPF 43, and an LPF 44 are provided in front of the rendering processing unit 41.
[0131] Moreover, the playback signal generation unit 101 is provided with N HPFs 42, N BPFs 43, and N LPFs 44 in order to perform filtering (bandwidth limiting) on the object signals of the N objects that are input to the rendering processing unit 41. That is, an HPF 42, a BPF 43, and an LPF 44 are provided for each object.
[0132] Therefore, each of the HPFs 42-1 to 42-N performs filtering processing on the object signals of each of the N pieces of object data supplied, and supplies the resulting object signals containing only high-frequency components to the rendering processing unit 41-1. Note that the HPFs 42-1 to 42-N perform the same filtering processing (bandwidth limiting processing) as the HPF 42 in the playback signal generation unit 31.
[0133] Similarly, each of the BPFs 43-1 to 43-N performs filtering processing on the object signals of each of the N pieces of object data supplied, and supplies the resulting object signals containing only mid-frequency components to the rendering processing unit 41-2. The BPFs 43-1 to 43-N perform the same filtering processing (bandwidth limiting processing) as the BPF 43 in the playback signal generation unit 31.
[0134] Each of the LPFs 44-1 to 44-N performs filtering processing on the object signals of each of the N pieces of object data supplied, and supplies the resulting object signals containing only low-frequency components to the rendering processing unit 41-3. The LPFs 44-1 to 44-N perform the same filtering processing (bandwidth limiting processing) as the LPF 44 in the playback signal generation unit 31.
[0135] In this way, while the audio reproduction system 11 shown in FIG. 2 has an HPF 42, BPF 43, and LPF 44 provided for each speaker 51, the audio reproduction system 81 has an HPF 42, BPF 43, and LPF 44 provided for each object.
[0136] In this example, since the content is made up of N objects, the audio playback system 81 is provided with N HPFs 42, N BPFs 43, and N LPFs 44.
[0137] In this example, as in the audio playback system 11, the N HPFs 42 have the same frequency characteristics, but these N HPFs 42 may be filters (HPFs) with different characteristics from each other, or no HPF 42 may be provided upstream of the rendering processing unit 41-1.
[0138] Similarly, the N BPFs 43 are assumed to have the same characteristics (frequency characteristics), but these BPFs 43 may have different characteristics from each other, or no BPF 43 may be provided in front of the rendering processing unit 41-2.
[0139] Furthermore, although the N LPFs 44 are assumed to have the same characteristics (frequency characteristics), these LPFs 44 may have different characteristics from each other, or no LPF 44 may be provided in front of the rendering processing unit 41-3.
[0140] <Explanation of the regeneration process> Next, the playback process performed by the audio playback system 81 will be described with reference to the flowchart of FIG.
[0141] In step S41, each of the HPFs 42-1 to 42-N performs filtering processing using an HPF on each of the supplied object signals of the N objects, and supplies the resulting band-limited object signals to the rendering processing unit 41-1.
[0142] In step S42, the rendering processing unit 41-1 performs rendering processing for the high-band speaker 51-1 based on the metadata of each of the N objects supplied and the N object signals supplied from the HPF 42-1 to HPF 42-N.
[0143] For example, in step S42, the same processing as in step S11 in Fig. 4 is performed. The rendering processing unit 41-1 supplies speaker playback signals corresponding to each speaker 51-1 obtained by the rendering processing to the D / A conversion units 32-1-1 to 32-1-Nt.
[0144] The D / A conversion unit 32-1 D / A converts the speaker playback signal supplied from the rendering processing unit 41-1 and supplies it to the amplification unit 33-1, and the amplification unit 33-1 amplifies the speaker playback signal supplied from the D / A conversion unit 32-1 and supplies it to the speaker 51-1.
[0145] In step S43, each of the BPFs 43-1 to 43-N performs a filtering process using a BPF on each of the supplied object signals of the N objects, and supplies the resulting band-limited object signals to the rendering processing unit 41-2.
[0146] In step S44, the rendering processing unit 41-2 performs rendering processing for the mid-band speaker 51-2 based on the metadata of each of the N objects supplied and the N object signals supplied from the BPFs 43-1 to 43-N.
[0147] For example, in step S44, the same processing as step S13 in Fig. 4 is performed. The rendering processing unit 41-2 supplies speaker playback signals corresponding to each speaker 51-2 obtained by the rendering processing to the D / A conversion units 32-2-1 to 32-2-Ns.
[0148] The D / A conversion unit 32-2 D / A converts the speaker playback signal supplied from the rendering processing unit 41-2 and supplies it to the amplification unit 33-2, and the amplification unit 33-2 amplifies the speaker playback signal supplied from the D / A conversion unit 32-2 and supplies it to the speaker 51-2.
[0149] In step S45, each of the LPFs 44-1 to 44-N performs filtering processing using an LPF on each of the supplied object signals of the N objects, and supplies the resulting band-limited object signals to the rendering processing unit 41-3.
[0150] In step S46, the rendering processing unit 41-3 performs rendering processing for the low-band speaker 51-3 based on the metadata of each of the N objects supplied and the N object signals supplied from the LPFs 44-1 to 44-N.
[0151] For example, in step S46, the same process as step S15 in Fig. 4 is performed. The rendering processing unit 41-3 supplies speaker playback signals corresponding to each speaker 51-3 obtained by the rendering process to the D / A conversion units 32-3-1 to 32-3-Nw.
[0152] The D / A conversion unit 32-3 D / A converts the speaker playback signal supplied from the rendering processing unit 41-3 and supplies it to the amplification unit 33-3, and the amplification unit 33-3 amplifies the speaker playback signal supplied from the D / A conversion unit 32-3 and supplies it to the speaker 51-3.
[0153] Once rendering processing has been performed on the speaker layout for each playback band in this manner, processing in step S47 is then performed and the playback processing ends. However, since the processing in step S47 is the same as the processing in step S17 in FIG. 4, its description will be omitted.
[0154] In this way, the audio playback system 81 performs filtering processing for each object, then performs rendering processing for each speaker layout of each of the multiple playback bands, and plays back the content. In this way, it is possible to suppress deterioration in sound quality caused by the playback band of the speaker 51, and to play back audio with higher sound quality.
[0155] A configuration such as audio playback system 81 that performs filtering processing before rendering processing can reduce the amount of processing compared to audio playback system 11, especially when the number of objects that make up the content (number of objects N) is small.
[0156] For example, assume that the amount of filtering processing is the same for the HPF 42, BPF 43, and LPF 44. In such a case, the amount of filtering processing (number of processes) required in the audio playback system 81 is the number of objects N × 3, where "3" is the number of rendering processing units 41.
[0157] On the other hand, in the audio reproduction system 11, the filtering process is performed the number of times equal to the total number (Nt+Ns+Nw) of speakers 51 that make up the speaker system 22.
[0158] Therefore, when the number of objects N×3 is smaller than the total number of speakers 51 (Nt+Ns+Nw), the configuration of audio playback system 81 can reduce the number of filtering processes (number of processing times) compared to audio playback system 11, and as a result, the overall processing volume can be kept low.
[0159] Third Embodiment <Example of audio playback system configuration> Incidentally, whether performing the filtering process before or after the rendering process reduces the amount of processing required depends on the number of objects N, the total number of speakers 51, and the number of types (playback bands) of speakers 51 (the number of rendering processing units 41).
[0160] Therefore, whether the filtering process is performed before or after the rendering process may be switched based on a determination criterion based on the number of objects N and the total number of speakers 51, for example.
[0161] In such a case, the audio playback system may be configured, for example, as shown in Figure 7. In Figure 7, parts corresponding to those in Figure 2 or Figure 5 are given the same reference numerals, and their explanation will be omitted where appropriate.
[0162] The audio reproduction system 131 shown in FIG.
[0163] The sound processing device 141 also includes a selection unit 151, a playback signal generation unit 31, a playback signal generation unit 101, D / A conversion units 32-1-1 through 32-3-Nw, and amplification units 33-1-1 through 33-3-Nw.
[0164] The reproduction signal generating section 31 has the same configuration as in FIG. 2, and the reproduction signal generating section 101 has the same configuration as in FIG.
[0165] In this example, object data for each of N objects is input to the selection unit 151. Based on the number of objects N and the total number of speakers 51, the selection unit 151 selects either the playback signal generation unit 31 or the playback signal generation unit 101 as the output destination of the object data, and outputs the object data to the selected output destination.
[0166] In other words, the selection unit 151 selects, for each object, whether to have the playback signal generation unit 31 perform rendering processing and then perform band limiting processing, or to have the playback signal generation unit 101 perform band limiting processing and then perform rendering processing.
[0167] Therefore, in the audio playback system 131, either the playback signal generation unit 31 or the playback signal generation unit 101 generates a speaker playback signal based on the object data, and supplies the speaker playback signal to the D / A conversion unit 32.
[0168] <Explanation of the regeneration process> Next, the playback process by the audio playback system 131 will be described with reference to the flowchart in Fig. 8. This playback process starts when the selection unit 151 is supplied with object data of N objects that make up the content.
[0169] In step S71, the selection unit 151 determines whether or not to perform filtering processing before rendering processing based on the number N of supplied object data, the total number of speakers 51, and the number of playback bands (the number of rendering processing units 41). That is, the selection unit 151 selects an output destination for the supplied object data. Note that in this example, the number of playback bands, i.e., the number of rendering processing units 41, is "3."
[0170] For example, if the number of objects N×3 is smaller than the total number of speakers 51 (Nt+Ns+Nw), the selection unit 151 determines that filtering processing is to be performed first.
[0171] On the other hand, for example, when the number of objects N×3 is equal to or greater than the total number of speakers 51 (Nt+Ns+Nw), the selection unit 151 determines that the filtering process is to be performed after the rendering process.
[0172] If it is determined in step S71 that filtering processing is to be performed first, the selection unit 151 selects the playback signal generation unit 101 as the output destination of the supplied object data, and then the process proceeds to step S72.
[0173] In this case, the selection unit 151 supplies the object signal of the supplied object data to the HPF 42, BPF 43, and LPF 44 of the playback signal generation unit 101, and also supplies the metadata of the object data to the rendering processing unit 41 of the playback signal generation unit 101.
[0174] When the object data is supplied to the playback signal generation unit 101 in this way, the processes of steps S72 to S77 are performed, but these processes are the same as the processes of steps S41 to S46 in Fig. 6, so their description will be omitted. After these processes are performed, a speaker playback signal is supplied to the speaker 51.
[0175] On the other hand, if it is determined in step S71 that the filtering process will be performed later, the selection unit 151 selects the playback signal generation unit 31 as the output destination of the supplied object data, and then the process proceeds to step S78.
[0176] In this case, the selection unit 151 supplies the supplied object data, that is, the object signal and metadata, to the rendering processing unit 41 of the playback signal generation unit 31 .
[0177] After the object data is supplied to the playback signal generation unit 31, the processes of steps S78 to S83 are performed, but these processes are the same as the processes of steps S11 to S16 in Fig. 4, so a description thereof will be omitted. After these processes are performed, a speaker playback signal is supplied to the speaker 51.
[0178] After the processing of step S77 or step S83 is performed, the processing of step S84 is then performed.
[0179] That is, in step S84, all the speakers 51 constituting the speaker system 22 output sounds based on the speaker playback signals supplied from the amplifier 33, and the playback process ends.
[0180] In this way, the audio reproduction system 131 selects the reproduction signal generation unit 31 or the reproduction signal generation unit 101 that requires less processing and performs filtering processing and rendering processing, based on the number of objects N and the total number of speakers 51. In other words, it is possible to switch between the reproduction signal generation unit 31 and the reproduction signal generation unit 101 to perform rendering processing and filtering processing, depending on the number of objects N and the total number of speakers 51.
[0181] In this way, audio can be reproduced with higher sound quality with a smaller amount of processing. Note that the switching (selection) of whether the rendering process and filtering process should be performed by the reproduction signal generation unit 31 or the reproduction signal generation unit 101 may be performed for each frame, etc.
[0182] In particular, in the reproduction signal generation unit 31, band limitation on the speaker reproduction signal according to the speaker layout for each reproduction band is effective when the number of objects N is large. On the other hand, in the reproduction signal generation unit 101, band limitation on the object signal according to the speaker layout for each reproduction band is effective when the number of objects N is small.
[0183] <Fourth embodiment> <Example of audio playback system configuration> Furthermore, the speaker layout for reproducing the sound of an object may be switched depending on the content of the object, that is, the characteristics of the object such as the type of sound source of the object or the characteristics of the object signal.
[0184] In such a case, the audio reproduction system may be configured as shown in Fig. 9. In Fig. 9, the same reference numerals are used to designate parts that correspond to those in Fig. 2, and the description thereof will be omitted where appropriate.
[0185] The audio reproduction system 181 shown in FIG. 9 includes a sound processing device 191 and a speaker system 192.
[0186] The sound processing device 191 has a playback signal generation unit 201, D / A conversion units 32-1-1 to 32-1-Nt, D / A conversion units 32-3-1 to 32-3-Nw, amplification units 33-1-1 to 33-1-Nt, and amplification units 33-3-1 to 33-3-Nw.
[0187] The playback signal generating unit 201 also includes a determining unit 211, a switching unit 212, a rendering processing unit 41-1, and a rendering processing unit 41-3.
[0188] The speaker system 192 includes speakers 51-1-1 through 51-1-Nt and speakers 51-3-1 through 51-3-Nw.
[0189] For example, the reproduction band of the speaker 51-1 can be partially overlapped with the reproduction band of the speaker 51-3, that is, the speaker 51-1 and the speaker 51-3 can have a partial common reproduction band.
[0190] Furthermore, the reproduction signal generation unit 201 is not provided with a filter that functions as a band limiting processing unit such as the HPF 42. Furthermore, the speaker system 192 is provided with a speaker 51-1 which is a tweeter and a speaker 51-3 which is a woofer, but is not provided with a speaker 51-2 which is a squawker. Note that, like the speaker system 22 described above, the speaker system 192 may also be provided with a speaker 51-2 which is a squawker.
[0191] The determination unit 211 is supplied with object data of each of the N objects.
[0192] The determination unit 211 performs a determination process for each object based on the object signal and metadata contained in the supplied object data to determine which rendering processing unit 41 should perform the rendering process, i.e., which speaker layout should be used for playback.
[0193] For example, the determination unit 211 determines (decides) for each object whether rendering processing should be performed only by the rendering processing unit 41-1, only by the rendering processing unit 41-3, or by both the rendering processing units 41-1 and 41-3. At this time, the determination can be made using at least one of information about the object, such as an object signal and metadata.
[0194] The determination unit 211 supplies the supplied object data to the switching unit 212, and controls the switching unit 212 based on the result of the determination process to supply the object data to the rendering processing unit 41 according to the result of the determination process.
[0195] For example, in the determination process, it may be determined for each object which reproduction band should be rendered into the speaker layout based on the frequency characteristics of the object signal as the characteristics of the object.
[0196] In such a case, for example, the determination unit 211 performs frequency analysis on the supplied object signal using FFT (Fast Fourier Transform) or the like, and from the information indicating the frequency characteristics obtained as a result, determines (decides) which playback band to render to the speaker layout, i.e., which rendering processing unit 41 to use to perform the rendering processing.
[0197] Specifically, for example, when the object signal contains only low-frequency components, rendering processing can be performed only by the rendering processing unit 41-3.
[0198] For example, in the audio reproduction system 11, each object is rendered by a rendering processor 41 that supports all reproduction bands. However, if the object signal contains only low-frequency components, there is no degradation in sound quality even if rendering is performed only by the rendering processor 41-3.
[0199] In the audio playback system 181, for example, an object signal containing only low-frequency components is rendered only by the rendering processing unit 41-3 corresponding to the low frequency band, thereby reducing the amount of processing without causing degradation in sound quality.
[0200] Furthermore, for example, if an object signal contains both low-frequency components and high-frequency components, rendering processing can be performed by both the rendering processing unit 41-1 and the rendering processing unit 41-3.
[0201] Additionally, information about the object may be included, for example in metadata.
[0202] Specifically, it is assumed that metadata contains sound source type information indicating the type of sound source an object is, such as a musical instrument such as a guitar, or vocals.
[0203] In such a case, for example, the determining unit 211 determines (decides) which rendering processing unit 41 should perform the rendering processing, based on the sound source type information included in the metadata.
[0204] In this case, if the object is a sound source that contains a lot of high-frequency components, such as a hi-hat, the rendering process for that object can be performed by the rendering process unit 41-1 that targets high-frequency components. Note that it may be predetermined which rendering process unit 41 will perform the rendering process for an object of a given sound source type. Also, the sound source type of an object may be identified from the file name of the object signal, etc.
[0205] Alternatively, for example, a content creator may specify in advance which object should be rendered by which rendering processing unit 41, and the specification information indicating the result of the specification may be included in the metadata as information about the object.
[0206] In such a case, the determination unit 211 determines (decides) which rendering processing unit 41 should perform rendering processing on the object based on the designation information included in the metadata. Note that the designation information may be supplied to the determination unit 211 separately from the object data.
[0207] The switching unit 212, under the control of the determining unit 211, switches the output destination of the object data supplied from the determining unit 211 for each object.
[0208] That is, under the control of the determination unit 211, the switching unit 212 supplies the object data to the rendering processing unit 41-1, to the rendering processing unit 41-3, or to both the rendering processing unit 41-1 and the rendering processing unit 41-3.
[0209] <Explanation of the regeneration process> Next, the playback process by the audio playback system 181 will be described with reference to the flowchart in Fig. 10. This playback process starts when the determination unit 211 is supplied with object data of N objects that make up the content.
[0210] In step S111, the determination unit 211 performs a determination process for each object based on the supplied object data.
[0211] For example, in the determination process, it is determined which rendering processing unit 41 corresponding to which playback band should perform the rendering process, based on at least the object signal and metadata. The determination unit 211 supplies the supplied object data to the switching unit 212, and controls the output of the object data by the switching unit 212 based on the result of the determination process.
[0212] In step S112, the switching unit 212, under the control of the determining unit 211, supplies the object data supplied from the determining unit 211 in accordance with the result of the determination process.
[0213] That is, the switching unit 212 supplies the object data supplied from the determination unit 211 to the rendering processing unit 41-1, the rendering processing unit 41-3, or both the rendering processing unit 41-1 and the rendering processing unit 41-3 for each object.
[0214] In step S113, the rendering processing unit 41-1 performs rendering processing for the high-band speaker 51-1 based on the object data supplied from the switching unit 212, and supplies the resulting speaker playback signal to the speaker 51-1 via the D / A conversion unit 32-1 and the amplification unit 33-1.
[0215] In step S114, the rendering processing unit 41-3 performs rendering processing for the low-band speaker 51-3 based on the object data supplied from the switching unit 212, and supplies the resulting speaker playback signal to the speaker 51-3 via the D / A conversion unit 32-3 and the amplification unit 33-3.
[0216] For example, in steps S113 and S114, the same processes as in steps S11 and S15 in FIG. 4 are performed.
[0217] In step S115, all the speakers 51 constituting the speaker system 192 output sounds based on the speaker playback signals supplied from the amplifier 33, and the playback process ends.
[0218] In this example, sounds are output from a high-band speaker 51-1 and a low-band speaker 51-3, and the sounds of N objects in the content are reproduced.
[0219] In this way, the audio playback system 181 determines which rendering processing unit 41 corresponding to the playback band should perform the processing, based on at least one of the object signal and information about the object, such as metadata, and performs the rendering processing according to the result of the determination.
[0220] In this way, it is possible to selectively perform rendering processing in the rendering processing unit 41 that corresponds to the appropriate reproduction band, thereby enabling audio reproduction with higher sound quality.
[0221] In this example, by switching (selecting) the speaker layout for each playback band to be subjected to rendering processing according to the main frequency band components of the object signal, it is possible to minimize the increase in processing volume due to multiple rendering processes. In other words, it is possible to omit rendering processes for unnecessary playback bands and reduce the processing volume.
[0222] Fifth Embodiment <Example of audio playback system configuration> By the way, subwoofers are sometimes added to reinforce the low frequencies during audio playback, and a technique called bass management is sometimes used.
[0223] Bass management involves filtering the low-frequency components from the main speaker signal and routing the filtered signals to one or more subwoofers.
[0224] However, when multiple subwoofers are used, the same low frequency components are generally reproduced by all the subwoofers, which impairs the sense of object positioning.
[0225] To avoid this loss of sense of positioning, it is also possible to route the low-frequency components of each main speaker to different subwoofers, so that the subwoofer that reproduces the low-frequency components changes depending on the direction of the object's position. However, in such cases, the behavior of the entire system, including routing, is up to the design, which can be complex and difficult.
[0226] In contrast, with this technology, rendering is performed for each of several playback bands, and content is played back using a speaker layout for each of those playback bands, enabling bass management that can suppress a decrease in the sense of object positioning without requiring complex design.
[0227] Furthermore, depending on the content, an audio signal of an LFE (Low Frequency Effect) channel for a subwoofer (hereinafter also referred to as an LFE channel signal) may be prepared in advance. In such a case, the present technology may appropriately adjust the gain of the LFE channel signal and add it to the speaker playback signal of the subwoofer.
[0228] In this way, when an LFE channel signal is prepared in advance in the content and bass management is also performed, the audio playback system may look like that shown in FIG. 11, for example.
[0229] The audio reproduction system 241 shown in FIG. 11 includes a sound processing device 251 and a speaker system 252, and reproduces object-based audio content based on supplied object data.
[0230] In this example, the content data consists of object data of N objects and a channel-based LFE channel signal. In this case, since the LFE channel signal is a channel-based audio signal, metadata including position information and the like is not supplied. Furthermore, the number of objects N can be any number.
[0231] The sound processing device 251 has a playback signal generation unit 261, D / A conversion units 271-1-1 through 271-2-Nsw, and amplification units 272-1-1 through 272-2-Nsw.
[0232] The playback signal generation unit 261 also includes a rendering processing unit 281-1, a rendering processing unit 281-2, HPFs 282-1 to 282-Nls, and LPFs 283-1 to 283-Nsw.
[0233] The speaker system 252 has speakers 291-1-1 to 291-1-Nls and speakers 291-2-1 to 291-2-Nsw that have different reproduction bands.
[0234] Hereinafter, when there is no need to particularly distinguish between the speakers 291-1-1 to 291-1-Nls, they will also be simply referred to as speaker 291-1, and when there is no need to particularly distinguish between the speakers 291-2-1 to 291-2-Nsw, they will also be simply referred to as speaker 291-2.
[0235] In the following description, when there is no need to particularly distinguish between the speaker 291-1 and the speaker 291-2, they will also be simply referred to as the speaker 291.
[0236] In this example, the Nls speakers 291-1 constituting the speaker system 252 are speakers called wideband loudspeakers, which have a wide band (wideband) as their reproduction band, mainly ranging from relatively low to high frequencies. In the speaker system 252, these Nls speakers 291-1 form one wideband speaker layout.
[0237] Furthermore, each of the Nsw speakers 291-2 constituting the speaker system 252 is a speaker called a subwoofer for reinforcing low frequencies, having a low-frequency reproduction band of, for example, about 100 Hz or less. In the speaker system 252, these Nsw speakers 291-2 form one speaker layout for the low frequency band.
[0238] The rendering processing unit 281-1 and the rendering processing unit 281-2 are each supplied with object data of N objects that make up the content.
[0239] The rendering processing units 281-1 and 281-2 perform rendering processing such as VBAP based on the object signal and metadata that constitute the supplied object data. That is, the rendering processing units 281-1 and 281-2 perform processing similar to that in the rendering processing unit 41.
[0240] For example, the rendering processing unit 281-1 generates speaker playback signals for each object, with the speakers 291-1-1 to 291-1-Nls as the output destinations, respectively. Then, the speaker playback signals for each object generated for the same speaker 291-1 are added together to form a final speaker playback signal.
[0241] In particular, when VBAP is performed as the rendering process, the rendering processing unit 281-1 uses a mesh formed by Nls speakers 291-1.
[0242] The rendering processing unit 281-1 supplies the final speaker playback signals generated for the speakers 291-1-1 through 291-1-Nls to the HPFs 282-1 through 282-Nls.
[0243] Similarly to the rendering processing unit 281-1, the rendering processing unit 281-2 generates speaker playback signals for each speaker 291-2, with the speakers 291-2-1 to 291-2-Nsw serving as final output destinations, respectively. In particular, when VBAP is performed as the rendering processing, the rendering processing unit 281-2 uses a mesh formed by the Nsw speakers 291-2.
[0244] Furthermore, the rendering processing unit 281-2 is supplied with an LFE channel signal.
[0245] Generally, since an LFE channel signal does not have localization information (position information), the rendering processing unit 281-2 does not perform rendering processing such as VBAP, but rather multiplies the LFE channel signal by a certain coefficient so that the signal is distributed to all speakers 291-2 and then outputs it.
[0246] That is, for each speaker 291-2, the rendering processing unit 281-2 adds a signal obtained by adjusting the gain of the LFE channel signal using a predetermined coefficient to the speaker playback signal corresponding to the speaker 291-2 obtained by the rendering processing, to obtain a final speaker playback signal. At this time, the coefficient used in the gain adjustment is, for example, (1 / Nsw) 1 / 2 etc.
[0247] The rendering processing unit 281-2 supplies the final speaker playback signals generated for the speakers 291-2-1 through 291-2-Nsw to the LPFs 283-1 through 283-Nsw.
[0248] Hereinafter, when there is no need to particularly distinguish between the rendering processing units 281-1 and 281-2, they will also be simply referred to as the rendering processing units 281.
[0249] The HPFs 282-1 to 282-Nls are HPFs that pass frequency components in a frequency band that includes at least the reproduction band of the speaker 291-1, that is, a relatively wide predetermined band.
[0250] HPF282-1 to HPF282-Nls perform filtering processing on the speaker playback signal supplied from the rendering processing unit 281-1, and supply the resulting speaker playback signal consisting of frequency components in a predetermined band to D / A conversion units 271-1-1 to 271-1-Nls.
[0251] In the following, when there is no need to particularly distinguish between the HPFs 282-1 to 282-Nls, they will also be simply referred to as HPFs 282. Like the HPF 42 shown in Fig. 2, this HPF 282 also functions as a band limiting processing unit that performs band limiting processing in accordance with the playback band of the speaker 291-1.
[0252] The LPFs 283-1 to 283-Nsw are LPFs that pass frequency components in a frequency band that includes at least the reproduction band of the speaker 291-2, that is, in a band of, for example, about 100 Hz or less.
[0253] The LPFs 283-1 through 283-Nsw perform filtering processing on the speaker playback signal supplied from the rendering processing unit 281-2, and supply the resulting speaker playback signal consisting of low-band frequency components to the D / A conversion units 271-2-1 through 271-2-Nsw.
[0254] In the following description, when there is no need to particularly distinguish between LPF 283-1 to LPF 283-Nsw, they will also be simply referred to as LPF 283. Like LPF 44 shown in Fig. 2, this LPF 283 also functions as a band limiting processing unit that performs band limiting processing in accordance with the reproduction band of speaker 291-2.
[0255] The D / A conversion units 271-1-1 through 271-1-Nls D / A convert the speaker playback signals supplied from the HPFs 282-1 through 282-Nls, and supply the resulting analog speaker playback signals to the amplification units 272-1-1 through 272-1-Nls.
[0256] Hereinafter, when there is no need to particularly distinguish between the D / A conversion units 271-1-1 to 271-1-Nls, they will also be simply referred to as D / A conversion unit 271-1.
[0257] The D / A conversion units 271-2-1 through 271-2-Nsw D / A convert the speaker reproduction signals supplied from the LPFs 283-1 through 283-Nsw, and supply the analog speaker reproduction signals obtained as a result to the amplification units 272-2-1 through 272-2-Nsw.
[0258] Hereinafter, when there is no need to particularly distinguish between the D / A conversion units 271-2-1 to 271-2-Nsw, they will also be simply referred to as D / A conversion units 271-2. Also, hereinafter, when there is no need to particularly distinguish between the D / A conversion units 271-1 and 271-2, they will also be simply referred to as D / A conversion units 271.
[0259] The amplifiers 272-1-1 through 272-1-Nls amplify the speaker playback signals supplied from the D / A converters 271-1-1 through 271-1-Nls, and supply the amplified signals to the speakers 291-1-1 through 291-1-Nls.
[0260] The amplifiers 272-2-1 through 272-2-Nsw amplify the speaker playback signals supplied from the D / A converters 271-2-1 through 271-2-Nsw, and supply the amplified signals to the speakers 291-2-1 through 291-2-Nsw.
[0261] In the following, when there is no need to particularly distinguish between the amplifier units 272-1-1 to 272-1-Nls, they will simply be referred to as amplifier unit 272-1, and when there is no need to particularly distinguish between the amplifier units 272-2-1 to 272-2-Nsw, they will simply be referred to as amplifier unit 272-2.
[0262] Furthermore, hereinafter, when there is no need to particularly distinguish between the amplifier unit 272-1 and the amplifier unit 272-2, they will also be simply referred to as amplifier unit 272.
[0263] The speakers 291-1-1 through 291-1-Nls output sounds based on the speaker playback signals supplied from the amplifiers 272-1-1 through 272-1-Nls.
[0264] The speakers 291-2-1 through 291-2-Nsw output sounds based on the speaker playback signals supplied from the amplifiers 272-2-1 through 272-2-Nsw.
[0265] In this way, the speaker system 252 is configured from a plurality of speakers 291 having different reproduction bands. That is, a plurality of speakers 291 having different reproduction bands are arranged around a listener who listens to content.
[0266] Although an example in which the speaker system 252 is provided separately from the sound processing device 251 will be described here, the speaker system 252 may be provided in the sound processing device 251 .
[0267] The frequency characteristics of the HPF 282 and LPF 283, which function as band limiting processing units, i.e., the limiting band (pass band), are as shown in Fig. 12. In Fig. 12, the horizontal axis represents frequency (Hz) and the vertical axis represents sound pressure level (dB).
[0268] In FIG. 12, a polygonal line L21 indicates the frequency characteristics of the HPF 282, and a polygonal line L22 indicates the frequency characteristics of the LPF 283.
[0269] As can be seen from the broken line L21, the HPF 282 performs high-pass filtering that passes components in a frequency band higher than that of the LPF 283, i.e., a wide frequency band of about 100 Hz or higher. In contrast, as can be seen from the broken line L22, the LPF 283 performs low-pass filtering that passes components in a frequency band lower than that of the HPF 282, i.e., low-frequency components of about 100 Hz or lower. Here, the pass bands of the HPF 282 and LPF 283 cross over each other, but the pass bands of the HPF 282 and LPF 283 do not necessarily have to cross over each other.
[0270] In the audio reproduction system 241, the Nls HPFs 282 are assumed to have the same characteristics (frequency characteristics), but these Nls HPFs 282 may be filters (HPFs) having different characteristics from each other. Also, the HPF 282 may not be provided between the rendering processing unit 281-1 and the speaker 291-1.
[0271] Similarly, the Nsw LPFs 283 are assumed to have the same characteristics (frequency characteristics), but these LPFs 283 may have different characteristics from each other, or no LPF 283 may be provided between the rendering processing unit 281-2 and the speaker 291-2.
[0272] <Explanation of the regeneration process> Next, the playback process by the audio playback system 241 will be described with reference to the flowchart of FIG.
[0273] In step S141, the rendering processing unit 281-1 performs rendering processing for the wideband speaker 291-1 based on the supplied N pieces of object data, and supplies the resulting speaker playback signal to the HPF 282. For example, in step S141, processing similar to step S11 in FIG.
[0274] In step S142, the HPF 282 performs filtering processing (band limiting processing) using the HPF on the speaker playback signal supplied from the rendering processing unit 281-1.
[0275] The HPF 282 supplies the band-limited speaker playback signal obtained by the filtering process to the speaker 291-1 via the D / A conversion unit 271-1 and the amplification unit 272-1.
[0276] In step S143, the rendering processing unit 281-2 performs rendering processing for the low-band speaker 291-2 based on the supplied N pieces of object data. For example, in step S143, processing similar to step S15 in FIG.
[0277] In step S144, the rendering processing unit 281-2 adjusts the gain of the supplied LFE channel signal using a predetermined coefficient, adds the adjusted signal to the speaker reproduction signal, and supplies the resulting final speaker reproduction signal to the LPF 283.
[0278] In step S145, the LPF 283 performs filtering processing (band limiting processing) using the LPF on the speaker playback signal supplied from the rendering processing unit 281-2.
[0279] The LPF 283 supplies the band-limited speaker playback signal obtained by the filtering process to the speaker 291-2 via the D / A conversion unit 271-2 and the amplification unit 272-2.
[0280] In the sound processing device 251, bass management is realized by the processes of steps S143 and S144.
[0281] In particular, in this example, the rendering processing unit 281-2 performs rendering processing for the low band, so that a decrease in the sense of localization of an object can be easily suppressed without requiring a complex design.
[0282] In step S146, all the speakers 291 that make up the speaker system 252 output sounds based on the speaker playback signals supplied from the amplifier 272, and the playback process ends.
[0283] In this way, the audio playback system 241 performs rendering processing for each playback band of the speaker 291, i.e., for each speaker layout of multiple playback bands, and also adjusts the gain of the LFE channel signal and adds it to the low-band speaker playback signal.
[0284] In this way, optimal rendering according to the object's metadata can be achieved in the audio playback system 241, even when multiple subwoofers (speakers 291-2) are used to reinforce the low frequencies. This suppresses deterioration in sound quality due to the playback band of the speaker 291, and easily suppresses deterioration in the sense of object positioning without requiring a complex design, enabling audio playback with higher sound quality.
[0285] <Example of computer configuration> The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the programs constituting the software are installed on a computer. Here, the term "computer" includes computers built into dedicated hardware, and general-purpose personal computers, for example, that can execute various functions by installing various programs.
[0286] FIG. 14 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by a program.
[0287] In the computer, a CPU (Central Processing Unit) 501, a ROM (Read Only Memory) 502, and a RAM (Random Access Memory) 503 are interconnected by a bus 504.
[0288] An input / output interface 505 is further connected to the bus 504. An input unit 506, an output unit 507, a recording unit 508, a communication unit 509, and a drive 510 are connected to the input / output interface 505.
[0289] The input unit 506 includes a keyboard, a mouse, a microphone, an image sensor, etc. The output unit 507 includes a display, a speaker, etc. The recording unit 508 includes a hard disk, a nonvolatile memory, etc. The communication unit 509 includes a network interface, etc. The drive 510 drives a removable recording medium 511 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0290] In a computer configured as described above, the CPU 501 performs the above-described series of processes by, for example, loading a program recorded in the recording unit 508 into the RAM 503 via the input / output interface 505 and the bus 504 and executing it.
[0291] The program executed by the computer (CPU 501) can be provided by being recorded on a removable recording medium 511 such as a package medium, for example. The program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.
[0292] In a computer, a program can be installed in the recording unit 508 via the input / output interface 505 by inserting a removable recording medium 511 into the drive 510. The program can also be received by the communication unit 509 via a wired or wireless transmission medium and installed in the recording unit 508. Alternatively, the program can be installed in the ROM 502 or the recording unit 508 in advance.
[0293] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.
[0294] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present technology.
[0295] For example, this technology can be configured as cloud computing, in which a single function is shared and processed collaboratively by multiple devices via a network.
[0296] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by multiple devices.
[0297] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.
[0298] Furthermore, the present technology can also be configured as follows.
[0299] (1) a first rendering processing unit that performs rendering processing based on the audio signal and generates a first output audio signal for outputting sound from a plurality of first speakers; a second rendering processing unit that performs rendering processing based on the audio signal and generates second output audio signals for outputting sound from a plurality of second speakers having a reproduction band different from that of the first speakers; An acoustic processing device comprising: (2) a first band limitation processing unit that performs band limitation processing on the first output audio signal in accordance with a reproduction band of the first speaker; a second band limitation processing unit that performs band limitation processing on the second output audio signal in accordance with a reproduction band of the second speaker; The sound processing device according to (1) further comprises: (3) a third band limitation processing unit that performs band limitation processing on the audio signal in accordance with a reproduction band of the first speaker; a third rendering processing unit that performs rendering processing based on the first band-limited signal obtained by the band-limiting processing by the third band-limiting processing unit, and generates a third output audio signal for outputting sound from a plurality of the first speakers; a fourth band limitation processing unit that performs band limitation processing on the audio signal in accordance with a reproduction band of the second speaker; a fourth rendering processing unit that performs rendering processing based on the second band-limited signal obtained by the band-limiting processing by the fourth band-limiting processing unit, and generates a fourth output audio signal for outputting sound from a plurality of the second speakers; causing the third band-limiting processing unit and the fourth band-limiting processing unit to perform band-limiting processing, and causing the third rendering processing unit and the fourth rendering processing unit to perform rendering processing; or The first rendering processing unit and the second rendering processing unit are caused to perform rendering processing, and the first band limiting processing unit and the second band limiting processing unit are caused to perform band limiting processing. A selection section to select The sound processing device according to (2) further comprises: (4) The selection unit performs the selection based on the number of the audio signals and the total number of the first speakers and the second speakers. The sound processing device according to (3). (5) a first band limiting processing unit that performs band limiting processing on the audio signal in accordance with a reproduction band of the first speaker; a second band limitation processing unit that performs band limitation processing on the audio signal in accordance with a reproduction band of the second speaker; Furthermore, the first rendering processing unit performs rendering processing based on a first band-limited signal obtained by the band-limiting processing by the first band-limiting processing unit; The second rendering processing unit performs rendering processing based on a second band-limited signal obtained by the band-limiting processing by the second band-limiting processing unit. The sound processing device according to (1). (6) The audio processing device further includes a determination unit that determines, for each audio signal, whether to cause the first rendering processing unit, the second rendering processing unit, or both the first rendering processing unit and the second rendering processing unit to perform rendering processing based on the audio signal, based on at least one of the audio signal and information related to the audio signal. The sound processing device according to (1), (2), or (5). (7) The determination unit makes the determination based on the frequency characteristics of the audio signal. (6) The sound processing device according to (6). (8) The determination unit makes the determination based on information indicating a sound source type of the audio signal. The sound processing device according to (6) or (7). (9) the audio signal is an object signal of an audio object, The first rendering processing unit and the second rendering processing unit perform rendering processing based on the audio signal and metadata of the audio signal. An audio processing device according to any one of (1) to (8). (10) The metadata includes location information indicating the location of the audio object. The sound processing device according to (9). (11) The position information is information indicating the relative position of the audio object with respect to a predetermined listening position. The sound processing device according to (10). (12) The second rendering processing unit adds the second output audio signal obtained by the rendering processing and a channel-based audio signal to obtain a final second output audio signal. The sound processing device according to any one of (9) to (11). (13) The channel-based audio signal is an audio signal of an LFE channel. The sound processing device according to (12). (14) The first rendering processing unit and the second rendering processing unit perform processing using VBAP as rendering processing. The sound processing device according to any one of (1) to (13). (15) The plurality of first speakers and the plurality of second speakers are further included. A sound processing device according to any one of (1) to (14). (16) The sound processing device performing a rendering process based on the audio signal to generate a first output audio signal for outputting sound from a plurality of first speakers; A rendering process is performed based on the audio signal, and a second output audio signal is generated for outputting sound from a plurality of second speakers having a reproduction band different from that of the first speakers. Acoustic processing methods. (17) performing a rendering process based on the audio signal to generate a first output audio signal for outputting sound from a plurality of first speakers; A rendering process is performed based on the audio signal, and a second output audio signal is generated for outputting sound from a plurality of second speakers having a reproduction band different from that of the first speakers. A program that causes a computer to execute a process that includes steps. [Explanation of symbols]
[0300] 11 Audio playback system, 21 Sound processing device, 22 Speaker system, 41-1 to 41-3, 41 Rendering processing unit, 42-1 to 42-Nt, 42 HPF, 43-1 to 43-Ns, 43 BPF, 44-1 to 44-Nw, 44 LPF, 151 Selection unit, 211 Determination unit
Claims
1. a first rendering processing unit that performs rendering processing based on the audio signal and generates a first output audio signal for outputting sound from a plurality of first speakers; a second rendering processing unit that performs rendering processing based on the audio signal and generates second output audio signals for outputting sound from a plurality of second speakers having a reproduction band different from that of the first speakers; Equipped with the first rendering processing unit performs VBAP using a mesh formed by the plurality of first speakers as a rendering process; The second rendering processing unit performs VBAP using a mesh formed by the plurality of second speakers as a rendering process. Sound processing equipment.
2. a first band limitation processing unit that performs band limitation processing on the first output audio signal in accordance with a reproduction band of the first speaker; a second band limitation processing unit that performs band limitation processing on the second output audio signal in accordance with a reproduction band of the second speaker; The sound processing device of claim 1 , further comprising:
3. a third band limiting processor that performs band limiting processing on the audio signal in accordance with a reproduction band of the first speaker; a third rendering processing unit that performs rendering processing based on the first band-limited signal obtained by the band-limiting processing by the third band-limiting processing unit, and generates a third output audio signal for outputting sound from the plurality of first speakers; a fourth band limitation processing unit that performs band limitation processing on the audio signal in accordance with a reproduction band of the second speaker; a fourth rendering processing unit that performs rendering processing based on the second band-limited signal obtained by the band-limiting processing by the fourth band-limiting processing unit, and generates a fourth output audio signal for outputting sound from the plurality of second speakers; causing the third band-limiting processing unit and the fourth band-limiting processing unit to perform band-limiting processing, and causing the third rendering processing unit and the fourth rendering processing unit to perform rendering processing; or The first rendering processing unit and the second rendering processing unit perform rendering processing, and the first band limiting processing unit and the second band limiting processing unit perform band limiting processing. A selection section to select The sound processing device according to claim 2 , further comprising:
4. The selection unit performs the selection based on the number of the audio signals and the total number of the first speakers and the second speakers. The sound processing device according to claim 3 .
5. a first band limiting processing unit that performs band limiting processing on the audio signal in accordance with a reproduction band of the first speaker; a second band limitation processing unit that performs band limitation processing on the audio signal in accordance with a reproduction band of the second speaker; Furthermore, the first rendering processing unit performs rendering processing based on a first band-limited signal obtained by the band-limiting processing by the first band-limiting processing unit; The second rendering processing unit performs rendering processing based on a second band-limited signal obtained by the band-limiting processing by the second band-limiting processing unit. The sound processing device according to claim 1 .
6. The audio processing device further includes a determination unit that determines, for each audio signal, whether to cause the first rendering processing unit, the second rendering processing unit, or both the first rendering processing unit and the second rendering processing unit to perform rendering processing based on the audio signal, based on at least one of the audio signal and information related to the audio signal. The sound processing device according to claim 1 .
7. The determination unit makes the determination based on the frequency characteristics of the audio signal. The sound processing device according to claim 6 .
8. The determination unit makes the determination based on information indicating a sound source type of the audio signal. The sound processing device according to claim 6 .
9. the audio signal is an object signal of an audio object, The first rendering processing unit and the second rendering processing unit perform rendering processing based on the audio signal and metadata of the audio signal. The sound processing device according to claim 1 .
10. The metadata includes location information indicating the location of the audio object. The sound processing device according to claim 9 .
11. The position information is information indicating the relative position of the audio object with respect to a predetermined listening position. The sound processing device according to claim 10.
12. The second rendering processing unit adds the second output audio signal obtained by the rendering processing and a channel-based audio signal to obtain a final second output audio signal. The sound processing device according to claim 9 .
13. The channel-based audio signal is an audio signal of an LFE channel. The sound processing device according to claim 12.
14. The plurality of first speakers and the plurality of second speakers are further included. The sound processing device according to claim 1 .
15. The sound processing device performing a rendering process based on the audio signal to generate a first output audio signal for outputting sound from a plurality of first speakers; performing a rendering process based on the audio signal to generate a second output audio signal for outputting sound from a plurality of second speakers having a reproduction band different from that of the first speakers; performing VBAP using a mesh formed by the plurality of first speakers as a rendering process for generating the first output audio signal; A VBAP process is performed using a mesh formed by the plurality of second speakers as a rendering process for generating the second output audio signal. Acoustic processing methods.
16. performing a rendering process based on the audio signal to generate a first output audio signal for outputting sound from a plurality of first speakers; A rendering process is performed based on the audio signal, and a second output audio signal is generated for outputting sound from a plurality of second speakers having a reproduction band different from that of the first speakers. Have the computer execute the process, performing VBAP using a mesh formed by the plurality of first speakers as a rendering process for generating the first output audio signal; A VBAP process is performed using a mesh formed by the plurality of second speakers as a rendering process for generating the second output audio signal. program.
Citation Information
Patent Citations
IEC23008-3
Stereoscopic sound reproduction equipment, stereophonic sound reproduction method, and computer program
JP2009077379A
Audio system and its operation method
JP2011529658A
Bass management for object-based audio
JP2018527825A
Audio signal processing method using generating virtual object
US20160066118A1