Audio processing apparatus and method, and program

By outputting audio from four or more speakers and calculating gain coefficients based on positional relationships, the method stabilizes sound image localization and expands the sweet spot, addressing the instability issues in conventional VBAP methods.

JP7704244B2Active Publication Date: 2025-07-08SONY GROUP CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024044571
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2013-04-26
Filing Date
2024-03-21
Publication Date
2025-07-08
Estimated Expiration
2034-04-11

AI Technical Summary

Technical Problem

Conventional sound image localization methods, such as VBAP, can result in unstable sound image localization and a narrow sweet spot due to the use of fewer speakers, causing the sound image to move in directions different from the user's movement, especially when using three-dimensional VBAP with three speakers.

Method used

The method involves outputting audio from four or more speakers, calculating gain coefficients based on the positional relationships between the speakers and the sound image localization position, using combinations of two or three speakers to stabilize the sound image localization and expand the sweet spot.

Benefits of technology

This approach enhances the stability of sound image localization and expands the range of the sweet spot by ensuring the sound image moves in the same direction as the user's movement, providing a more stable and consistent listening experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007704244000004
    Figure 0007704244000004
  • Figure 0007704244000005
    Figure 0007704244000005
  • Figure 0007704244000006
    Figure 0007704244000006
Patent Text Reader

Abstract

To enable more stable localization of a sound image.SOLUTION: A virtual speaker is assumed to exist on a lower side among sides of a tetragon having its corners formed with four speakers surrounding a target sound image position on a spherical plane. Three-dimensional VBAP is performed with respect to the virtual speaker and the two speakers located at an upper right and an upper left, to calculate respectively gains of the two speakers at the upper right and the upper left and the virtual speaker, the gains being to be used for fixing a sound image at the target sound image position. Further, two-dimensional VBAP is performed with respect to the lower right and lower left speakers to calculate gains of the lower right and lower left speakers, the gains being to be used for fixing a sound image at the position of the virtual speaker. Values obtained by multiplying these gains by the gain of the virtual speaker are set as the gains of the lower right and lower left speakers for fixing a sound image at the target sound image position. The present technology can be applied to sound processing apparatuses.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present technology relates to an acoustic processing apparatus and method, and a program, and particularly to an acoustic processing apparatus and method, and a program that can make the localization of sound images more stable.

Background Art

[0002] Conventionally, VBAP (Vector Base Amplitude Pannning) is known as a technology for controlling the localization of sound images using a plurality of speakers (see, for example, Non-Patent Document 1).

[0003] In VBAP, the target sound image localization position is represented by the linear sum of vectors directed in the directions of two or three speakers around that localization position. Then, the coefficients multiplied by each vector in that linear sum are used as the gains of the sound emitted from each speaker, and gain adjustment is performed so that the sound image localizes at the target position.

Prior Art Documents

Non-Patent Documents

[0004]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, with the above-described technology, although it is possible to localize the sound image at the target position, depending on the localization position, the localization of the sound image may become unstable.

[0006] For example, in 3D VBAP that performs VBAP using three speakers, depending on the localization position of the target sound image, among the three speakers, sound may be output only from two speakers, and the remaining one speaker may be controlled so that no sound is output.

[0007] In such a case, if the user moves while listening to the sound, the sound image may move in a direction different from the moving direction, and it may be perceived that the localization of the sound image is unstable. When the localization of the sound image becomes unstable in this way, the range of the sweet spot, which is the optimal viewing position, becomes narrow.

[0008] The present technology has been made in view of such a situation, and enables more stable localization of the sound image.

Means for Solving the Problems

[0009] An acoustic processing apparatus according to one aspect of the present technology includes Audio object signals for outputting sound an acquisition unit that acquires, and a plurality of sound output units located near the sound image localization position For combinations of two or three of them, for each of the plurality of combinations, by calculating the gain coefficient of the audio output unit based on the positional relationship between the audio output unit and the audio-visual localization position, a calculation unit that calculates the final gain coefficients of the plurality of audio output units is provided, the plurality of Based on the final gain coefficients of the audio output unit, sound is output from the plurality of audio output units.

[0010] An acoustic processing method or program according to one aspect of the present technology includes Audio object signals for outputting sound acquiring, and a plurality of sound output units located near the sound image localization position For combinations of two or three of them, for each of the plurality of combinations, by calculating the gain coefficient of the audio output unit based on the positional relationship between the audio output unit and the audio-visual localization position, the final gain coefficients of the plurality of audio output units are calculated, the plurality of Based on the final gain coefficients of the audio output unit, sound is output from the plurality of audio output units steps.

[0011] In one aspect of the present technology, Audio object signals for outputting sound is acquired, a plurality of sound output units located near the sound image localization position For combinations of two or three of them, for each of the plurality of combinations, by calculating the gain coefficient of the audio output unit based on the positional relationship between the audio output unit and the audio-visual localization position, the final gain coefficients of the plurality of audio output units are calculated are, and the plurality of Based on the final gain coefficients of the audio output unit, sound from the plurality of audio output units are output.

Advantages of the Invention

[0012] According to one aspect of the present technology, the localization of the sound image can be made more stable.

Brief Description of the Drawings

[0013]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Embodiments for Carrying Out the Invention

[0014] Hereinafter, embodiments to which this technology is applied will be described with reference to the drawings.

[0015] 〈The First Embodiment〉 〈Regarding the Outline of This Technology〉 First, with reference to FIGS. 1 to 8, the outline of the present technology will be described. In FIGS. 1 to 8, corresponding parts are denoted by the same reference numerals, and the description thereof will be omitted as appropriate.

[0016] For example, as shown in FIG. 1, assume that a user U11 who views content such as a moving image with audio or music listens to two-channel audio output from two speakers SP1 and speaker SP2 as the audio of the content.

[0017] In such a case, consider localizing the sound image at the position of the virtual sound source VSP1 using the position information of the two speakers SP1 and SP2 that output the audio of each channel.

[0018] For example, taking the position of the user U11's head as the origin O, in the two-dimensional coordinate system where the vertical and horizontal directions in the figure are the x-axis direction and the y-axis direction, the position of the virtual sound source VSP1 is represented by a vector P starting from the origin O.

[0019] Since the vector P is a two-dimensional vector, the vector P can be represented by the linear sum of vectors L1 and L2 starting from the origin O and facing the directions of the positions of the speakers SP1 and SP2 respectively. That is, the vector P can be represented by the following equation (1) using the vectors L1 and L2.

[0020]

Equation

[0021] By calculating the coefficients g1 and g2 multiplied by the vectors L1 and L2 in Equation (1) and using these coefficients g1 and g2 as the gains of the audio output from the speakers SP1 and SP2 respectively, the sound image can be localized at the position of the virtual sound source VSP1. That is, the sound image can be localized at the position indicated by the vector P.

[0022] In this way, the method of obtaining the coefficients g1 and g2 using the position information of the two speakers SP1 and SP2 and controlling the sound image localization position is called 2D VBAP.

[0023] In the example of FIG. 1, the sound image can be localized at any position on the arc AR11 connecting the speakers SP1 and SP2. Here, the arc AR11 is a part of a circle centered at the origin O and passing through the positions of the speakers SP1 and SP2.

[0024] Since the vector P is a two-dimensional vector, when the angle formed by the vectors L1 and L2 is greater than 0 degrees and less than 180 degrees, the coefficients g1 and g2 used as gains can be uniquely obtained. The calculation method of these coefficients g1 and g2 is described in detail in Non-Patent Document 1 mentioned above.

[0025] On the other hand, when trying to reproduce three-channel audio, for example, as shown in FIG. 2, the number of speakers for outputting the audio becomes three.

[0026] In the example of FIG. 2, the audio of each channel is output from the three speakers SP1, SP2, and SP3.

[0027] Even in such a case, the gains of the audio of each channel output from the speakers SP1 to SP3, that is, the coefficients obtained as gains, are only three, and the concept is the same as that of the 2D VBAP described above.

[0028] That is, when trying to localize the sound image at the position of the virtual sound source VSP2, in a three-dimensional coordinate system with the position of the user U11's head as the origin O, the position of the virtual sound source VSP2 is represented by a three-dimensional vector P starting from the origin O.

[0029] Also, when three-dimensional vectors that start from the origin O and point in the directions of the positions of the respective speakers SP1 to SP3 are defined as vectors L1 to L3, vector P can be expressed as a linear sum of vectors L1 to L3 as shown in the following equation (2).

[0030]

Equation

[0031] If the coefficients g1 to g3 multiplied by vectors L1 to L3 in equation (2) are calculated and these coefficients g1 to g3 are used as the gains of the audio signals output from the respective speakers SP1 to SP3, the sound image can be localized at the position of the virtual sound source VSP2.

[0032] The method of obtaining the coefficients g1 to g3 using the position information of the three speakers SP1 to SP3 and controlling the localization position of the sound image is called three-dimensional VBAP.

[0033] In the example of FIG. 2, the sound image can be localized at any position within the triangular region TR11 on the spherical surface that includes the positions of the speakers SP1, SP2, and SP3. Here, the region TR11 is a region on the surface of a sphere centered at the origin O and including the respective positions of the speakers SP1 to SP3, and is a triangular region on the spherical surface surrounded by the speakers SP1 to SP3.

[0034] By using such three-dimensional VBAP, the sound image can be localized at any position in space.

[0035] For example, as shown in FIG. 3, if the number of speakers for outputting audio is increased and a plurality of regions corresponding to the triangular region TR11 shown in FIG. 2 are provided in space, the sound image can be localized at any position on those regions.

[0036] In the example shown in FIG. 3, five speakers SP1 to SP5 are arranged, and the audio of each channel is output from these speakers SP1 to SP5. Here, the speakers SP1 to SP5 are arranged on the spherical surface centered on the origin O at the position of the head of the user U11.

[0037] In this case, taking the origin O as the starting point, and regarding the three-dimensional vectors pointing in the directions of the positions of the respective speakers SP1 to SP5 as vectors L1 to L5, a calculation similar to the calculation of solving the above-described formula (2) is performed, and the gain of the audio output from each speaker is obtained.

[0038] Here, among the regions on the spherical surface centered on the origin O, the triangular region surrounded by the speakers SP1, SP4, and SP5 is defined as the region TR21. Similarly, among the regions on the spherical surface centered on the origin O, the triangular region surrounded by the speakers SP3, SP4, and SP5 is defined as the region TR22, and the triangular region surrounded by the speakers SP2, SP3, and SP5 is defined as the region TR23.

[0039] These regions TR21 to TR23 are regions corresponding to the region TR11 shown in FIG. 2. Now, assuming that the three-dimensional vector indicating the position where the sound image is to be localized is the vector P, in the example of FIG. 3, the vector P indicates a position on the region TR21.

[0040] Therefore, in this example, a calculation similar to the calculation of solving the formula (2) is performed using the vectors L1, L4, and L5 indicating the positions of the speakers SP1, SP4, and SP5, and the gains of the audio output from each of the speakers SP1, SP4, and SP5 are calculated. Also, in this case, the gains of the audio output from the other speakers SP2 and SP3 are set to 0. That is, no audio is output from these speakers SP2 and SP3.

[0041] If five speakers SP1 to SP5 are arranged in space in this way, it becomes possible to localize sound images at any position on the area composed of areas TR21 to TR23.

[0042] By the way, as shown in FIG. 4, four speakers SP1 to SP4 are arranged in space, and it is assumed that the sound image is localized at the position of a virtual sound source VSP3 at the center positions of those speakers SP1 to SP4.

[0043] In the example of FIG. 4, the speakers SP1 to SP4 are arranged on the surface of a sphere centered on an origin O (not shown), and a triangular area on the surface thereof that is surrounded by the speakers SP1 to SP3 is area TR31. Also, a triangular area on the surface of the sphere centered on the origin O that is surrounded by the speakers SP2 to SP4 is area TR32.

[0044] And the virtual sound source VSP3 is located on the lower right side edge of area TR31. Also, the virtual sound source VSP3 is located on the upper left side edge of area TR32.

[0045] Therefore, in this case, 3D VBAP may be performed for the speakers SP1 to SP3, or 3D VBAP may be performed for the speakers SP2 to SP4. In either case, the calculation results of 3D VBAP are the same, and gains are obtained such that sound is output only from two speakers SP2 and SP3, and no sound is output from the remaining speakers SP1 and SP4.

[0046] In 3D VBAP, when the position where the sound image is to be localized is on the boundary line of a triangular area on the spherical surface connecting three speakers, that is, on the side of the triangle on the spherical surface, sound is output only from two speakers located at both ends of that side.

[0047] When audio is output from only two speakers, i.e., speaker SP2 and speaker SP3, as shown in FIG. 5 for example, assume that a user U11 at the sweet spot, which is the optimal viewing position, moves to the left in the figure as indicated by arrow A11.

[0048] Then, since the user U11's head approaches speaker SP3, the audio output from this speaker SP3 becomes louder, so the user U11 perceives that the virtual sound source VSP3, i.e., the sound image, has moved to the lower left in the figure as indicated by arrow A12.

[0049] In 3D VBAP, when audio is output from only two speakers as shown in FIG. 5, if the user U11 moves a little from the sweet spot, the sound image moves in a direction perpendicular to the moving direction of the user U11. In such a case, the user U11 perceives that the sound image has moved in a direction different from their moving direction, resulting in a sense of discomfort. That is, for the user U11, the localization of the sound image is perceived as unstable, and the range of the sweet spot becomes narrower.

[0050] Therefore, in this technology, different from the above-described VBAP, by outputting audio from more than three speakers, i.e., four or more speakers, the localization of the sound image is made more stable, thereby making the range of the sweet spot wider.

[0051] Note that the number of speakers for outputting audio may be any number as long as it is four or more. Hereinafter, the case of outputting audio from four speakers will be continued as an example for explanation.

[0052] For example, similar to the example shown in FIG. 4, assume that the sound image is localized at the position of the virtual sound source VSP3 at the center of the four speakers SP1 to SP4.

[0053] In such a case, in the present technology, two or three speakers are selected to form one combination, and VBAP is performed for a plurality of different combinations, and the gains of the voices output from the four speakers SP1 to SP4 are calculated.

[0054] Therefore, in the present technology, for example, as shown in FIG. 6, voices are output from all four speakers SP1 to SP4.

[0055] In such a case, in FIG. 6, as shown by arrow A21, even if the user U11 moves leftward from the sweet spot in the figure, the position of the virtual sound source VSP3, that is, the sound image localization position, only moves leftward in the figure as shown by arrow A22. That is, as in the example shown in FIG. 5, the sound image does not move downward, that is, in a direction perpendicular to the moving direction of the user U11, but only moves in the same direction as the moving direction of the user U11.

[0056] This is because when the user U11 moves leftward, the user U11 gets closer to the speaker SP3, and since the speaker SP1 is also located above the speaker SP3. In this case, the sound reaches the ears of the user U11 from both the upper left side and the lower left side as seen from the user U11, so it is difficult to perceive that the sound image has moved downward in the figure.

[0057] Therefore, compared with the conventional VBAP method, the sound image localization can be made more stable, and as a result, the range of the sweet spot can be expanded.

[0058] Next, the control of sound image localization according to the present technology will be described more specifically.

[0059] In the present technology, a vector indicating the position where the sound image is to be localized is expressed as a vector P starting from the origin O (not shown) of the three-dimensional coordinate system, and the vector P is expressed by the following formula (3).

[0060]

Equation

[0061] In Equation (3), vectors L1 to L4 are in the vicinity of the sound localization position of the sound image, and represent the directions of the positions of speakers SP1 to SP4 arranged so as to surround the sound image localization position, which are three-dimensional vectors. Also, g1 to g4 represent coefficients that will be the gains of the audio of each channel to be output from speakers SP1 to SP4 and are to be obtained hereinafter.

[0062] In Equation (3), vector P is represented by the linear sum of the four vectors L1 to L4. Here, since vector P is a three-dimensional vector, the four coefficients g1 to g4 cannot be uniquely determined.

[0063] Therefore, in this technology, each coefficient g1 to g4 that becomes a gain is calculated by the following method.

[0064] Now, assume that the sound image is localized at the central position of the quadrilateral on the spherical surface surrounded by the four speakers SP1 to SP4 shown in FIG. 4, that is, the position of the virtual sound source VSP3.

[0065] Here, first, any one side of the quadrilateral on the spherical surface with speakers SP1 to SP4 as vertices is selected, and it is assumed that there is a virtual speaker (hereinafter referred to as a virtual speaker) on that side.

[0066] For example, as shown in FIG. 7, assume that among the quadrilaterals on the spherical surface with speakers SP1 to SP4 as vertices, the side connecting speakers SP3 and SP4 located at the lower left and lower right in the figure is selected. And, for example, it is assumed that there is a virtual speaker VSP' at the intersection position of the perpendicular line dropped from the position of the virtual sound source VSP3 to the side connecting speakers SP3 and SP4.

[0067] Subsequently, for a total of three speakers, namely this virtual speaker VSP' and speakers SP1 and SP2 located in the upper left and upper right in the figure, 3D VBAP is performed. That is, by solving an equation similar to the above-described equation (2), coefficients g1, which is the gain of the sound output from each of speaker SP1, speaker SP2, and virtual speaker VSP', coefficient g2, and coefficient g' are obtained.

[0068] In FIG. 7, a vector P is represented by the linear sum of three vectors starting from the origin O, that is, a vector L1 pointing in the direction of speaker SP1, a vector L2 pointing in the direction of speaker SP2, and a vector L' pointing in the direction of virtual speaker VSP'. That is, P = g1L1 + g2L2 + g'L'.

[0069] Here, in order to localize the sound image at the position of the virtual sound source VSP3, sound must be output from the virtual speaker VSP' with a gain g', but the virtual speaker VSP' does not actually exist. Therefore, in this technology, as shown in FIG. 8, two speakers SP3 and SP4 located at both ends of the sides of the quadrilateral where the virtual speaker VSP' is located are used to localize the sound image at the position of the virtual speaker VSP', thereby realizing the virtual speaker VSP'.

[0070] Specifically, for two speakers SP3 and SP4 located at both ends of the side on the spherical surface where the virtual speaker VSP' is located, 2D VBAP is performed. That is, by solving an equation similar to the above-described equation (1), coefficients g3' and g4', which are the gains of the sound output from each of speaker SP3 and speaker SP4, are calculated.

[0071] In the example of FIG. 8, a vector L' pointing in the direction of the virtual speaker VSP' is represented by the linear sum of a vector L3 pointing in the direction of speaker SP3 and a vector L4 pointing in the direction of speaker SP4. That is, L' = g3'L3 + g4'L4.

[0072] Then, the value g'g3' obtained by multiplying the obtained coefficient g3' by the coefficient g' is set as the gain of the sound to be output from the speaker SP3, and the value g'g4' obtained by multiplying the coefficient g4' by the coefficient g' is set as the gain of the sound to be output from the speaker SP4. As a result, the virtual speaker VSP' that outputs sound with the gain g' is realized by the speakers SP3 and SP4.

[0073] Here, the value of g'g3' used as the gain value is the value of the coefficient g3 in the above-described formula (3), and the value of g'g4' used as the gain value is the value of the coefficient g4 in the above-described formula (3).

[0074] If the non-zero values g1, g2, g'g3', and g'g4' obtained as described above are set as the gains of the sound of each channel output from the speakers SP1 to SP4, sound can be output from the four speakers to localize the sound image at the target position.

[0075] By outputting sound from the four speakers in this way to localize the sound image, the localization of the sound image can be made more stable than when localizing the sound image by the conventional VBAP method, and thereby the range of the sweet spot can be expanded.

[0076] <Configuration Example of Audio Processing Device> Next, a specific embodiment to which the present technology described above is applied will be described. FIG. 9 is a diagram showing a configuration example of an embodiment of an audio processing device to which the present technology is applied.

[0077] The audio processing device 11 generates N-channel (where N ≥ 5) audio signals by performing gain adjustment for each channel on the monaural audio signal supplied from the outside, and supplies the audio signals to the speakers 12-1 to 12-N corresponding to each of the N channels.

[0078] Speakers 12-1 to 12-N output the audio of each channel based on the audio signal supplied from the audio processing device 11. That is, speakers 12-1 to 12-N are audio output units that serve as sound sources for outputting the audio of each channel. Hereinafter, when there is no particular need to distinguish speakers 12-1 to 12-N, they may simply be referred to as speaker 12. In FIG. 9, speaker 12 is configured not to be included in the audio processing device 11, but speaker 12 may be included in the audio processing device 11. Further, each part constituting the audio processing device 11 and the speaker 12 may be provided separately in, for example, several devices, etc., to form an audio processing system consisting of each part of the audio processing device 11 and the speaker 12.

[0079] Speaker 12 is arranged so as to surround the position where the user is assumed to be located when viewing content or the like (hereinafter, also simply referred to as the user's position). For example, each speaker 12 is arranged at a position on the surface of a sphere centered on the user's position. In other words, each speaker 12 is arranged at a position equidistant from the user. Also, the supply of the audio signal from the audio processing device 11 to the speaker 12 may be performed by wire or wirelessly.

[0080] The audio processing device 11 is composed of a speaker selection unit 21, a gain calculation unit 22, a gain determination unit 23, a gain output unit 24, and a gain adjustment unit 25.

[0081] To the audio processing device 11, for example, an audio signal of audio picked up by a microphone attached to an object such as a moving object and the position information of the object are supplied.

[0082] Based on the position information of the object supplied from the outside, the speaker selection unit 21 specifies the position (hereinafter, also referred to as the target sound image position) where the sound image of the audio emitted from the object should be localized in the space where the speaker 12 is arranged, and supplies the specification result to the gain calculation unit 22.

[0083] Further, based on the target sound image position, the speaker selection unit 21 selects four speakers 12 from among the N speakers 12 as the speakers 12 to output voice, and supplies selection information indicating the selection result to the gain calculation unit 22, the gain determination unit 23, and the gain output unit 24.

[0084] Based on the selection information supplied from the speaker selection unit 21 and the target sound image position, the gain calculation unit 22 calculates the gain of the speaker 12 to be processed, and supplies it to the gain output unit 24. Based on the selection information supplied from the speaker selection unit 21, the gain determination unit 23 determines the gain of the speakers 12 not being processed, and supplies it to the gain output unit 24. For example, the gain of the speakers 12 not being processed is set to "0". That is, control is performed so that the voice of the object is not output from the speakers 12 not being processed.

[0085] The gain output unit 24 supplies the N gains supplied from the gain calculation unit 22 and the gain determination unit 23 to the gain adjustment unit 25. At this time, based on the selection information supplied from the speaker selection unit 21, the gain output unit 24 determines the supply destination within the gain adjustment unit 25 for each of the N gains supplied from the gain calculation unit 22 and the gain determination unit 23.

[0086] Based on each gain supplied from the gain output unit 24, the gain adjustment unit 25 performs gain adjustment on the voice signal of the object supplied from the outside, and supplies the voice signals of each of the N channels obtained as a result to the speaker 12 to output voice.

[0087] The gain adjustment unit 25 includes amplification units 31-1 to 31-N. Based on the gain supplied from the gain output unit 24, the amplification units 31-1 to 31-N perform gain adjustment on the voice signal supplied from the outside, and supply the voice signals obtained as a result to the speakers 12-1 to 12-N.

[0088] Incidentally, hereinafter, when it is not necessary to distinguish the amplification units 31-1 to 31-N individually, they are simply referred to as the amplification unit 31.

[0089] <Example configuration of gain calculation unit> Further, the gain calculation unit 22 shown in FIG. 9 is configured as shown in FIG. 10, for example.

[0090] The gain calculation unit 22 shown in FIG. 10 is composed of a virtual speaker positioning unit 61, a three-dimensional gain calculation unit 62, a two-dimensional gain calculation unit 63, a multiplication unit 64, and a multiplication unit 65.

[0091] The virtual speaker positioning unit 61 determines the position of the virtual speaker based on the information indicating the target sound image position and the selection information supplied from the speaker selection unit 21. The virtual speaker positioning unit 61 supplies the information indicating the target sound image position, the selection information, and the information indicating the position of the virtual speaker to the three-dimensional gain calculation unit 62, and supplies the selection information and the information indicating the position of the virtual speaker to the two-dimensional gain calculation unit 63.

[0092] The three-dimensional gain calculation unit 62 performs three-dimensional VBAP for two of the speakers 12 to be processed and the virtual speaker based on each information supplied from the virtual speaker positioning unit 61. Then, the three-dimensional gain calculation unit 62 supplies the gains of the two speakers 12 obtained by the three-dimensional VBAP to the gain output unit 24, and supplies the gain of the virtual speaker to the multiplication unit 64 and the multiplication unit 65.

[0093] The two-dimensional gain calculation unit 63 performs two-dimensional VBAP for two of the speakers 12 to be processed based on each information supplied from the virtual speaker positioning unit 61, and supplies the gains of the speakers 12 obtained as a result to the multiplication unit 64 and the multiplication unit 65.

[0094] The multiplication unit 64 multiplies the gain supplied from the two-dimensional gain calculation unit 63 by the gain supplied from the three-dimensional gain calculation unit 62 to obtain the final gain of the speaker 12, and supplies it to the gain output unit 24. The multiplication unit 65 multiplies the gain supplied from the two-dimensional gain calculation unit 63 by the gain supplied from the three-dimensional gain calculation unit 62 to obtain the final gain of the speaker 12, and supplies it to the gain output unit 24.

[0095] <Explanation of sound image localization control processing> Incidentally, when the position information of the object and the audio signal are supplied to the audio processing device 11 and the output of the object's voice is instructed, the audio processing device 11 starts the sound image localization control processing to output the object's voice and localize the sound image at an appropriate position.

[0096] Hereinafter, with reference to the flowchart of FIG. 11, the sound image localization control processing by the audio processing device 11 will be described.

[0097] In step S11, the speaker selection unit 21 selects the speaker 12 to be processed based on the position information of the object supplied from the outside.

[0098] Specifically, for example, the speaker selection unit 21 specifies the target sound image position based on the position information of the object, and among the N speakers 12, selects the four speakers 12 that are near the target sound image position and arranged so as to surround the target sound image position as the speakers 12 to be processed.

[0099] For example, when the position of the virtual sound source VSP3 shown in FIG. 7 is set as the target sound image position, the speakers 12 corresponding to the four speakers SP1 to SP4 surrounding the virtual sound source VSP3 are selected as the speakers 12 to be processed.

[0100] The speaker selection unit 21 supplies information indicating the target sound image position to the virtual speaker position determination unit 61, and supplies selection information indicating the four speakers 12 to be processed to the virtual speaker position determination unit 61, the gain determination unit 23, and the gain output unit 24.

[0101] In step S12, the virtual speaker position determination unit 61 determines the position of the virtual speaker based on the information indicating the target sound image position and the selection information supplied from the speaker selection unit 21. For example, similar to the example shown in FIG. 7, the position of the intersection of the side on the spherical surface connecting the speakers 12 located at the lower left and lower right as viewed from the user among the speakers 12 to be processed and the perpendicular line dropped from the target sound image position to that side is set as the position of the virtual speaker.

[0102] When the position of the virtual speaker is determined, the virtual speaker position determination unit 61 supplies the information indicating the target sound image position, the selection information, and the information indicating the position of the virtual speaker to the three-dimensional gain calculation unit 62, and supplies the selection information and the information indicating the position of the virtual speaker to the two-dimensional gain calculation unit 63.

[0103] Note that the position of the virtual speaker may be set to any position as long as it is on the side of the quadrilateral on the spherical surface having the four speakers 12 to be processed as each vertex. Further, even when there are five or more speakers 12 to be processed, an arbitrary position on the side of the polygon on the spherical surface having those speakers 12 as each vertex may be set as the position of the virtual speaker.

[0104] In step S13, the three-dimensional gain calculation unit 62 calculates the gain for the virtual speaker and the two speakers 12 to be processed based on the information indicating the target sound image position, the selection information, and the information indicating the position of the virtual speaker supplied from the virtual speaker position determination unit 61.

[0105] Specifically, the three-dimensional gain calculation unit 62 sets the three-dimensional vector indicating the target sound image position as vector P, and the three-dimensional vector directed toward the virtual speaker as vector L'. Further, the three-dimensional gain calculation unit 62 sets the vector directed toward the speaker 12 having the same positional relationship as the speaker SP1 shown in FIG. 7 among the speakers 12 to be processed as vector L1, and the vector directed toward the speaker 12 having the same positional relationship as the speaker SP2 as vector L2.

[0106] Then, the three-dimensional gain calculation unit 62 obtains an expression representing the vector P as a linear combination of the vector L’, the vector L1, and the vector L2, and solves the expression to calculate the coefficients g’, g1, and g2 of the vector L’, the vector L1, and the vector L2 as gains. That is, an operation similar to the operation of solving the above-described equation (2) is performed.

[0107] The three-dimensional gain calculation unit 62 supplies the coefficients g1 and g2 of the speaker 12 in the same positional relationship as the speakers SP1 and SP2 obtained as a result of the calculation to the gain output unit 24 as the gains of the sound output from those speakers 12.

[0108] Also, the three-dimensional gain calculation unit 62 supplies the coefficient g’ of the virtual speaker obtained as a result of the calculation to the multiplication units 64 and 65 as the gain of the sound output from the virtual speaker.

[0109] In step S14, the two-dimensional gain calculation unit 63 calculates the gains for the two speakers 12 to be processed based on the selection information supplied from the virtual speaker positioning unit 61 and the information indicating the position of the virtual speaker.

[0110] Specifically, the two-dimensional gain calculation unit 63 sets the three-dimensional vector indicating the position of the virtual speaker as the vector L’. Also, the two-dimensional gain calculation unit 63 sets the vector facing the speaker 12 in the same positional relationship as the speaker SP3 shown in FIG. 8 among the speakers 12 to be processed as the vector L3, and the vector facing the speaker 12 in the same positional relationship as the speaker SP4 as the vector L4.

[0111] Then, the two-dimensional gain calculation unit 63 obtains an expression representing the vector L’ as a linear combination of the vector L3 and the vector L4, and solves the expression to calculate the coefficients g3’ and g4’ of the vector L3 and the vector L4 as gains. That is, an operation similar to the operation of solving the above-described equation (1) is performed.

[0112] The two-dimensional gain calculation unit 63 supplies the coefficients g3’ and g4’ of the speaker 12 in the same positional relationship as the speakers SP3 and SP4 obtained as a result of the calculation to the multiplication unit 64 and the multiplication unit 65 as the gains of the sound output from those speakers 12.

[0113] In step S15, the multiplication unit 64 and the multiplication unit 65 multiply the gains g3’ and g4’ supplied from the two-dimensional gain calculation unit 63 by the gain g’ of the virtual speaker supplied from the three-dimensional gain calculation unit 62, and supply the result to the gain output unit 24.

[0114] Therefore, as the final gain of the speaker 12 in the same positional relationship as the speaker SP3 in FIG. 8 among the four speakers 12 being processed, g3 = g’g3’ is supplied to the gain output unit 24. Similarly, as the final gain of the speaker 12 in the same positional relationship as the speaker SP4 in FIG. 8 among the four speakers 12 being processed, g4 = g’g4’ is supplied to the gain output unit 24.

[0115] In step S16, the gain determination unit 23 determines the gains of the speakers 12 that are not the processing targets based on the selection information supplied from the speaker selection unit 21, and supplies the result to the gain output unit 24. For example, the gains of all the speakers 12 that are not the processing targets are set to “0”.

[0116] When the gain output unit 24 is supplied with the gains g1, g2, g’g3’, and g’g4’ from the gain calculation unit 22 and the gain “0” from the gain determination unit 23, the gain output unit 24 supplies those gains to the amplifier unit 31 of the gain adjustment unit 25 based on the selection information from the speaker selection unit 21.

[0117] Specifically, the gain output unit 24 supplies the gain g1, the gain g2, the gain g'g3', and the gain g'g4' to the amplification unit 31 that supplies an audio signal to each speaker 12 to be processed, that is, each speaker 12 corresponding to each of the speakers SP1 to SP4 in FIG. 7. For example, when the speaker 12 corresponding to the speaker SP1 is the speaker 12-1, the gain output unit 24 supplies the gain g1 to the amplification unit 31-1.

[0118] Also, the gain output unit 24 supplies the gain "0" supplied from the gain determination unit 23 to the amplification unit 31 that supplies an audio signal to the speaker 12 that is not the processing target.

[0119] In step S17, the amplification unit 31 of the gain adjustment unit 25 adjusts the gain of the audio signal of the object supplied from the outside based on the gain supplied from the gain output unit 24, supplies the obtained audio signal to the speaker 12, and outputs the sound.

[0120] Each speaker 12 outputs sound based on the audio signal supplied from the amplification unit 31. More specifically, sound is output only from the four speakers 12 that are the processing targets. Thereby, the sound image can be localized at the target position. When sound is output from the speaker 12, the sound image localization control process ends.

[0121] As described above, the audio processing device 11 selects four speakers 12 to be processed from the position information of the object, and performs VBAP on combinations of two or three speakers among those speakers 12 and the virtual speaker. Then, the audio processing device 11 adjusts the gain of the audio signal based on the gain of each speaker 12 to be processed obtained by performing VBAP on a plurality of different combinations.

[0122] As a result, sound is output from the four speakers 12 located around the target sound image position, and the localization of the sound image can be made more stable. As a result, the range of the sweet spot can be further expanded.

[0123] <Second Embodiment> <Calculation of Gain> In the above, among the five speakers including the virtual speaker, two or three speakers are selected to form a combination of one speaker, and the gain of the speaker 12 to be processed is calculated by performing VBAP for a plurality of combinations. However, in this technology, it is also possible to calculate the gain by selecting a plurality of combinations from the four speakers 12 to be processed without defining a virtual speaker and performing VBAP for each of those combinations.

[0124] In such a case, for example, as shown in FIG. 12, the number of times VBAP should be performed changes depending on the target sound image position. In FIG. 12, the parts corresponding to those in FIG. 7 are denoted by the same reference numerals, and the description thereof will be omitted as appropriate.

[0125] For example, when the position of the virtual sound source, that is, the target sound image position, is at the position indicated by the arrow Q11, the position indicated by the arrow Q11 is within the triangular region surrounded by the speakers SP1, SP2, and SP4 on the spherical surface. Therefore, for the set of speakers (hereinafter also referred to as the first set) consisting of the speakers SP1, SP2, and SP4, if 3D VBAP is performed, the gains of the voices output from the three speakers SP1, SP2, and SP4 can be obtained.

[0126] On the other hand, the position indicated by the arrow Q11 is also a position within the triangular region surrounded by the speakers SP2, SP3, and SP4 on the spherical surface. Therefore, for the set of speakers (hereinafter also referred to as the second set) consisting of the speakers SP2, SP3, and SP4, if 3D VBAP is performed, the gains of the voices output from the three speakers SP2, SP3, and SP4 can be obtained.

[0127] Here, in the first group and the second group, if the gain of the speaker that was not used in each group is set to "0", in this example, a total of two types of gains can be obtained as the gains of each of the four speakers SP1 to SP4 in the first group and the second group.

[0128] Therefore, for each speaker, the sum of the gains of the speakers obtained in the first group and the second group is obtained as the gain sum. For example, if the gain of speaker SP1 obtained for the first group is g1(1) and the gain of speaker SP1 obtained for the second group is g1(2), the gain sum g s1 of speaker SP1 is s1 g = g1(1) + g1(2).

[0129] Here, since speaker SP1 is not included in the combination of the second group, g1(2) becomes 0, but since speaker SP1 is included in the combination of the speakers in the first group, g1(1) becomes a non-zero value. Eventually, the gain sum g s1 of speaker SP1 does not become 0. The same applies to the gain sums of the other speakers SP2 to SP4.

[0130] When the gain sum of each speaker is obtained in this way, the value obtained by normalizing the gain sum of each speaker with the sum of the squares of those gain sums can be used as the final gain of those speakers, or more specifically, the gain of the sound output from the speakers.

[0131] When the gains of each of the speakers SP1 to SP4 are obtained in this way, a non-zero gain is always obtained. Therefore, sound can be output from each of the four speakers SP1 to SP4, and the sound image can be localized at a desired position.

[0132] Note that hereinafter, the gain of speaker SPk (where 1 ≤ k ≤ 4) obtained for the m-th group (where 1 ≤ m ≤ 4) will be represented as g k (m). Also, the gain sum of speaker SPk (where 1 ≤ k ≤ 4) will be represented as g sk .

[0133] Furthermore, when the target sound image position is at the position indicated by arrow Q12, that is, on the spherical surface, at the intersection position of the line connecting speaker SP2 and speaker SP3 and the line connecting speaker SP1 and speaker SP4, there are four combinations of three speakers.

[0134] That is, a combination of speaker SP1, speaker SP2, and speaker SP3 (hereinafter referred to as the first set), and a combination of speaker SP1, speaker SP2, and speaker SP4 (hereinafter referred to as the second set) can be considered. In addition, a combination of speaker SP1, speaker SP3, and speaker SP4 (hereinafter referred to as the third set), and a combination of speaker SP2, speaker SP3, and speaker SP4 (hereinafter referred to as the fourth set) can be considered.

[0135] In this case, for each combination from the first set to the fourth set, three-dimensional VBAP may be performed for each to obtain the gain of each speaker. Then, taking the sum of the four gains obtained for the same speaker as the gain sum, and using the square sum of the four gain sums obtained for each speaker, the value obtained by normalizing the gain sum of each speaker can be used as the final gain of those speakers.

[0136] Note that when the target sound image position is at the position indicated by arrow Q12, if the quadrilateral on the spherical surface composed of speakers SP1 to SP4 is a rectangle or the like, for example, the same calculation result can be obtained as three-dimensional VBAP for the first set and the fourth set. Therefore, in such a case, by performing three-dimensional VBAP for two appropriate combinations such as the first set and the second set, the gain of each speaker can be obtained. However, when the quadrilateral on the spherical surface composed of speakers SP1 to SP4 is an asymmetric quadrilateral that is not a rectangle or the like, it is necessary to perform three-dimensional VBAP for each of the four combinations.

[0137] <Example of the configuration of the gain calculation unit> As described above, when selecting a plurality of combinations from the four speakers 12 to be processed without determining virtual speakers and performing VBAP for each of those combinations to calculate gains, the gain calculation unit 22 shown in FIG. 9 is configured as shown in FIG. 13, for example.

[0138] The gain calculation unit 22 shown in FIG. 13 is composed of a selection unit 91, a three-dimensional gain calculation unit 92-1, a three-dimensional gain calculation unit 92-2, a three-dimensional gain calculation unit 92-3, a three-dimensional gain calculation unit 92-4, and an addition unit 93.

[0139] Based on the information indicating the target sound image position and the selection information supplied from the speaker selection unit 21, the selection unit 91 determines a combination of three speakers 12 that surround the target sound image position from among the four speakers 12 to be processed. The selection unit 91 supplies the information indicating the combination of speakers 12 and the information indicating the target sound image position to the three-dimensional gain calculation units 92-1 to 92-4.

[0140] The three-dimensional gain calculation units 92-1 to 92-4 perform three-dimensional VBAP based on the information indicating the combination of speakers 12 and the information indicating the target sound image position supplied from the selection unit 91, and supply the gains of the respective speakers 12 obtained as a result to the addition unit 93. Hereinafter, when there is no particular need to distinguish between the three-dimensional gain calculation units 92-1 to 92-4, they are also simply referred to as the three-dimensional gain calculation unit 92.

[0141] The addition unit 93 obtains a gain sum based on the gains of the respective speakers 12 to be processed supplied from the three-dimensional gain calculation units 92-1 to 92-4, and further calculates the final gain of each speaker 12 to be processed by normalizing those gain sums, and supplies it to the gain output unit 24.

[0142] <Explanation of Sound Image Localization Control Processing> Next, with reference to the flowchart of FIG. 14, the sound image localization control processing performed when the gain calculation unit 22 has the configuration shown in FIG. 13 will be described.

[0143] Note that the process of step S41 is the same as the process of step S11 in FIG. 11, so the description thereof is omitted.

[0144] In step S42, the selection unit 91 determines a combination of speakers 12 based on the information indicating the target sound image position and the selection information supplied from the speaker selection unit 21, and supplies the information indicating the combination of speakers 12 and the information indicating the target sound image position to the three-dimensional gain calculation unit 92.

[0145] For example, when the target sound image position is at the position indicated by the arrow Q11 shown in FIG. 12, a combination (first set) of speakers 12 composed of three speakers 12 corresponding to the speaker SP1, the speaker SP2, and the speaker SP4 is determined. Also, a combination (second set) of speakers 12 composed of three speakers 12 corresponding to the speaker SP2, the speaker SP3, and the speaker SP4 is determined.

[0146] In this case, for example, the selection unit 91 supplies the information indicating the combination of the first set of speakers 12 and the information indicating the target sound image position to the three-dimensional gain calculation unit 92-1, and supplies the information indicating the combination of the second set of speakers 12 and the information indicating the target sound image position to the three-dimensional gain calculation unit 92-2. Also, in this case, information indicating a combination of speakers 12 and the like is not supplied to the three-dimensional gain calculation unit 92-3 and the three-dimensional gain calculation unit 92-4, and the three-dimensional VBAP calculation is not performed by the three-dimensional gain calculation unit 92-3 and the three-dimensional gain calculation unit 92-4.

[0147] In step S43, the three-dimensional gain calculation unit 92 calculates the gain of each speaker 12 to be processed for the combination of speakers 12 based on the information indicating the combination of speakers 12 and the information indicating the target sound image position supplied from the selection unit 91, and supplies it to the addition unit 93.

[0148] Specifically, the three-dimensional gain calculation unit 92 performs the same processing as step S13 in FIG. 11 described above for the three speakers 12 indicated by the information indicating the combination of the speakers 12, and obtains the gain of each speaker 12. That is, the same operation as the operation of solving the above-described formula (2) is performed. Further, among the four speakers 12 to be processed, the gain of the remaining one speaker 12 that is not the three speakers 12 indicated by the information indicating the combination of the speakers 12 is set to "0".

[0149] For example, when two combinations of the first set and the second set are obtained in step S42, the three-dimensional gain calculation unit 92-1 calculates the gain of each speaker 12 by three-dimensional VBAP for the first set. Further, the three-dimensional gain calculation unit 92-2 calculates the gain of each speaker 12 by three-dimensional VBAP for the second set.

[0150] Specifically, it is assumed that a combination of three speakers 12 corresponding to the speaker SP1, the speaker SP2, and the speaker SP4 shown in FIG. 12 is determined as the first set. In this case, in the three-dimensional gain calculation unit 92-1, the gain g1(1) of the speaker 12 corresponding to the speaker SP1, the gain g2(1) of the speaker 12 corresponding to the speaker SP2, and the gain g4(1) of the speaker 12 corresponding to the speaker SP4 are calculated. Further, the gain g3(1) of the speaker 12 corresponding to the speaker SP3 is set to "0".

[0151] In step S44, the addition unit 93 calculates the final gain of the speaker 12 to be processed based on the gain of each speaker 12 supplied from the three-dimensional gain calculation unit 92, and supplies it to the gain output unit 24.

[0152] For example, the addition unit 93 obtains the sum of the gains g1(1), g1(2), g1(3), and g1(4) of the speaker 12 corresponding to the speaker SP1 supplied from the three-dimensional gain calculation unit 92, thereby obtaining the gain sum g of the speaker 12 s1is calculated. Similarly, the adder 93 calculates the sum of gains g of the speaker 12 corresponding to the speaker SP2 s2 , the sum of gains g of the speaker 12 corresponding to the speaker SP3 s3 , and the sum of gains g of the speaker 12 corresponding to the speaker SP4 s4 as well.

[0153] Then, the adder 93 normalizes the sum of gains g of the speaker 12 corresponding to the speaker SP1 s1 by the sum of squares of the sum of gains g s1 to s4 to obtain the final gain g1 (coefficient g1) of the speaker 12 corresponding to the speaker SP1. Also, the adder 93 obtains the final gains g2 to g4 of the speaker 12 corresponding to the speakers SP2 to SP4 by the same calculation.

[0154] After the gain of the speaker 12 to be processed is obtained in this way, the processes of step S45 and step S46 are then performed, and the sound image localization control process ends. However, since these processes are the same as the processes of step S16 and step S17 in FIG. 11, the description thereof is omitted.

[0155] As described above, the audio processing device 11 selects four speakers 12 to be processed from the position information of the object, and performs VBAP on combinations of speakers 12 composed of three of those speakers 12. Then, the audio processing device 11 obtains the sum of the gains of the same speaker 12 obtained by performing VBAP on a plurality of different combinations, thereby obtaining the final gain of each speaker 12 to be processed and performing gain adjustment on the audio signal.

[0156] As a result, sound is output from the four speakers 12 located around the target sound image position, and the sound image localization can be made more stable. As a result, the range of the sweet spot can be further expanded.

[0157] In this embodiment, an example has been described in which the four speakers 12 surrounding the target sound image position are the speakers 12 to be processed. However, the number of speakers 12 to be processed may be four or more.

[0158] For example, when five speakers 12 are selected as the speakers 12 to be processed, among these five speakers 12, a set of speakers 12 consisting of any three speakers 12 surrounding the target sound image position is selected as one combination.

[0159] Specifically, as shown in FIG. 15, assume that the speakers 12 corresponding to the five speakers SP1 to SP5 are selected as the speakers 12 to be processed, and the target sound image position is the position indicated by the arrow Q21.

[0160] In this case, as the first set, a combination consisting of speaker SP1, speaker SP2, and speaker SP3 is selected, and as the second set, a combination consisting of speaker SP1, speaker SP2, and speaker SP4 is selected. Also, as the third set, a combination consisting of speaker SP1, speaker SP2, and speaker SP5 is selected.

[0161] Then, for these first to third sets, the gain of each speaker is obtained, and the final gain is calculated from the sum of the gains of each speaker. That is, for the first to third sets, the process of step S43 in FIG. 14 is performed, and then the processes of steps S44 to S46 are performed.

[0162] In this way, even when five or more speakers 12 are selected as the speakers 12 to be processed, it is possible to output sound from all the speakers 12 to be processed to localize the sound image.

[0163] Incidentally, the above-described series of processes can be executed by hardware or by software. When the series of processes is executed by software, the program constituting the software is installed in a computer. Here, the computer includes a computer incorporated in dedicated hardware, or a general-purpose computer, for example, that can execute various functions by installing various programs.

[0164] FIG. 16 is a block diagram showing a configuration example of the hardware of a computer that executes the above-described series of processes by a program.

[0165] In the computer, the CPU 801, ROM 802, and RAM 803 are interconnected by a bus 804.

[0166] Further connected to the bus 804 is an input / output interface 805. Connected to the input / output interface 805 are an input unit 806, an output unit 807, a recording unit 808, a communication unit 809, and a drive 810.

[0167] The input unit 806 includes a keyboard, a mouse, a microphone, an imaging element, etc. The output unit 807 includes a display, a speaker, etc. The recording unit 808 includes a hard disk, a non-volatile memory, etc. The communication unit 809 includes a network interface, etc. The drive 810 drives a removable medium 811 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.

[0168] In the computer configured as described above, the CPU 801 loads and executes, for example, the program recorded in the recording unit 808 via the input / output interface 805 and the bus 804 into the RAM 803, whereby the above-described series of processes is performed.

[0169] The program executed by the computer (CPU 801) can be recorded and provided on a removable medium 811 such as a package medium or the like. Also, the program can be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.

[0170] In the computer, the program can be installed in the recording unit 808 via the input / output interface 805 by attaching the removable medium 811 to the drive 810. Also, the program can be received by the communication unit 809 via a wired or wireless transmission medium and installed in the recording unit 808. Additionally, the program can be pre-installed in the ROM 802 or the recording unit 808.

[0171] Note that the program executed by the computer may be a program in which processing is performed in time series in accordance with the order described in this specification, or may be a program in which processing is performed in parallel or at a necessary timing such as when a call is made.

[0172] Also, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible without departing from the gist of the present technology.

[0173] For example, the present technology can adopt a cloud computing configuration in which one function is shared and jointly processed by a plurality of devices via a network.

[0174] Also, each step described in the above flowchart can be executed by one device or can be shared and executed by a plurality of devices.

[0175] Furthermore, when a plurality of processes are included in one step, the plurality of processes included in that one step can be executed by one device or can be shared and executed by a plurality of devices.

[0176] Furthermore, the present technology can also be configured as follows.

[0177] [1] For combinations of two or three of the four or more audio output units located near the target audio localization position, for each of a plurality of different combinations, by calculating the gain of the audio output from the audio output units based on the positional relationship of the audio output units, a gain calculation unit for obtaining the output gain of the audio output from the four or more audio output units for localizing the audio at the audio localization position, a gain adjustment unit that adjusts the gain of the audio output from the audio output units based on the output gain; An audio processing apparatus comprising: [2] At least four or more of the output gain values are non-zero values The audio processing apparatus according to [1]. [3] The gain calculation unit includes: A first gain calculation unit that calculates the output gain of the virtual audio output unit and the two audio output units based on the positional relationship between the virtual audio output unit, the two audio output units, and the audio localization position; A second gain calculation unit that calculates the gain of the other two audio output units for localizing the audio at the position of the virtual audio output unit based on the positional relationship between the other two audio output units different from the two audio output units and the virtual audio output unit; An arithmetic unit that calculates the output gain of the other two audio output units based on the gain of the other two audio output units and the output gain of the virtual audio output unit; Comprising The audio processing apparatus according to [1] or [2]. [4] The arithmetic unit calculates the output gain of the other two audio output units by multiplying the gain of the other two audio output units by the output gain of the virtual audio output unit. The audio processing apparatus according to [3]. [5] The position of the virtual voice output unit is determined to be located on the side of a polygon having the four or more voice output units as vertices. The voice processing apparatus according to [3] or [4]. [6] The gain calculation unit a virtual gain calculation unit that calculates the output gain of the three voice output units based on the positional relationship between the three voice output units and the sound image localization position; an arithmetic unit that calculates the final output gain of the voice output unit based on the output gains calculated by the plurality of virtual gain calculation units that calculate the output gain for different combinations; and includes The voice processing apparatus according to [1] or [2]. [7] The arithmetic unit calculates the final output gain of the voice output unit by obtaining the sum of the output gains obtained for the same voice output unit. The voice processing apparatus according to [6].

Explanation of Signs

[0178] 11 Voice processing apparatus, 12-1 to 12-N, 12 Speakers, 21 Speaker selection unit, 22 Gain calculation unit, 25 Gain adjustment unit, 61 Virtual speaker positioning unit, 62 Three-dimensional gain calculation unit, 63 Two-dimensional gain calculation unit, 64 Multiplication unit, 65 Multiplication unit, 91 Selection unit, 92-1 to 92-4, 92 Three-dimensional gain calculation unit, 93 Addition unit

Claims

1. An acquisition unit that acquires an audio object signal for outputting sound; For each combination of two or three of a plurality of sound output units located near the sound image localization position, a calculation unit that calculates the final gain coefficients of the plurality of sound output units by calculating the gain coefficients of the sound output units based on the positional relationship between the sound output units and the sound image localization position Comprising: Based on the final gain coefficients of the plurality of sound output units, sound is output from the plurality of sound output units An acoustic processing apparatus.

2. An acoustic processing apparatus Acquires an audio object signal for outputting sound, For each combination of two or three of a plurality of sound output units located near the sound image localization position, calculates the final gain coefficients of the plurality of sound output units by calculating the gain coefficients of the sound output units based on the positional relationship between the sound output units and the sound image localization position, Based on the final gain coefficients of the plurality of sound output units, sound is output from the plurality of sound output units An acoustic processing method.

3. A program that causes a computer to execute a process including steps of acquiring an audio object signal for outputting sound, for each combination of two or three of a plurality of sound output units located near the sound image localization position, calculating the final gain coefficients of the plurality of sound output units by calculating the gain coefficients of the sound output units based on the positional relationship between the sound output units and the sound image localization position, and outputting sound from the plurality of sound output units based on the final gain coefficients of the plurality of sound output units.

Citation Information

Patent Citations

  • Sound field control device

    JP2009038641A

  • Acoustic device and program

    JP2010041190A

  • Method and apparatus for decoding audio sound field representation for audio playback.

    JP2013524564A

  • System and tools for enhanced 3D audio authoring and rendering

    WO2013006330A2