Voice processing device, information processing method, and program

By employing four or more speakers and calculating gains through virtual speaker positioning, the sound image localization method stabilizes the audio experience, ensuring the sound image aligns with user movement and expands the sweet spot.

JP7704257B2Active Publication Date: 2025-07-08SONY GROUP CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024091181
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2013-04-26
Filing Date
2024-06-05
Publication Date
2025-07-08
Estimated Expiration
2034-04-11

AI Technical Summary

Technical Problem

Conventional sound image localization methods using VBAP can result in unstable sound image positioning, particularly when users move, leading to a narrow sweet spot and discomfort due to the sound image moving in directions different from the user's movement.

Method used

The use of four or more speakers, with virtual speaker positioning and gain calculations through VBAP for multiple speaker combinations, ensures stable sound image localization by outputting sound from specific speaker pairs to maintain alignment with user movement.

Benefits of technology

This approach enhances sound image stability and expands the sweet spot, providing a more consistent and comfortable listening experience by ensuring the sound image moves in the same direction as the user, regardless of their position changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007704257000004
    Figure 0007704257000004
  • Figure 0007704257000005
    Figure 0007704257000005
  • Figure 0007704257000006
    Figure 0007704257000006
Patent Text Reader

Abstract

To stabilize a location of a sound image.SOLUTION: In a side of quadrangle on a spherical surface in which four speakers surrounding a target sound image position are set as a peak, it is assumed that the virtual speaker is positioned onto the side positioned at a lower side. By performing a three-dimension VBAP from the virtual speaker and the two speakers on an upper right and an upper left, each gain of the two speaker on the upper right and the upper left for positioning a sound image to the target sound image position and the virtual speaker is calculated. Further, by performing a secondary VBAP of the speakers in a lower right and a lower left, each gain of the speakers on the lower right and the lower left for positioning the sound image to the virtual speaker is calculated. The gain obtained by multiplying the gain of the virtual speaker to the gain is a gain of the speaker in the lower right and the lower left for positioning the sound image to the target sound image position. The present technique is adapted to the voice processing device.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present technology relates to an audio processing apparatus, an information processing method, and a program, and particularly to an audio processing apparatus, an information processing method, and a program that can make the localization of sound images more stable.

Background Art

[0002] Conventionally, as a technology for controlling the localization of sound images using a plurality of speakers, VBAP (Vector Base Amplitude Pannning) is known (see, for example, Non-Patent Document 1).

[0003] In VBAP, the target sound image localization position is expressed as a linear sum of vectors directed toward the directions of two or three speakers around the localization position. Then, the coefficients multiplied by each vector in the linear sum are used as the gains of the audio output from each speaker, and gain adjustment is performed so that the sound image is localized at the target position.

Prior Art Documents

Non-Patent Documents

[0004]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, with the above-described technology, although it is possible to localize the sound image at the target position, depending on the localization position, the localization of the sound image may become unstable.

[0006] For example, in three-dimensional VBAP that performs VBAP using three speakers, depending on the localization position of the target sound image, depending on the localization position of the target sound image, among the three speakers, sound may be output only from two of the three speakers, and the remaining one speaker may be controlled so that no sound is output.

[0007] In such a case, if the user moves while listening to the sound, the sound image may move in a direction different from the moving direction, and it may be perceived that the localization of the sound image is unstable. When the localization of the sound image becomes unstable in this way, the range of the sweet spot, which is the optimal viewing position, becomes narrow.

[0008] The present technology has been made in view of such a situation, and enables more stable localization of the sound image.

Means for Solving the Problem

[0009] The audio processing device according to one aspect of the present technology includes Positions of Four Speakers a virtual speaker position determination unit that determines the position of a virtual speaker based on Based on the target sound image position a gain calculation unit that calculates the gains of two of the four speakers and the virtual speaker using VBAP, and a gain adjustment unit that performs gain adjustment of the sound output from at least two of the speakers based on the gain of the virtual speaker.

[0010] The information processing method or program according to one aspect of the present technology Positions of Four Speakers determines the position of a virtual speaker based on Based on the target sound image position calculates the gains of two of the four speakers and the virtual speaker using VBAP, and includes a step of performing gain adjustment of the sound output from at least two of the speakers based on the gain of the virtual speaker.

[0011] In one aspect of the present technology, Positions of Four Speakers the position of a virtual speaker is determined based on Based on the target sound image positionThe gains of two of the four speakers and the virtual speaker are calculated using VBAP, and based on the gain of the virtual speaker, gain adjustment of the audio output from at least two of the speakers is performed.

Advantages of the Invention

[0012] According to one aspect of the present technology, the sound image localization can be made more stable.

Brief Description of the Drawings

[0013]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Embodiments for Carrying Out the Invention

[0014] Hereinafter, embodiments to which the present technology is applied will be described with reference to the drawings.

[0015] 〈First Embodiment〉 〈Regarding the Outline of the Present Technology〉 First, with reference to FIGS. 1 to 8, the outline of the present technology will be described. In FIGS. 1 to 8, corresponding parts are denoted by the same reference numerals, and the description thereof will be omitted as appropriate.

[0016] For example, as shown in FIG. 1, assume that a user U11 who views content such as a moving image with sound or music listens to two-channel audio output from two speakers SP1 and SP2 as the audio of the content.

[0017] In such a case, consider localizing the sound image at the position of the virtual sound source VSP1 using the position information of the two speakers SP1 and SP2 that output the audio of each channel.

[0018] For example, taking the position of the user U11's head as the origin O, in the two-dimensional coordinate system where the vertical and horizontal directions in the figure are the x-axis direction and the y-axis direction, the position of the virtual sound source VSP1 is represented by a vector P starting from the origin O.

[0019] Since the vector P is a two-dimensional vector, the vector P can be represented by the linear sum of vectors L1 and L2 starting from the origin O and pointing in the directions of the positions of the speakers SP1 and SP2, respectively. That is, the vector P can be represented by the following equation (1) using the vectors L1 and L2.

[0020]

Equation

[0021] In formula (1), if the coefficients g1 and g2 multiplied by the vector L1 and the vector L2 are calculated and these coefficients g1 and g2 are used as the gains of the sound output from each of the speakers SP1 and SP2, the sound image can be localized at the position of the virtual sound source VSP1. That is, the sound image can be localized at the position indicated by the vector P.

[0022] In this way, the method of obtaining the coefficients g1 and g2 using the position information of the two speakers SP1 and SP2 and controlling the localization position of the sound image is called two-dimensional VBAP.

[0023] In the example of FIG. 1, the sound image can be localized at any position on the arc AR11 connecting the speakers SP1 and SP2. Here, the arc AR11 is a part of a circle centered on the origin O and passing through the positions of the speakers SP1 and SP2.

[0024] Since the vector P is a two-dimensional vector, when the angle formed by the vector L1 and the vector L2 is greater than 0 degrees and less than 180 degrees, the coefficients g1 and g2 used as the gains are uniquely determined. The calculation method of these coefficients g1 and g2 is described in detail in Non-Patent Document 1 mentioned above.

[0025] On the other hand, when attempting to reproduce three-channel sound, for example, as shown in FIG. 2, the number of speakers that output sound becomes three.

[0026] In the example of FIG. 2, the sound of each channel is output from the three speakers SP1, SP2, and SP3.

[0027] Even in such a case, the gain of the audio of each channel output from the speakers SP1 to SP3, that is, the coefficients obtained as gains, are only three, and the concept is the same as that of the two-dimensional VBAP described above.

[0028] That is, when trying to localize the sound image at the position of the virtual sound source VSP2, in a three-dimensional coordinate system with the position of the head of the user U11 as the origin O, the position of the virtual sound source VSP2 is represented by a three-dimensional vector P starting from the origin O.

[0029] Also, assuming that the three-dimensional vectors starting from the origin O and pointing in the directions of the positions of the speakers SP1 to SP3 are vectors L1 to L3, the vector P can be represented as a linear sum of the vectors L1 to L3 as shown in the following equation (2).

[0030]

Equation

[0031] If the coefficients g1 to g3 multiplied by the vectors L1 to L3 in Equation (2) are calculated and these coefficients g1 to g3 are used as the gains of the audio output from each of the speakers SP1 to SP3, the sound image can be localized at the position of the virtual sound source VSP2.

[0032] In this way, the method of obtaining the coefficients g1 to g3 using the position information of the three speakers SP1 to SP3 and controlling the localization position of the sound image is called three-dimensional VBAP.

[0033] In the example of FIG. 2, the sound image can be localized at any position within the triangular region TR11 on the spherical surface including the positions of the speakers SP1, SP2, and SP3. Here, the region TR11 is a region on the surface of the sphere centered at the origin O and including the positions of the speakers SP1 to SP3, and is a triangular region on the spherical surface surrounded by the speakers SP1 to SP3.

[0034] By using such 3D VBAP, it becomes possible to localize the audio-visual content at any position in space.

[0035] For example, as shown in FIG. 3, if the number of speakers for outputting audio is increased and a plurality of regions corresponding to the triangular region TR11 shown in FIG. 2 are provided in space, the audio-visual content can be localized at any position on those regions.

[0036] In the example shown in FIG. 3, five speakers SP1 to SP5 are arranged, and the audio of each channel is output from those speakers SP1 to SP5. Here, the speakers SP1 to SP5 are arranged on the spherical surface centered at the origin O at the position of the head of the user U11.

[0037] In this case, taking the origin O as the starting point and the three-dimensional vectors pointing in the directions of the positions of the respective speakers SP1 to SP5 as vectors L1 to L5, a calculation similar to the calculation for solving the above-described formula (2) is performed to obtain the gain of the audio output from each speaker.

[0038] Here, among the regions on the spherical surface centered at the origin O, the triangular region surrounded by the speakers SP1, SP4, and SP5 is defined as the region TR21. Similarly, among the regions on the spherical surface centered at the origin O, the triangular region surrounded by the speakers SP3, SP4, and SP5 is defined as the region TR22, and the triangular region surrounded by the speakers SP2, SP3, and SP5 is defined as the region TR23.

[0039] These regions TR21 to TR23 are regions corresponding to the region TR11 shown in FIG. 2. Now, if the three-dimensional vector indicating the position where the audio-visual content is to be localized is taken as the vector P, in the example of FIG. 3, the vector P indicates a position on the region TR21.

[0040] Therefore, in this example, vectors L1, L4, and L5 indicating the positions of speaker SP1, speaker SP4, and speaker SP5 are used to perform calculations similar to those for solving equation (2), and the gains of the sound outputs from each of speaker SP1, speaker SP4, and speaker SP5 are calculated. Also, in this case, the gains of the sound outputs from the other speakers SP2 and SP3 are set to 0. That is, no sound is output from these speakers SP2 and SP3.

[0041] By arranging five speakers SP1 to SP5 in space in this way, it becomes possible to localize the sound image at any position on the area composed of areas TR21 to TR23.

[0042] Incidentally, as shown in FIG. 4, four speakers SP1 to SP4 are arranged in space, and it is assumed that the sound image is localized at the position of a virtual sound source VSP3 at the center position of those speakers SP1 to SP4.

[0043] In the example of FIG. 4, speakers SP1 to SP4 are arranged on the surface of a sphere centered on an origin O (not shown), and a triangular area on that surface surrounded by speakers SP1 to SP3 is area TR31. Also, a triangular area on the surface of the sphere centered on the origin O surrounded by speakers SP2 to SP4 is area TR32.

[0044] And the virtual sound source VSP3 is located on the lower right side of the side of area TR31. Also, the virtual sound source VSP3 is located on the upper left side of the side of area TR32.

[0045] Therefore, in this case, 3D VBAP can be performed on speakers SP1 to SP3, or 3D VBAP can be performed on speakers SP2 to SP4. In either case, the calculation results of 3D VBAP are the same, and a gain can be obtained such that audio is output only from the two speakers SP2 and SP3, and no audio is output from the remaining speakers SP1 and SP4.

[0046] In 3D VBAP, when the position where the sound image is to be localized is on the boundary line of the triangular region on the spherical surface connecting three speakers, that is, on the side of the triangle on the spherical surface, audio is output only from the two speakers located at both ends of that side.

[0047] When audio is output only from the two speakers SP2 and SP3 in this way, for example, as shown in FIG. 5, assume that user U11 at the sweet spot, which is the optimal viewing position, moves to the left in the figure as indicated by arrow A11.

[0048] Then, since the head of user U11 approaches speaker SP3, the audio output from this speaker SP3 becomes louder, so user U11 perceives that the virtual sound source VSP3, that is, the sound image, has moved to the lower left in the figure as indicated by arrow A12.

[0049] In 3D VBAP, when audio is output from only two speakers as shown in FIG. 5, if user U11 moves a little from the sweet spot, the sound image moves in a direction perpendicular to the moving direction of user U11. In such a case, user U11 perceives that the sound image has moved in a direction different from their moving direction, resulting in a sense of discomfort. That is, for user U11, the localization of the sound image is perceived as unstable, and the range of the sweet spot becomes narrower.

[0050] Therefore, in this technology, unlike the above-described VBAP, by outputting sound from more than three speakers, that is, four or more speakers, the sound image localization is made more stable, thereby making the range of the sweet spot wider.

[0051] Note that the number of speakers for outputting sound may be any number as long as it is four or more. Hereinafter, the case of outputting sound from four speakers will be continued as an example for explanation.

[0052] For example, similar to the example shown in FIG. 4, it is assumed that the sound image is localized at the position of the virtual sound source VSP3 at the center position of the four speakers SP1 to SP4.

[0053] In such a case, in this technology, two or three speakers are selected to form one combination, and VBAP is performed for a plurality of different combinations to calculate the gain of the sound output from the four speakers SP1 to SP4.

[0054] Therefore, in this technology, for example, as shown in FIG. 6, sound is output from all four speakers SP1 to SP4.

[0055] In such a case, in FIG. 6, as shown by the arrow A21, even if the user U11 moves leftward from the sweet spot in the figure, the position of the virtual sound source VSP3, that is, the sound image localization position, only moves leftward in the figure as shown by the arrow A22. That is, unlike the example shown in FIG. 5, the sound image does not move downward, that is, in a direction perpendicular to the moving direction of the user U11, but only moves in the same direction as the moving direction of the user U11.

[0056] This is because when the user U11 moves leftward, the user U11 gets closer to the speaker SP3, and since the speaker SP1 is also located above the speaker SP3. In this case, sound reaches the ears of the user U11 from both the upper left side and the lower left side as seen from the user U11, so it becomes difficult to perceive that the sound image has moved downward in the figure.

[0057] Therefore, compared with the conventional VBAP method, the sound image localization can be made more stable, and as a result, the range of the sweet spot can be expanded.

[0058] Next, the control of sound image localization according to the present technology will be described more specifically.

[0059] In the present technology, a vector indicating the position where the sound image is to be localized is expressed by the following formula (3) as a vector P starting from the origin O (not shown) of the three-dimensional coordinate system.

[0060]

Equation

[0061] In Equation (3), vectors L1 to L4 are three-dimensional vectors that are in the vicinity of the sound image localization position and point in the directions of the positions of speakers SP1 to SP4 arranged so as to surround the sound image localization position. Also, g1 to g4 represent coefficients that are the gains of the audio of each channel to be output from speakers SP1 to SP4 and are to be obtained hereinafter.

[0062] In Equation (3), vector P is represented by the linear sum of the four vectors L1 to L4. Here, since vector P is a three-dimensional vector, the four coefficients g1 to g4 cannot be uniquely determined.

[0063] Therefore, in the present technology, the coefficients g1 to g4 that are gains are calculated by the following method.

[0064] Now, assume that the sound image is localized at the center position of the quadrilateral on the spherical surface surrounded by the four speakers SP1 to SP4 shown in FIG. 4, that is, the position of the virtual sound source VSP3.

[0065] Here, first, any one side of the quadrilateral on the spherical surface with speakers SP1 to SP4 as vertices is selected, and it is assumed that there is a virtual speaker (hereinafter referred to as a virtual speaker) on that side.

[0066] For example, as shown in FIG. 7, among the quadrilaterals on the spherical surface with speakers SP1 to SP4 as vertices, it is assumed that the side connecting speakers SP3 and SP4 located at the lower left and lower right in the figure is selected. Then, for example, from the position of the virtual sound source VSP3, it is assumed that there is a virtual speaker VSP' at the intersection position of the perpendicular line dropped to the side connecting speakers SP3 and SP4.

[0067] Subsequently, 3D VBAP is performed for a total of three speakers, namely this virtual speaker VSP' and speakers SP1 and SP2 located at the upper left and upper right in the figure. That is, by solving an equation similar to the above-described equation (2), the coefficients g1, g2, and g' that are the gains of the sound reproduced from speakers SP1, SP2, and the virtual speaker VSP' are obtained.

[0068] In FIG. 7, the vector P is represented by the linear sum of three vectors starting from the origin O, that is, the vector L1 pointing in the direction of speaker SP1, the vector L2 pointing in the direction of speaker SP2, and the vector L' pointing in the direction of the virtual speaker VSP'. That is, P = g1L1 + g2L2 + g'L'.

[0069] Here, in order to localize the sound image at the position of the virtual sound source VSP3, the sound must be output from the virtual speaker VSP' with the gain g'. However, the virtual speaker VSP' does not actually exist. Therefore, in this technique, as shown in FIG. 8, the virtual speaker VSP' is realized by localizing the sound image at the position of the virtual speaker VSP' using two speakers SP3 and SP4 located at both ends of the side of the quadrilateral where the virtual speaker VSP' is located.

[0070] Specifically, for the two speakers SP3 and SP4 located at both ends of the side on the spherical surface where the virtual speaker VSP’ is located, two-dimensional VBAP is performed. That is, by solving an equation similar to the above-described equation (1), coefficients g3’ and g4’ that are the gains of the sound output from each of the speaker SP3 and the speaker SP4 are calculated.

[0071] In the example of FIG. 8, the vector L’ pointing in the direction of the virtual speaker VSP’ is represented by the linear sum of the vector L3 pointing in the direction of the speaker SP3 and the vector L4 pointing in the direction of the speaker SP4. That is, L’ = g3’L3 + g4’L4.

[0072] Then, the value g’g3’ obtained by multiplying the obtained coefficient g3’ by the coefficient g’ is set as the gain of the sound output from the speaker SP3, and the value g’g4’ obtained by multiplying the coefficient g4’ by the coefficient g’ is set as the gain of the sound output from the speaker SP4. As a result, the virtual speaker VSP’ that outputs sound with the gain g’ is realized by the speaker SP3 and the speaker SP4.

[0073] Here, the value of g’g3’ that is the gain value becomes the value of the coefficient g3 in the above-described equation (3), and the value of g’g4’ that is the gain value becomes the value of the coefficient g4 in the above-described equation (3).

[0074] If the non-zero values g1, g2, g’g3’, and g’g4’ obtained as described above are set as the gains of the sound of each channel output from the speakers SP1 to SP4, it is possible to output sound from the four speakers and localize the sound image at the target position.

[0075] By outputting sound from the four speakers in this way and localizing the sound image, it is possible to make the localization of the sound image more stable than when localizing the sound image by the conventional VBAP method, and thereby expand the range of the sweet spot.

[0076] <Example of the configuration of the audio processing device> Next, a specific embodiment to which the present technology described above is applied will be described. FIG. 9 is a diagram showing a configuration example of an embodiment of an audio processing apparatus to which the present technology is applied.

[0077] The audio processing apparatus 11 generates an N-channel (where N ≧ 5) audio signal by performing gain adjustment for each channel on the monaural audio signal supplied from the outside, and supplies the audio signal to speakers 12-1 to 12-N corresponding to each of the N channels.

[0078] The speakers 12-1 to 12-N output the audio of each channel based on the audio signal supplied from the audio processing apparatus 11. That is, the speakers 12-1 to 12-N are audio output units that serve as sound sources for outputting the audio of each channel. Hereinafter, when it is not particularly necessary to distinguish the speakers 12-1 to 12-N, they may be simply referred to as speakers 12. In FIG. 9, the speaker 12 is configured not to be included in the audio processing apparatus 11, but the speaker 12 may be included in the audio processing apparatus 11. Further, each part constituting the audio processing apparatus 11 and the speaker 12 may be provided separately in, for example, several devices, so as to form an audio processing system including each part of the audio processing apparatus 11 and the speaker 12.

[0079] The speaker 12 is arranged so as to surround the position where the user is assumed to be located when viewing content or the like (hereinafter, also simply referred to as the user's position). For example, each speaker 12 is arranged at a position on the surface of a sphere centered on the user's position. In other words, each speaker 12 is arranged at a position equidistant from the user. Also, the supply of the audio signal from the audio processing apparatus 11 to the speaker 12 may be performed by wire or wirelessly.

[0080] The audio processing apparatus 11 is composed of a speaker selection unit 21, a gain calculation unit 22, a gain determination unit 23, a gain output unit 24, and a gain adjustment unit 25.

[0081] The voice processing device 11 is supplied with a voice signal of voice picked up by a microphone attached to an object such as a moving object, and position information of the object.

[0082] Based on the position information of the object supplied from the outside, the speaker selection unit 21 specifies a position (hereinafter also referred to as a target sound image position) where the sound image of the voice emitted from the object should be localized in the space where the speaker 12 is arranged, and supplies the specified result to the gain calculation unit 22.

[0083] Also, based on the target sound image position, the speaker selection unit 21 selects four speakers 12 from among the N speakers 12 as the speakers 12 to output voice, and supplies selection information indicating the selection result to the gain calculation unit 22, the gain determination unit 23, and the gain output unit 24.

[0084] Based on the selection information supplied from the speaker selection unit 21 and the target sound image position, the gain calculation unit 22 calculates the gain of the speaker 12 to be processed and supplies it to the gain output unit 24. The gain determination unit 23 determines the gain of the speakers 12 that are not the processing targets based on the selection information supplied from the speaker selection unit 21, and supplies it to the gain output unit 24. For example, the gain of the speakers 12 that are not the processing targets is set to "0". That is, control is performed so that the voice of the object is not output from the speakers 12 that are not the processing targets.

[0085] The gain output unit 24 supplies the N gains supplied from the gain calculation unit 22 and the gain determination unit 23 to the gain adjustment unit 25. At this time, based on the selection information supplied from the speaker selection unit 21, the gain output unit 24 determines the supply destinations in the gain adjustment unit 25 of the N gains supplied from the gain calculation unit 22 and the gain determination unit 23.

[0086] Based on each gain supplied from the gain output unit 24, the gain adjustment unit 25 performs gain adjustment on the audio signal of the object supplied from the outside, and supplies the resulting audio signals of each of the N channels to the speaker 12 to output sound.

[0087] The gain adjustment unit 25 includes amplification units 31-1 to 31-N. The amplification units 31-1 to 31-N perform gain adjustment on the audio signal supplied from the outside based on the gain supplied from the gain output unit 24, and supply the resulting audio signals to the speakers 12-1 to 12-N.

[0088] Hereinafter, when it is not necessary to individually distinguish the amplification units 31-1 to 31-N, they are simply also referred to as the amplification unit 31.

[0089] <Example configuration of the gain calculation unit> Also, the gain calculation unit 22 shown in FIG. 9 is configured as shown in FIG. 10, for example.

[0090] The gain calculation unit 22 shown in FIG. 10 is composed of a virtual speaker positioning unit 61, a three-dimensional gain calculation unit 62, a two-dimensional gain calculation unit 63, a multiplication unit 64, and a multiplication unit 65.

[0091] The virtual speaker positioning unit 61 determines the position of the virtual speaker based on the information indicating the target sound image position and the selection information supplied from the speaker selection unit 21. The virtual speaker positioning unit 61 supplies the information indicating the target sound image position, the selection information, and the information indicating the position of the virtual speaker to the three-dimensional gain calculation unit 62, and supplies the selection information and the information indicating the position of the virtual speaker to the two-dimensional gain calculation unit 63.

[0092] Based on each piece of information supplied from the virtual speaker positioning unit 61, the three-dimensional gain calculation unit 62 performs three-dimensional VBAP for two of the speakers 12 to be processed and the virtual speaker. Then, the three-dimensional gain calculation unit 62 supplies the gains of the two speakers 12 obtained by the three-dimensional VBAP to the gain output unit 24, and supplies the gain of the virtual speaker to the multiplication unit 64 and the multiplication unit 65.

[0093] Based on each piece of information supplied from the virtual speaker positioning unit 61, the two-dimensional gain calculation unit 63 performs two-dimensional VBAP for two of the speakers 12 to be processed, and supplies the gains of the speakers 12 obtained as a result to the multiplication unit 64 and the multiplication unit 65.

[0094] The multiplication unit 64 multiplies the gain supplied from the two-dimensional gain calculation unit 63 by the gain supplied from the three-dimensional gain calculation unit 62 to obtain the final gain of the speaker 12, and supplies it to the gain output unit 24. The multiplication unit 65 multiplies the gain supplied from the two-dimensional gain calculation unit 63 by the gain supplied from the three-dimensional gain calculation unit 62 to obtain the final gain of the speaker 12, and supplies it to the gain output unit 24.

[0095] <Explanation of sound image localization control processing> Incidentally, when the position information of the object and the audio signal are supplied to the audio processing device 11 and the output of the audio of the object is instructed, the audio processing device 11 starts the sound image localization control processing, outputs the audio of the object, and localizes the sound image at an appropriate position.

[0096] Hereinafter, with reference to the flowchart of FIG. 11, the sound image localization control processing by the audio processing device 11 will be described.

[0097] In step S11, the speaker selection unit 21 selects the speaker 12 to be processed based on the position information of the object supplied from the outside.

[0098] Specifically, for example, the speaker selection unit 21 identifies the target sound image position based on the position information of the object, and among the N speakers 12, the four speakers 12 that are near the target sound image position and are arranged to surround the target sound image position are set as the speakers 12 to be processed.

[0099] For example, when the position of the virtual sound source VSP3 shown in FIG. 7 is set as the target sound image position, the speakers 12 corresponding to the four speakers SP1 to SP4 surrounding the virtual sound source VSP3 are selected as the speakers 12 to be processed.

[0100] The speaker selection unit 21 supplies information indicating the target sound image position to the virtual speaker position determination unit 61, and supplies selection information indicating the four speakers 12 to be processed to the virtual speaker position determination unit 61, the gain determination unit 23, and the gain output unit 24.

[0101] In step S12, the virtual speaker position determination unit 61 determines the position of the virtual speaker based on the information indicating the target sound image position and the selection information supplied from the speaker selection unit 21. For example, similar to the example shown in FIG. 7, the position of the intersection of the side on the spherical surface connecting the speakers 12 located at the lower left and lower right as viewed by the user among the speakers 12 to be processed and the perpendicular line dropped from the target sound image position to that side is set as the position of the virtual speaker.

[0102] When the position of the virtual speaker is determined, the virtual speaker position determination unit 61 supplies the information indicating the target sound image position, the selection information, and the information indicating the position of the virtual speaker to the three-dimensional gain calculation unit 62, and supplies the selection information and the information indicating the position of the virtual speaker to the two-dimensional gain calculation unit 63.

[0103] Note that the position of the virtual speaker may be any position on the side of the quadrilateral on the spherical surface with the four speakers 12 to be processed as each vertex. Also, even when there are five or more speakers 12 to be processed, any position on the side of the polygon on the spherical surface with those speakers 12 as each vertex may be set as the position of the virtual speaker.

[0104] In step S13, the three-dimensional gain calculation unit 62 calculates the gains for the virtual speaker and the two speakers 12 to be processed based on the information indicating the target sound image position, the selection information, and the information indicating the position of the virtual speaker, which are supplied from the virtual speaker positioning unit 61.

[0105] Specifically, when the three-dimensional vector indicating the target sound image position is vector P and the three-dimensional vector facing the virtual speaker is vector L', the three-dimensional gain calculation unit 62 sets the vector facing the speaker 12 having the same positional relationship as the speaker SP1 shown in FIG. 7 among the speakers 12 to be processed as vector L1, and the vector facing the speaker 12 having the same positional relationship as the speaker SP2 as vector L2.

[0106] Then, the three-dimensional gain calculation unit 62 obtains an expression representing vector P as a linear sum of vector L', vector L1, and vector L2, and by solving that expression, calculates the coefficients g', g1, and g2 of vector L', vector L1, and vector L2 as gains. That is, an operation similar to the operation of solving the above-described expression (2) is performed.

[0107] The three-dimensional gain calculation unit 62 supplies the coefficients g1 and g2 of the speakers 12 having the same positional relationship as the speakers SP1 and SP2 obtained as a result of the calculation to the gain output unit 24 as the gains of the sound output from those speakers 12.

[0108] Also, the three-dimensional gain calculation unit 62 supplies the coefficient g' of the virtual speaker obtained as a result of the calculation to the multiplication units 64 and 65 as the gain of the sound output from the virtual speaker.

[0109] In step S14, the two-dimensional gain calculation unit 63 calculates the gains for the two speakers 12 to be processed based on the selection information supplied from the virtual speaker positioning unit 61 and the information indicating the position of the virtual speaker.

[0110] Specifically, the two-dimensional gain calculation unit 63 designates the three-dimensional vector indicating the position of the virtual speaker as vector L'. Further, among the speakers 12 to be processed, the two-dimensional gain calculation unit 63 designates the vector facing the speaker 12 in the same positional relationship as the speaker SP3 shown in FIG. 8 as vector L3, and the vector facing the speaker 12 in the same positional relationship as the speaker SP4 as vector L4.

[0111] Then, the two-dimensional gain calculation unit 63 obtains an expression representing vector L' as a linear sum of vector L3 and vector L4, and by solving that expression, calculates the coefficients g3' and g4' of vector L3 and vector L4 as gains. That is, an operation similar to the operation of solving the above-described equation (1) is performed.

[0112] The two-dimensional gain calculation unit 63 supplies the coefficients g3' and g4' of the speakers 12 in the same positional relationship as the speakers SP3 and SP4 obtained as a result of the calculation to the multiplication unit 64 and the multiplication unit 65 as the gains of the audio output from those speakers 12.

[0113] In step S15, the multiplication unit 64 and the multiplication unit 65 multiply the gains g3' and g4' supplied from the two-dimensional gain calculation unit 63 by the gain g' of the virtual speaker supplied from the three-dimensional gain calculation unit 62, and supply the result to the gain output unit 24.

[0114] Therefore, among the four speakers 12 to be processed, g3 = g'g3' is supplied to the gain output unit 24 as the final gain of the speaker 12 in the same positional relationship as the speaker SP3 in FIG. 8. Similarly, among the four speakers 12 to be processed, g4 = g'g4' is supplied to the gain output unit 24 as the final gain of the speaker 12 in the same positional relationship as the speaker SP4 in FIG. 8.

[0115] In step S16, the gain determination unit 23 determines the gain of the speaker 12 that is not the processing target based on the selection information supplied from the speaker selection unit 21, and supplies it to the gain output unit 24. For example, the gain of all speakers 12 that are not the processing target is set to "0".

[0116] When the gain output unit 24 is supplied with the gains g1, g2, g'g3', and g'g4' from the gain calculation unit 22 and the gain "0" from the gain determination unit 23, the gain output unit 24 supplies those gains to the amplification unit 31 of the gain adjustment unit 25 based on the selection information from the speaker selection unit 21.

[0117] Specifically, the gain output unit 24 supplies the gains g1, g2, g'g3', and g'g4' to the amplification unit 31 that supplies the audio signal to each speaker 12 that is the processing target, that is, the speaker 12 corresponding to each of the speakers SP1 to SP4 in FIG. 7. For example, when the speaker 12 corresponding to the speaker SP1 is the speaker 12-1, the gain output unit 24 supplies the gain g1 to the amplification unit 31-1.

[0118] In addition, the gain output unit 24 supplies the gain "0" supplied from the gain determination unit 23 to the amplification unit 31 that supplies the audio signal to the speaker 12 that is not the processing target.

[0119] In step S17, the amplification unit 31 of the gain adjustment unit 25 performs gain adjustment on the audio signal of the object supplied from the outside based on the gain supplied from the gain output unit 24, supplies the resulting audio signal to the speaker 12, and outputs the sound.

[0120] Each speaker 12 outputs sound based on the audio signal supplied from the amplification unit 31. More specifically, sound is output only from the four speakers 12 that are the processing target. Thereby, the sound image can be localized at the target position. When sound is output from the speaker 12, the sound image localization control process ends.

[0121] As described above, the audio processing device 11 selects four speakers 12 to be processed from the position information of the object, and performs VBAP for combinations of two or three speakers out of these speakers 12 and the virtual speaker. Then, the audio processing device 11 adjusts the gain of the audio signal based on the gain of each speaker 12 to be processed obtained by performing VBAP for a plurality of different combinations.

[0122] As a result, audio is output from the four speakers 12 located around the target sound image position, and the localization of the sound image can be made more stable. As a result, the range of the sweet spot can be further expanded.

[0123] <Second Embodiment> <Regarding Calculation of Gain> In the above, an example of calculating the gain of the speaker 12 to be processed by selecting two or three speakers out of five speakers including the virtual speaker as one speaker combination and performing VBAP for a plurality of combinations has been described. However, in this technology, it is also possible to calculate the gain by selecting a plurality of combinations from the four speakers 12 to be processed without defining the virtual speaker and performing VBAP for each of these combinations.

[0124] In such a case, for example, as shown in FIG. 12, the number of times VBAP should be performed changes depending on the target sound image position. In FIG. 12, parts corresponding to those in FIG. 7 are denoted by the same reference numerals, and the description thereof is omitted as appropriate.

[0125] For example, when the position of the virtual sound source, that is, the target sound image position, is at the position indicated by arrow Q11, the position indicated by arrow Q11 is within the triangular region surrounded by speakers SP1, SP2, and SP4 on the spherical surface. Therefore, for the set of speakers (hereinafter also referred to as the first set) consisting of speakers SP1, SP2, and SP4, if 3D VBAP is performed, the gains of the voices output from the three speakers SP1, SP2, and SP4 can be obtained.

[0126] On the other hand, the position indicated by arrow Q11 is also a position within the triangular region surrounded by speakers SP2, SP3, and SP4 on the spherical surface. Therefore, for the set of speakers (hereinafter also referred to as the second set) consisting of speakers SP2, SP3, and SP4, if 3D VBAP is performed, the gains of the voices output from the three speakers SP2, SP3, and SP4 can be obtained.

[0127] Here, in the first set and the second set, if the gains of the speakers not used in each are set to "0", in this example, a total of two sets of gains can be obtained as the gains of the four speakers SP1 to SP4 in the first set and the second set.

[0128] Therefore, for each speaker, the sum of the gains of the speakers obtained in the first set and the second set is obtained as the gain sum. For example, if the gain of speaker SP1 obtained for the first set is g1(1) and the gain of speaker SP1 obtained for the second set is g1(2), the gain sum g s1 of speaker SP1 is s1 g = g1(1) + g1(2).

[0129] Here, since speaker SP1 is not included in the combination of the second set, g1(2) becomes 0, but since speaker SP1 is included in the combination of the speakers in the first set, g1(1) is not a value of 0. Eventually, the gain sum g s1It will not be zero. The same applies to the sum of the gains of the other speakers SP2 to SP4.

[0130] When the sum of the gains of each speaker is obtained in this way, the value obtained by normalizing the sum of the gains of each speaker with the sum of the squares of those gains may be used as the final gain of those speakers, or more specifically, the gain of the sound output from the speakers.

[0131] When the gains of each of the speakers SP1 to SP4 are obtained in this way, a non-zero gain will always be obtained. Therefore, sound can be output from each of the four speakers SP1 to SP4, and the sound image can be localized at the desired position.

[0132] Hereinafter, the gain of the speaker SPk (where 1 ≤ k ≤ 4) obtained for the m-th set (where 1 ≤ m ≤ 4) will be denoted as g k (m). Also, the sum of the gains of the speaker SPk (where 1 ≤ k ≤ 4) will be denoted as g sk and.

[0133] Furthermore, when the target sound image position is at the intersection position of the line connecting the speakers SP2 and SP3 and the line connecting the speakers SP1 and SP4 on the sphere, that is, at the position indicated by the arrow Q12, there are four combinations of three speakers.

[0134] That is, a combination of the speakers SP1, SP2, and SP3 (hereinafter referred to as the first set) and a combination of the speakers SP1, SP2, and SP4 (hereinafter referred to as the second set) can be considered. In addition, a combination of the speakers SP1, SP3, and SP4 (hereinafter referred to as the third set) and a combination of the speakers SP2, SP3, and SP4 (hereinafter referred to as the fourth set) can be considered.

[0135] In this case, for each combination from the first group to the fourth group, 3D VBAP may be performed respectively to obtain the gain of each speaker. Then, the sum of the four gains obtained for the same speaker is defined as the gain sum, and the value obtained by normalizing the gain sum of the four speakers for each speaker with the sum of squares of the four gain sums is set as the final gain of those speakers.

[0136] In addition, when the target sound image position is at the position indicated by the arrow Q12 and the quadrilateral on the spherical surface composed of the speakers SP1 to SP4 is a rectangle or the like, for example, the same calculation result can be obtained as 3D VBAP for the first group and the fourth group. Therefore, in such a case, 3D VBAP may be performed for appropriate two combinations such as the first group and the second group to obtain the gain of each speaker. However, when the quadrilateral on the spherical surface composed of the speakers SP1 to SP4 is an asymmetric quadrilateral that is not a rectangle or the like, it is necessary to perform 3D VBAP for each of the four combinations.

[0137] <Example Configuration of Gain Calculation Unit> As described above, when a plurality of combinations are selected from the four speakers 12 to be processed without defining virtual speakers, and VBAP is performed for each of those combinations to calculate the gain, the gain calculation unit 22 shown in FIG. 9 is configured as shown in FIG. 13, for example.

[0138] The gain calculation unit 22 shown in FIG. 13 includes a selection unit 91, a 3D gain calculation unit 92-1, a 3D gain calculation unit 92-2, a 3D gain calculation unit 92-3, a 3D gain calculation unit 92-4, and an addition unit 93.

[0139] The selection unit 91 determines a combination of three speakers 12 surrounding the target sound image position from among the four speakers 12 to be processed based on the information indicating the target sound image position and the selection information supplied from the speaker selection unit 21. The selection unit 91 supplies the information indicating the combination of the speakers 12 and the information indicating the target sound image position to the 3D gain calculation units 92-1 to 92-4.

[0140] The three-dimensional gain calculation units 92-1 to 92-4 perform three-dimensional VBAP based on the information indicating the combination of the speakers 12 supplied from the selection unit 91 and the information indicating the target sound image position, and supply the gains of the respective speakers 12 obtained as a result to the addition unit 93. Hereinafter, when it is not particularly necessary to distinguish between the three-dimensional gain calculation units 92-1 to 92-4, they are simply also referred to as the three-dimensional gain calculation unit 92.

[0141] The addition unit 93 obtains a gain sum based on the gains of the respective speakers 12 to be processed supplied from the three-dimensional gain calculation units 92-1 to 92-4, and further calculates the final gain of each speaker 12 to be processed by normalizing the gain sum, and supplies it to the gain output unit 24.

[0142] <Explanation of sound image localization control processing> Next, with reference to the flowchart of FIG. 14, the sound image localization control processing performed when the gain calculation unit 22 has the configuration shown in FIG. 13 will be described.

[0143] Note that the processing in step S41 is the same as the processing in step S11 of FIG. 11, and thus the description thereof is omitted.

[0144] In step S42, the selection unit 91 determines a combination of the speakers 12 based on the information indicating the target sound image position and the selection information supplied from the speaker selection unit 21, and supplies the information indicating the combination of the speakers 12 and the information indicating the target sound image position to the three-dimensional gain calculation unit 92.

[0145] For example, when the target sound image position is at the position indicated by the arrow Q11 shown in FIG. 12, a combination of the speakers 12 (first set) composed of three speakers 12 corresponding to the speaker SP1, the speaker SP2, and the speaker SP4 is determined. Also, a combination of the speakers 12 (second set) composed of three speakers 12 corresponding to the speaker SP2, the speaker SP3, and the speaker SP4 is determined.

[0146] In this case, for example, the selection unit 91 supplies information indicating the combination of the first set of speakers 12 and information indicating the target sound image position to the three-dimensional gain calculation unit 92-1, and supplies information indicating the combination of the second set of speakers 12 and information indicating the target sound image position to the three-dimensional gain calculation unit 92-2. Also, in this case, information indicating the combination of the speakers 12 and the like is not supplied to the three-dimensional gain calculation unit 92-3 and the three-dimensional gain calculation unit 92-4, and the three-dimensional gain calculation unit 92-3 and the three-dimensional gain calculation unit 92-4 do not perform the calculation of three-dimensional VBAP.

[0147] In step S43, the three-dimensional gain calculation unit 92 calculates the gain of each speaker 12 to be processed for the combination of the speakers 12 based on the information indicating the combination of the speakers 12 supplied from the selection unit 91 and the information indicating the target sound image position, and supplies it to the addition unit 93.

[0148] Specifically, the three-dimensional gain calculation unit 92 performs the same processing as step S13 in FIG. 11 described above for the three speakers 12 indicated by the information indicating the combination of the speakers 12 to obtain the gain of each speaker 12. That is, the same operation as the operation of solving the above-described formula (2) is performed. Also, among the four speakers 12 to be processed, the gain of the remaining one speaker 12 that is not the three speakers 12 indicated by the information indicating the combination of the speakers 12 is set to "0".

[0149] For example, when two combinations of the first set and the second set are obtained in step S42, the three-dimensional gain calculation unit 92-1 calculates the gain of each speaker 12 by three-dimensional VBAP for the first set. Also, the three-dimensional gain calculation unit 92-2 calculates the gain of each speaker 12 by three-dimensional VBAP for the second set.

[0150] Specifically, assume that a combination of three speakers 12 corresponding to the speakers SP1, SP2, and SP4 shown in FIG. 12 is determined as the first group. In this case, in the three-dimensional gain calculation unit 92-1, the gain g1(1) of the speaker 12 corresponding to the speaker SP1, the gain g2(1) of the speaker 12 corresponding to the speaker SP2, and the gain g4(1) of the speaker 12 corresponding to the speaker SP4 are calculated. Also, the gain g3(1) of the speaker 12 corresponding to the speaker SP3 is set to "0".

[0151] In step S44, the addition unit 93 calculates the final gain of the speaker 12 to be processed based on the gains of the respective speakers 12 supplied from the three-dimensional gain calculation unit 92, and supplies it to the gain output unit 24.

[0152] For example, the addition unit 93 obtains the sum of the gains g1(1), g1(2), g1(3), and g1(4) of the speaker 12 corresponding to the speaker SP1 supplied from the three-dimensional gain calculation unit 92, thereby obtaining the gain sum g s1 of that speaker 12. Similarly, the addition unit 93 calculates the gain sum g s2 of the speaker 12 corresponding to the speaker SP2, the gain sum g s3 of the speaker 12 corresponding to the speaker SP3, and the gain sum g s4 of the speaker 12 corresponding to the speaker SP4.

[0153] Then, the addition unit 93 normalizes the gain sum g s1 of the speaker 12 corresponding to the speaker SP1 by the sum of the squares of the gain sums g s1 to g s4 to obtain the final gain g1 (coefficient g1) of the speaker 12 corresponding to the speaker SP1. Also, the addition unit 93 obtains the final gains g2 to g4 of the speakers 12 corresponding to the speakers SP2 to SP4 by the same calculation.

[0154] In this way, when the gain of the speaker 12 to be processed is obtained, the processes of step S45 and step S46 are then performed, and the sound image localization control process ends. However, since these processes are the same as the processes of step S16 and step S17 in FIG. 11, the description thereof will be omitted.

[0155] As described above, the audio processing device 11 selects four speakers 12 to be processed from the position information of the object, and performs VBAP on combinations of speakers 12 composed of three of these speakers 12. Then, the audio processing device 11 obtains the sum of the gains of the same speakers 12 obtained by performing VBAP on a plurality of different combinations, thereby obtaining the final gain of each speaker 12 to be processed, and performing gain adjustment on the audio signal.

[0156] As a result, sound is output from the four speakers 12 located around the target sound image position, and the sound image localization can be made more stable. As a result, the range of the sweet spot can be further expanded.

[0157] In this embodiment, an example in which the four speakers 12 surrounding the target sound image position are the speakers 12 to be processed has been described. However, the number of speakers 12 to be processed may be four or more.

[0158] For example, when five speakers 12 are selected as the speakers 12 to be processed, among these five speakers 12, a set of speakers 12 composed of any three speakers 12 surrounding the target sound image position is selected as one combination.

[0159] Specifically, as shown in FIG. 15, assume that the speakers 12 corresponding to the five speakers SP1 to SP5 are selected as the speakers 12 to be processed, and the target sound image position is the position indicated by the arrow Q21.

[0160] In this case, as the first set, a combination consisting of speaker SP1, speaker SP2, and speaker SP3 is selected, and as the second set, a combination consisting of speaker SP1, speaker SP2, and speaker SP4 is selected. Also, as the third set, a combination consisting of speaker SP1, speaker SP2, and speaker SP5 is selected.

[0161] Then, for these first to third sets, the gain of each speaker is obtained, and the final gain is calculated from the sum of the gains of each speaker. That is, for the first to third sets, the process of step S43 in FIG. 14 is performed, and then the processes of steps S44 to S46 are performed.

[0162] Thus, even when selecting five or more speakers 12 as the speakers 12 to be processed, it is possible to output sound from all the speakers 12 to be processed and localize the sound image.

[0163] By the way, the above-described series of processes can be executed by hardware or by software. When the series of processes is executed by software, the program constituting the software is installed in a computer. Here, the computer includes a computer incorporated in dedicated hardware, and, for example, a general-purpose computer that can execute various functions by installing various programs.

[0164] FIG. 16 is a block diagram showing a configuration example of the hardware of a computer that executes the above-described series of processes by a program.

[0165] In a computer, the CPU 801, ROM 802, and RAM 803 are interconnected by a bus 804.

[0166] The bus 804 is further connected to an input / output interface 805. The input / output interface 805 is connected to an input unit 806, an output unit 807, a recording unit 808, a communication unit 809, and a drive 810.

[0167] The input unit 806 includes a keyboard, a mouse, a microphone, an imaging device, etc. The output unit 807 includes a display, a speaker, etc. The recording unit 808 includes a hard disk, a non-volatile memory, etc. The communication unit 809 includes a network interface, etc. The drive 810 drives a removable medium 811 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.

[0168] In the computer configured as described above, for example, the CPU 801 loads and executes a program recorded in the recording unit 808 via the input / output interface 805 and the bus 804 into the RAM 803, thereby performing the above-described series of processes.

[0169] The program executed by the computer (CPU 801) can be recorded and provided, for example, on a removable medium 811 as a package medium or the like. Further, the program can be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.

[0170] In the computer, the program can be installed in the recording unit 808 via the input / output interface 805 by mounting the removable medium 811 on the drive 810. Further, the program can be received by the communication unit 809 via a wired or wireless transmission medium and installed in the recording unit 808. Additionally, the program can be installed in advance in the ROM 802 or the recording unit 808.

[0171] Note that the program executed by the computer may be a program that is processed in time series according to the order described in this specification, or a program that is processed in parallel or at a necessary timing such as when a call is made.

[0172] Also, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible without departing from the gist of the present technology.

[0173] For example, the present technology can adopt a cloud computing configuration in which one function is shared and jointly processed by a plurality of devices via a network.

[0174] In addition, each step described in the above flowchart can be executed by one device or shared and executed by a plurality of devices.

[0175] Furthermore, when a plurality of processes are included in one step, the plurality of processes included in that one step can be executed by one device or shared and executed by a plurality of devices.

[0176] Furthermore, the present technology can also have the following configuration.

[0177] [1] Four or more voice output units, For combinations of two or three of the four or more voice output units located near the target audio-visual localization position, for each of the plurality of different combinations, by calculating the gain of the voice output from the voice output units based on the positional relationship of the voice output units, a gain calculation unit that obtains the output gain of the voice output from the four or more voice output units for localizing the audio-visual at the audio-visual localization position, A gain adjustment unit that adjusts the gain of the voice output from the voice output units based on the output gain A voice processing apparatus comprising. [2] At least four of the values of the output gain are non-zero values. The audio processing device according to [1]. [3] The gain calculation unit Based on the positional relationship between the virtual audio output unit, the two audio output units, and the audio image localization position, a first gain calculation unit that calculates the output gain of the virtual audio output unit and the two audio output units; Based on the positional relationship between the other two audio output units different from the two audio output units and the virtual audio output unit, a second gain calculation unit that calculates the gain of the other two audio output units for localizing an audio image at the position of the virtual audio output unit; An arithmetic unit that calculates the output gain of the other two audio output units based on the gain of the other two audio output units and the output gain of the virtual audio output unit Comprising The audio processing device according to [1] or [2]. [4] The arithmetic unit calculates the output gain of the other two audio output units by multiplying the gain of the other two audio output units by the output gain of the virtual audio output unit. The audio processing device according to [3]. [5] The position of the virtual audio output unit is determined to be located on the side of a polygon having the four or more audio output units as vertices. The audio processing device according to [3] or [4]. [6] The gain calculation unit Based on the positional relationship between the three audio output units and the audio image localization position, a temporary gain calculation unit that calculates the output gain of the three audio output units; An arithmetic unit that calculates the final output gain of the audio output unit based on the output gain calculated by a plurality of the temporary gain calculation units that calculate the output gain for different combinations; Comprising The audio processing device according to [1] or [2]. [7] The calculation unit calculates the final output gain of the voice output unit by obtaining the sum of the output gains obtained for the same voice output unit. The voice processing device according to [6].

Explanation of Signs

[0178] 11 Voice processing device, 12-1 to 12-N, 12 Speakers, 21 Speaker selection unit, 22 Gain calculation unit, 25 Gain adjustment unit, 61 Virtual speaker positioning unit, 62 3D gain calculation unit, 63 2D gain calculation unit, 64 Multiplication unit, 65 Multiplication unit, 91 Selection unit, 92-1 to 92-4, 92 3D gain calculation unit, 93 Addition unit

Claims

1. A virtual speaker position determination unit that determines the position of a virtual speaker based on the positions of four speakers, a gain calculation unit that calculates the gains of two of the four speakers and the virtual speaker based on a target sound image position using VBAP, and a gain adjustment unit that performs gain adjustment on the sound output from at least two of the speakers based on the gain of the virtual speaker A voice processing apparatus comprising:

2. A voice processing apparatus, determines the position of a virtual speaker based on the positions of four speakers, calculates the gains of two of the four speakers and the virtual speaker based on a target sound image position using VBAP, and performs gain adjustment on the sound output from at least two of the speakers based on the gain of the virtual speaker An information processing method.

3. Determine the position of a virtual speaker based on the positions of four speakers, calculate the gains of two of the four speakers and the virtual speaker based on a target sound image position using VBAP, and perform gain adjustment on the sound output from at least two of the speakers based on the gain of the virtual speaker A program for causing a computer to execute a process including steps.

Citation Information

Patent Citations

  • Speaker system and video display device

    JP2007134939A

  • Sound field control device

    JP2009038641A

  • Acoustic device and program

    JP2010041190A

  • Apparatus and method for calculating the driving coefficient of a speaker in a speaker system with respect to an audio signal associated with a virtual sound source.

    JP2013510481A

  • Method and device for decoding an audio soundfield representation for audio playback

    WO2011117399A1