Voice processing device, information processing method, and program

By selecting four speakers and employing virtual speaker positioning and gain calculations, the method stabilizes sound image localization and expands the sweet spot, addressing the instability issues in conventional techniques.

JP2025128392APending Publication Date: 2025-09-02SONY GROUP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025105423
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2013-04-26
Filing Date
2025-06-23
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

Conventional sound image localization techniques, such as VBAP, can result in unstable sound image positioning and a narrow sweet spot when the user moves, leading to an uncomfortable listening experience.

Method used

An audio processing device and method that selects four speakers and utilizes virtual speaker positioning and gain calculations using VBAP to stabilize sound image localization, ensuring sound is output from multiple speakers even when the user moves.

Benefits of technology

The proposed method achieves more stable sound image localization and expands the range of the sweet spot, providing a consistent listening experience regardless of user movement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025128392000001_ABST
    Figure 2025128392000001_ABST
Patent Text Reader

Abstract

To stabilize a location of a sound image.SOLUTION: In a side of quadrangle on a spherical surface in which four speakers surrounding a target sound image position are set as a peak, it is assumed that the virtual speaker is positioned onto the side positioned at a lower side. By performing a three-dimension VBAP from the virtual speaker and the two speakers on an upper right and an upper left, each gain of the two speaker on the upper right and the upper left for positioning a sound image to the target sound image position and the virtual speaker is calculated. Further, by performing a secondary VBAP of the speakers in a lower right and a lower left, each gain of the speakers on the lower right and the lower left for positioning the sound image to the virtual speaker is calculated. The gain obtained by multiplying the gain of the virtual speaker to the gain is a gain of the speaker in the lower right and the lower left for positioning the sound image to the target sound image position. The present technique is adapted to the voice processing device.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present technology relates to an audio processing device, an information processing method, and a program, and more particularly to an audio processing device, an information processing method, and a program that enable more stable localization of a sound image. [Background technology]

[0002] BACKGROUND ART Conventionally, VBAP (Vector Base Amplitude Panning) is known as a technique for controlling the localization of a sound image using a plurality of speakers (see, for example, Non-Patent Document 1).

[0003] In VBAP, the localization position of the target sound image is expressed as a linear sum of vectors pointing in the directions of two or three speakers surrounding that localization position. The coefficients multiplied by each vector in the linear sum are then used as the gain of the sound output from each speaker to adjust the gain, so that the sound image is localized at the target position. [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] Ville Pulkki, “Virtual Sound Source Positioning Using Vector Base Amplitude Panning”, Journal of AES, vol.45, no.6, pp.456-466, 1997 Summary of the Invention [Problem to be solved by the invention]

[0005] However, although the above-mentioned techniques can localize a sound image at a target position, the localization of the sound image may become unstable depending on the localization position.

[0006] For example, in a three-dimensional VBAP using three speakers, depending on the localization position of the target sound image, sound may be output from only two of the three speakers, and no sound may be output from the remaining speaker.

[0007] In such cases, if the user moves while listening to audio, the sound image may move in a direction different from the direction of the user's movement, causing the sound image to be perceived as unstable. When the sound image becomes unstable in this way, the range of the sweet spot, which is the optimal listening position, becomes narrow.

[0008] The present technology has been made in view of such circumstances, and is intended to make it possible to make the localization of a sound image more stable. [Means for solving the problem]

[0009] An audio processing device according to one aspect of the present technology includes a speaker selection unit that selects four speakers to be processed, a virtual speaker position determination unit that determines the position of a virtual speaker based on a selection result by the speaker selection unit, a gain calculation unit that calculates gains for two of the four speakers and the virtual speaker using VBAP, and a gain adjustment unit that adjusts the gain of audio to be output from at least two of the speakers based on the gain of the virtual speaker.

[0010] An information processing method or program according to one aspect of the present technology includes steps of selecting four speakers to be processed, determining the position of a virtual speaker based on the selection results of the four speakers, calculating gains for two of the four speakers and the virtual speaker using VBAP, and adjusting the gain of audio to be output from at least two of the speakers based on the gain of the virtual speaker.

[0011] In one aspect of the present technology, four speakers to be processed are selected, the position of a virtual speaker is determined based on the selection results of the four speakers, the gains of two of the four speakers and the virtual speaker are calculated using VBAP, and gain adjustment of the audio to be output from at least two of the speakers is performed based on the gain of the virtual speaker. [Effects of the Invention]

[0012] According to one aspect of the present technology, it is possible to make the localization of a sound image more stable. [Brief explanation of the drawings]

[0013] [Figure 1] FIG. 1 is a diagram illustrating a two-dimensional VBAP. [Figure 2] FIG. 1 is a diagram illustrating a three-dimensional VBAP. [Figure 3] FIG. 1 is a diagram illustrating speaker placement. [Figure 4] FIG. 10 is a diagram illustrating a gain calculation method when four speakers are arranged. [Figure 5] FIG. 10 is a diagram illustrating movement of a sound image. [Figure 6] 10A and 10B are diagrams illustrating movement of a sound image when the present technology is applied. [Figure 7] 10A and 10B are diagrams illustrating calculation of a gain according to the present technology. [Figure 8] 10A and 10B are diagrams illustrating calculation of a gain according to the present technology. [Figure 9] FIG. 1 illustrates an example of the configuration of a voice processing device. [Figure 10] FIG. 2 illustrates an example of the configuration of a gain calculation unit. [Figure 11] 10 is a flowchart illustrating a sound image localization control process. [Figure 12] 10A and 10B are diagrams illustrating another method for calculating the speaker gain. [Figure 13] FIG. 10 is a diagram illustrating another example of the configuration of the gain calculation unit. [Figure 14] 10 is a flowchart illustrating a sound image localization control process. [Figure 15] FIG. 10 is a diagram illustrating a method for calculating the gain of a speaker. [Figure 16] FIG. 1 illustrates an example of the configuration of a computer. DETAILED DESCRIPTION OF THE INVENTION

[0014] Hereinafter, embodiments to which the present technology is applied will be described with reference to the drawings.

[0015] First Embodiment Overview of this technology First, an overview of the present technology will be described with reference to Figures 1 to 8. In Figures 1 to 8, corresponding parts are denoted by the same reference numerals, and descriptions thereof will be omitted where appropriate.

[0016] For example, as shown in FIG. 1, assume that a user U11 viewing content such as audio-accompanying moving images or music is listening to two-channel audio output from two speakers SP1 and SP2 as the audio of the content.

[0017] In such a case, it is considered that the sound image is localized at the position of the virtual sound source VSP1 using the position information of the two speakers SP1 and SP2 that output the sound of each channel.

[0018] For example, the position of the head of the user U11 is set as the origin O, and the position of the virtual sound source VSP1 in a two-dimensional coordinate system in which the vertical and horizontal directions are the x-axis and y-axis directions in the figure is represented by a vector P starting from the origin O.

[0019] Because vector P is a two-dimensional vector, it can be expressed as the linear sum of vectors L1 and L2, which start from origin O and point in the directions of the positions of speakers SP1 and SP2, respectively. That is, vector P can be expressed by the following equation (1) using vectors L1 and L2.

[0020]

number

[0021] By calculating the coefficients g1 and g2 by which the vectors L1 and L2 are multiplied in equation (1) and setting these coefficients g1 and g2 as the gains of the sounds output from the speakers SP1 and SP2, respectively, it is possible to localize the sound image at the position of the virtual sound source VSP1. In other words, the sound image can be localized at the position indicated by the vector P.

[0022] This method of determining the coefficients g1 and g2 using the position information of the two speakers SP1 and SP2 and controlling the localization position of the sound image is called two-dimensional VBAP.

[0023] In the example of Fig. 1, the sound image can be localized at any position on the arc AR11 connecting the speakers SP1 and SP2. Here, the arc AR11 is a part of a circle that has the origin O as its center and passes through the positions of the speakers SP1 and SP2.

[0024] Since the vector P is a two-dimensional vector, when the angle between the vectors L1 and L2 is greater than 0 degrees and less than 180 degrees, the coefficients g1 and g2, which are used as gains, are uniquely determined. The method for calculating these coefficients g1 and g2 is described in detail in the above-mentioned Non-Patent Document 1.

[0025] On the other hand, when trying to play back three-channel audio, the number of speakers that output the audio becomes three, as shown in FIG.

[0026] In the example of FIG. 2, the audio of each channel is output from three speakers: a speaker SP1, a speaker SP2, and a speaker SP3.

[0027] In such a case, the concept is the same as that of the two-dimensional VBAP described above, except that the gains of the audio of each channel output from the speakers SP1 to SP3, that is, the coefficients calculated as gains, become three.

[0028] That is, when trying to localize a sound image at the position of the virtual sound source VSP2, in a three-dimensional coordinate system with the position of the head of the user U11 as the origin O, the position of the virtual sound source VSP2 is represented by a three-dimensional vector P with the origin O as the starting point.

[0029] Furthermore, if the three-dimensional vectors starting from the origin O and pointing in the direction of the positions of each speaker SP1 to SP3 are vectors L1 to L3, then vector P can be expressed as a linear sum of vectors L1 to L3, as shown in the following equation (2).

[0030]

number

[0031] By calculating the coefficients g1 to g3 by which vectors L1 to L3 are multiplied in equation (2) and using these coefficients g1 to g3 as the gains of the sound output from speakers SP1 to SP3, respectively, the sound image can be localized at the position of virtual sound source VSP2.

[0032] The method of determining coefficients g1 to g3 using the position information of the three speakers SP1 to SP3 in this way and controlling the localization position of the sound image is called three-dimensional VBAP.

[0033] 2, the sound image can be localized at any position within a triangular region TR11 on a sphere that includes the positions of speakers SP1, SP2, and SP3. Here, region TR11 is a region on the surface of a sphere that is centered at the origin O and includes the positions of speakers SP1 to SP3, and is a triangular region on the sphere that is surrounded by speakers SP1 to SP3.

[0034] Using such a 3D VBAP, it becomes possible to localize a sound image at any position in space.

[0035] For example, as shown in Figure 3, by increasing the number of speakers that output sound and creating multiple areas in space that correspond to the triangular area TR11 shown in Figure 2, it is possible to position the sound image at any position within those areas.

[0036] 3, five speakers SP1 to SP5 are arranged, and audio for each channel is output from these speakers SP1 to SP5. Here, the speakers SP1 to SP5 are arranged on a spherical surface centered at an origin O, which is located at the position of the head of the user U11.

[0037] In this case, the origin O is used as the starting point, and three-dimensional vectors pointing in the direction of the positions of each speaker SP1 to SP5 are defined as vectors L1 to L5, and a calculation similar to that for solving equation (2) described above is performed to find the gain of the sound output from each speaker.

[0038] Here, among the regions on the spherical surface centered on the origin O, the triangular region surrounded by speakers SP1, SP4, and SP5 is defined as region TR21. Similarly, among the regions on the spherical surface centered on the origin O, the triangular region surrounded by speakers SP3, SP4, and SP5 is defined as region TR22, and the triangular region surrounded by speakers SP2, SP3, and SP5 is defined as region TR23.

[0039] These regions TR21 to TR23 correspond to the region TR11 shown in Fig. 2. If a three-dimensional vector indicating the position where a sound image is desired to be localized is denoted by vector P, then in the example of Fig. 3, vector P indicates a position on region TR21.

[0040] Therefore, in this example, vectors L1, L4, and L5 indicating the positions of speakers SP1, SP4, and SP5 are used to perform a calculation similar to that for solving equation (2), and the gain of the sound output from each of speakers SP1, SP4, and SP5 is calculated. In this case, the gain of the sound output from other speakers SP2 and SP3 is set to 0. In other words, no sound is output from speakers SP2 and SP3.

[0041] By arranging the five speakers SP1 to SP5 in the space in this way, it is possible to localize a sound image at any position in the area consisting of the areas TR21 to TR23.

[0042] Now, suppose that four speakers SP1 to SP4 are arranged in a space as shown in FIG. 4, and a sound image is to be localized at the position of a virtual sound source VSP3 located at the center of the speakers SP1 to SP4.

[0043] 4, speakers SP1 to SP4 are arranged on the surface of a sphere centered at an origin O (not shown), and the triangular area on that surface surrounded by speakers SP1 to SP3 is area TR31. Also, the triangular area on the surface of the sphere centered at the origin O surrounded by speakers SP2 to SP4 is area TR32.

[0044] The virtual sound source VSP3 is located on the lower right side of the region TR31. The virtual sound source VSP3 is also located on the upper left side of the region TR32.

[0045] Therefore, in this case, it is sufficient to perform three-dimensional VBAP for speakers SP1 to SP3, or to perform three-dimensional VBAP for speakers SP2 to SP4. In either case, the calculation result of three-dimensional VBAP will be the same, and gains will be determined such that sound is output only from the two speakers, SP2 and SP3, and no sound is output from the remaining speakers, SP1 and SP4.

[0046] In 3D VBAP, if the position where you want to localize the sound image is on the boundary line of a triangular area on the sphere connecting the three speakers, that is, on a side of the triangle on the sphere, sound will only be output from the two speakers located at both ends of that side.

[0047] In this case where sound is output only from two speakers, SP2 and SP3, suppose that user U11, who is in the sweet spot, the optimal listening position, moves to the left in the figure, as shown by arrow A11, as shown in FIG.

[0048] As a result, the head of user U11 moves closer to speaker SP3, causing the sound output from speaker SP3 to sound louder, and user U11 perceives the virtual sound source VSP3, i.e., the sound image, as having moved to the bottom left in the figure, as shown by arrow A12.

[0049] In 3D VBAP, when sound is output from only two speakers as shown in Figure 5, even if the user U11 moves slightly from the sweet spot, the sound image moves in a direction perpendicular to the direction of movement of the user U11. In such a case, the user U11 perceives the sound image as if it has moved in a direction different from the direction of movement of the user U11, which creates an uncomfortable feeling. In other words, the user U11 perceives the positioning of the sound image as unstable, and the range of the sweet spot becomes narrower.

[0050] Therefore, unlike the above-mentioned VBAP, this technology outputs sound from more than three speakers, i.e., four or more speakers, thereby making the positioning of the sound image more stable and thereby widening the range of the sweet spot.

[0051] The number of speakers that output sound may be any number equal to or greater than four, but the following description will be given taking as an example a case where sound is output from four speakers.

[0052] For example, as in the example shown in FIG. 4, it is assumed that a sound image is localized at the position of a virtual sound source VSP3 located at the center of the four speakers SP1 to SP4.

[0053] In such a case, in the present technology, two or three speakers are selected to form one combination, and VBAP is performed for multiple different combinations to calculate the gain of the audio output from the four speakers SP1 to SP4.

[0054] Therefore, in the present technology, for example, as shown in FIG. 6, sounds are output from all four speakers SP1 to SP4.

[0055] In such a case, even if the user U11 moves leftward from the sweet spot as indicated by arrow A21 in Fig. 6, the position of the virtual sound source VSP3, i.e., the localization position of the sound image, simply moves leftward as indicated by arrow A22 in the figure. In other words, unlike the example shown in Fig. 5, the sound image does not move downward, i.e., in a direction perpendicular to the direction of movement of the user U11, but moves only in the same direction as the direction of movement of the user U11.

[0056] This is because when user U11 moves leftward, he or she approaches speaker SP3, but speaker SP1 is located above speaker SP3. In this case, sounds reach user U11's ears from both the upper left and lower left sides as seen from user U11, making it difficult for the sound image to be perceived as having moved downward in the figure.

[0057] Therefore, compared to the conventional VBAP method, the sound image can be localized more stably, and as a result, the range of the sweet spot can be expanded.

[0058] Next, the control of sound image localization according to the present technology will be described more specifically.

[0059] In this technology, a vector indicating the position where a sound image is to be localized is defined as a vector P starting from an origin O (not shown) of a three-dimensional coordinate system, and the vector P is expressed by the following equation (3).

[0060]

number

[0061] In equation (3), vectors L1 to L4 are three-dimensional vectors that point toward the positions of speakers SP1 to SP4, which are located near the sound image localization position and surround the sound image localization position. Also, g1 to g4 represent coefficients that will be used as gains for the sounds of each channel to be output from speakers SP1 to SP4.

[0062] In equation (3), vector P is expressed as a linear sum of four vectors L1 to L4. Here, vector P is a three-dimensional vector, and therefore the four coefficients g1 to g4 cannot be uniquely determined.

[0063] Therefore, in the present technology, the coefficients g1 to g4 that become the gains are calculated by the following method.

[0064] Now, it is assumed that a sound image is to be localized at the center position of a quadrangle on a spherical surface surrounded by the four speakers SP1 to SP4 shown in FIG. 4, that is, at the position of the virtual sound source VSP3.

[0065] First, any one side of a quadrangle on the spherical surface having the speakers SP1 to SP4 as its vertices is selected, and it is assumed that a virtual speaker (hereinafter referred to as a virtual speaker) is located on that side.

[0066] For example, as shown in Fig. 7, let us assume that the side connecting speakers SP3 and SP4 located at the bottom left and bottom right in the figure is selected from a quadrangle on a spherical surface with speakers SP1 to SP4 as vertices. Then, for example, virtual speaker VSP' is assumed to be located at the intersection of a perpendicular line drawn from the position of virtual sound source VSP3 to the side connecting speakers SP3 and SP4.

[0067] Next, three-dimensional VBAP is performed for this virtual speaker VSP' and the speakers SP1 and SP2 located at the top left and top right in the figure, for a total of three speakers. That is, by solving an equation similar to the above-mentioned equation (2), coefficients g1, g2, and g', which are the gains of the sounds output from speaker SP1, speaker SP2, and virtual speaker VSP', respectively, are found.

[0068] 7, vector P is expressed as a linear sum of three vectors starting from origin O: vector L1 pointing in the direction of speaker SP1, vector L2 pointing in the direction of speaker SP2, and vector L' pointing in the direction of virtual speaker VSP'. That is, P = g1L1 + g2L2 + g'L'.

[0069] Here, in order to localize a sound image at the position of the virtual sound source VSP3, sound must be output from the virtual speaker VSP' with a gain g', but the virtual speaker VSP' does not actually exist. Therefore, in this technology, as shown in Fig. 8, the virtual speaker VSP' is realized by localizing a sound image at the position of the virtual speaker VSP' using two speakers SP3 and SP4 located at both ends of the sides of the rectangle in which the virtual speaker VSP' is located.

[0070] Specifically, two-dimensional VBAP is performed for the two speakers SP3 and SP4 located at both ends of the side of the sphere on which the virtual speaker VSP' is located. That is, by solving an equation similar to the above-mentioned equation (1), coefficients g3' and g4', which are the gains of the sounds output from the speakers SP3 and SP4, respectively, are calculated.

[0071] 8, the vector L' pointing in the direction of the virtual speaker VSP' is expressed as the linear sum of the vector L3 pointing in the direction of the speaker SP3 and the vector L4 pointing in the direction of the speaker SP4, i.e., L'=g3'L3+g4'L4.

[0072] The value g'g3' obtained by multiplying the coefficient g3' thus obtained by the coefficient g' is set as the gain of the sound to be output from the speaker SP3, and the value g'g4' obtained by multiplying the coefficient g4' by the coefficient g' is set as the gain of the sound to be output from the speaker SP4. In this way, the speakers SP3 and SP4 realize a virtual speaker VSP' that outputs sound with a gain g'.

[0073] Here, the value of g'g3', which is the gain value, becomes the value of the coefficient g3 in the above-mentioned equation (3), and the value of g'g4', which is the gain value, becomes the value of the coefficient g4 in the above-mentioned equation (3).

[0074] If the non-zero values ​​g1, g2, g'g3', and g'g4' obtained in this manner are used as the gains of the audio for each channel output from speakers SP1 to SP4, audio can be output from the four speakers and the sound image can be localized at the target position.

[0075] By outputting sound from four speakers in this way and localizing the sound image, the sound image can be localized more stably than when localizing the sound image using the conventional VBAP method, thereby expanding the range of the sweet spot.

[0076] <Configuration example of voice processing device> Next, a specific embodiment to which the present technology described above is applied will be described. Fig. 9 is a diagram showing an example of the configuration of an embodiment of a voice processing device to which the present technology is applied.

[0077] The audio processing device 11 generates audio signals of N channels (where N≧5) by performing gain adjustment for each channel on a monaural audio signal supplied from an external device, and supplies the audio signals to speakers 12-1 to 12-N corresponding to each of the N channels.

[0078] The speakers 12-1 to 12-N output audio for each channel based on an audio signal supplied from the audio processing device 11. That is, the speakers 12-1 to 12-N are audio output units that serve as sound sources that output audio for each channel. Hereinafter, when there is no need to particularly distinguish between the speakers 12-1 to 12-N, they will also be simply referred to as speakers 12. Although FIG. 9 shows a configuration in which the speaker 12 is not included in the audio processing device 11, the speaker 12 may be included in the audio processing device 11. Furthermore, the respective units constituting the audio processing device 11 and the speaker 12 may be provided in, for example, several devices, thereby forming an audio processing system comprising the respective units of the audio processing device 11 and the speaker 12.

[0079] The speakers 12 are arranged so as to surround a position where the user is expected to be located when viewing content, etc. (hereinafter simply referred to as the user's position). For example, the speakers 12 are arranged at positions on the surface of a sphere centered on the user's position. In other words, the speakers 12 are arranged at positions equidistant from the user. Furthermore, the audio signals may be supplied from the audio processing device 11 to the speakers 12 via a wired or wireless connection.

[0080] The audio processing device 11 includes a speaker selection unit 21 , a gain calculation unit 22 , a gain determination unit 23 , a gain output unit 24 , and a gain adjustment unit 25 .

[0081] The audio processing device 11 is supplied with an audio signal of an audio sound picked up by a microphone attached to an object such as a moving object, and position information of the object.

[0082] The speaker selection unit 21 identifies a position (hereinafter also referred to as a target sound image position) where the sound image of the sound emitted from the object should be localized in the space in which the speakers 12 are placed, based on the position information of the object supplied from outside, and supplies the identification result to the gain calculation unit 22.

[0083] In addition, the speaker selection unit 21 selects four speakers 12 from among the N speakers 12 based on the target sound image position, as the speakers 12 to be processed, from which audio should be output, and supplies selection information indicating the selection result to the gain calculation unit 22, the gain determination unit 23, and the gain output unit 24.

[0084] The gain calculation unit 22 calculates the gain of the speaker 12 to be processed based on the selection information supplied from the speaker selection unit 21 and the target sound image position, and supplies it to the gain output unit 24. The gain determination unit 23 determines the gain of the speaker 12 that is not the processing target based on the selection information supplied from the speaker selection unit 21, and supplies it to the gain output unit 24. For example, the gain of the speaker 12 that is not the processing target is set to "0". In other words, the speaker 12 that is not the processing target is controlled so that the sound of the object is not output.

[0085] The gain output unit 24 supplies the N gains supplied from the gain calculation unit 22 and the gain determination unit 23 to the gain adjustment unit 25. At this time, the gain output unit 24 determines, based on the selection information supplied from the speaker selection unit 21, the supply destinations within the gain adjustment unit 25 of the N gains supplied from the gain calculation unit 22 and the gain determination unit 23.

[0086] The gain adjustment unit 25 performs gain adjustment on the audio signals of the externally supplied objects based on the gains supplied from the gain output unit 24, and supplies the resulting audio signals of each of the N channels to the speaker 12 to output audio.

[0087] The gain adjustment unit 25 includes amplifiers 31-1 to 31-N. The amplifiers 31-1 to 31-N adjust the gain of the audio signal supplied from the outside based on the gain supplied from the gain output unit 24, and supply the audio signal obtained as a result to the speakers 12-1 to 12-N.

[0088] In the following description, when there is no need to distinguish between the amplifiers 31-1 to 31-N, they will also be simply referred to as amplifiers 31.

[0089] <Configuration example of gain calculation unit> The gain calculation section 22 shown in FIG. 9 is configured as shown in FIG. 10, for example.

[0090] The gain calculation unit 22 shown in FIG. 10 includes a virtual speaker position determination unit 61, a three-dimensional gain calculation unit 62, a two-dimensional gain calculation unit 63, a multiplication unit 64, and a multiplication unit 65.

[0091] The virtual speaker position determination unit 61 determines the positions of the virtual speakers based on the information indicating the target sound image position and the selection information supplied from the speaker selection unit 21. The virtual speaker position determination unit 61 supplies the information indicating the target sound image position, the selection information, and the information indicating the positions of the virtual speakers to a three-dimensional gain calculation unit 62, and supplies the selection information and the information indicating the positions of the virtual speakers to a two-dimensional gain calculation unit 63.

[0092] The three-dimensional gain calculation unit 62 performs three-dimensional VBAP on two of the speakers 12 to be processed and on the virtual speaker, based on the information supplied from the virtual speaker position determination unit 61. Then, the three-dimensional gain calculation unit 62 supplies the gains of the two speakers 12 obtained by the three-dimensional VBAP to the gain output unit 24, and supplies the gain of the virtual speaker to the multiplication units 64 and 65.

[0093] The two-dimensional gain calculation unit 63 performs two-dimensional VBAP on two of the speakers 12 to be processed based on the information supplied from the virtual speaker position determination unit 61, and supplies the resulting gains of the speakers 12 to the multiplication units 64 and 65.

[0094] The multiplication unit 64 multiplies the gain supplied from the two-dimensional gain calculation unit 63 by the gain supplied from the three-dimensional gain calculation unit 62 to obtain the final gain of the speaker 12, and supplies the result to the gain output unit 24. The multiplication unit 65 multiplies the gain supplied from the two-dimensional gain calculation unit 63 by the gain supplied from the three-dimensional gain calculation unit 62 to obtain the final gain of the speaker 12, and supplies the result to the gain output unit 24.

[0095] <Explanation of sound image localization control processing> When the position information and audio signal of an object are supplied to the audio processing device 11 and an instruction to output the audio of the object is given, the audio processing device 11 starts a sound image localization control process, outputs the audio of the object, and localizes the sound image at an appropriate position.

[0096] The sound image localization control process performed by the sound processing device 11 will be described below with reference to the flowchart of FIG.

[0097] In step S11, the speaker selection unit 21 selects the speaker 12 to be processed based on the position information of the object supplied from the outside.

[0098] Specifically, for example, the speaker selection unit 21 identifies the target sound image position based on the position information of the object, and selects four speakers 12 out of the N speakers 12 that are located near the target sound image position and are arranged to surround the target sound image position as the speakers 12 to be processed.

[0099] For example, when the position of the virtual sound source VSP3 shown in FIG. 7 is set as the target sound image position, the speakers 12 corresponding to the four speakers SP1 to SP4 surrounding the virtual sound source VSP3 are selected as the speakers 12 to be processed.

[0100] The speaker selection unit 21 supplies information indicating the target sound image position to the virtual speaker position determination unit 61, and also supplies selection information indicating the four speakers 12 to be processed to the virtual speaker position determination unit 61, the gain determination unit 23, and the gain output unit 24.

[0101] In step S12, the virtual speaker position determination unit 61 determines the positions of the virtual speakers based on the information indicating the target sound image position and the selection information supplied from the speaker selection unit 21. For example, as in the example shown in Fig. 7, the positions of the virtual speakers are determined to be the positions of the intersections of the sides on the spherical surface connecting the speakers 12 located at the lower left and lower right as seen from the user among the speakers 12 to be processed, and the perpendicular lines drawn from the target sound image position to the sides.

[0102] Once the positions of the virtual speakers are determined, the virtual speaker position determination unit 61 supplies information indicating the target sound image position, selection information, and information indicating the positions of the virtual speakers to a three-dimensional gain calculation unit 62, and also supplies the selection information and information indicating the positions of the virtual speakers to a two-dimensional gain calculation unit 63.

[0103] The positions of the virtual speakers may be any positions on the sides of a quadrangle on the spherical surface, with the four speakers 12 being the processing targets as vertices. Even if there are five or more speakers 12 being the processing targets, the positions of the virtual speakers may be any positions on the sides of a polygon on the spherical surface, with the speakers 12 being vertices.

[0104] In step S13, the three-dimensional gain calculation unit 62 calculates gains for the virtual speaker and the two speakers 12 to be processed based on the information indicating the target sound image position, the selection information, and the information indicating the positions of the virtual speakers supplied from the virtual speaker position determination unit 61.

[0105] Specifically, the three-dimensional gain calculation unit 62 defines a three-dimensional vector indicating the target sound image position as vector P, and defines a three-dimensional vector pointing toward the virtual speaker as vector L'. Furthermore, the three-dimensional gain calculation unit 62 defines, among the speakers 12 being processed, a vector pointing toward the speaker 12 that is in the same positional relationship as speaker SP1 shown in Fig. 7 as vector L1, and a vector pointing toward the speaker 12 that is in the same positional relationship as speaker SP2 as vector L2.

[0106] Then, the three-dimensional gain calculation unit 62 finds an equation that expresses the vector P as a linear sum of the vectors L', L1, and L2, and solves the equation to calculate the coefficients g', g1, and g2 of the vectors L', L1, and L2 as gains. That is, a calculation similar to the calculation for solving the above-mentioned equation (2) is performed.

[0107] The three-dimensional gain calculation unit 62 supplies the coefficients g1 and g2 of the speakers 12 in the same positional relationship as the speakers SP1 and SP2 obtained as a result of the calculation to the gain output unit 24 as the gains of the sounds to be output from those speakers 12.

[0108] Furthermore, the three-dimensional gain calculation unit 62 supplies the coefficient g' of the virtual speaker obtained as a result of the calculation to the multiplication units 64 and 65 as the gain of the sound to be output from the virtual speaker.

[0109] In step S14, the two-dimensional gain calculation unit 63 calculates gains for the two speakers 12 to be processed, based on the selection information supplied from the virtual speaker position determination unit 61 and the information indicating the positions of the virtual speakers.

[0110] Specifically, the two-dimensional gain calculation unit 63 sets a three-dimensional vector indicating the position of the virtual speaker as vector L'. Furthermore, the two-dimensional gain calculation unit 63 sets a vector pointing toward the speaker 12 that is the processing target and has the same positional relationship as the speaker SP3 shown in Fig. 8 as vector L3, and a vector pointing toward the speaker 12 that is the same positional relationship as the speaker SP4 shown in Fig. 8 as vector L4.

[0111] Then, the two-dimensional gain calculation unit 63 finds an equation that expresses the vector L' as a linear sum of the vectors L3 and L4, and solves the equation to calculate the coefficients g3' and g4' of the vectors L3 and L4 as gains. That is, a calculation similar to the calculation for solving the above-mentioned equation (1) is performed.

[0112] The two-dimensional gain calculation unit 63 supplies the coefficients g3' and g4' of the speakers 12 that are in the same positional relationship as the speakers SP3 and SP4 obtained as a result of the calculation to the multiplication units 64 and 65 as the gains of the audio to be output from those speakers 12.

[0113] In step S15, the multiplication units 64 and 65 multiply the gains g3' and g4' supplied from the two-dimensional gain calculation unit 63 by the gain g' of the virtual speaker supplied from the three-dimensional gain calculation unit 62, and supply the result to the gain output unit 24.

[0114] Therefore, of the four speakers 12 being processed, g3=g'g3' is supplied to the gain output unit 24 as the final gain of the speaker 12 that is in the same positional relationship as the speaker SP3 in Fig. 8. Similarly, of the four speakers 12 being processed, g4=g'g4' is supplied to the gain output unit 24 as the final gain of the speaker 12 that is in the same positional relationship as the speaker SP4 in Fig. 8.

[0115] In step S16, the gain determination unit 23 determines the gains of the speakers 12 that are not the processing targets based on the selection information supplied from the speaker selection unit 21, and supplies the determined gains to the gain output unit 24. For example, the gains of all the speakers 12 that are not the processing targets are set to "0".

[0116] When the gain output unit 24 is supplied with the gain g1, gain g2, gain g'g3', and gain g'g4' from the gain calculation unit 22 and the gain "0" from the gain determination unit 23, the gain output unit 24 supplies those gains to the amplification unit 31 of the gain adjustment unit 25 based on the selection information from the speaker selection unit 21.

[0117] Specifically, the gain output unit 24 supplies gains g1, g2, g'g3', and g'g4' to the amplifier unit 31 that supplies audio signals to the speakers 12 corresponding to the respective speakers 12 to be processed, that is, the speakers SP1 to SP4 in Fig. 7. For example, when the speaker 12 corresponding to the speaker SP1 is the speaker 12-1, the gain output unit 24 supplies the gain g1 to the amplifier unit 31-1.

[0118] Moreover, the gain output unit 24 supplies the gain "0" supplied from the gain determination unit 23 to the amplifier unit 31 that supplies the audio signal to the speaker 12 that is not the processing target.

[0119] In step S17, the amplifier 31 of the gain adjustment unit 25 adjusts the gain of the audio signal of the object supplied from outside based on the gain supplied from the gain output unit 24, and supplies the resulting audio signal to the speaker 12 to output audio.

[0120] Each speaker 12 outputs sound based on the audio signal supplied from the amplifier 31. More specifically, sound is output from only the four speakers 12 that are the processing targets. This allows the sound image to be localized at the target position. When sound is output from the speakers 12, the sound image localization control process ends.

[0121] In this way, the audio processing device 11 selects four speakers 12 to be processed from the object position information, and performs VBAP for combinations of those speakers 12 and two or three of the virtual speakers.The audio processing device 11 then adjusts the gain of the audio signal based on the gain of each speaker 12 to be processed, which is obtained by performing VBAP for a plurality of different combinations.

[0122] This allows sound to be output from the four speakers 12 positioned around the target sound image position, making it possible to more stabilize the positioning of the sound image, and as a result, the range of the sweet spot can be further expanded.

[0123] Second Embodiment <Calculating gain> Note that the above description has been given of an example in which two or three speakers are selected from five speakers including a virtual speaker to form one speaker combination, and VBAP is performed for the plurality of combinations to calculate the gain of the speaker 12 to be processed. However, with the present technology, it is also possible to calculate the gain by selecting a plurality of combinations from the four speakers 12 to be processed without specifying a virtual speaker, and performing VBAP for each of the combinations.

[0124] In such a case, the number of times VBAP should be performed varies depending on the target sound image position, as shown in Fig. 12. Note that in Fig. 12, the same reference numerals are used to designate parts corresponding to those in Fig. 7, and the description thereof will be omitted where appropriate.

[0125] For example, if the position of the virtual sound source, that is, the target sound image position, is at the position indicated by arrow Q11, the position indicated by arrow Q11 is within a triangular area on the spherical surface surrounded by speakers SP1, SP2, and SP4. Therefore, if three-dimensional VBAP is performed on the speaker group consisting of speakers SP1, SP2, and SP4 (hereinafter also referred to as the first group), the gains of the sounds output from the three speakers, speaker SP1, speaker SP2, and speaker SP4, can be determined.

[0126] On the other hand, the position indicated by arrow Q11 is also a position within the triangular area on the spherical surface surrounded by speakers SP2, SP3, and SP4. Therefore, if three-dimensional VBAP is performed on the speaker group consisting of speakers SP2, SP3, and SP4 (hereinafter also referred to as the second group), the gain of the sound output from the three speakers, speaker SP2, speaker SP3, and speaker SP4, can be found.

[0127] Here, if the gains of the speakers that are not used in the first and second groups are set to "0", then in this example, a total of two gains can be obtained for each of the four speakers SP1 to SP4 in the first and second groups.

[0128] Therefore, for each speaker, the sum of the gains of the speakers obtained in the first and second groups is calculated as the gain sum. For example, if the gain of speaker SP1 obtained in the first group is g1(1) and the gain of speaker SP1 obtained in the second group is g1(2), then the gain sum of speaker SP1 is g s1 is the gain sum g s1 =g1(1)+g1(2).

[0129] Here, the second set of speaker combinations does not include speaker SP1, so g1(2) is 0, but the first set of speaker combinations includes speaker SP1, so g1(1) is a non-zero value. In the end, the gain sum g of speaker SP1 is s1does not become 0. The same applies to the sum of the gains of the other speakers SP2 to SP4.

[0130] Once the gain sum of each speaker is calculated in this way, the value obtained by normalizing the gain sum of each speaker by the sum of the squares of those gain sums can be used as the final gain of those speakers, or more specifically, the gain of the sound output from the speakers.

[0131] By calculating the gain of each speaker SP1 to SP4 in this way, a gain that is not zero is always obtained, so that sound can be output from each of the four speakers SP1 to SP4 and the sound image can be localized at the desired position.

[0132] In the following, the gain of the speaker SPk (where 1≦k≦4) obtained for the mth group (where 1≦m≦4) is referred to as g k The sum of the gains of the speakers SPk (where 1≦k≦4) is expressed as g sk This will be expressed as follows.

[0133] Furthermore, if the target sound image position is at the position indicated by arrow Q12, that is, on the spherical surface, at the intersection of the line connecting speakers SP2 and SP3 and the line connecting speakers SP1 and SP4, there are four possible combinations of the three speakers.

[0134] That is, possible combinations are a combination of speakers SP1, SP2, and SP3 (hereinafter referred to as a first group), and a combination of speakers SP1, SP2, and SP4 (hereinafter referred to as a second group).In addition, possible combinations are a combination of speakers SP1, SP3, and SP4 (hereinafter referred to as a third group), and a combination of speakers SP2, SP3, and SP4 (hereinafter referred to as a fourth group).

[0135] In this case, 3D VBAP can be performed on each of the first through fourth combinations to find the gain for each speaker. The sum of the four gains found for the same speaker is then taken as the gain sum, and the value obtained by normalizing the gain sum for each speaker by the sum of the squares of the four gain sums found for each speaker is taken as the final gain for those speakers.

[0136] If the target sound image position is at the position indicated by arrow Q12 and the quadrangle on the spherical surface formed by speakers SP1 to SP4 is a rectangle or the like, the same calculation result will be obtained for, for example, the first and fourth pairs of speakers using three-dimensional VBAP. Therefore, in such a case, the gain of each speaker can be obtained by performing three-dimensional VBAP for two appropriate combinations, such as the first and second pairs. However, if the quadrangle on the spherical surface formed by speakers SP1 to SP4 is an asymmetric quadrangle other than a rectangle or the like, it is necessary to perform three-dimensional VBAP for each of the four combinations.

[0137] <Configuration example of gain calculation unit> As described above, when multiple combinations are selected from the four speakers 12 to be processed without determining a virtual speaker, and VBAP is performed for each combination to calculate a gain, the gain calculation unit 22 shown in FIG. 9 is configured, for example, as shown in FIG. 13.

[0138] The gain calculation section 22 shown in FIG. 13 includes a selection section 91, a three-dimensional gain calculation section 92-1, a three-dimensional gain calculation section 92-2, a three-dimensional gain calculation section 92-3, a three-dimensional gain calculation section 92-4, and an addition section 93.

[0139] The selection unit 91 determines a combination of three speakers 12 surrounding the target sound image position from among the four speakers 12 targeted for processing, based on the information indicating the target sound image position and the selection information supplied from the speaker selection unit 21. The selection unit 91 supplies the information indicating the combination of the speakers 12 and the information indicating the target sound image position to three-dimensional gain calculation units 92-1 to 92-4.

[0140] The three-dimensional gain calculation units 92-1 to 92-4 perform three-dimensional VBAP based on the information indicating the combination of the speakers 12 and the information indicating the target sound image position supplied from the selection unit 91, and supply the resulting gains of each speaker 12 to the addition unit 93. Note that hereinafter, when there is no need to particularly distinguish between the three-dimensional gain calculation units 92-1 to 92-4, they will also be simply referred to as three-dimensional gain calculation units 92.

[0141] The adder 93 calculates a gain sum based on the gains of each speaker 12 to be processed supplied from the three-dimensional gain calculation units 92-1 to 92-4, and further normalizes the gain sums to calculate the final gain of each speaker 12 to be processed, and supplies the resulting gain to the gain output unit 24.

[0142] <Explanation of sound image localization control processing> Next, with reference to the flowchart of FIG. 14, a sound image localization control process performed when the gain calculation section 22 has the configuration shown in FIG. 13 will be described.

[0143] The process of step S41 is the same as the process of step S11 in FIG. 11, and therefore a description thereof will be omitted.

[0144] In step S42, the selection unit 91 determines a combination of speakers 12 based on the information indicating the target sound image position and the selection information supplied from the speaker selection unit 21, and supplies the information indicating the combination of speakers 12 and the information indicating the target sound image position to the three-dimensional gain calculation unit 92.

[0145] For example, when the target sound image position is at the position indicated by the arrow Q11 in Fig. 12, a combination (first set) of speakers 12 consisting of three speakers 12 corresponding to the speakers SP1, SP2, and SP4 is determined. Also, a combination (second set) of speakers 12 consisting of three speakers 12 corresponding to the speakers SP2, SP3, and SP4 is determined.

[0146] In this case, for example, the selection unit 91 supplies information indicating a first combination of speakers 12 and information indicating a target sound image position to the three-dimensional gain calculation unit 92-1, and supplies information indicating a second combination of speakers 12 and information indicating a target sound image position to the three-dimensional gain calculation unit 92-2. In this case, the three-dimensional gain calculation unit 92-3 and the three-dimensional gain calculation unit 92-4 are not supplied with information indicating the combination of speakers 12, and the three-dimensional gain calculation unit 92-3 and the three-dimensional gain calculation unit 92-4 do not calculate the three-dimensional VBAP.

[0147] In step S43, the three-dimensional gain calculation unit 92 calculates the gain of each speaker 12 to be processed for the combination of speakers 12 based on the information indicating the combination of speakers 12 supplied from the selection unit 91 and the information indicating the target sound image position, and supplies the calculated gain to the addition unit 93.

[0148] Specifically, the three-dimensional gain calculation unit 92 performs the same process as step S13 in Fig. 11 described above for the three speakers 12 indicated by the information indicating the combination of the speakers 12 to determine the gain of each speaker 12. That is, a calculation similar to the calculation for solving the above-described equation (2) is performed. Furthermore, the gain of the remaining speaker 12, which is not one of the three speakers 12 indicated by the information indicating the combination of the speakers 12, among the four speakers 12 to be processed, is set to "0".

[0149] For example, when two combinations, a first set and a second set, are obtained in step S42, the three-dimensional gain calculation unit 92-1 calculates the gain of each speaker 12 for the first set using three-dimensional VBAP, and the three-dimensional gain calculation unit 92-2 calculates the gain of each speaker 12 for the second set using three-dimensional VBAP.

[0150] Specifically, it is assumed that a first set of three speakers 12 corresponding to speakers SP1, SP2, and SP4 shown in Fig. 12 has been determined. In this case, the three-dimensional gain calculation unit 92-1 calculates a gain g1(1) of the speaker 12 corresponding to speaker SP1, a gain g2(1) of the speaker 12 corresponding to speaker SP2, and a gain g4(1) of the speaker 12 corresponding to speaker SP4. In addition, the gain g3(1) of the speaker 12 corresponding to speaker SP3 is set to "0".

[0151] In step S44, the adder 93 calculates the final gain of the target speaker 12 based on the gain of each speaker 12 supplied from the three-dimensional gain calculator 92, and supplies the calculated final gain to the gain output unit .

[0152] For example, the adder 93 calculates the sum of the gains g1(1), g1(2), g1(3), and g1(4) of the speaker 12 corresponding to the speaker SP1 supplied from the three-dimensional gain calculator 92, to calculate the gain sum g1(1), g1(2), g1(3), and g1(4) of the speaker 12. s1 Similarly, the adder 93 calculates the gain sum g of the speaker 12 corresponding to the speaker SP2. s2 , the gain sum g of speaker 12 corresponding to speaker SP3 s3 , and the gain sum g of speaker 12 corresponding to speaker SP4 s4 Also calculate.

[0153] Then, the adder 93 calculates the gain sum g of the speaker 12 corresponding to the speaker SP1. s1 The gain sum g s1 to gain sum g s4 The adder 93 also calculates the final gains g2 to g4 of the speakers 12 corresponding to the speakers SP2 to SP4 by performing a similar calculation.

[0154] Once the gain of the speaker 12 to be processed is determined in this manner, the processes of steps S45 and S46 are performed, and the sound image localization control process ends. However, since these processes are similar to the processes of steps S16 and S17 in FIG. 11, their description will be omitted.

[0155] In this way, the audio processing device 11 selects four speakers 12 to be processed from the object position information, and performs VBAP on a combination of speakers 12 consisting of three speakers 12 out of the selected speakers 12. The audio processing device 11 then calculates the sum of the gains of the same speakers 12 obtained by performing VBAP on a plurality of different combinations, thereby calculating the final gain of each speaker 12 to be processed, and performs gain adjustment of the audio signal.

[0156] This allows sound to be output from the four speakers 12 positioned around the target sound image position, making it possible to more stabilize the positioning of the sound image, and as a result, the range of the sweet spot can be further expanded.

[0157] In this embodiment, an example has been described in which the four speakers 12 surrounding the target sound image position are the speakers 12 to be processed, but the number of speakers 12 to be processed may be four or more.

[0158] For example, if five speakers 12 are selected as the speakers 12 to be processed, a set of speakers 12 consisting of any three speakers 12 that surround the target sound image position among those five speakers 12 is selected as one combination.

[0159] Specifically, it is assumed that the speakers 12 corresponding to five speakers SP1 to SP5 as shown in FIG. 15 are selected as the speakers 12 to be processed, and the target sound image position is set to the position indicated by the arrow Q21.

[0160] In this case, a combination of speakers SP1, SP2, and SP3 is selected as the first set, a combination of speakers SP1, SP2, and SP4 is selected as the second set, and a combination of speakers SP1, SP2, and SP5 is selected as the third set.

[0161] Then, the gain of each speaker is found for these first to third groups, and the final gain is calculated from the sum of the gains of each speaker. That is, the process of step S43 in Fig. 14 is performed for the first to third groups, and then the processes of steps S44 to S46 are performed.

[0162] In this way, even when five or more speakers 12 are selected as the speakers 12 to be processed, sound can be output from all the speakers 12 to be processed and the sound image can be localized.

[0163] The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the programs constituting the software are installed on a computer. Here, the term "computer" includes computers built into dedicated hardware, and general-purpose computers that can execute various functions by installing various programs.

[0164] FIG. 16 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes using a program.

[0165] In the computer, a CPU 801 , a ROM 802 , and a RAM 803 are interconnected by a bus 804 .

[0166] An input / output interface 805 is further connected to the bus 804. An input unit 806, an output unit 807, a recording unit 808, a communication unit 809, and a drive 810 are connected to the input / output interface 805.

[0167] The input unit 806 includes a keyboard, a mouse, a microphone, an image sensor, etc. The output unit 807 includes a display, a speaker, etc. The recording unit 808 includes a hard disk, a nonvolatile memory, etc. The communication unit 809 includes a network interface, etc. The drive 810 drives removable media 811 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.

[0168] In a computer configured as described above, the CPU 801 performs the above-described series of processes by, for example, loading a program recorded in the recording unit 808 into the RAM 803 via the input / output interface 805 and the bus 804 and executing it.

[0169] The program executed by the computer (CPU 801) can be provided by being recorded on removable media 811 such as package media, for example. The program can also be provided via wired or wireless transmission media such as a local area network, the Internet, or digital satellite broadcasting.

[0170] In a computer, the program can be installed in the recording unit 808 via the input / output interface 805 by inserting the removable medium 811 into the drive 810. The program can also be received by the communication unit 809 via a wired or wireless transmission medium and installed in the recording unit 808. Alternatively, the program can be installed in the ROM 802 or the recording unit 808 in advance.

[0171] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.

[0172] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present technology.

[0173] For example, this technology can be configured as cloud computing, in which a single function is shared and processed collaboratively by multiple devices via a network.

[0174] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by multiple devices.

[0175] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.

[0176] Furthermore, the present technology can also be configured as follows.

[0177] [1] 4 or more audio outputs; a gain calculation unit that calculates a gain of the sound to be output from the four or more audio output units for each of a plurality of different combinations of two or three audio output units among the four or more audio output units located near a target sound image localization position, based on a positional relationship of the audio output units, thereby determining an output gain of the sound to be output from the four or more audio output units in order to localize a sound image at the sound image localization position; a gain adjustment unit that adjusts the gain of the sound output from the sound output unit based on the output gain; An audio processing device comprising: [2] The output gain value is at least 4 and is not 0. [1] The voice processing device according to [1]. [3] The gain calculation unit a first gain calculation unit that calculates the output gains of the virtual audio output unit and the two audio output units based on a positional relationship between the virtual audio output unit, the two audio output units, and the sound image localization position; a second gain calculation unit that calculates gains of the other two audio output units for localizing a sound image at the position of the virtual audio output unit, based on a positional relationship between the other two audio output units different from the two audio output units and the virtual audio output unit; a calculation unit that calculates the output gains of the other two audio output units based on gains of the other two audio output units and the output gain of the virtual audio output unit; Equipped with [1] or [2]. [4] The calculation unit calculates the output gains of the other two audio output units by multiplying the gains of the other two audio output units by the output gain of the virtual audio output unit. [3] The audio processing device according to [3]. [5] The positions of the virtual audio output units are determined so as to be located on the sides of a polygon having the four or more audio output units as vertices. [3] or [4], the voice processing device. [6] The gain calculation unit a temporary gain calculation unit that calculates the output gains of the three audio output units based on a positional relationship between the three audio output units and the sound image localization position; a calculation unit that calculates a final output gain of the audio output unit based on the output gains calculated by a plurality of temporary gain calculation units that calculate the output gains for the different combinations; Equipped with [1] or [2]. [7] The calculation unit calculates the final output gain of the audio output unit by calculating the sum of the output gains calculated for the same audio output unit. [6] The audio processing device according to [6]. [Explanation of symbols]

[0178] 11 audio processing device, 12-1 to 12-N, 12 speakers, 21 speaker selection unit, 22 gain calculation unit, 25 gain adjustment unit, 61 virtual speaker position determination unit, 62 three-dimensional gain calculation unit, 63 two-dimensional gain calculation unit, 64 multiplication unit, 65 multiplication unit, 91 selection unit, 92-1 to 92-4, 92 three-dimensional gain calculation unit, 93 addition unit

Claims

1. a virtual speaker position determination unit that determines the position of a virtual speaker based on the positions of the four speakers; a gain calculation unit that calculates gains of two of the four speakers and the virtual speaker based on a target sound image position using a three-dimensional VBAP; a gain adjustment unit that adjusts the gain of the sound to be output from at least two of the speakers based on the gain of the virtual speaker; An audio processing device comprising:

2. The audio processing device determining a position of a virtual speaker based on the positions of the four speakers; Calculating gains of two of the four speakers and the virtual speaker based on a target sound image position using a three-dimensional VBAP; adjusting the gain of the sound output from at least two of the speakers based on the gain of the virtual speaker; An information processing method including:

3. determining a position of a virtual speaker based on the positions of the four speakers; Calculating gains of two of the four speakers and the virtual speaker based on a target sound image position using a three-dimensional VBAP; adjusting the gain of the sound output from at least two of the speakers based on the gain of the virtual speaker; A program that causes a computer to execute a process including the above.