Information processing program, information processing method, and information processing apparatus

The information processing program encodes sound data using specific coefficients for the r and u channels to address sound image deviation issues, ensuring precise sound localization by adjusting coefficients based on the virtual sound source's position, thus improving sound reproduction accuracy.

JP2026083866APending Publication Date: 2026-05-20KOEI TECMO GAMES CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
KOEI TECMO GAMES CO LTD
Filing Date
2024-11-08
Publication Date
2026-05-20

AI Technical Summary

Technical Problem

Existing techniques for forming a three-dimensional sound field using multiple speakers around a listener can cause the position of sound images to deviate from the target position, especially when sound is radiated asymmetrically or from non-symmetrical directions.

Method used

An information processing program that encodes sound data using specific coefficients for each channel, including the r and u channels, to strengthen signals in particular plane directions, and adjusts coefficients based on the position of a virtual sound source relative to a listener within a set area, using a combination of first, second, and third coefficients to suppress sound image deviation.

Benefits of technology

The method effectively suppresses sound image deviation from the target position, ensuring accurate sound localization by adjusting coefficients to account for the virtual sound source's position and speaker arrangement, thereby enhancing the precision of sound reproduction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026083866000001_ABST
    Figure 2026083866000001_ABST
Patent Text Reader

Abstract

The objective is to obtain an information processing program, an information processing method, and an information processing device that can suppress the deviation of the sound image's position from the target position. [Solution] The information processing program is a scene-based encoding method that encodes sound data using coefficients for each of a plurality of channels, including a specific channel that has the property of strengthening signals in a particular plane direction. In this encoding method, when a virtual sound source is located within a set area of ​​a virtual three-dimensional space based on a virtual listener, the program causes the computer to perform a process of encoding sound data to be played back from the virtual sound source using a first coefficient obtained by combining the coefficients of the specific channel corresponding to each of a plurality of positions on a specific plane facing each other across the listener.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing program, an information processing method, and an information processing apparatus.

Background Art

[0002] Patent Document 1 discloses a technique for reproducing a three-dimensional sound field such that a listener appears to be present at the installation location of an ambisonics microphone by performing signal processing on a signal picked up using an ambisonics microphone, using a plurality of speakers.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] By the way, a technique for forming a three-dimensional sound field that realizes sound images including not only the planar direction but also the height direction using a plurality of speakers arranged around a listener is known. This technique is also called a surround system. Here, for example, when sound is radiated from above the listener toward the listener and not from the listener's feet toward the listener, or when sound is radiated in an asymmetric direction toward the listener, the position of the sound image may deviate from the target position.

[0005] An object of the present disclosure is to provide an information processing program, an information processing method, and an information processing apparatus capable of suppressing the deviation of the position of a sound image from the target position. <00​​​​​The first embodiment of the information processing program is a scene-based encoding method in which sound data is encoded using coefficients for each of a plurality of channels, including a specific channel having the property of strengthening signals in a particular plane direction, and in the case where a virtual sound source is located within a set area of ​​a virtual three-dimensional space with respect to a virtual listener, the computer is made to perform a process of encoding sound data to be played back from the virtual sound source using a first coefficient obtained by combining the coefficients of the specific channel corresponding to each of a plurality of positions on a specific plane facing each other with respect to the listener.

[0007] The second embodiment of the information processing program causes the computer to perform a process of encoding the sound data using the position of the virtual sound source and a second coefficient calculated based on a predetermined mathematical formula corresponding to each of the multiple channels, when the virtual sound source is located outside the setting area.

[0008] The third embodiment of the information processing program causes the computer to perform a process of encoding the sound data using a third coefficient obtained by combining the first coefficient and the second coefficient, when the virtual sound source is located within the setting area of ​​the information processing program of the second embodiment.

[0009] The fourth aspect of the information processing program is an information processing program of the third aspect, wherein the third coefficient is a composite coefficient such that the ratio of the first coefficient increases as the position of the virtual sound source approaches the position of the virtual listener.

[0010] The information processing program of the fifth embodiment is an information processing program of the third or fourth embodiment in which the third coefficient is equal to the first coefficient when the position of the virtual sound source coincides with the position of the virtual listener.

[0011] The sixth embodiment of the information processing program causes a computer to perform a process in which the information processing program of any one of the first to fifth embodiments amplifies the coefficient of the specific channel among the first coefficients.

[0012] The seventh embodiment of the information processing program causes the computer to perform a process to attenuate the amplified first coefficient corresponding to each of the multiple channels when the sound pressure of the sound output from the speaker using sound data encoded according to the amplified first coefficient exceeds a threshold in the information processing program of the sixth embodiment.

[0013] The information processing program of the eighth embodiment is an information processing program of any one of the first to seventh embodiments, wherein the encoding method is an ambisonic encoding method of order 2 or higher, and the specific plane is a horizontal plane.

[0014] The information processing program of the ninth embodiment is an information processing program of the eighth embodiment in which the specific channels are the r channel and the u channel.

[0015] The information processing program of the tenth embodiment is an information processing program of any one embodiment from the first to the ninth embodiment, wherein the multiple positions on a specific plane facing each other with the listener in between are the position directly in front of the listener and the position directly behind the listener.

[0016] The eleventh aspect of the information processing method is a scene-based encoding method in which sound data is encoded using coefficients for each of a plurality of channels, including a specific channel having the property of strengthening signals in a particular plane direction, wherein when a virtual sound source is located within a set area of ​​a virtual three-dimensional space with respect to a virtual listener, the computer performs a process of encoding sound data to be played back from the virtual sound source using a first coefficient obtained by combining the coefficients of the specific channel corresponding to each of a plurality of positions on a specific plane facing each other with respect to the listener.

[0017] The information processing apparatus according to the twelfth aspect includes a processor, and in an encoding method that is a scene-based encoding method and includes a plurality of channels each having a specific channel having a property of strengthening a signal in a specific plane direction, the processor encodes audio data using coefficients of each of the plurality of channels. When a virtual sound source is located within a set area of a virtual three-dimensional space based on a virtual listener, audio data to be reproduced from the virtual sound source is encoded using a first coefficient obtained by synthesizing coefficients of the specific channels corresponding to a plurality of positions on a specific plane facing each other with the listener interposed therebetween.

[0018] The information processing apparatus according to the thirteenth aspect is the information processing apparatus according to the twelfth aspect, and the plurality of speakers from which sound based on the audio data is output includes speakers arranged so that sound reaches the user from directions that are not symmetric with the user interposed therebetween.

Advantages of the Invention

[0019] <0 motivo>According to the present disclosure, it is possible to suppress the position of the sound image from deviating from the target position.

Brief Description of the Drawings

[0020] [Figure 1] It is a block diagram showing an example of the hardware configuration of an information processing apparatus. <oo00081> [Figure 2] It is a conceptual diagram showing an example of the installation position of speakers. [Figure 3] It is a diagram for explaining a virtual sound source and a virtual listener in a virtual three-dimensional space in a game. [Figure 4] It is a conceptual diagram visually representing the B-format components of second-order ambisonics. [Figure 5] It is a diagram for explaining that the sound image floats above the player's head. [Figure 6] It is a block diagram showing an example of the functional configuration of an information processing apparatus. [Figure 7] It is a diagram for explaining a set area. [Figure 8] This is a diagram for explaining the derivation process of the first coefficient. [Figure 9] This is a flowchart showing an example of sound reproduction processing. [Figure 10] This is a diagram for explaining that the deviation of the position of the sound image is reduced. [Figure 11] This is a diagram for explaining the position of the sound image when the r channel and the u channel are amplified. [Figure 12] This is a diagram for explaining the position of the sound image when there is a lower layer speaker. MODE FOR CARRYING OUT THE INVENTION

[0021] Hereinafter, with reference to the drawings, a mode example for implementing the technology of the present disclosure will be described in detail.

[0022] First, referring to FIG. 1, the hardware configuration of the information processing apparatus 10 according to the present embodiment will be described. The information processing apparatus 10 is, for example, a home game machine, a portable game machine, a business game machine, a smartphone, a tablet terminal, a personal computer, or the like. In the present embodiment, as an example, an example in which a home game machine is applied as the information processing apparatus 10 will be described. The information processing apparatus 10 is an example of a computer.

[0023] As shown in FIG. 1, the information processing apparatus 10 includes a CPU (Central Processing Unit) 11, a memory 12, a storage 13, an external I / F (InterFace) 14, a communication I / F 15, and an input I / F 16. Each of the CPU 11, the memory 12, the storage 13, the external I / F 14, the communication I / F 15, and the input I / F 16 is communicably connected to each other via a bus 20.

[0024] The CPU 11 is a central processing unit that executes various programs and controls various parts. The CPU 11 is an example of a processor. The memory 12 temporarily stores programs or data as a working area. The storage 13 consists of a storage device such as an HDD (Hard Disk Drive), SSD (Solid State Drive), or flash memory, and stores various programs and various data.

[0025] The storage device 13 stores an information processing program 30 for running a predetermined game on the information processing device 10. The CPU 11 reads the information processing program 30 from the storage device 13 and executes it using the memory 12 as a working area. The information processing program 30 is not limited to being stored in the storage device 13, but may also be stored on a recording medium such as an optical disc, USB (Universal Serial Bus) memory, or SD memory card. Furthermore, the information processing program 30 may be downloadable to the information processing device 10 via the communication interface 15. In addition, the information processing program 30 can be provided as a program product. A program product includes all forms of products for providing programs. For example, a program product includes programs provided via a network such as the Internet, and non-temporary computer-readable recording media such as optical discs on which programs are stored.

[0026] The external I / F 14 is an interface for connecting various external devices to the information processing device 10. In this embodiment, multiple speakers 21 and displays 22 are connected to the external I / F 14.

[0027] Multiple speakers 21 output sound based on sound data described later. At least one of the multiple speakers 21 may be integrated with the information processing device 10 or integrated with the display 22. The display 22 is, for example, a liquid crystal display or an organic EL (Electro-Luminescence) display, and displays various types of information. The display 22 may have an integrated touch panel. The display 22 may also be integrated with the information processing device 10.

[0028] Communication I / F15 is an interface for connecting the information processing device 10 to a network. Communication I / F15 may use, for example, wired communication standards such as Ethernet (registered trademark) or FDDI (Fiber Distributed Data Interface), or wireless communication standards such as 4G, 5G, or Wi-Fi (registered trademark).

[0029] The input interface 16 is an interface for connecting the input device 23 to the information processing device 10. The input device 23 is a game controller, mouse, or keyboard, which has operation buttons and directional keys, and is used for various types of input. The user operates the game using the input device 23. Operation information indicating the content of the input operations performed by the user using the input device 23 is stored in the memory 12. The input device 23 may be integrated with the information processing device 10. Alternatively, the input device 23 may be detachable from the information processing device 10. Furthermore, there may be one or more input devices 23. Also, the input device 23 may be a touch panel integrated with the display 22.

[0030] The information processing device 10 according to this embodiment provides a game in which a player, who is a user of the information processing device 10, operates a character (hereinafter referred to as "player character") in a virtual three-dimensional space within the game. A game is a collection of activities and rules for playing or competing. A game is played, for example, by a player using strategies and techniques to achieve a specific objective. Games are played to achieve various objectives, such as competitive objectives like winning a championship, combat objectives like defeating an enemy, educational objectives like learning, and narrative objectives like completing the progression of a scenario. A game may be in a competitive format or a non-competitive format.

[0031] As an example, as shown in Figure 2, in this embodiment, multiple speakers 21 are arranged at different height levels. Player P is the listener who actually hears the sound emitted from the speakers 21. Hereinafter, when distinguishing between speakers 21 arranged at different height levels, the speaker 21 arranged at the first height level will be referred to as speaker 21A, and the speaker 21 arranged at the second height level, which is higher than the first height level, will be referred to as speaker 21B. Speakers 21 arranged at the same height level may differ in height within an acceptable range.

[0032] Speakers 21A are, for example, front speakers and rear speakers, and are installed around the player P in the horizontal direction. Speakers 21A radiate sound horizontally toward the player P. Speakers 21A include speakers that radiate sound from the front of the player P toward the player P, and from the back of the player P toward the player P. That is, speakers 21A include speakers positioned so that sound reaches the player P from the front and back directions, which are symmetrical directions with respect to the player P. Speakers 21A also include speakers that radiate sound from the right of the player P toward the player P, and from the left of the player P toward the player P. That is, speakers positioned so that sound reaches the player P from the left and right directions, which are symmetrical directions with respect to the player P.

[0033] Speaker 21B is, for example, a height speaker and is installed above the player P's head. Speaker 21B radiates sound downwards from above the player P's head. However, the installation position of speaker 21B is not limited to above the player P's head, as long as it radiates sound downwards from above the player P's head. For example, speaker 21B may be installed with its sound-radiating surface facing the ceiling, and the sound is reflected off the ceiling to radiate sound downwards from above the player P's head.

[0034] In this embodiment, the multiple speakers 21 do not include speakers that radiate sound upwards from the feet towards the player P. That is, the multiple speakers 21 include speakers arranged so that sound reaches the player P from directions that are not symmetrical with respect to the player P. The number of speakers 21A and speakers 21B are not limited to the example shown in Figure 2. Also, the placement of the speakers 21 is not limited to the example shown in Figure 2.

[0035] As an example, as shown in Figure 3, a virtual sound source S and a virtual listener L that hears the sound emitted from the sound source S exist in the virtual 3D space of the game. In this embodiment, we will describe an example in which a 3D orthogonal coordinate system is applied as the coordinate system of the virtual 3D space, with the position of listener L as the origin, the front-to-back direction of listener L as the X-axis direction, the left-to-right direction of listener L as the Y-axis direction, and the up-to-down direction of listener L as the Z-axis direction. This is because, in this embodiment, the channel order of the MaxN normalized Fuma (Furse-Malham) method is used in the ambisonics described later, and this is aligned with the axis directions of the X, Y, and Z channels of that channel order. The X, Y, and Z axis directions are not limited to the example shown in Figure 3. For example, the left-to-right direction of listener L may be the X-axis direction, the up-to-down direction of listener L may be the Y-axis direction, and the front-to-back direction of listener L may be the Z-axis direction. Listener L is, for example, the player character. The position of listener L is, for example, the position inside listener L's head. Sound source S is the location where sounds such as sound effects are generated and played in the game. For example, if the sound played from sound source S is the footsteps of an enemy character walking, it is the location where the enemy character's feet touch the ground. The sound data representing the sound played from sound source S is stored in storage 13, for example. In this embodiment, the direction that listener L is facing is defined as the forward direction of listener L. The backward, left-right, and up-down directions of listener L are also directions relative to the forward direction of listener L.

[0036] In the following, the coordinates of the sound source S are denoted as (Px, Py, Pz), the azimuth angle as Az, and the elevation angle as El. The azimuth angle is defined as a counterclockwise angle with the front of the listener L set to 0°. The elevation angle is defined as a positive value for the upper hemisphere with the horizontal plane set to 0°. In this embodiment, the explanation is given using the example of a third-person perspective game in which the listener L is displayed on the display 22, but the disclosed technology is not limited to this embodiment. For example, the game in which the information processing device 10 is provided may be a first-person perspective game in which the listener L is not displayed on the display 22.

[0037] In the environment described above, Ambisonics exists as a spatial audio technology that reproduces a three-dimensional sound field as if a player P in real space were located at the position of a listener L in virtual space, by controlling the sound output from multiple speakers 21. The Ambisonics encoding method is an example of a scene-based encoding method. Furthermore, the Ambisonics encoding method is an encoding method that stores sound data as directional sound data representing a position on the surface of a 360° sphere, regardless of the configuration such as the number and position of the speakers that play the sound data. In the Ambisonics encoding method, sound data is encoded by coefficients for each of the multiple channels. For example, in the second-order Ambisonics encoding method, the coefficients for the nine channels shown in equation (1) below are used.

[0038]

number

[0039] In the second-order ambisonics encoding method used in this embodiment, among the nine channels, the r channel and the u channel are included as examples of specific channels that have the property of strengthening signals in a particular plane direction. The r channel and the u channel have the property of strengthening signals in the direction of the horizontal plane, which is an example of a specific plane. Specifically, as shown in Figure 4 as an example, the r channel has the property of strengthening the signal of the speaker 21 in the horizontal plane direction when the coefficient is a negative value, and the u channel has the property of strengthening the signal of the speaker 21 in the front-to-back direction when the coefficient is a positive value. That is, the specific plane is, for example, the XY plane corresponding to the black area in Figure 4 for the r channel. Also, the position opposite the listener L is, for example, the front-to-back position corresponding to the white area in Figure 4 for the u channel. Figure 4 is a diagram that visually represents the B-format components of second-order ambisonics. In the example in Figure 4, the white spheres represent the signal when the coefficient is a positive value, and the black spheres represent the signal when the coefficient is a negative value.

[0040] As mentioned above, in equation (1), Az is the azimuth angle of the virtual sound source S in the virtual 3D space of the game, and El is the elevation angle of the azimuth angle of the virtual sound source S. When the position of the sound source S and the position of the listener L coincide, that is, when the sound is localized inside the head, the coordinates of the sound source S become (0,0,0), and assuming no directionality, the equation becomes only for the w channel out of the 9 channels, as shown in equation (2) below.

[0041]

number

[0042] As shown in Figure 5, in an environment with speaker 21B, if the position of sound source S and the position of listener L coincide, decoding the sound data encoded using equation (2) and outputting it from speakers 21A and 21B will cause the sound image to appear to float above the player P's head. In other words, in this case, the position of the sound image will be shifted from the target position. Therefore, the information processing device 10 according to this embodiment has a function to suppress the shift of the sound image position from the target position.

[0043] Next, the functional configuration of the information processing device 10 will be described with reference to Figure 6. As shown in Figure 6, the information processing device 10 includes an acquisition unit 40, an output unit 42, and a playback control unit 44. The CPU 11 executes the information processing program 30, thereby enabling the acquisition unit 40, the output unit 42, and the playback control unit 44 to function.

[0044] The acquisition unit 40 acquires sound data representing the sound to be played from the virtual sound source S from the storage 13, according to the progress of the game.

[0045] As shown in Figure 7, the derivation unit 42 determines whether the virtual sound source S is located within a set region R in a virtual three-dimensional space relative to the virtual listener L. Figure 7 shows an example where the sound source S is located within the set region R. In this embodiment, the set region R is described as an area enclosed by the surface of a sphere with radius r centered on the listener L. The radius r is set by converting it to a distance in the virtual three-dimensional space, for example, according to the distance between a typical player P and each speaker 21. In this embodiment, the radius r is normalized to 1.

[0046] The derivation unit 42 derives a first coefficient by combining the coefficients of specific channels corresponding to each of several positions on a specific plane facing each other with respect to the listener L, when the sound source S is located within the setting region R. Specifically, in this case, the derivation unit 42 derives a first coefficient by combining the coefficients of the r channel and the u channel corresponding to each of several positions on the horizontal plane. As an example, as shown in Figure 8, in this embodiment, the case in which the position P1 directly in front of the listener L and the position P2 directly behind the listener L in the XY plane are applied as several positions on the horizontal plane facing each other with respect to the listener L will be explained as an example. In this embodiment, positions P1 and P2 are positions that are equal in distance from the listener L and are positions on the boundary of the setting region R. That is, the coordinates of position P1 are represented as (1,0,0) and the coordinates of position P2 are represented as (-1,0,0).

[0047] The derivation unit 42 derives the coefficients of nine channels corresponding to positions P1 and P2, respectively, based on the coordinates of positions P1 and P2 and equation (1). Then, according to equation (3), the derivation unit 42 combines the derived coefficients corresponding to positions P1 and P2 by adding them together and then dividing by 2. In this way, the derivation unit 42 derives the first coefficient.

[0048]

number

[0049] In the first coefficient synthesized by equation (3), the x, y, and z channels become 0. That is, in a first-order ambisonic encoding scheme, only the w channel remains. In this embodiment, a second-order ambisonic encoding scheme is used, so in the first coefficient, the coefficients of the r and u channels are synthesized and remain.

[0050] The derivation unit 42 amplifies the coefficients of the r channel and u channel, which are examples of specific channels, from the first coefficients according to equation (4) below. F1 and F2 in equation (4) are values ​​for amplifying the coefficients of the r channel and u channel, and are values ​​greater than 1. Note that F1 and F2 may be the same value or different values.

[0051]

number

[0052] The derivation unit 42, when the sound pressure of the sound output from the speaker 21 using the sound data encoded according to the first amplified coefficients of equation (4) exceeds a threshold, attenuates the first amplified coefficients corresponding to each of the multiple channels. Specifically, in this case, the derivation unit 42 attenuates the first coefficients corresponding to each of the nine channels according to the following equation (5).

[0053]

number

[0054] In equation (5), Att is a value that attenuates the coefficient of each of the nine channels, and is a value less than 1.

[0055] The derivation unit 42 derives a second coefficient corresponding to multiple channels based on the position of the sound source S and equation (1), which is a predetermined mathematical formula corresponding to each of the multiple channels. Specifically, the derivation unit 42 derives a second coefficient corresponding to nine channels by substituting the azimuth angle Az and elevation angle El, which correspond to the coordinates representing the position of the sound source S, into equation (1).

[0056] The derivation unit 42 derives a third coefficient by combining the first coefficient and the second coefficient when the sound source S is located within the set region R. Specifically, the derivation unit 42 derives the third coefficient according to the following equation (6). In equation (6), C1 represents the first coefficient, C2 represents the second coefficient, and C3 represents the third coefficient.

[0057]

number

[0058] Here, the derivation unit 42 derives the Ratio in equation (6) according to the following equation (7).

[0059]

number

[0060] In equation (7), d represents the distance between listener L and sound source S, and r represents the radius of the setting region R. In equation (7), the closer the sound source S is to the position of listener L, the smaller d becomes and the smaller the Ratio becomes. That is, the third coefficient is a composite coefficient such that the proportion of the first coefficient increases as the position of sound source S is closer to the position of listener L. Also, when the position of sound source S coincides with the position of listener L, the Ratio becomes 0, and the third coefficient becomes equal to the first coefficient. The derivation unit 42 can determine whether the aforementioned sound source S is located within the setting region R by whether the Ratio is less than 1.

[0061] When the sound source S is located within the set area R, the playback control unit 44 encodes the sound data acquired by the acquisition unit 40 using a third coefficient derived by the derivation unit 42. The playback control unit 44 then decodes the sound data encoded using the third coefficient and controls the output from each speaker 21. As a result, the playback control unit 44 reproduces the sound represented by the sound data from each speaker 21.

[0062] If the sound source S is located outside the set area R, the playback control unit 44 encodes the sound data acquired by the acquisition unit 40 using a second coefficient calculated by the derivation unit 42. The playback control unit 44 then decodes the sound data encoded using the second coefficient and controls the output from each speaker 21. As a result, the playback control unit 44 reproduces the sound represented by the sound data from each speaker 21.

[0063] Next, the operation of the information processing device 10 will be explained with reference to Figure 9. The CPU 11 executes the information processing program 30, thereby performing the sound playback process shown in Figure 9.

[0064] In step S10 of Figure 9, the acquisition unit 40 acquires sound data representing the sound to be played from the virtual sound source S from the storage 13 according to the progress of the game. In step S12, the derivation unit 42 derives a second coefficient corresponding to multiple channels based on the position of the sound source S and equation (1), which is a predetermined mathematical formula corresponding to each of the multiple channels. In step S14, the derivation unit 42 determines whether or not the sound source S is located within the set area R. If this determination is positive, the process proceeds to step S16.

[0065] In step S16, the derivation unit 42 derives a first coefficient by combining the coefficients of the r channel and u channel corresponding to each of several positions on the horizontal plane opposite each other with respect to the listener L. In step S18, the derivation unit 42 amplifies the coefficients of the r channel and u channel from the first coefficient derived in step S16.

[0066] In step S20, the derivation unit 42 uses the sound data encoded according to the first coefficient after amplification in step S18 to determine whether the sound pressure of the sound output from the speaker 21 exceeds a threshold. If this determination is positive, the process moves to step S22. In step S22, the derivation unit 42 attenuates the first coefficient after amplification corresponding to each of the multiple channels. Once the process in step S22 is completed, the process moves to step S24. If the determination in step S20 is negative, step S22 is not executed, and the process moves to step S24.

[0067] In step S24, the derivation unit 42 derives a third coefficient by combining the first coefficient obtained through the above processing with the second coefficient derived in step S12. In step S26, the playback control unit 44 encodes the sound data acquired in step S10 using the third coefficient derived in step S24. When the processing in step S26 is completed, the process moves on to step S30.

[0068] If the determination in step S14 is negative, the process proceeds to step S28. In step S28, the playback control unit 44 encodes the sound data acquired in step S10 using the second coefficient calculated in step S12. When the process in step S28 is completed, the process proceeds to step S30.

[0069] In step S30, the playback control unit 44 decodes the sound data encoded in step S26 or step S28 and controls the output from each speaker 21. When the processing in step S30 is completed, the sound playback process ends.

[0070] As explained above, this embodiment makes it possible to suppress the shift in the sound image position from the target position. As an example, as shown in Figure 10, when the position of the sound source S and the position of the listener L coincide, when the sound data encoded using the position of the sound source S and equation (2), which is an equation for only the w channel, is decoded and output from the speaker 21, the position of the sound image shifts in the height direction. In the example in Figure 10, the darker the color inside the circle representing the speaker 21, the higher the output level. Also, in the example in Figure 10, the position of the sound image is shown as a value calculated as a deviation in the X, Y, and Z axes after normalizing the energy of the output level of the speaker 21 when using the encoding formula of 5th order ambisonics, synthesizing it as a vector, and converting it back to amplitude. Also, in the example in Figure 10, for the sake of simplicity, the azimuth angles of speaker 21A are set to 45°, 90°, and 135°, and the elevation angle of speaker 21B is set to 135°.

[0071] As shown in the dashed rectangle in Figure 10, in this embodiment, the sound data encoded using the first coefficient is decoded and output from the speaker 21, so the shift in the height direction of the sound image position is reduced compared to when only the w channel formula is used. Also, in this embodiment, when the position of the sound source S is directly in front outside the setting area R, the sound is generally localized forward, and the position of the sound image does not shift significantly in the height direction. Also, in this embodiment, when the position of the sound source S is directly to the side outside the setting area R, the sound is generally localized to the side, and the position of the sound image does not shift significantly in the height direction.

[0072] Furthermore, as shown in the area enclosed by the dashed rectangle in Figure 11, in this embodiment, the coefficients of the r channel and u channel are amplified, which further reduces the shift in the height direction of the sound image position. In the example in Figure 11, the sound image position is shown when the coefficients of the r channel and u channel are not amplified (indicated as 1x in Figure 11), and the sound image position is shown when the coefficients of the r channel and u channel are amplified by 2x, 3x, and 4x, respectively.

[0073] Furthermore, as shown in Figure 12, in an environment where, in addition to speaker 21B, speaker 21C is positioned to radiate sound upward from the feet towards player P, so that sound reaches player P from the up and down directions which are symmetrical with respect to player P, it can be seen that the method of the above embodiment does not adversely affect the position of the sound image.

[0074] In the above embodiment, the case in which the derivation unit 42 combines the coefficients of the two channels, the r channel and the u channel, was described, but the disclosed technology is not limited to this embodiment. For example, the derivation unit 42 may combine the coefficients of only the r channel among the r channel and the u channel.

[0075] Furthermore, although the above embodiment describes a case in which the derivation unit 42 amplifies the coefficients of the r channel and u channel among the first coefficients, the disclosed technology is not limited to this embodiment. For example, the derivation unit 42 may be configured to amplify the coefficients of the r channel and u channel among the third coefficients.

[0076] Furthermore, although the above embodiment describes a case in which sound data is encoded using a second-order ambisonics encoding scheme, the disclosed technology is not limited to this embodiment. For example, a configuration in which sound data is encoded using a third-order or higher ambisonics encoding scheme may also be used.

[0077] Furthermore, although the above embodiment describes the case in which a MaxN normalized FuMa channel order is used in ambisonics, the disclosed technology is not limited to this embodiment. For example, an SN3D normalized ACN (Ambisonic Channel Number) channel order may be used in ambisonics. In this case, the r channel of the FuMa system corresponds to the 7th channel of the ACN system, and the u channel of the FuMa system corresponds to the 9th channel of the ACN system.

[0078] Furthermore, the various processes that the CPU 11 reads and executes in the above embodiment may be executed by various processors other than the CPU. Examples of such processors include PLDs (Programmable Logic Devices) such as FPGAs (Field-Programmable Gate Arrays) whose circuit configuration can be changed after manufacturing, and dedicated electrical circuits that are processors with circuit configurations specifically designed to execute specific processes, such as ASICs (Application Specific Integrated Circuits). In addition, the various processes may be executed by one of these various processors, or by a combination of two or more processors of the same or different types (for example, multiple FPGAs, and a combination of a CPU and an FPGA). More specifically, the hardware structure of these various processors is an electrical circuit that combines circuit elements such as semiconductor elements. [Explanation of Symbols]

[0079] 10 Information Processing Devices 11 CPU 30 Information Processing Programs 40 Acquisition Department 42 Derivation part 44 Regeneration Control Unit

Claims

1. In a scene-based encoding method, in which sound data is encoded using coefficients for each of multiple channels, including a specific channel that has the property of strengthening signals in a particular plane direction, When a virtual sound source is located within a set area of ​​a virtual three-dimensional space relative to a virtual listener, the sound data to be played back from the virtual sound source is encoded using a first coefficient obtained by combining the coefficients of the specific channel corresponding to each of multiple positions on a specific plane facing the listener. An information processing program that causes a computer to perform a task.

2. If the virtual sound source is located outside the setting area, the sound data is encoded using the position of the virtual sound source and a second coefficient calculated based on a predetermined formula corresponding to each of the multiple channels. An information processing program according to claim 1 for causing a computer to perform processing.

3. If the virtual sound source is located within the setting area, the sound data is encoded using a third coefficient obtained by combining the first coefficient and the second coefficient. An information processing program according to claim 2 for causing a computer to perform processing.

4. The third coefficient is a composite coefficient such that the proportion of the first coefficient increases as the position of the virtual sound source is closer to the position of the virtual listener. The information processing program according to claim 3.

5. The third coefficient is equal to the first coefficient when the position of the virtual sound source coincides with the position of the virtual listener. The information processing program according to claim 3 or claim 4.

6. Of the first coefficients, the coefficient of the specific channel is amplified. An information processing program according to claim 1 or claim 2 for causing a computer to perform processing.

7. If the sound pressure of the sound output from the speaker using the sound data encoded according to the amplified first coefficient exceeds a threshold, the amplified first coefficient corresponding to each of the multiple channels is attenuated. An information processing program according to claim 6 for causing a computer to perform processing.

8. The encoding method is an ambisonic encoding method of order 2 or higher. The aforementioned specific plane is a horizontal plane. The information processing program according to claim 1 or claim 2.

9. The aforementioned specific channels are the r channel and the u channel. The information processing program according to claim 8.

10. The multiple positions on a specific plane facing each other with the listener in between are the position directly in front of the listener and the position directly behind the listener. The information processing program according to claim 1 or claim 2.

11. In a scene-based encoding method, in which sound data is encoded using coefficients for each of multiple channels, including a specific channel that has the property of strengthening signals in a particular plane direction, When a virtual sound source is located within a set area of ​​a virtual three-dimensional space relative to a virtual listener, the sound data to be played back from the virtual sound source is encoded using a first coefficient obtained by combining the coefficients of the specific channel corresponding to each of multiple positions on a specific plane facing the listener. An information processing method in which a computer performs the processing.

12. The processor comprises, In a scene-based encoding method, in which sound data is encoded using coefficients for each of multiple channels, including a specific channel that has the property of strengthening signals in a particular plane direction, When a virtual sound source is located within a set area of ​​a virtual three-dimensional space relative to a virtual listener, the sound data to be played back from the virtual sound source is encoded using a first coefficient obtained by combining the coefficients of the specific channel corresponding to each of multiple positions on a specific plane facing the listener. Information processing device.

13. Multiple speakers that output sound based on the aforementioned sound data include speakers arranged so that the sound reaches the user from a direction that is not symmetrical with respect to the user. The information processing apparatus according to claim 12.