Information processing device and method, playback device and method, and program
The described technology addresses the time-consuming nature of 3D Audio content creation by determining gain correction values based on listener direction and auditory characteristics, enabling faster and easier production of high-quality 3D Audio content.
Patent Information
- Application Number
- JP2024104305
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-04-11
- Filing Date
- 2024-06-27
- Publication Date
- 2025-10-22
- Estimated Expiration
- 2040-03-27
AI Technical Summary
The production of high-quality 3D Audio content is time-consuming due to the high dimensionality of object position information and the need for precise gain correction, and there is a scarcity of such content and creators, necessitating easier gain correction methods.
An information processing device and method that determines gain correction values based on the direction of audio objects relative to the listener, using auditory characteristics to adjust gain values for 3D Audio content creation, including interpolation for missing data points and user interface tools for manual adjustment.
Facilitates the creation of high-quality 3D Audio content more efficiently by simplifying gain correction processes, reducing time and effort in content production.
Smart Images

Figure 0007758108000004 
Figure 0007758108000005 
Figure 0007758108000006
Abstract
Description
[Technical Field]
[0001] The present technology relates to an information processing device and method, a playback device and method, and a program, and particularly to an information processing device and method, a playback device and method, and a program that enable gain correction to be performed more easily. [Background technology]
[0002] Conventionally, the Moving Picture Experts Group (MPEG)-H 3D Audio standard is known (see, for example, Non-Patent Document 1 and Non-Patent Document 2).
[0003] 3D Audio, which is handled by standards such as MPEG-H 3D Audio, can reproduce the direction, distance, and spread of sound in three dimensions, making it possible to reproduce audio with a more realistic feel than conventional stereo playback. [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] ISO / IEC 23008-3, MPEG-H 3D Audio [Non-patent document 2] ISO / IEC 23008-3:2015 / AMENDMENT3, MPEG-H 3D Audio Phase 2 Summary of the Invention [Problem to be solved by the invention]
[0005] However, with 3D Audio, the time cost of producing content (3D Audio content) is high.
[0006] For example, in 3D Audio, the number of dimensions of object position information, i.e., the position information of the sound source, is higher than in stereo (3D Audio is three-dimensional, while stereo is two-dimensional).As a result, with 3D Audio, the time cost is high, especially when determining the parameters that make up the metadata for each object, such as the horizontal and vertical angles that indicate the object's position, distance, and gain for the object.
[0007] Furthermore, compared to stereo content, there are far fewer 3D Audio content available, both in terms of content and creators, which means there is currently very little high-quality 3D Audio content.
[0008] On the other hand, as a characteristic of hearing, the perceived loudness of a sound differs depending on the direction from which the sound comes. That is, even if the sound is from the same object, the perceived loudness of the sound will differ depending on whether the object is in front of or to the side of the listener, or above or below the listener. Therefore, gain correction that takes such hearing characteristics into account is necessary.
[0009] For these reasons, it is desirable to make it easier to perform gain correction, thereby enabling the creation of 3D Audio content of sufficient quality in a short amount of time.
[0010] The present technology has been made in view of such circumstances, and makes it possible to perform gain correction more easily. [Means for solving the problem]
[0011] An information processing device according to a first aspect of the present technology includes a gain correction value determination unit that determines a correction value for a gain value for performing gain correction on an audio signal of an audio object according to a direction of the audio object as seen by a listener.
[0012] An information processing method or program according to a first aspect of the present technology includes determining a gain correction value for gain-correcting an audio signal of an audio object according to a direction of the audio object as seen by a listener.
[0013] In a first aspect of the present technology, a gain correction value for performing gain correction on an audio signal of an audio object is determined according to a direction of the audio object as seen by a listener.
[0014] A playback device according to a second aspect of the present technology includes a gain correction unit that determines a gain correction value for gain correcting an audio signal of an audio object based on position information indicating the position of the audio object, the correction value corresponding to the direction of the audio object as seen by a listener, and performs the gain correction of the audio signal based on the gain value corrected by the correction value; and a renderer processing unit that performs rendering processing based on the audio signal obtained by the gain correction, and generates playback signals of multiple channels for playing back the sound of the audio object.
[0015] A playback method or program according to a second aspect of the present technology includes the steps of: determining a gain correction value for gain correcting an audio signal of an audio object based on position information indicating the position of the audio object, the correction value corresponding to the direction of the audio object as seen by a listener; performing the gain correction of the audio signal based on the gain value corrected by the correction value; performing a rendering process based on the audio signal obtained by the gain correction; and generating playback signals of multiple channels for reproducing the sound of the audio object.
[0016] In a second aspect of the present technology, a gain correction value for gain correcting an audio signal of an audio object is determined based on position information indicating the position of the audio object, the correction value corresponding to the direction of the audio object as seen by a listener, the gain correction of the audio signal is performed based on the gain value corrected by the correction value, a rendering process is performed based on the audio signal obtained by the gain correction, and playback signals of multiple channels for reproducing the sound of the audio object are generated. [Brief explanation of the drawings]
[0017] [Figure 1] FIG. 1 is a diagram illustrating auditory characteristics with respect to the direction from which a sound comes. [Figure 2] FIG. 1 is a diagram illustrating auditory characteristics with respect to the direction from which a sound comes. [Figure 3] FIG. 1 is a diagram illustrating auditory characteristics with respect to the direction from which a sound comes. [Figure 4] FIG. 1 illustrates an example of the configuration of an information processing device. [Figure 5] FIG. 10 is a diagram illustrating an example of an auditory characteristic table. [Figure 6] FIG. 10 is a diagram illustrating an example of an auditory characteristic table. [Figure 7] 10 is a flowchart illustrating a gain value determination process. [Figure 8] FIG. 10 is a diagram showing an example of a display screen of a content production tool. [Figure 9] FIG. 10 is a diagram showing an example of a display screen of a content production tool. [Figure 10] FIG. 10 is a diagram showing an example of a display screen of a content production tool. [Figure 11] FIG. 10 is a diagram showing an example of a display screen of a content production tool. [Figure 12] FIG. 1 illustrates an example of the configuration of an information processing device. [Figure 13] 10 is a flowchart illustrating a table generation process. [Figure 14] FIG. 1 illustrates an example of the configuration of a voice processing device. [Figure 15] 10 is a flowchart illustrating a reproduction signal generation process. [Figure 16] FIG. 10 is a diagram illustrating an example of an auditory characteristic table. [Figure 17] FIG. 10 is a diagram illustrating an example of syntax of gain auditory characteristic information. [Figure 18] FIG. 1 illustrates an example of the configuration of a voice processing device. [Figure 19] FIG. 1 illustrates an example of the configuration of a computer. DETAILED DESCRIPTION OF THE INVENTION
[0018] Hereinafter, embodiments to which the present technology is applied will be described with reference to the drawings.
[0019] First Embodiment About this technology This technology makes it easier to perform gain correction by determining the gain correction value according to the direction of the object as seen by the listener, thereby making it possible to create 3D Audio content of sufficiently high quality more easily, i.e., in a shorter amount of time.
[0020] In particular, the present technology has the following features (F1) to (F5).
[0021] Feature (F1): Determine the gain correction value of the object according to the three-dimensional auditory characteristics for the sound image localization position. Feature (F2): When the auditory characteristics are given by a table or the like, the gain correction value for a localization position without data is calculated by interpolation or the like based on the gain correction values of adjacent positions. Feature (F3): In automatic mixing, gain information is determined from separately determined position information. Feature (F4): Provides a user interface to set and adjust gain correction values for object positions Feature (F5): Apply gain correction values according to the 3D auditory characteristics when the object position changes relative to the listening position.
[0022] First, the determination of gain parameters based on the three-dimensional auditory characteristics of humans will be explained.
[0023] Figure 1 shows the amount of gain correction applied to pink noise so that the perceived loudness of a sound when played from different directions is the same, based on the perceived loudness of the sound when played directly in front of the listener. In other words, Figure 1 shows the human hearing characteristics in the horizontal direction.
[0024] In FIG. 1, the vertical axis indicates the gain correction amount, and the horizontal axis indicates the azimuth value (horizontal angle), which is the angle in the horizontal direction indicating the position of the sound source as seen by the listener.
[0025] For example, the Azimuth value indicating the direction directly ahead from the listener's perspective is 0 degrees, the Azimuth value indicating the direction directly to the side from the listener's perspective, i.e., to the side, is ±90 degrees, and the Azimuth value indicating the direction directly behind the listener, i.e., directly behind, is 180 degrees. In particular, the left direction from the listener's perspective is the positive Azimuth value.
[0026] In addition, in Figure 1, the vertical position when the pink noise is played is the same height as the listener. In other words, if the vertical angle indicating the vertical (elevation) position of the sound source as seen by the listener is defined as the Elevation value, Figure 1 shows an example where the Elevation value is 0 degrees. Note that the upward direction as seen by the listener is the positive direction for the Elevation value.
[0027] This example shows the average gain correction amount for each Azimuth value obtained from the results of an experiment conducted on multiple listeners. In particular, the range indicated by the dotted line for each Azimuth value indicates the 95% confidence interval.
[0028] For example, when playing pink noise to the side (Azimuth value = ±90 degrees, Elevation value = 0 degrees), by lowering the gain slightly, the listener will perceive the sound as being at the same volume as when the pink noise is played directly in front of them.
[0029] Also, for example, when playing pink noise from behind (Azimuth value = 180 degrees, Elevation value = 0 degrees), by increasing the gain slightly, the listener will perceive the sound as being at the same volume as when the pink noise is played directly in front of them.
[0030] That is, for a certain object sound source, if the localization position of the object sound source is to the side of the listener, the gain of the sound of the object sound source is slightly lowered, and if the localization position of the object sound source is behind the listener, the gain of the sound of the object sound source is slightly increased, so that the listener feels as if they are hearing the sound at the same volume.
[0031] Furthermore, as shown in Figures 2 and 3, even if the Azimuth value is the same, if the Elevation value changes, the way the listener hears the sound also changes.
[0032] 2 and 3, the vertical axis represents the gain correction amount, and the horizontal axis represents the azimuth value (horizontal angle) that indicates the sound source position as seen by the listener. In addition, in Figures 2 and 3, the range indicated by the dotted line for each azimuth value represents the 95% confidence interval.
[0033] FIG. 2 shows the gain correction amount for each azimuth value when the elevation value is 30 degrees.
[0034] Figure 2 shows that when the sound source is located higher than the listener, the sound will be perceived as quieter when the sound source is in front of, behind, or diagonally behind the listener, and slightly louder when the sound source is diagonally in front of the listener.
[0035] Similarly, FIG. 3 shows the gain correction amount for each Azimuth value when the Elevation value is −30 degrees.
[0036] From Figure 3, we can see that when the sound source is located lower than the listener, the sound will be perceived as louder when the sound source is directly in front of or diagonally in front of the listener, and as quieter when the sound source is directly behind or diagonally behind the listener.
[0037] From the above-described auditory characteristics with respect to the direction from which the sound comes, it can be seen that appropriate gain correction can be performed more easily by determining the amount of gain correction for the object sound source based on position information indicating the position of the object sound source and the auditory characteristics of the listener.
[0038] <Configuration example of information processing device> FIG. 4 is a diagram showing an example of the configuration of an embodiment of an information processing device to which the present technology is applied.
[0039] The information processing device 11 shown in FIG. 4 functions as a gain determination device that determines a gain value for gain correction of an audio signal for reproducing the sound of an audio object (hereinafter simply referred to as an object) that constitutes 3D Audio content.
[0040] Such an information processing device 11 is provided in, for example, an editing device that mixes audio signals that make up 3D Audio content.
[0041] The information processing device 11 includes a gain correction value determination unit 21 and an auditory characteristic table storage unit 22 .
[0042] The gain correction value determination unit 21 is supplied with position information and initial gain values as metadata of objects that make up the 3D Audio content.
[0043] Here, the position information of an object is information that indicates the position of the object as seen from a reference position in three-dimensional space, and here the position information consists of an Azimuth value, an Elevation value, and a Radius value. In this example, the position of the listener is the reference position.
[0044] The azimuth value and elevation value are angles indicating the horizontal and vertical positions of the object as seen by a listener (user) at a reference position, and these azimuth values and elevation values are the same as those in Figures 1 to 3.
[0045] The Radius value is the distance (radius) from the listener at the reference position in three-dimensional space to the object.
[0046] It can be said that the position information consisting of such Azimuth value, Elevation value, and Radius value indicates the localized position of the sound image of the sound of the object.
[0047] Furthermore, the initial gain value included in the metadata supplied to the gain correction value determination unit 21 is a gain value for gain correction of the audio signal of the object, i.e., an initial value of the gain information, and this initial gain value is determined, for example, by the creator of the 3D Audio content, etc. For simplicity of explanation, the initial gain value is assumed to be 1.0.
[0048] The gain correction value determination unit 21 determines a gain correction value indicating the amount of gain correction for correcting the initial gain value of the object based on the position information as metadata supplied and the auditory characteristic table stored in the auditory characteristic table storage unit 22.
[0049] In addition, the gain correction value determination unit 21 corrects the supplied initial gain value based on the determined gain correction value, and uses the resulting gain value as information indicating the final gain correction amount for gain correcting the audio signal of the object.
[0050] In other words, the gain correction value determination unit 21 determines the gain value of the audio signal by determining a gain correction value according to the direction of the object (the direction from which the sound comes) as seen by the listener, which is indicated by the position information. The gain value determined in this way and the supplied position information are output to a subsequent stage as the final metadata of the object.
[0051] The auditory characteristic table storage unit 22 stores an auditory characteristic table, and supplies the gain correction value indicated by the auditory characteristic table to the gain correction value determination unit 21 as needed.
[0052] Here, the auditory characteristics table is a table in which the direction from which sound comes from an object that is a sound source to a listener, that is, the direction of the sound source as seen by the listener, is associated with a gain correction value corresponding to that direction.
[0053] More specifically, the auditory characteristics table is a table in which the relative positional relationship between the sound source and the listener is associated with a gain correction value corresponding to that positional relationship.
[0054] The gain correction values indicated by the auditory characteristics table are determined according to the human auditory characteristics relative to the direction from which the sound comes, as shown in Figures 1 to 3, and are gain correction amounts that keep the auditory loudness of the sound constant regardless of the direction from which the sound comes.
[0055] In other words, if the audio signal of an object is gain-corrected using a gain value obtained by correcting the initial gain value using the gain correction value indicated in the auditory characteristics table, the sound of the same object will be heard at the same volume regardless of the object's position.
[0056] FIG. 5 shows an example of the auditory characteristics table.
[0057] In the example shown in FIG. 5, the gain correction value is associated with the position of the object determined by the Azimuth value, Elevation value, and Radius value, that is, the direction of the object.
[0058] In particular, in this example, all Elevation and Radius values are 0 and 1.0, which assumes that the object is vertically positioned at the same height as the listener and that the distance from the listener to the object remains constant.
[0059] In the example of Figure 5, when the object that is the sound source is behind the listener, for example when the Azimuth value is 180 degrees, the gain correction value is larger than when the object is in front of the listener, for example when the Azimuth value is 0 degrees or 30 degrees.
[0060] On the other hand, when the object that is the sound source is to the side of the listener, for example when the Azimuth value is 90 degrees, the gain correction value is smaller than when the object is in front of the listener.
[0061] Furthermore, a specific example of correction of the initial gain value by the gain correction value determination unit 21 when the auditory characteristic table storage unit 22 stores the auditory characteristic table shown in FIG. 5 will be described.
[0062] For example, if the Azimuth value, Elevation value, and Radius value indicating the position of the object are 90 degrees, 0 degrees, and 1.0 m, respectively, the gain correction value corresponding to the object position is −0.52 dB from FIG.
[0063] Therefore, the gain correction value determination unit 21 performs the calculation of the following equation (1) based on the gain correction value "-0.52 dB" read out from the auditory characteristics table and the initial gain value "1.0", and obtains the gain value "0.94".
[0064]
number
[0065] Similarly, if the Azimuth value, Elevation value, and Radius value indicating the object position are -150 degrees, 0 degrees, and 1.0 m, respectively, the gain correction value corresponding to the object position is 0.51 dB from FIG.
[0066] Therefore, the gain correction value determination unit 21 performs the calculation of the following equation (2) based on the gain correction value "0.51 dB" read out from the auditory characteristics table and the initial gain value "1.0", and obtains the gain value "1.06".
[0067]
number
[0068] 5, an example has been described in which gain correction values determined based on two-dimensional auditory characteristics in which only the horizontal direction is taken into consideration are used. That is, an example has been described in which an auditory characteristics table generated based on two-dimensional auditory characteristics (hereinafter also referred to as a two-dimensional auditory characteristics table) is used.
[0069] However, the initial gain value may be corrected using a gain correction value determined based on three-dimensional auditory characteristics that take into account not only horizontal but also vertical characteristics.
[0070] In such a case, for example, the auditory characteristics table shown in FIG. 6 can be used.
[0071] In the example shown in FIG. 6, the gain correction value is associated with the position of the object determined by the Azimuth value, the Elevation value, and the Radius value, that is, the direction of the object.
[0072] In particular, in this example, the Radius value is set to 1.0 for all combinations of Azimuth and Elevation values.
[0073] Hereinafter, the auditory characteristic table generated based on the three-dimensional auditory characteristic for the sound arrival direction as shown in FIG. 6 will also be referred to as a three-dimensional auditory characteristic table.
[0074] Here, a specific example of correction of the initial gain value by the gain correction value determination unit 21 when the auditory characteristic table storage unit 22 stores the auditory characteristic table shown in FIG. 6 will be described.
[0075] For example, if the Azimuth value, Elevation value, and Radius value indicating the position of the object are 60 degrees, 30 degrees, and 1.0 m, respectively, then from FIG. 6, the gain correction value corresponding to the object position is −0.07 dB.
[0076] Therefore, the gain correction value determination unit 21 performs the calculation of the following equation (3) based on the gain correction value "-0.07 dB" read out from the auditory characteristics table and the initial gain value "1.0", and obtains the gain value "0.99".
[0077]
number
[0078] In the specific example of gain value calculation described above, gain correction values based on auditory characteristics determined for the position (direction) of an object are prepared in advance. That is, the example has been described in which gain correction values corresponding to object position information are stored in the auditory characteristics table.
[0079] However, the position of the object is not necessarily at a position in the auditory characteristics table where the corresponding gain correction value is stored.
[0080] Specifically, for example, it is assumed that the auditory characteristic table shown in FIG. 6 is stored in the auditory characteristic table storage unit 22, and the Azimuth value, Elevation value, and Radius value as position information are −120 degrees, 15 degrees, and 1.0 m.
[0081] In this case, the auditory characteristics table of FIG. 6 does not store gain correction values corresponding to the Azimuth value "-120", the Elevation value "15", and the Radius value "1.0".
[0082] Therefore, if the auditory characteristics table does not contain a gain correction value corresponding to the position indicated by the position information, the gain correction value determination unit 21 may calculate the gain correction value for the desired position by interpolation processing or the like using data (gain correction values) for multiple positions adjacent to the position indicated by the position information for which corresponding gain correction values exist.
[0083] In other words, if a gain correction value corresponding to the direction (position) of an object as seen by the listener is not stored in the auditory characteristics table, that gain correction value may be obtained by interpolation processing or the like based on gain correction values corresponding to other directions of the object as seen by the listener.
[0084] For example, one of the methods for interpolating gain correction values is VBAP (Vector Base Amplitude Panning).
[0085] VBAP is used to determine, for each object, gain values for multiple speakers in a playback environment from the object's metadata.
[0086] Here, by replacing the multiple speakers in the playback environment with multiple gain correction values, it is possible to calculate the gain correction value at a desired position.
[0087] Specifically, meshes are divided into a plurality of positions in a three-dimensional space for which gain correction values are prepared. For example, if gain correction values are prepared for three positions in a three-dimensional space, a triangular area with the three positions as vertices is defined as one mesh.
[0088] When the three-dimensional space is divided into a plurality of meshes in this manner, a desired position for which a gain correction value is to be obtained is set as a focus position, and a mesh that includes the focus position is identified.
[0089] In addition, coefficients are calculated to be multiplied by the position vectors indicating the three vertex positions when the position vector indicating the target position is expressed by multiplying and adding the position vectors indicating the three vertex positions that make up the identified mesh.
[0090] Then, each of the three coefficients thus obtained is multiplied by the gain correction value of each of the three vertex positions of the mesh containing the position of interest, and the sum of the gain correction values multiplied by the coefficients is calculated as the gain correction value of the position of interest.
[0091] Specifically, it is assumed that the position vectors indicating the three vertex positions of the mesh containing the target position are P1 to P3, and the gain correction values of these vertex positions are G1 to G3.
[0092] In this case, it is assumed that the position vector indicating the position of interest is expressed as g1P1+g2P2+g3P3, and the gain correction value of the position of interest is g1G1+g2G2+g3G3.
[0093] The method of interpolating the gain correction value is not limited to interpolation using VBAP, and any other method may be used.
[0094] For example, among positions where gain correction values exist in the auditory characteristics table, the average value of gain correction values at N positions (for example, N=5) near the target position may be used as the gain correction value of the target position.
[0095] Furthermore, for example, among positions in the auditory characteristics table where gain correction values exist, the gain correction value at the position closest to the target position may be used as the gain correction value at the target position.
[0096] Furthermore, although the example in which the gain correction value is calculated in decibel values has been described here, the gain correction value may be calculated in linear values. In such a case, even when the gain correction value is calculated in linear values by interpolation using VBAP, the gain correction value at any position can be obtained by the same calculation as in the case of the decibel values described above.
[0097] In addition, this technology can also be applied to determining position information as metadata for an object, that is, Azimuth value, Elevation value, and Radius value, based on the type, priority, sound pressure, sound pitch, etc. of the object.
[0098] In this case, the gain correction value is determined based on the position information determined based on the type and priority of the object, for example, and a three-dimensional auditory characteristic table prepared in advance.
[0099] <Description of Gain Value Determination Process> Next, a description will be given of the operation of the information processing device 11. That is, the gain value determination process performed by the information processing device 11 will be described below with reference to the flowchart of FIG.
[0100] In step S11, the gain correction value determination unit 21 acquires metadata from an external source.
[0101] That is, the gain correction value determination unit 21 acquires, as metadata, position information including an Azimuth value, an Elevation value, and a Radius value, and an initial gain value.
[0102] In step S12, the gain correction value determination unit 21 determines a gain correction value based on the position information acquired in step S11 and the auditory characteristic table held in the auditory characteristic table holding unit 22.
[0103] That is, the gain correction value determination unit 21 reads out the gain correction values associated with the Azimuth value, Elevation value, and Radius value that constitute the acquired position information from the auditory characteristics table, and sets the read-out gain correction value as the determined gain correction value.
[0104] In step S13, the gain correction value determination unit 21 determines a gain value based on the initial gain value acquired in step S11 and the gain correction value determined in step S12.
[0105] That is, the gain correction value determination unit 21 performs a calculation similar to that of equation (1) based on the initial gain value and the gain correction value, and corrects the initial gain value with the gain correction value, thereby obtaining the gain value.
[0106] Once the gain value is determined in this manner, the gain correction value determination unit 21 outputs the determined gain value to a subsequent stage, and the gain value determination process ends. The output gain value is used for gain correction (gain adjustment) of the audio signal in the subsequent stage.
[0107] In this manner, the information processing device 11 determines the gain correction value using the auditory characteristics table, and determines the gain value by correcting the initial gain value with the gain correction value.
[0108] This makes it possible to perform gain correction more easily, which in turn makes it possible to create 3D Audio content of sufficiently high quality more easily, i.e., in a shorter time.
[0109] Second Embodiment About the user interface Furthermore, according to the present technology, it is possible to provide a user interface for setting and adjusting the gain correction value described above.
[0110] For example, this technology can be applied to a 3D audio content creation tool that determines the position of objects based on user input or automatically.
[0111] Specifically, in a 3D audio content production tool, for example, a user interface (display screen) shown in Figure 8 can be used to set and adjust a gain correction value (gain value) based on the auditory characteristics of the listener in relation to the direction of an object.
[0112] In the example shown in FIG. 8, the display screen of the 3D audio content production tool is provided with a pull-down box BX11 for selecting a desired auditory characteristic from a plurality of different auditory characteristics that are pre-set.
[0113] In this example, multiple two-dimensional hearing characteristics are prepared in advance, such as male hearing characteristics, female hearing characteristics, and the user's personal hearing characteristics, and the user can select the desired hearing characteristic by operating the pull-down box BX11.
[0114] When the user selects an auditory characteristic, the gain correction value for each Azimuth value corresponding to the auditory characteristic selected by the user is displayed in the gain correction value display area R11 provided below the pull-down box BX11 in the figure.
[0115] In particular, in the gain correction value display region R11, the vertical axis indicates the gain correction value, and the horizontal axis indicates the azimuth value.
[0116] Furthermore, curve L11 indicates the gain correction value for negative Azimuth values, that is, for each Azimuth value in the right direction as seen from the listener, and curve L12 indicates the gain correction value for each Azimuth value in the left direction as seen from the listener.
[0117] By looking at such a gain correction value display area R11, the user can intuitively and instantly grasp the gain correction value for each Azimuth value.
[0118] Furthermore, below the gain correction value display area R11 in the figure, there is provided a slider display area R12 displaying a slider or the like for adjusting the gain correction value displayed in the gain correction value display area R11.
[0119] In the slider display area R12, for each Azimuth value for which the user can adjust the gain correction value, a number indicating the Azimuth value, a scale indicating the gain correction value, and a slider for adjusting the gain correction value are displayed.
[0120] For example, the slider SD11 is for adjusting the gain correction value when the Azimuth value is 30 degrees, and the user can specify a desired value as the adjusted gain correction value by moving the slider SD11 up and down.
[0121] When the gain correction value is adjusted by the slider SD11, the display in the gain correction value display region R11 is updated in accordance with the adjustment. That is, here, the curve L12 changes in accordance with the operation of the slider SD11.
[0122] In this way, in the example shown in FIG. 8, it is possible to independently adjust the gain correction values in each direction on the right side as seen by the listener and the gain correction values in each direction on the left side as seen by the listener.
[0123] In particular, in this example, by selecting any one of a plurality of pre-prepared auditory characteristics, it is possible to specify a gain correction value corresponding to the desired auditory characteristic, i.e., an auditory characteristic table, and then by operating the slider, it is possible to further adjust the gain correction value corresponding to the selected auditory characteristic.
[0124] For example, since the pre-prepared hearing characteristics are average, the user can adjust the gain correction value by operating a slider to suit the user's individual hearing characteristics. Also, by operating the slider to adjust the gain correction value, it becomes possible to make adjustments according to the user's intentions, such as by applying a larger gain correction to objects in the background to emphasize them.
[0125] In this way, the gain correction value for each Azimuth value is set and adjusted, and when, for example, a save button (not shown) is operated, a two-dimensional auditory characteristic table is generated in which the gain correction value displayed in the gain correction value display area R11 corresponds to each Azimuth value.
[0126] 8, an example has been described in which the gain correction values are different between the right and left sides as seen by the listener, that is, the gain correction values are asymmetric. However, the gain correction values may be symmetric.
[0127] In such a case, the gain correction value is set or adjusted as shown in Fig. 9. In Fig. 9, the same reference numerals are used to designate parts corresponding to those in Fig. 8, and the description thereof will be omitted as appropriate.
[0128] FIG. 9 shows the display screen of a 3D audio content production tool, and in this example, a pull-down box BX11, a gain correction value display area R21, and a slider display area R22 are displayed on the display screen.
[0129] The gain correction value display area R21 displays the gain correction value for each Azimuth value, similar to the gain correction value display area R11 in Figure 8, but in this case, since the gain correction value is common to both the left and right directions, only one curve indicating the gain correction value is displayed.
[0130] For example, the average value of the gain correction values in the left and right directions can be set as a common gain correction value for both the left and right. In this case, for example, the average value of the gain correction values for the Azimuth values of 90 degrees and −90 degrees in the example of Fig. 8 is set as a common gain correction value for the Azimuth values of ±90 degrees in the example of Fig. 9.
[0131] In addition, the slider display area R22 displays a slider or the like for adjusting the gain correction value displayed in the gain correction value display area R21.
[0132] For example, in this example, the user can adjust a common gain correction value for Azimuth values of ±30 degrees by moving the slider SD21 up or down.
[0133] Furthermore, for example, the gain correction value at each Azimuth value may be adjusted for each Elevation value as shown in Fig. 10. Note that in Fig. 10, the same reference numerals are used to designate parts corresponding to those in Fig. 8, and descriptions thereof will be omitted where appropriate.
[0134] Figure 10 shows the display screen of a 3D audio content production tool, and in this example, the display screen displays a pull-down box BX11, gain correction value display areas R31 to R33, and slider display areas R34 to R36.
[0135] In the example shown in FIG. 10, the gain correction values are symmetrical, similar to the example shown in FIG.
[0136] The gain correction value display area R31 displays the gain correction values for each Azimuth value when the Elevation value is 30 degrees, and the user can adjust these gain correction values by operating the sliders displayed in the slider display area R34.
[0137] Similarly, the gain correction value display area R32 displays the gain correction values for each Azimuth value when the Elevation value is 0 degrees, and the user can adjust these gain correction values by operating the sliders displayed in the slider display area R35.
[0138] In addition, the gain correction value display area R33 displays the gain correction values for each Azimuth value when the Elevation value is -30 degrees, and the user can adjust these gain correction values by operating the sliders displayed in the slider display area R36.
[0139] In this way, the gain correction value for each Azimuth value is set and adjusted, and when, for example, a save button (not shown) is operated, a three-dimensional auditory characteristic table is generated in which the gain correction value is associated with the Elevation value and the Azimuth value.
[0140] Furthermore, as another example of the display screen of a 3D audio content production tool, a radar chart-type gain correction value display area may be provided as shown in Fig. 11. Note that in Fig. 11, parts corresponding to those in Fig. 10 are given the same reference numerals, and their explanation will be omitted as appropriate.
[0141] 11, the display screen displays a pull-down box BX11, gain correction value display areas R41 to R43, and slider display areas R34 to R36. In this example, the gain correction values are symmetrical, similar to the example shown in FIG.
[0142] The gain correction value display area R41 displays the gain correction values for each Azimuth value when the Elevation value is 30 degrees, and the user can adjust these gain correction values by operating the sliders displayed in the slider display area R34.
[0143] In particular, in the gain correction value display area R41, each item on the radar chart is an Azimuth value, so the user can instantly grasp not only each direction (Azimuth value) and the gain correction value for that direction, but also the relative difference in gain correction value between each direction.
[0144] Like the gain correction value display region R41, the gain correction value display region R42 displays the gain correction value for each azimuth value when the elevation value is 0 degrees. Moreover, the gain correction value display region R43 displays the gain correction value for each azimuth value when the elevation value is -30 degrees.
[0145] <Configuration example of information processing device> Next, an information processing device that generates an auditory characteristic table using the 3D audio content production tool described with reference to FIG. 8 and the like will be described.
[0146] Such an information processing device is configured, for example, as shown in FIG.
[0147] An information processing device 51 shown in FIG. 12 implements a content production tool, and causes a display device 52 to display the display screen of the content production tool.
[0148] The information processing device 51 includes an input unit 61 , an auditory characteristic table generating unit 62 , an auditory characteristic table holding unit 63 , and a display control unit 64 .
[0149] The input unit 61 is made up of, for example, a mouse, a keyboard, a switch, a button, a touch panel, etc., and supplies an input signal according to a user's operation to the auditory characteristic table generation unit 62 .
[0150] The auditory characteristic table generation unit 62 generates a new auditory characteristic table based on the input signal supplied from the input unit 61 and the auditory characteristic table of preset auditory characteristics stored in the auditory characteristic table storage unit 63, and supplies it to the auditory characteristic table storage unit 63.
[0151] Furthermore, the auditory characteristic table generating unit 62 instructs the display control unit 64 to update the display on the display screen of the display device 52 as appropriate when generating the auditory characteristic table.
[0152] The auditory characteristic table holding unit 63 holds an auditory characteristic table of pre-set auditory characteristics, supplies the auditory characteristic table to the auditory characteristic table generation unit 62 as appropriate, and holds the auditory characteristic table supplied from the auditory characteristic table generation unit 62.
[0153] The display control unit 64 controls the display of the display screen by the display device 52 in accordance with instructions from the auditory characteristic table generating unit 62 .
[0154] The input unit 61, the auditory characteristic table generating unit 62, and the display control unit 64 shown in FIG. 12 may be provided in the information processing device 11 shown in FIG.
[0155] <Explanation of table generation process> Next, the operation of the information processing device 51 will be described.
[0156] That is, the table generation process performed by the information processing device 51 will be described below with reference to the flowchart of FIG.
[0157] In step S41, the display control unit 64, in response to an instruction from the auditory characteristic table generating unit 62, causes the display device 52 to display a display screen of the content production tool.
[0158] Specifically, for example, the display control unit 64 causes the display device 52 to display the display screens shown in FIG. 8, FIG. 9, FIG. 10, FIG. 11, and the like.
[0159] At this time, for example, if the user operates the input unit 61 to select a preset auditory characteristic, the auditory characteristic table generation unit 62 reads out an auditory characteristic table corresponding to the auditory characteristic selected by the user from the auditory characteristic table storage unit 63 in accordance with the input signal supplied from the input unit 61.
[0160] Then, the auditory characteristic table generating unit 62 instructs the display control unit 64 to display a gain correction value display area so that the gain correction value for each Azimuth value indicated by the read auditory characteristic table is displayed on the display device 52. In response to the instruction from the auditory characteristic table generating unit 62, the display control unit 64 causes the gain correction value display area to be displayed on the display screen of the display device 52.
[0161] When the display screen of the content production tool is displayed on the display device 52, the user operates the input unit 61 as appropriate to operate the slider or the like displayed in the slider display area to instruct a change (adjustment) of the gain correction value.
[0162] Then, in step S 42 , the auditory characteristic table generating unit 62 generates an auditory characteristic table in accordance with the input signal supplied from the input unit 61 .
[0163] That is, the auditory characteristic table generating unit 62 generates a new auditory characteristic table by changing the auditory characteristic table read out from the auditory characteristic table holding unit 63 in accordance with the input signal supplied from the input unit 61. That is, the preset auditory characteristic table is changed (updated) in accordance with the operation of a slider or the like displayed in the slider display area.
[0164] In this way, when the gain correction value for each Azimuth value is adjusted (changed) in response to the operation of a slider or the like and a new auditory characteristic table is generated, the auditory characteristic table generation unit 62 instructs the display control unit 64 to update the display of the gain correction value display area in accordance with the new auditory characteristic table.
[0165] In step S43, the display control unit 64 controls the display device 52 in accordance with an instruction from the auditory characteristic table generating unit 62, and performs display according to the newly generated auditory characteristic table.
[0166] Specifically, the display control unit 64 updates the display of the gain correction value display area on the display screen of the display device 52 in accordance with the newly generated auditory characteristics table.
[0167] In step S44, the auditory characteristic table generating unit 62 determines, based on the input signal supplied from the input unit 61, whether or not to end the process.
[0168] For example, the auditory characteristic table generation unit 62 determines to terminate processing when a signal instructing the saving of the auditory characteristic table is supplied as an input signal by the user operating the input unit 61 and operating a save button or the like displayed on the display device 52.
[0169] If it is determined in step S44 that the process is not yet finished, the process returns to step S42, and the above-described process is repeated.
[0170] On the other hand, if it is determined in step S44 that the process is to be ended, the process proceeds to step S45.
[0171] In step S45, the auditory characteristic table generating unit 62 supplies the auditory characteristic table obtained in the last step S42 to the auditory characteristic table holding unit 63 as a newly generated auditory characteristic table, and causes the auditory characteristic table holding unit 63 to hold it.
[0172] When the auditory characteristic table is stored in the auditory characteristic table storage unit 63, the table generation process ends.
[0173] In this manner, the information processing device 51 displays the display screen of the content production tool on the display device 52 and adjusts the gain correction value in response to the user's operation, thereby generating a new auditory characteristics table.
[0174] This allows users to easily and intuitively obtain an auditory characteristic table that matches their desired auditory characteristics, enabling them to more easily create 3D Audio content of sufficiently high quality in a short amount of time.
[0175] Third Embodiment <Configuration example of voice processing device> Furthermore, for example, in free viewpoint content, the listener's position in three-dimensional space can be freely moved, and as the listener moves, the relative positional relationship between the object and the listener in three-dimensional space also changes.
[0176] In this way, when the listener's position can be freely moved, a technique has been proposed in which the sound source position is corrected in accordance with changes in the listener's position, and rendering processing is performed based on the corrected position information obtained as a result (see, for example, International Publication No. 2015 / 107926).
[0177] This technology can also be applied to playback devices that play back such free viewpoint content. In such cases, gain correction is performed using not only the corrected position information but also the above-mentioned three-dimensional auditory characteristics.
[0178] Fig. 14 is a diagram showing an example of the configuration of an embodiment of an audio processing device to which the present technology is applied, functioning as a playback device for playing back free viewpoint content. Note that in Fig. 14, parts corresponding to those in Fig. 4 are assigned the same reference numerals, and descriptions thereof will be omitted as appropriate.
[0179] The audio processing device 91 shown in FIG. 14 includes an input unit 121, a position information correction unit 122, a gain / frequency characteristic correction unit 123, an auditory characteristic table storage unit 22, a spatial acoustic characteristic addition unit 124, a rendering processing unit 125, and a convolution processing unit 126.
[0180] For each object, audio signals and metadata of the audio signals are supplied as audio information of the content to be played back to the audio processing device 91. Note that, although an example in which audio signals and metadata of two objects are supplied to the information processing device 91 will be described in Fig. 14, the number of objects is not limited to this and may be any number.
[0181] Here, the metadata supplied to the audio processing device 91 is the position information and gain initial value of the object.
[0182] The position information is made up of the Azimuth, Elevation, and Radius values mentioned above, and indicates the position of the object as seen from a reference position in three-dimensional space, i.e., the localization position of the sound from the object. Note that hereinafter, the reference position in three-dimensional space will also be referred to as the standard listening position.
[0183] The input unit 121 includes a mouse, buttons, a touch panel, etc., and when operated by a user, outputs a signal corresponding to the operation. For example, the input unit 121 receives input of an expected listening position by the user, and supplies expected listening position information indicating the expected listening position input by the user to the position information correction unit 122 and the spatial acoustic characteristic addition unit 124.
[0184] Here, the assumed listening position is the listening position of the sounds that make up the content in the virtual sound field that is to be reproduced. Therefore, it can be said that the assumed listening position indicates the position after changing (correcting) the predetermined standard listening position.
[0185] The position information correction unit 122 corrects the position information as metadata of the object supplied from the outside, based on the assumed listening position information supplied from the input unit 121 and directional information indicating the orientation of the listener supplied from the outside.
[0186] The position information correcting unit 122 supplies the corrected position information obtained by correcting the position information to the gain / frequency characteristic correcting unit 123 and the rendering processing unit 125 .
[0187] The direction information can be obtained, for example, from a gyro sensor provided on the head of the user (listener). The corrected position information is information indicating the position of the object as seen by the listener who is at the assumed listening position and facing the direction indicated by the direction information, that is, the localized position of the sound of the object.
[0188] The gain / frequency characteristic correction unit 123 performs gain correction and frequency characteristic correction of the audio signal of the externally supplied object based on the corrected position information supplied from the position information correction unit 122, the auditory characteristic table stored in the auditory characteristic table storage unit 22, and metadata supplied from the outside.
[0189] The gain / frequency characteristic correction unit 123 supplies the audio signal obtained by the gain correction and frequency characteristic correction to the spatial acoustic characteristic addition unit 124.
[0190] The spatial acoustic characteristic adding unit 124 adds spatial acoustic characteristics to the audio signal supplied from the gain / frequency characteristic correcting unit 123 based on the assumed listening position information supplied from the input unit 121 and the position information of the object supplied from outside, and supplies the audio signal to the renderer processing unit 125.
[0191] The rendering processing unit 125 performs a rendering process, i.e., a mapping process, on the audio signal supplied from the spatial acoustic characteristic adding unit 124 based on the corrected position information supplied from the position information correcting unit 122, and generates playback signals for M channels, where M is two or more.
[0192] That is, M-channel playback signals are generated from the audio signals of each object. The renderer processing unit 125 supplies the generated M-channel playback signals to the convolution processing unit 126.
[0193] The M-channel playback signals obtained in this way are audio signals that, when played back through M virtual speakers (M-channel speakers), reproduce the sounds output from each object, heard at the assumed listening position in the virtual sound field that you want to reproduce.
[0194] The convolution processing unit 126 performs convolution processing on the M-channel playback signals supplied from the rendering processing unit 125, and generates and outputs two-channel playback signals.
[0195] That is, in this example, the device on the content playback side is a headphone, and the convolution processing unit 126 generates and outputs a playback signal to be played back by two speakers (drivers) provided in the headphone.
[0196] <Description of playback signal generation process> Next, the operation of the audio processing device 91 will be described.
[0197] That is, the playback signal generation process performed by the audio processing device 91 will be described below with reference to the flowchart of FIG.
[0198] In step S71, the input unit 121 receives input of an assumed listening position.
[0199] When the user operates the input unit 121 to input an assumed listening position, the input unit 121 supplies assumed listening position information indicating the assumed listening position to the position information correcting unit 122 and the spatial acoustic characteristic adding unit 124 .
[0200] In step S72, the position information correcting unit 122 calculates corrected position information based on the assumed listening position information supplied from the input unit 121 and the position information and direction information of the object supplied from the outside.
[0201] The position information correcting unit 122 supplies the corrected position information obtained for each object to the gain / frequency characteristic correcting unit 123 and the rendering processing unit 125 .
[0202] In step S73, the gain / frequency characteristic correction unit 123 performs gain correction and frequency characteristic correction of the audio signal of the object supplied from outside based on the corrected position information supplied from the position information correction unit 122, the metadata supplied from outside, and the auditory characteristic table held in the auditory characteristic table holding unit 22.
[0203] Specifically, for example, the gain / frequency characteristic correction unit 123 reads out, from the auditory characteristic table, the gain correction values associated with the Azimuth value, Elevation value, and Radius value that constitute the correction position information.
[0204] In addition, the gain / frequency characteristic correction unit 123 corrects the gain correction value by multiplying the gain correction value by the ratio between the Radius value of the position information supplied as metadata and the Radius value of the corrected position information, and corrects the initial gain value using the resulting gain correction value to obtain the gain value.
[0205] As a result, gain correction according to the direction of the object as seen from the assumed listening position and gain correction according to the distance from the assumed listening position to the object are realized by gain correction using gain values.
[0206] Furthermore, the gain / frequency characteristic correction unit 123 selects a filter coefficient based on the Radius value of the position information supplied as metadata and the Radius value of the corrected position information.
[0207] The filter coefficients selected in this way are used in filter processing to achieve the desired frequency characteristic correction. More specifically, the filter coefficients are used to reproduce the characteristic of high-frequency components of sound from an object being attenuated by the walls and ceiling of a virtual sound field to be reproduced, depending on the distance from the assumed listening position to the object.
[0208] The gain / frequency characteristic correction unit 123 performs gain correction and filter processing on the audio signal of the object based on the filter coefficient and gain value obtained as described above, thereby achieving gain correction and frequency characteristic correction.
[0209] The gain / frequency characteristic correction unit 123 supplies the audio signal of each object obtained by the gain correction and frequency characteristic correction to the spatial acoustic characteristic addition unit 124.
[0210] In step S74, the spatial acoustic characteristic adding unit 124 adds spatial acoustic characteristics to the audio signal supplied from the gain / frequency characteristic correcting unit 123 based on the assumed listening position information supplied from the input unit 121 and the position information of the object supplied from outside, and supplies the audio signal to the renderer processing unit 125.
[0211] For example, the spatial acoustic characteristic adding unit 124 adds spatial acoustic characteristics by performing multi-tap delay processing, comb filter processing, or all-pass filter processing on the audio signal based on the delay amount and gain amount determined from the object position information and the assumed listening position information. As a result, for example, early reflections, reverberation characteristics, etc. are added to the audio signal as spatial acoustic characteristics.
[0212] In step S75, the rendering processing unit 125 performs mapping processing on the audio signal supplied from the spatial acoustic characteristic adding unit 124 based on the corrected position information supplied from the position information correcting unit 122, thereby generating an M-channel playback signal and supplying it to the convolution processing unit 126.
[0213] For example, in the process of step S75, the playback signal is generated by VBAP, but any other method may be used to generate the M-channel playback signal.
[0214] In step S76, the convolution processing unit 126 generates and outputs a two-channel playback signal by performing convolution processing on the M-channel playback signal supplied from the rendering processing unit 125. For example, BRIR (Binaural Room Impulse Response) processing is performed as the convolution processing.
[0215] When the two-channel playback signals are generated and output, the playback signal generation process ends.
[0216] In this way, the audio processing device 91 calculates corrected position information based on the expected listening position information, and performs gain correction and frequency characteristic correction on the audio signal of each object, or adds spatial acoustic characteristics, based on the obtained corrected position information and expected listening position information.
[0217] This makes it easier to perform appropriate gain correction and frequency characteristic correction. It also makes it possible to realistically reproduce how the sound output from each object sounds at any assumed listening position. This allows users to freely specify the listening position according to their preferences when playing content, enabling audio playback with a higher degree of freedom.
[0218] In step S73, in addition to performing gain correction and frequency characteristic correction according to the distance from the expected listening position to the object based on the correction position information, gain correction based on three-dimensional auditory characteristics is also performed using an auditory characteristic table.
[0219] At this time, the auditory characteristics table used in step S73 is, for example, the one shown in FIG.
[0220] The auditory characteristics table shown in FIG. 16 is obtained by inverting the signs of the gain correction values in the auditory characteristics table shown in FIG.
[0221] By correcting the initial gain value using such an auditory characteristic table, it is possible to reproduce the phenomenon in which the auditory loudness of a sound changes depending on the direction from which the sound from the same object (sound source) comes, thereby achieving a more realistic sound field reproduction.
[0222] On the other hand, depending on the reproduction conditions, more appropriate gain correction may be achieved by using the auditory characteristics table shown in FIG. 6 rather than the auditory characteristics table shown in FIG.
[0223] That is, for example, let us consider a case where content is not played back using headphones, but rather played back using real speakers arranged in a three-dimensional space.
[0224] In this case, in the audio processing device 91, the M-channel playback signals obtained by the renderer processing unit 125 are supplied to speakers corresponding to the M channels, and the sound of the content is reproduced.
[0225] When content is played back using such real speakers, the sound of the actual sound source, that is, the sound of the object, is played back at the position of the object as seen from the assumed listening position.
[0226] Therefore, there is no need for gain correction that reproduces the phenomenon where the perceived loudness of a sound changes depending on the direction from which the sound is coming. In fact, there are cases where it is not desirable to change the perceived loudness of a sound so as not to change the volume balance.
[0227] In such a case, in step S73, a gain correction value is determined using the auditory characteristics table shown in Fig. 6, and the initial gain value is corrected using the determined gain correction value. In this way, gain correction is performed so that the audible loudness of the sound becomes constant regardless of the direction of the object.
[0228] <Modification 1 of the third embodiment> <On the code transmission of gain auditory characteristic information> Incidentally, audio signals, metadata, etc. may be encoded and transmitted as an encoded bitstream.
[0229] In such a case, for example, in the gain / frequency characteristic correction unit 123, gain auditory characteristic information including flag information indicating whether or not to perform gain correction using an auditory characteristic table can be transmitted by the coded bit stream.
[0230] In this case, the gain auditory characteristic information can include not only flag information but also an auditory characteristic table and index information indicating an auditory characteristic table to be used for gain correction among a plurality of auditory characteristic tables.
[0231] The syntax of such gain auditory characteristic information can be, for example, as shown in FIG.
[0232] In the example of FIG. 17, the characters "numGainAuditoryPropertyTables" indicate the number of auditory property tables transmitted by the coded bitstream, that is, the number of auditory property tables included in the gain auditory property information.
[0233] Furthermore, the characters "numElements[i]" indicate the number of elements that make up the i-th auditory characteristic table included in the gain auditory characteristic information.
[0234] The elements here refer to the Azimuth value, Elevation value, Radius value, and gain correction value that are associated with each other.
[0235] Furthermore, the letters "azimuth[i][n]", "elevation[i][n]", and "radius[i][n]" indicate the Azimuth value, Elevation value, and Radius value that make up the nth element of the ith auditory characteristic table.
[0236] In other words, azimuth[i][n], elevation[i][n], and radius[i][n] indicate the direction from which the sound is coming from the object that is the sound source, i.e., the horizontal angle, vertical angle, and distance (radius) that indicate the object's position.
[0237] Furthermore, the characters "gainCompensValue[i][n]" indicate the gain compensation value constituting the nth element of the i-th auditory characteristic table, i.e., the gain compensation value for the position (direction) indicated by azimuth[i][n], elevation[i][n], and radius[i][n].
[0238] Furthermore, the characters "hasGainCompensObjects" are flag information indicating whether or not there are any objects for which gain compensation using the auditory characteristics table is performed.
[0239] In addition, the characters "num_objects" indicate the number of objects (number of objects) that make up the content, and this number of objects, num_objects, is assumed to be transmitted to the device that plays back the content, i.e., the audio processing device, separately from the gain auditory characteristic information.
[0240] If the value of the flag information hasGainCompensObjects indicates that there is an object that performs gain compensation using the auditory characteristic table, the gain auditory characteristic information contains flag information indicated by the characters "isGainCompensObject[o]" for the number of objects, num_objects.
[0241] The flag information isGainCompensObject[o] indicates whether or not gain compensation using the auditory characteristics table is to be performed on the o-th object.
[0242] Furthermore, if the value of the flag information isGainCompensObject[o] is a value indicating that gain compensation is to be performed using an auditory characteristic table, the gain auditory characteristic information includes an index indicated by the characters "applyTableIndex[o]".
[0243] This index applyTableIndex[o] is information indicating the auditory characteristic table to be used when performing gain correction on the oth object.
[0244] For example, when the number of auditory property tables, numGainAuditoryPropertyTables, is 0, no auditory property table is transmitted, and the gain auditory property information does not include the index applyTableIndex[o]. That is, the index applyTableIndex[o] is not transmitted.
[0245] In such a case, for example, the auditory characteristic table stored in the auditory characteristic table storage unit 22 may be used to perform gain correction, or gain correction may not be performed.
[0246] <Configuration example of voice processing device> When the above-described gain auditory characteristic information is transmitted by an encoded bit stream, the audio processing device may be configured as shown in Fig. 18. In Fig. 18, the same reference numerals are used to designate parts corresponding to those in Fig. 14, and their explanation will be omitted as appropriate.
[0247] The audio processing device 151 shown in FIG. 18 includes an input unit 121, a position information correction unit 122, a gain / frequency characteristic correction unit 123, an auditory characteristic table storage unit 22, a spatial acoustic characteristic addition unit 124, a rendering processing unit 125, and a convolution processing unit 126.
[0248] The configuration of the audio processing device 151 is the same as the configuration of the audio processing device 91 shown in Figure 14, but differs from the audio processing device 91 in that an audio characteristic table, etc. read from the gain audio characteristic information extracted from the encoded bitstream is supplied to the gain / frequency characteristic correction unit 123.
[0249] That is, in the audio processing device 151, the gain / frequency characteristic correction unit 123 is supplied with an auditory characteristic table read from the gain auditory characteristic information, flag information hasGainCompensObjects, flag information isGainCompensObject[o], index applyTableIndex[o], and the like.
[0250] In the audio processing device 151, basically, the playback signal generation process described with reference to FIG. 15 is performed.
[0251] However, in step S73, if the number of auditory characteristic tables numGainAuditoryPropertyTables is 0, that is, if no auditory characteristic table is supplied from the outside, the gain / frequency characteristic correction unit 123 performs gain correction using the auditory characteristic table stored in the auditory characteristic table storage unit 22.
[0252] On the other hand, when an auditory characteristic table is supplied from the outside, the gain / frequency characteristic correction unit 123 performs gain correction using the supplied auditory characteristic table.
[0253] Specifically, the gain / frequency characteristic correction unit 123 performs gain correction on the o-th object using the auditory characteristic table indicated by the index applyTableIndex[o] from among a plurality of auditory characteristic tables supplied from the outside.
[0254] However, the gain / frequency characteristic correction unit 123 does not perform gain correction using the auditory characteristic table for an object whose flag information isGainCompensObject[o] has a value indicating that gain correction using the auditory characteristic table is not to be performed.
[0255] That is, in the gain / frequency characteristic correction unit 123, when flag information isGainCompensObject[o] having a value indicating that gain correction is to be performed using an auditory characteristic table is supplied, gain correction is performed using the auditory characteristic table indicated by the index applyTableIndex[o].
[0256] Also, for example, if the value of flag information hasGainCompensObjects indicates that there is no object for which gain compensation using the auditory characteristics table is to be performed, the gain / frequency characteristic compensation unit 123 does not perform gain compensation using the auditory characteristics table for the object.
[0257] As described above, this technology makes it possible to easily determine the gain information, i.e., the gain value, of each object in 3D mixing of object audio, playback of free viewpoint content, etc. This makes it possible to perform gain correction more easily.
[0258] Furthermore, according to the present technology, it is possible to appropriately correct a change in the volume perceived by the listener that accompanies a change in the relative positional relationship between the listener and an object (sound source) when the listening position is changed.
[0259] <Example of computer configuration> The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the programs constituting the software are installed on a computer. Here, the term "computer" includes computers built into dedicated hardware, and general-purpose personal computers, for example, that can execute various functions by installing various programs.
[0260] FIG. 19 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes using a program.
[0261] In the computer, a CPU (Central Processing Unit) 501, a ROM (Read Only Memory) 502, and a RAM (Random Access Memory) 503 are interconnected by a bus 504.
[0262] An input / output interface 505 is further connected to the bus 504. An input unit 506, an output unit 507, a recording unit 508, a communication unit 509, and a drive 510 are connected to the input / output interface 505.
[0263] The input unit 506 includes a keyboard, a mouse, a microphone, an image sensor, etc. The output unit 507 includes a display, a speaker, etc. The recording unit 508 includes a hard disk, a non-volatile memory, etc. The communication unit 509 includes a network interface, etc. The drive 510 drives a removable recording medium 511 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0264] In a computer configured as described above, the CPU 501 performs the above-described series of processes by, for example, loading a program recorded in the recording unit 508 into the RAM 503 via the input / output interface 505 and the bus 504 and executing it.
[0265] The program executed by the computer (CPU 501) can be provided by being recorded on a removable recording medium 511 such as a package medium, for example. The program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.
[0266] In a computer, a program can be installed in the recording unit 508 via the input / output interface 505 by inserting a removable recording medium 511 into the drive 510. The program can also be received by the communication unit 509 via a wired or wireless transmission medium and installed in the recording unit 508. Alternatively, the program can be installed in the ROM 502 or the recording unit 508 in advance.
[0267] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.
[0268] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present technology.
[0269] For example, this technology can be configured as cloud computing, in which a single function is shared and processed collaboratively by multiple devices via a network.
[0270] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by multiple devices.
[0271] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.
[0272] Furthermore, the present technology can also be configured as follows.
[0273] (1) a gain correction value determination unit that determines a gain correction value for gain-correcting an audio signal of an audio object according to a direction of the audio object as seen by a listener; Information processing device. (2) The gain correction value determination unit determines the correction value based on the three-dimensional hearing characteristics of the listener with respect to the direction from which the sound comes. The information processing device described in (1). (3) The gain correction value determination unit determines the correction value based on the orientation of the listener. An information processing device according to (1) or (2). (4) The gain correction value determiner determines the correction value so that when the audio object is located behind the listener, the correction value is larger than when the audio object is located in front of the listener. An information processing device according to any one of (1) to (3). (5) The gain correction value determiner determines the correction value so that the correction value is smaller when the audio object is located to the side of the listener than when the audio object is located in front of the listener. An information processing device according to any one of (1) to (4). (6) The gain correction value determination unit determines the correction value according to the predetermined direction by performing an interpolation process based on the correction values according to other directions. An information processing device according to any one of (1) to (5). (7) The gain correction value determination unit performs VBAP as the interpolation processing. (6) An information processing device according to the present invention. (8) The gain correction value determination unit determines the correction value in a linear value or a decibel value. (7) An information processing device according to (7). (9) The information processing device A gain correction value for correcting the gain of an audio signal of an audio object is determined according to a direction of the audio object as seen from a listener. Information processing methods. (10) A gain correction value for correcting the gain of an audio signal of an audio object is determined according to a direction of the audio object as seen from a listener. A program that causes a computer to execute a process that includes steps. (11) a gain correction unit that determines a gain correction value for correcting a gain of an audio signal of an audio object based on position information indicating a position of the audio object, the correction value corresponding to a direction of the audio object as seen by a listener, and performs the gain correction of the audio signal based on the gain value corrected by the correction value; a renderer processing unit that performs rendering processing based on the audio signal obtained by the gain correction and generates playback signals of a plurality of channels for playing back the sound of the audio object; A playback device comprising: (12) The gain correction unit corrects the gain value included in the metadata of the audio signal using the correction value. The playback device according to (11). (13) When a flag indicating that the gain value is to be corrected is supplied, the gain correction unit corrects the gain value using the correction value. A playback device according to (11) or (12). (14) The gain correction unit determines the correction value using the table indicated by the supplied index from among a plurality of tables in which the direction of the audio object as seen from the listener is associated with the correction value. (13) A playback device according to (13). (15) a position information correction unit that corrects the position information included in the metadata of the audio signal based on information indicating the position of the listener; The gain correction unit determines the correction value based on the corrected position information. A playback device according to any one of (11) to (14). (16) The position information correcting unit corrects the position information based on information indicating the position of the listener and direction information indicating the orientation of the listener. (15) A playback device according to (15). (17) The playback device determining a gain correction value for gain-correcting an audio signal of the audio object based on position information indicating the position of the audio object, the gain correction value corresponding to the direction of the audio object as seen from the listener; performing the gain correction of the audio signal based on the gain value corrected by the correction value; Rendering is performed based on the audio signal obtained by the gain correction, and a reproduction signal of a plurality of channels for reproducing the sound of the audio object is generated. How to play. (18) determining a gain correction value for gain-correcting an audio signal of the audio object based on position information indicating the position of the audio object, the gain correction value corresponding to the direction of the audio object as seen from the listener; performing the gain correction of the audio signal based on the gain value corrected by the correction value; Rendering is performed based on the audio signal obtained by the gain correction, and a reproduction signal of a plurality of channels for reproducing the sound of the audio object is generated. A program that causes a computer to execute a process that includes steps. [Explanation of symbols]
[0274] 11 information processing device, 21 gain correction value determination unit, 22 auditory characteristic table storage unit, 62 auditory characteristic table generation unit, 64 display control unit, 122 position information correction unit, 123 gain / frequency characteristic correction unit
Claims
1. a gain correction value determination unit that determines a gain correction value for gain-correcting an audio signal of an audio object according to a direction of the audio object as seen by a listener; The gain correction value determination unit determines the correction value according to the predetermined direction by performing an interpolation process based on the correction values according to other directions. Information processing device.
2. The gain correction value determination unit determines the correction value based on three-dimensional hearing characteristics of the listener with respect to the direction from which the sound comes. The information processing device according to claim 1 .
3. The gain correction value determination unit determines the correction value based on the orientation of the listener. The information processing device according to claim 1 .
4. The gain correction value determination unit performs VBAP as the interpolation processing. The information processing device according to claim 1 .
5. The gain correction value determination unit determines the correction value in a linear value or a decibel value. The information processing device according to claim 4 .
6. The information processing device determining a gain correction value for gain-correcting an audio signal of the audio object according to a direction of the audio object as seen by a listener; The correction value according to the predetermined direction is determined by performing an interpolation process based on the correction values according to other directions. Information processing methods.
7. A gain correction value for correcting the gain of an audio signal of an audio object is determined according to a direction of the audio object as seen from a listener. causing a computer to execute a process including the steps, The correction value according to the predetermined direction is determined by performing an interpolation process based on the correction values according to other directions. program.
8. a gain correction unit that determines a gain correction value for correcting a gain of an audio signal of an audio object based on position information indicating a position of the audio object, the correction value corresponding to a direction of the audio object as seen by a listener, and performs the gain correction of the audio signal based on the gain value corrected by the correction value; a renderer processing unit that performs rendering processing based on the audio signal obtained by the gain correction and generates playback signals of a plurality of channels for playing back the sound of the audio object; Equipped with The gain correction unit determines the correction value using the table indicated by the supplied index from among a plurality of tables in which the direction of the audio object as seen from the listener is associated with the correction value. playback device.
9. The gain correction unit corrects the gain value included in the metadata of the audio signal using the correction value. The playback device according to claim 8.
10. When a flag indicating that the gain value is to be corrected is supplied, the gain correction unit corrects the gain value using the correction value. The playback device according to claim 8.
11. a position information correction unit that corrects the position information included in the metadata of the audio signal based on information indicating the position of the listener; The gain correction unit determines the correction value based on the corrected position information. The playback device according to claim 8.
12. The position information correcting unit corrects the position information based on information indicating the position of the listener and direction information indicating the orientation of the listener. The playback device according to claim 11.
13. The playback device determining a gain correction value for gain-correcting an audio signal of the audio object based on position information indicating the position of the audio object, the gain correction value corresponding to the direction of the audio object as seen from the listener; performing the gain correction of the audio signal based on the gain value corrected by the correction value; Rendering is performed based on the audio signal obtained by the gain correction, and a reproduction signal of a plurality of channels for reproducing the sound of the audio object is generated. including the steps The correction value is determined using the table indicated by the supplied index from among a plurality of tables in which the direction of the audio object as seen from the listener is associated with the correction value. How to play.
14. determining a gain correction value for gain-correcting an audio signal of the audio object based on position information indicating the position of the audio object, the gain correction value corresponding to the direction of the audio object as seen from the listener; performing the gain correction of the audio signal based on the gain value corrected by the correction value; Rendering is performed based on the audio signal obtained by the gain correction, and a reproduction signal of a plurality of channels for reproducing the sound of the audio object is generated. causing a computer to execute a process including the steps, The correction value is determined using the table indicated by the supplied index from among a plurality of tables in which the direction of the audio object as seen from the listener is associated with the correction value. program.
Citation Information
Patent Citations
IEC23008-3
IEC23008-3,
Speaker device
JP2015126359A
Transmission device, transmission method, receiving device, and receiving method
JP2018116299A
JPP7513020B