Information processing device, method, and program

The described technology addresses high computational demands in virtual audio reproduction by selecting directional filters based on sound source characteristics and listener distance, enabling efficient and continuous sound localization in virtual spaces.

WO2025177809A1PCT designated stage Publication Date: 2025-08-28SONY GROUP CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/003375
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-19
Filing Date
2025-02-03
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Existing technologies for realistic audio reproduction in virtual spaces require significant computational resources due to the need to superimpose directivity equally for all object sound sources, leading to high calculation demands.

Method used

An information processing device and method that selects directional filters based on positional relationships between sound sources and listeners using parameters like directionality and distance, generating and encoding parameters to reduce computational load while maintaining realistic sound localization.

Benefits of technology

Realistic audio reproduction is achieved with reduced calculations by dynamically adapting directivity based on listener position, ensuring continuous sound localization and volume without incongruity, thus optimizing computational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025003375_28082025_PF_FP_ABST
    Figure JP2025003375_28082025_PF_FP_ABST
Patent Text Reader

Abstract

The present technology relates to an information processing device, method, and program that can realize audio playback having presence, with less computation. An information processing device is provided with a signal processing unit that selects, from among directional filters for reproducing the directivity of a sound source, a directional filter corresponding to the positional relationship between the sound source and a listener by means of a selection method determined on the basis of one or more parameters that represent the characteristics of the sound source, including at least a first parameter relating to the directivity of the sound source, and the distance in a virtual space from the sound source to the listener. The present technology can be applied to a decoder.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, method, and program

[0001] The present technology relates to an information processing device, method, and program, and more particularly to an information processing device, method, and program that enable realistic audio reproduction with less calculations.

[0002] When playing object audio in a virtual space, i.e., sound output from an object sound source called an audio object, it is possible to present a more immersive and realistic sound image by reproducing the directionality of the object sound source.

[0003] For example, a technology has been proposed in which the radiation characteristics (directivity) of each object sound source are maintained, and a directional filter based on the listener's ear position is selected according to the relative positional relationship between the object sound source and the listener, and the selected directional filter is superimposed on the signal of the object sound source (see, for example, Patent Document 1).

[0004] JP 2022-34268 A

[0005] However, the above-mentioned technology requires superimposing directivity equally for the left and right ears on all object sound sources present in the virtual space, which results in a huge amount of computational resources, i.e., a huge amount of calculation.

[0006] The present technology has been made in view of such circumstances, and aims to realize realistic audio reproduction with fewer calculations.

[0007] An information processing device according to a first aspect of the present technology includes a signal processing unit that selects a directional filter from among directional filters for reproducing the directionality of the sound source according to a positional relationship between the sound source and the listener, using a selection method determined from one or more parameters representing characteristics of the sound source, including at least a first parameter related to the directionality of the sound source, and a distance from the sound source to a listener in a virtual space.

[0008] An information processing method or program according to a first aspect of the present technology includes a step of selecting a directional filter corresponding to a positional relationship between the sound source and the listener from among directional filters for reproducing the directionality of the sound source, using a selection method determined from one or more parameters representing characteristics of the sound source, including at least a first parameter related to the directionality of the sound source, and a distance from the sound source to the listener in a virtual space.

[0009] In a first aspect of the present technology, a directional filter corresponding to the positional relationship between the sound source and the listener is selected from among directional filters for reproducing the directionality of the sound source using a selection method determined from one or more parameters representing characteristics of the sound source, including at least a first parameter related to the directionality of the sound source, and the distance from the sound source to the listener in a virtual space.

[0010] An information processing device according to a second aspect of the present technology includes a parameter generation unit that generates, based on directivity data indicating the directivity of a sound source placed in a virtual space, one or more parameters that represent characteristics of the sound source, including at least a first parameter related to the directivity of the sound source, and a communication unit that outputs a bitstream including the parameters.

[0011] In a second aspect of the present technology, based on directivity data indicating the directivity of a sound source placed in a virtual space, one or more parameters representing characteristics of the sound source, including at least a first parameter related to the directivity of the sound source, are generated, and a bitstream including the parameters is output.

[0012] FIG. 1 is a diagram illustrating an example of the configuration of an encoder. FIG. 2 is a diagram illustrating an example of directivity data. FIG. 3 is a flowchart illustrating encoding processing. FIG. 4 is a diagram illustrating an example of the configuration of a decoder. FIG. 5 is a diagram illustrating a transition curve. FIG. 6 is a diagram illustrating an example of a transition curve. FIG. 7 is a diagram illustrating a weighting coefficient. FIG. 8 is a diagram illustrating selection of a directional filter. FIG. 9 is a flowchart illustrating output audio data generation processing. FIG. 10 is a diagram illustrating selection of a transition curve when there are multiple directivity peaks. FIG. 11 is a diagram illustrating an example of obtaining a head radius. FIG. 12 is a diagram illustrating input of a head radius. FIG. 13 is a diagram illustrating input of a transition curve parameter. FIG. 14 is a diagram illustrating an example of the configuration of a computer.

[0013] Hereinafter, embodiments to which the present technology is applied will be described with reference to the drawings.

[0014] First Embodiment Example of Encoder Configuration The present technology relates to a technology for reproducing object audio having directionality in a virtual space. Hereinafter, an object sound source (audio object) will also be simply referred to as a sound source.

[0015] When playing directional object audio in a virtual space, i.e., sound output from an object sound source, reproducing the directionality of the sound source makes it possible to present a more realistic sound image with a sense of presence.

[0016] In such a case, for example, the directivity (directional characteristics) of the sound source is obtained in advance by measurement or simulation.

[0017] When the distance between the sound source and the listener in the virtual space is large enough, the directivity of the sound source can be approximated as omnidirectional, but as the listener approaches the sound source, the directivity of the sound source must be gradually adapted.

[0018] There are methods for reproducing the directivity of a sound source by appropriately applying the directivity of the sound source to the fixed positional relationship between the sound source and the listener. However, no method has been proposed that applies the directivity without creating a sense of incongruity while maintaining the sense of continuity of the sound localization and volume when the distance between the sound source and the listener changes dynamically.

[0019] Therefore, this technology makes it possible to smoothly switch the process from non-directional rendering to directional rendering using characteristics such as the sharpness of the sound source's directivity as a parameter.

[0020] FIG. 1 is a diagram showing an example of the configuration of an embodiment of an encoder to which the present technology is applied.

[0021] The encoder 11 is an information processing device such as a server that has a function as an encoding device, and generates and outputs a bit stream made up of various data for reproducing content.

[0022] The content may be, for example, audio content consisting of sounds from one or more sound sources (audio objects), or video content consisting of video and accompanying audio including sounds from one or more sound sources.

[0023] The encoder 11 includes a parameter generating unit 21 , a parameter encoding unit 22 , an audio encoding unit 23 , a multiplexing unit 24 , and a communication unit 25 .

[0024] The parameter generation unit 21 is supplied with directivity data indicating the directional characteristics (directivity) of each sound source placed in the virtual space, more specifically, for each type of sound source. For example, the directivity data may be data containing information on the amplitude and phase of the sound from the sound source for positions in each direction as seen from the sound source. Note that the directivity data may be prepared for each frequency.

[0025] The parameter generation unit 21 generates (extracts) one or more parameters representing the characteristics of the sound source from the supplied directional data based on a predetermined mathematical model, metadata of the sound source, signals supplied in response to input operations by a creator, etc., and supplies them to the parameter encoding unit 22.

[0026] For example, the parameter generating unit 21 generates a parameter κ indicating the sharpness of the directivity of the sound source and a parameter p indicating the priority, i.e., the importance (level of priority), of the sound source (audio object) as parameters indicating the characteristics of the sound source. The one or more parameters indicating the characteristics of the sound source include at least a parameter related to the directivity of the sound source, for example, the parameter κ indicating the sharpness of the directivity of the sound source.

[0027] For example, the value of parameter κ increases as the directivity of the sound source increases, and decreases as the directivity of the sound source approaches omnidirectionality.Furthermore, for example, the value of parameter p increases as the importance of the sound source increases.

[0028] In the following description, it is assumed that parameters κ and p are generated as parameters representing the characteristics of the sound source.

[0029] The parameter encoding unit 22 encodes the parameters κ and p supplied from the parameter generating unit 21 using a predetermined encoding method, and supplies the resulting encoded parameters of each excitation signal to the multiplexing unit 24 .

[0030] Audio data of each sound source is supplied to the audio encoding unit 23. The audio data of a sound source is data for reproducing the sound of that sound source.

[0031] The audio encoding unit 23 encodes the audio data of each sound source supplied thereto using a predetermined encoding method, and supplies the resulting encoded audio data to the multiplexing unit 24 .

[0032] The audio encoding unit 23 may also encode audio data of sounds other than the sound source, such as channel-based audio data, and the resulting encoded audio data may also be supplied to the multiplexing unit 24.

[0033] The multiplexing unit 24 multiplexes the encoding parameters supplied from the parameter encoding unit 22 and the encoded audio data supplied from the audio encoding unit 23 to generate a bit stream, and supplies the bit stream to the communication unit 25 .

[0034] The communication unit 25 transmits (outputs) the bit stream supplied from the multiplexing unit 24 to the destination of the content, that is, the information processing device that is the client.

[0035] Note that, when there is metadata for each sound source, i.e., metadata for the audio data of each sound source, the metadata may also be encoded by the parameter encoding unit 22 or the audio encoding unit 23, and the resulting encoded metadata may also be stored in the bitstream and transmitted (transmitted). In this case, the parameters κ and p may be transmitted as information constituting the metadata.

[0036] For example, metadata for a sound source may include at least one of the following: sound source position information indicating the position of the sound source in the virtual space; sound source priority information indicating the priority of the sound source; gain information indicating the gain of the audio data of the sound source; sound source type information indicating the type of sound source (e.g., instrument, vocals, etc.); and spread information indicating the degree of spread of the sound source. Additionally, information indicating the direction of the sound source in the virtual space may also be included in the metadata. Furthermore, for example, the priority indicated by the sound source priority information is expressed as a value from 0 to 7, and is set by the content creator, etc.

[0037] The directional data of each sound source or the directional filter generated from the directional data may also be coded by the parameter coding unit 22, and the resulting coded directional data or coded directional filter may also be stored in a bit stream and transmitted (transmitted).

[0038] The directional filter is a filter for each direction as seen from the sound source, and is used to reproduce the directivity of the sound source during audio playback.

[0039] By filtering using a directional filter, i.e., by superimposing a directional filter, the directional characteristics (directivity) of a sound source can be added to audio data. For example, a directional filter for a predetermined direction as seen from the sound source may have a gain value calculated from the amplitude of the directional data for that predetermined direction. Note that the directional filter may be the directional data itself, or may be modeled directional data, as described below.

[0040] Furthermore, the audio encoding unit 23 may be provided in a device separate from the encoder 11, and the encoding parameters and the encoded audio data may be transmitted to the client by separate devices. In other words, the audio encoding unit 23 does not have to be provided in the encoder 11.

[0041] <Regarding Parameters> Examples of the parameters κ and p will be described.

[0042] For example, the parameter generating unit 21 models the directivity data of the sound source using a mathematical model. More specifically, for example, the directivity data is modeled and expressed using one or more von Mises Fisher (vMF) distributions.

[0043] In this case, the directional data after modeling, i.e., the vMF distribution representing the directional data, is a vector γ 1 It can be expressed by parameters such as a mean vector (K) and a parameter concentration index κ (hereinafter also referred to as model parameters).

[0044] Figure 2 shows an example of directional data modeled using the vMF distribution.

[0045] In FIG. 2, the spheres indicated by arrows Q11 to Q14 represent the directivity data (vMF distribution) after modeling.

[0046] In particular, the vMF distribution indicated by arrow Q11 has a parameter concentration of κ = 3, the vMF distribution indicated by arrow Q12 has a parameter concentration of κ = 5, and the vMF distribution indicated by arrow Q13 has a parameter concentration of κ = 8. Furthermore, the vMF distribution indicated by arrow Q14 has a parameter concentration of κ = 10.

[0047] Each vMF distribution is a vector γ on the surface of the sphere. 1 The center of the sphere representing the vMF distribution corresponds to the position of the sound source, and the shade of color at each position on the surface of the sphere indicates the value of the directivity data for each position as seen from the center of the sphere (sound source).

[0048] In this example, it can be seen that the larger the parameter concentration κ, the sharper the shape of the directivity, that is, the shape of the distribution of values ​​on the surface of the sphere.

[0049] Among the model parameters that make up the vMF distribution, the value of the parameter concentration factor κ increases as the directivity of the sound source becomes sharper, so this parameter concentration factor κ is used as the parameter κ that indicates the sharpness of the directivity of the sound source.

[0050] The parameter concentration degree κ may be used as the parameter κ as it is, or a value calculated based on the parameter concentration degree κ may be used as the parameter κ.

[0051] Furthermore, when the directional data is represented by a distribution obtained by mixing multiple vMF distributions, the parameter concentration degree κ of each vMF distribution may be used as the parameter κ, or one or more of the multiple parameter concentration degrees κ may be used as the parameter κ.

[0052] In addition, the mathematical model used to generate (extract) the parameter κ from the directivity data of the sound source is not limited to the vMF distribution, and may be any model such as a Kent distribution, a complex Bingham distribution, a complex Watson distribution, etc. Furthermore, the directivity data may be modeled using a plurality of different types of distributions.

[0053] The parameter generating unit 21 generates (determines) a parameter p indicating the importance of the sound source based on, for example, metadata of the sound source, audio data of the sound source, and a signal supplied in response to an input operation by a creator or the like.

[0054] Specifically, for example, a creator or the like may directly input the importance for each sound source, and the value indicating the input importance may be set as the value of the parameter p. In this case, the parameter generation unit 21 acquires the importance of the sound source in response to an input operation by the creator or the like, and generates the parameter p based on the acquired importance.

[0055] Furthermore, for example, the parameter generating unit 21 may determine the value of the parameter p based on sound source type information, sound source position information, sound source priority information, gain information, spread information, and the like included in the metadata of the sound source.

[0056] As an example, if the content is music content, if the sound source type indicated by the sound source type information is vocal, the parameter p may be set to a large value, and if the sound source type is a reverb object, the parameter p may be set to a small value.

[0057] It is also possible to set the value of parameter p to be larger as the priority indicated by the sound source priority information is higher, or to set the value of parameter p to be larger as the gain value indicated by the gain information is larger.

[0058] Furthermore, it is also conceivable to make the value of parameter p larger the greater the spread of the sound source indicated by the spread information, i.e., the larger the sound source size, or to make the value of parameter p larger the greater the loudness (volume) of the sound based on the audio data of the sound source.

[0059] Alternatively, for example, the closer the position indicated by the sound source position information is to a specific position, the larger the value of parameter p may be, or the closer the position indicated by the sound source position information is to the center of the screen of the content video, the larger the value of parameter p may be.

[0060] In addition, when the position indicated by the sound source position information is outside the screen of the content video, the value of the parameter p may be made small, and when the position indicated by the sound source position information is inside the screen of the content video, the value of the parameter p may be made large.

[0061] Furthermore, although an example in which the parameter p is generated (determined) on the encoder 11 side will be described here, the parameter p may be generated on the client side, that is, on the decoder side, which will be described later.

[0062] <Description of Encoding Process> Next, a description will be given of the operation of the encoder 11. That is, the encoding process performed by the encoder 11 will be described below with reference to the flowchart of FIG.

[0063] In step S11 , the parameter generating unit 21 generates parameters representing the characteristics of the sound sources based on the supplied directivity data of each sound source, and supplies the generated parameters to the parameter encoding unit 22 .

[0064] For example, the parameter generation unit 21 models the directivity data using a mathematical model or the like, and some of the model parameters obtained as a result are used as the parameter κ indicating the sharpness of the directivity. For example, in the above example, the directivity data is modeled using a vMF distribution, and the parameter concentration κ, which is one of the model parameters, is used as the parameter κ. In other words, the parameter κ is extracted from the directivity data.

[0065] Furthermore, for example, the parameter generation unit 21 determines a parameter p indicating the importance of the sound source based on at least one of the sound source metadata, audio data, and signals supplied in response to input operations by a creator, etc., as described above.

[0066] In step S12, the parameter encoding unit 22 encodes the parameters κ and p supplied from the parameter generating unit 21, and supplies the resulting encoding parameters for each excitation signal to the multiplexing unit 24.

[0067] In step S13 , the audio encoding unit 23 encodes the audio data of each sound source supplied thereto, and supplies the resulting encoded audio data to the multiplexing unit 24 .

[0068] In step S14, the multiplexing unit 24 multiplexes the encoding parameters supplied from the parameter encoding unit 22 and the encoded audio data supplied from the audio encoding unit 23 to generate a bit stream, and supplies the bit stream to the communication unit 25.

[0069] In step S15, the communication unit 25 transmits the bit stream supplied from the multiplexing unit 24 to the client, and the encoding process ends.

[0070] In this manner, the encoder 11 generates and encodes the parameters κ and p that represent the characteristics of the sound source, and transmits (outputs) a bitstream that includes the encoding parameters.

[0071] In this way, the client (decoder) can select a directional filter by an appropriate method for each sound source using the parameters κ and p, and add directivity (reproduce directivity), thereby realizing realistic audio reproduction with fewer calculations.

[0072] <Configuration Example of Decoder> FIG. 4 is a diagram showing a configuration example of a decoder 61 that receives the bit stream transmitted by the encoder 11 and generates audio data for reproducing the content.

[0073] The decoder 61 is an information processing device such as a personal computer, a smartphone, a tablet, or a head mounted display (HMD), and functions as a decoding device. The decoder 61 may also function as a client that plays back content.

[0074] The decoder 61 is connected to a display unit 62 that displays various images such as video of the content, and an audio output unit 63 that outputs sound based on the audio data of the content.

[0075] For example, the display unit 62 is made up of a display etc., and the audio output unit 63 is made up of a speaker, headphones etc. The display unit 62 and the audio output unit 63 may be provided inside the decoder 61.

[0076] The decoder 61 includes a communication unit 71 , a demultiplexing unit 72 , an initialization unit 73 , an audio decoding unit 74 , a signal processing unit 75 , an input unit 76 , and a display control unit 77 .

[0077] The communication unit 71 receives the bit stream transmitted from the encoder 11 and supplies it to the demultiplexing unit 72 .

[0078] The demultiplexing unit 72 demultiplexes the bit stream supplied from the communication unit 71, extracts the coding parameters and the coded audio data, and supplies the extracted coding parameters to an initialization unit 73 and the coded audio data to an audio decoding unit 74.

[0079] The initialization unit 73 designs a transition curve for each excitation signal based on the encoding parameters supplied from the demultiplexing unit 72, i.e., the parameters representing the characteristics of the excitation signal, and supplies the resulting transition curve parameters to the signal processing unit 75.

[0080] The transition curve is a weighting coefficient curve (gain curve) for smoothly switching the selection method of the directional filter according to the distance between the sound source and the listener (user) in the virtual space. That is, the transition curve indicates the selection method (selection method) and weighting coefficient (gain value of the directional filter) of the directional filter according to the distance from the sound source in the virtual space.

[0081] The transition curve parameters are transition information consisting of a group of parameters for reproducing a transition curve, and are used to select a directional filter, more specifically, to select a directional filter selection method and to determine (select) a gain value to be used as a directional filter.

[0082] The initialization unit 73 includes a parameter decoding unit 91 and a transition curve design unit 92 .

[0083] The parameter decoding unit 91 decodes the encoding parameters supplied from the demultiplexing unit 72 and supplies the resulting parameters representing the characteristics of the excitation, that is, the parameters κ and p, to a transition curve design unit 92 .

[0084] The transition curve design unit 92 designs (generates) a transition curve based on the parameters κ and p supplied from the parameter decoding unit 91 , and supplies the transition curve parameters obtained as a result of the design to the signal processing unit 75 .

[0085] The audio decoding unit 74 decodes the encoded audio data supplied from the demultiplexing unit 72 and supplies the resulting audio data of the sound source to the signal processing unit 75 .

[0086] If the bitstream also contains encoded audio data of channel-based audio data, the channel-based audio data obtained by decoding in the audio decoding unit 74 is also supplied to the signal processing unit 75. The encoded audio data may be obtained from a device other than the encoder 11, or may be stored in the decoder 61 in advance.

[0087] The signal processing unit 75 generates audio data for reproducing the sound of the content including the sound of each sound source in the virtual space, based on the transition curve parameters supplied from the transition curve design unit 92 and the audio data of each sound source supplied from the audio decoding unit 74. Hereinafter, the audio data for reproducing the sound of the content will also be referred to as output audio data.

[0088] The signal processing unit 75 includes a relative distance calculation unit 93 , a directional filter superimposing unit 94 , and an HRTF (Head Related Transfer Function) superimposing unit 95 .

[0089] The relative distance calculation unit 93, the directional filter superimposition unit 94, and the HRTF superimposition unit 95 are supplied with sound source position / direction information indicating the position and direction of the sound source in the virtual space, and listener position / direction information indicating the position and direction of the listener (user) in the virtual space, as appropriate.

[0090] The sound source position / direction information may be acquired from the encoder 11 or a device on the network different from the encoder 11, or may be stored in the decoder 61 in advance by some method.

[0091] For example, if the multiplexing unit 24 of the encoder 11 stores the sound source position / direction information in the bit stream, the demultiplexing unit 72 of the decoder 61 extracts the sound source position / direction information from the bit stream and supplies it to the signal processing unit 75. Alternatively, the sound source position / direction information may be included in the metadata of the sound source. In such a case, for example, information consisting of the above-mentioned sound source position information included in the metadata and information indicating the direction of the sound source in virtual space is used as the sound source position / direction information.

[0092] In addition, the listener position and direction information may be obtained by a sensor or the like provided on a device worn on the listener's head, such as an HMD having a decoder 61, or may be input by the listener operating the input unit 76 or the like.

[0093] The relative distance calculation unit 93 calculates the distance between the sound source and the listener in the virtual space (hereinafter also referred to as relative distance x) based on the supplied sound source position / direction information and listener position / direction information, and supplies it to the directional filter superimposition unit 94.

[0094] The directional filter superimposition unit 94 selects a directional filter according to the transition curve based on the relative distance x supplied from the relative distance calculation unit 93, the transition curve parameters supplied from the transition curve design unit 92, and the supplied sound source position / direction information and listener position / direction information.

[0095] At this time, the directional filter superimposing unit 94 selects a weighting coefficient (gain value) according to the relative distance x, which is determined by the transition curve, as the final directional filter value, based on the relative distance x and the transition curve parameter.

[0096] For example, the directional filter superimposing unit 94 may store a directional filter in advance, or may acquire a directional filter from the encoder 11 or a device on the network other than the encoder 11. Note that directional data may be acquired, and a directional filter may be generated from the directional data, or a directional filter may be generated (restored) from the above-mentioned model parameters.

[0097] As an example, if the multiplexing unit 24 of the encoder 11 stores the encoding directional filter in the bitstream, the demultiplexing unit 72 of the decoder 61 extracts the encoding directional filter from the bitstream and supplies it to the parameter decoding unit 91.

[0098] The parameter decoding unit 91 decodes the supplied encoded directional filter and supplies the resulting directional filter to the directional filter superimposition unit 94 via the transition curve design unit 92. For example, the encoded directional filter may be an encoded directional filter, or may be an encoded model parameter for obtaining (restoring) directional data after modeling as a directional filter.

[0099] The directional filter superimposing unit 94 performs filtering on the audio data of each sound source supplied from the audio decoding unit 74 based on the directional filter selected for each sound source in accordance with the transition curve, and supplies the resulting audio data to the HRTF superimposing unit 95.

[0100] For example, the directional filter superimposing unit 94 generates audio data for each channel for each sound source through filtering. In particular, audio data for the L (left) channel and R (right) channel is generated here, but depending on the directional filter selection method based on the transition curve, the audio data for the L channel and the R channel may end up being the same.

[0101] The HRTF superimposing unit 95 holds an HRTF for each channel for each direction as seen from the listener, and selects an HRTF for each channel for each sound source based on the supplied sound source position direction information and listener position direction information.

[0102] The HRTF superimposing unit 95 superimposes the selected HRTF on the audio data for each channel of each sound source supplied from the directional filter superimposing unit 94. The HRTF superimposing unit 95 also generates output audio data for each channel by adding the audio data for the same channel after HRTF superimposition, and supplies the output audio data to the audio output unit 63 to play back the sound of the content. This output audio data is a binaural signal.

[0103] In the signal processing unit 75, a filtering process by the directional filter superimposing unit 94 and a process of superimposing an HRTF by the HRTF superimposing unit 95 are performed as rendering processes.

[0104] The input unit 76 consists of, for example, buttons, a mouse, a keyboard, a touch panel superimposed on the display unit 62, etc., and supplies signals according to operations by the listener (user) to the initialization unit 73, the signal processing unit 75, and the display control unit 77.

[0105] The display control unit 77 causes the display unit 62 to display various images in response to signals supplied from the input unit 76, etc.

[0106] <Design of Transition Curve> The design of the transition curve will be explained.

[0107] The transition curve is used to determine (select) the directional filter selection method and weighting coefficient, i.e., the gain value that becomes the directional filter value. The directional filter selection method (selection method), in other words, the type of directivity applied to the sound source, includes omnidirectional (omnidirectional method), head-centered directivity (head-centered directivity method), and ear-position-based directivity (ear-position-based directivity method).

[0108] The omnidirectional method is a selection method that does not apply a directional filter, assuming that the directivity (directional characteristics) of the sound source is omnidirectional, or more specifically, can be approximated to omnidirectionality. In other words, the omnidirectional method is a method that selects a directional filter with a value of 1.

[0109] The head-centered directivity method is a method of selecting a directional filter according to the directivity defined at the intersection of a line connecting the sound source center and the listener's head center with a virtual sphere on which the sound source directivity (directivity data) is defined. In other words, the head-centered directivity method is a method of selecting a directional filter based on the position of the listener's head center (using the position of the head center as a reference) assuming that the sound source directivity is head-centered directivity.

[0110] For example, a value of directivity data or a directional filter is associated with each position on the surface of a virtual sphere centered on the sound source. In other words, the directivity of a sound source based on the directivity data is represented by a virtual sphere centered on the sound source and values ​​at each position on the surface of the virtual sphere.

[0111] In the head-centered directivity method, a directional filter is selected that corresponds to a position in the direction of the center of the listener's head as viewed from the sound source (the center of the virtual sphere) among the positions on the surface of the virtual sphere.

[0112] In this case, the same directional filter is selected for the L channel and the R channel. In other words, the head-centered directivity is a single directional filter.

[0113] The ear-position-based directivity method is a method of selecting a directional filter according to the directivity defined at the intersection of two lines connecting the sound source center with the left and right ear positions of the listener, respectively, with a virtual sphere on which the sound source directivity (directivity data) is defined. In other words, the ear-position-based directivity method is a method of selecting a directional filter based on the left and right ear positions of the listener (using each ear position as a reference) assuming that the directivity of the sound source is the ear-position-based directivity.

[0114] In the ear-position-based directivity method, a directional filter is selected for each LR (left / right) channel, i.e., for each left and right ear. In other words, ear-position-based directivity is a stereo directional filter. For example, among the positions on the surface of the virtual sphere, the directional filter associated with the position in the direction of the listener's left ear as seen from the sound source (the center of the virtual sphere) is selected as the directional filter for the L channel.

[0115] In the transition curve, the directivity superimposed (applied) on the sound source is switched according to the distance (relative distance x) from the sound source to the listener in the virtual space, i.e., the selection method of the directional filter is switched.

[0116] Specifically, for example, as shown in FIG. 5, consider a case where a predetermined sound source SD11 and a listener U11 exist in a virtual space.

[0117] In this example, the area A that is sufficiently far from the sound source SD11, that is, the area A whose distance from the sound source SD11 is greater than a predetermined distance d1, is set as an omnidirectional (omnidirectional) area.

[0118] For example, when listener U11 is in area A, sound source SD11 is located far enough away from listener U11 that approximating the directivity of sound source SD11 to omnidirectionality makes almost no difference to the auditory sense when the sound from the sound source is reproduced. Therefore, when relative distance x is greater than distance d1, the omnidirectional method is selected. In other words, no directivity is superimposed on the sound from the sound source.

[0119] Area B that is some distance away from sound source SD11, that is, area B whose distance from sound source SD11 is equal to or less than a predetermined distance d2 and greater than a distance d3, is set as an area of ​​head-centered directivity (head-centered directivity method).

[0120] Moreover, the area where the distance from the sound source SD11 is equal to or less than the distance d1 and is greater than the distance d2 is defined as a transition area AB.

[0121] When the listener U11 is in region B or transition region AB, the sound source SD11 is located at a certain distance from the listener U11, so even if the directivity of the sound source SD11 is approximated to head-centered directivity, there is almost no audible difference when the sound from the sound source is reproduced.

[0122] That is, there is no significant difference in auditory sensation between the sound emitted from the sound source SD11 and reaching the right ear of the listener U11 as shown by the line L11, the sound emitted from the sound source SD11 and reaching the left ear of the listener U11 as shown by the line L12, and the sound emitted from the sound source SD11 and reaching the center of the listener U11's head as shown by the line L13. Therefore, when the relative distance x is greater than the distance d3 and equal to or less than the distance d1, the head-centered directivity method is selected. That is, the head-centered directivity is superimposed on the sound from the sound source.

[0123] When the listener U11 is in the transition region AB, a directional filter to be superimposed on the audio data of the sound source may be generated based on the directional filter selected by the omnidirectional method and the directional filter selected by the head-centered directivity method. Specifically, for example, the directional filter selected by each method (method) may be weighted and added according to the relative distance x to generate the directional filter to be finally used.

[0124] Area C including the position of the sound source SD11 and close to the sound source SD11, that is, area C whose distance from the sound source SD11 is equal to or less than a predetermined distance d4, is set as an area of ​​ear-position-based directivity (ear-position-based directivity method).

[0125] Moreover, the area whose distance from the sound source SD11 is equal to or less than distance d3 and greater than distance d4 is defined as a transition area BC.

[0126] When the listener U11 is in region C or transition region BC, the sound source SD11 is located close to the listener U11, so in order to reproduce the directionality of the sound source with sufficient accuracy, it is necessary to select appropriate directional filters for each ear.

[0127] Therefore, when the relative distance x is equal to or less than the distance d3, the ear-position-based directivity method is selected, i.e., the ear-position-based directivity is superimposed on the sound of the sound source.

[0128] When the listener U11 is in the transition region BC, a directional filter to be superimposed on the audio data of the sound source may be generated based on the directional filter selected by the ear-position-based directivity method and the directional filter selected by the head-centered directivity method. Specifically, for example, the directional filter selected by each method (method) may be weighted and added according to the relative distance x to generate the directional filter to be finally used.

[0129] As described above, when designing a transition curve, the area in the virtual space containing the sound source is divided into areas A to C according to the distance from the sound source, and transition areas AB and BC are provided between each area.

[0130] In this case, the start position (start point) and end position (end point) of each region, that is, the distances d1 to d4 that define each region, are calculated based on the parameters κ and p that represent the characteristics of the sound source.

[0131] For example, the transition curve design unit 92 calculates distances d1 to d4 to obtain the transition curve shown in Fig. 6. In Fig. 6, the horizontal axis indicates the distance from the sound source, and the vertical axis indicates the weighting coefficient (gain value) that becomes the value of the directional filter at each distance.

[0132] In the transition curve L21 shown in FIG. 6, the value on the vertical axis for each distance shown on the horizontal axis is the gain value of each frequency bin that is multiplied by the sound source as a directional filter.

[0133] For example, in area A, which is considered to be omnidirectional, the weighting coefficient is set to 1. In other words, since a directional filter is not used in area A, the weighting coefficient is essentially not defined. In area B, which is considered to be head-centered directivity, the weighting coefficient value is set to a constant value Dir_C, and in the transition area AB between areas A and B, the weighting coefficient value changes smoothly from Dir_C to 1 as the distance increases. The weighting coefficient Dir_C is the value of the directional filter selected by the head-centered directivity method.

[0134] In region C, which is the ear-position-based directivity, the weighting coefficients for the L and R channels are constant values, Dir_L and Dir_R, while in the transition region BC between regions B and C, the weighting coefficient values ​​change smoothly from Dir_L and Dir_R to Dir_C as the distance increases. In ear-position-based directivity (ear-position-based directivity method), weighting coefficients are determined for each of the L and R channels. The weighting coefficients Dir_L and Dir_R are the directional filter values ​​selected by the ear-position-based directivity method.

[0135] For example, the value of the weighting factor in each transition region is calculated by the transition curve design unit 92 based on distances d1 to d4, which are the start and end positions of each region.

[0136] For example, the transition curve design unit 92 calculates, as transition curve parameters, a group of parameters including distances d1 to d4 and parameters such as coefficients of functions for calculating weighting coefficients in each transition region.

[0137] In the decoder 61, a transition curve in which the weighting coefficients change continuously as described above is used, and a directional filter weighted according to the distance (relative distance x) between the sound source and the listener is superimposed on the audio data of each sound source.

[0138] In this way, audio reproduction can be realized that provides a sense of continuity in the sound received when the listener approaches or moves away from the sound source in the virtual space. In other words, by determining an appropriate weighting coefficient for each region (distance), the sense of continuity of the sound positioning, volume, etc. can be maintained, and the applied directivity can be switched appropriately without causing any sense of discomfort.

[0139] Furthermore, by switching the selection method of the directional filter for each region (distance) based on the transition curve, it is possible to reproduce the directivity of the sound source with fewer calculations and realize realistic audio reproduction.

[0140] For example, when the listener is sufficiently far from the sound source or when the sound source (audio object) is an omnidirectional object such as a dodecahedron speaker, the process of superimposing a directional filter (filter process) can be omitted. This reduces the computational resources, i.e., the amount of computation (processing amount). For example, when the sound source is omnidirectional, the parameter κ is small, and therefore the value of the distance d1 is also small.

[0141] Similarly, when the listener is at a certain distance from the sound source, by superimposing head-centered directivity on the sound source, it is possible to maintain the same level of directivity reproduction accuracy as in the case of ear-position-based directivity, while reducing the amount of calculation compared to the case of ear-position-based directivity.

[0142] Furthermore, for example, if the decoder 61 is a device with limited computing resources, such as a smartphone, the start distance of the directional processing, i.e., the distance d1, may be adjusted according to the detection result of the type of device, etc. This makes it possible to reduce processing costs according to the device.

[0143] Similarly, the starting distance (distance d1) of the directivity processing may be automatically adjusted depending on the remaining battery power of the device in which the decoder 61 is installed.

[0144] Here, a more specific example of the design of the transition curve will be described.

[0145] For example, it is assumed that the range of values ​​that the parameter κ can take is [0, 10], the range of values ​​that the parameter p can take is [0, 10], and the maximum value that the distance d1 can take is the maximum distance dm [m].

[0146] The listener may input the maximum distance dm by operating the input unit 76 on a UI (User Interface) displayed on the display unit 62 by the display control unit 77. In such a case, a signal indicating the maximum distance dm input by the listener's operation is supplied from the input unit 76 to the initialization unit 73.

[0147] The transition curve design unit 92 calculates the distance d1 by calculating the following equation (1) based on the parameter κ, the parameter p, and the maximum distance dm.

[0148]

[0149] In equation (1), max(p) is the maximum value that parameter p can take, and max(κ) is the maximum value that parameter κ can take. In this example, the larger the parameter κ is, the sharper the directionality of the sound source is, and the larger the parameter p is, the higher the importance of the sound source is, the larger the distance d1 becomes.

[0150] Also, for example, the maximum value of the width of the region from the position at distance d1 to the position at distance d2, i.e., the width of the transition region AB, is defined as dm2. The maximum value dm2 is defined as a value equal to or less than the distance d1. Note that the listener may input the maximum value dm2 within the range equal to or less than the distance d1 by operating the input unit 76 on the UI displayed on the display unit 62 by the display control unit 77, and a signal indicating the maximum value dm2 input by this operation may be supplied from the input unit 76 to the initialization unit 73.

[0151] The transition curve design unit 92 calculates the distance d2 by calculating the following equation (2) based on the parameter κ, the distance d1, and the maximum value dm2.

[0152]

[0153] The convergence angle at which the difference in sound after superimposition of a directional filter based on the ear position, i.e., a directional filter selected using the ear-position-based directivity method, and a directional filter based on the head center position, i.e., a directional filter selected using the head-centered directivity method, is perceptible is assumed to be 4 degrees.

[0154] The convergence angle here is the angle between a line connecting the center of the sound source and the listener's right ear and a line connecting the center of the sound source and the listener's left ear. For example, the angle between line L11 and line L12 shown in Figure 5 is the convergence angle.

[0155] In addition, the listener may input a convergence angle by operating the input unit 76 in response to the UI displayed on the display unit 62 by the display control unit 77, and a signal indicating the convergence angle input by this operation may be supplied from the input unit 76 to the initialization unit 73.

[0156] If the convergence angle is 4 degrees, and the distance from the center of the listener's head to the entrance of the ear canal is assumed to be 0.065 m, which is the average distance for an adult male, the distance between the sound source and the listener is 1.86 m, so the transition curve design unit 92 sets the distance d3 to 1.86 m. Note that if the distances d1 and d2 are 0 due to the parameters κ and p, the distances d3 and d4 are also set to 0.

[0157] Furthermore, the convergence angle at which directivity based on the ear position (ear-position-based directivity) becomes important perceptually, that is, the convergence angle at which the size of the listener's head cannot be ignored in terms of auditory sensation, is assumed to be 10 degrees.

[0158] Note that the listener may operate the input unit 76 to input the convergence angle for which the ear-position reference directivity is important to the UI displayed on the display unit 62 by the display control unit 77, and a signal indicating the convergence angle input by this operation may be supplied from the input unit 76 to the initialization unit 73.

[0159] If the convergence angle at which ear-position-based directivity becomes important is 10 degrees, and the distance from the center of the listener's head to the entrance of the ear canal is assumed to be 0.065 m, which is the average value for an adult male, the distance between the sound source and the listener is 0.74 m, so the transition curve design unit 92 sets the distance d4 to 0.74 m.

[0160] By the above calculations, distances d1 to d4 are found from parameters κ and p, and regions A, B, C, transition region AB, and transition region BC are determined.

[0161] The transition curve design unit 92 calculates (designs) the transition curve, that is, the weighting coefficient in each region, based on the distances d1 to d4, according to the following equation (3).

[0162]

[0163] In equation (3), x represents the distance between the sound source and the listener (relative distance x), and y represents a weighting coefficient.

[0164] When rendering processing is performed in the signal processing unit 75, i.e., when directional filter superposition processing is performed in the directional filter superposition unit 94, the relative distance x is input and the weighting coefficient that will become the final directional filter is obtained by the case classification shown in equation (3).

[0165] By using equation (3) like this, for example, a transition curve L21 shown in Fig. 7 can be obtained. The weighting coefficient y determined by equation (3) will be explained below with reference to Fig. 7. Note that in Fig. 7, parts corresponding to those in Fig. 6 are given the same reference numerals, and their explanation will be omitted as appropriate.

[0166] 7, the weighting coefficient y is set to 1.0 in the omnidirectional area A, and is constant as the gain value Dir_C of each frequency bin of the directional filter selected based on the sound source position and listener position in area B. In area C, the weighting coefficient y of the L channel is set to the gain value Dir_L of each frequency bin of the directional filter selected based on the sound source position and listener position, and the weighting coefficient y of the R channel is set to the gain value Dir_R of each frequency bin of the directional filter selected in a similar manner.

[0167] For example, the values ​​Dir_C, Dir_L, and Dir_R of the weighting coefficient y may be predetermined values ​​or may be values ​​calculated based on the distance d1 or the like.

[0168] In the transition region AB, the weighting coefficient y is set to y1. In equation (3), the weighting coefficient y=y1 is calculated based on the weighting coefficient Dir_C, the distance d1, the distance d2, and the relative distance x.

[0169] In the transition region BC, the weighting coefficient y of the L channel is set to y2, and the weighting coefficient y of the R channel is set to y3. In equation (3), the weighting coefficients y=y2 and y=y3 are calculated based on the weighting coefficient Dir_C, the weighting coefficients Dir_L and Dir_R, the distances d3 and d4, and the relative distance x.

[0170] <Selection of Directional Filter> Selection of directional filters in the head-centered directivity method and the ear-position-based directivity method will be described.

[0171] For example, when the distance between the sound source and the listener in the virtual space is sufficiently close, specifically when the relative distance x is equal to or less than distance d3 and the listener is in region C or transition region BC, directional filter selection is performed based on the ear position, i.e., selection is performed using the ear position-based directivity method.

[0172] At this time, it is necessary to convert coordinates from the center position of the listener's head to the ear position reference, and this conversion is performed using the distance b from the center position of the listener's head to one ear.

[0173] A specific method for selecting a directional filter will be described below with reference to FIG.

[0174] 8, position o indicates the central position of the sound source, and a directional filter is defined by a unit sphere S centered at this position o. That is, the unit sphere S corresponds to the above-mentioned virtual sphere, and a directional filter for each direction as seen from position o is associated with each position on the surface of the unit sphere S.

[0175] Position h indicates the center position of the listener's head, and position e indicates the position of the listener's left ear, more specifically, the position of the entrance to the ear canal.

[0176] In this example, the distance from position o to position h is the distance a [m] from the sound source (sound source center) to the listener (head center), and the distance from position h to position e is the distance b [m] from the center of the listener's head to their ear. This distance b can also be said to be the radius of the listener's head.

[0177] Now, let us represent positions on the surface of the unit sphere S using coordinates (polar coordinates) in a polar coordinate system with position o as the origin. In this case, positions on the surface of the unit sphere S are represented by an azimuth angle indicating the position in the left-right direction and an elevation angle indicating the position in the up-down direction, with the direction in which the sound source faces as viewed from position o being the front direction (azimuth angle 0 degrees, elevation angle 0 degrees). The front direction of the sound source is the direction in which the sound source faces (sound source orientation) indicated by the sound source position / orientation information.

[0178] For example, let the position q be the intersection of the line segment oh from the center position o of the unit sphere S to the center position h of the listener's head with the unit sphere S, and let the coordinates of position q be expressed as (θ, φ) using the azimuth angle θ and the elevation angle φ.

[0179] In this case, the directional filter superimposing unit 94 selects the directional filter associated with the position q indicated by the coordinates (θ, φ) in the head-centered directivity method as the directional filter to be superimposed on the audio data of the sound source. The value of the directional filter at this time is Dir_C.

[0180] Furthermore, for example, the position of the intersection between the line segment oe from the center position o of the unit sphere S to the position e of the listener's left ear and the unit sphere S is defined as position q', and the coordinates of position q' are expressed as (θ', φ') using the azimuth angle θ' and the elevation angle φ'.

[0181] In this case, the directional filter superimposing unit 94 selects the directional filter associated with the position q' indicated by the coordinates (θ', φ') in the ear-position-based directivity method as the directional filter for the left ear, i.e., the L channel, to be superimposed on the audio data of the sound source. The value of the directional filter at this time is Dir_L.

[0182] The coordinates (θ′, φ′) of the position q′ can be obtained as follows.

[0183] That is, if the angle formed by the line segments oh and oe is defined as angle α, the angle α can be obtained by calculating the following equation (4) based on the distances a and b.

[0184]

[0185] Furthermore, by calculating the following equation (5) based on the angle α and the coordinates (θ, φ) of the position q, the coordinates (θ', φ') of the position q' can be obtained.

[0186]

[0187] The position o of the center of the unit sphere S is the position indicated by the sound source position / direction information, and the position h of the center of the listener's head is the position indicated by the listener position / direction information.

[0188] Furthermore, for example, the distance b indicating the head radius may be a predetermined value, a value specified by the listener or creator, or a value measured by a sensor or the like possessed by a device such as an HMD in which the decoder 61 is provided.

[0189] The directional filter superimposition unit 94 determines the coordinates (θ, φ) of position q and the coordinates (θ', φ') of position q' based on the sound source position direction information, the listener position direction information, and the distance b, and can select an appropriate directional filter depending on the directional filter selection method.

[0190] For example, by superimposing a directional filter associated with position q' on the unit sphere S on the audio data of a sound source, it is possible to reproduce the sound that would be emitted from a directional sound source at position o and heard by the left ear of a listener. In other words, it is possible to reproduce the directivity of the sound source at position o.

[0191] <Description of Output Audio Data Generation Process> Next, a description will be given of the operation of the decoder 61. That is, the output audio data generation process by the decoder 61 will be described below with reference to the flowchart of FIG.

[0192] In step S 51 , the communication unit 71 receives the bit stream transmitted from the encoder 11 and supplies it to the demultiplexing unit 72 .

[0193] In step S52, the demultiplexing unit 72 demultiplexes the bit stream supplied from the communication unit 71, supplies the encoding parameters extracted by the demultiplexing to the parameter decoding unit 91, and supplies the encoded audio data to the audio decoding unit 74.

[0194] In step S 53 , the parameter decoding unit 91 decodes the encoding parameters supplied from the demultiplexing unit 72 , and supplies the resulting parameters κ and p of each excitation signal to the transition curve design unit 92 .

[0195] In step S54, the transition curve design unit 92 designs a transition curve for each sound source based on the parameters κ and p supplied from the parameter decoding unit 91, and supplies the resulting transition curve parameters to the directional filter superimposition unit 94.

[0196] For example, the transition curve design unit 92 calculates the above-mentioned formula (1) or formula (2) based on the parameters κ and p, the maximum distance dm, the maximum value dm2, and so on, to find the distances d1 to d4.

[0197] Furthermore, for example, the transition curve design unit 92 calculates weighting coefficients y1 to y3 (functions of y1 to y3) by appropriately calculating Equation (3) based on the distances d1 to d4, and defines a group of parameters including these weighting coefficients and the distances d1 to d4 as transition curve parameters. Note that the weighting coefficients y1 to y3 may be calculated by the directional filter superposition unit 94.

[0198] The processes of steps S53 and S54 do not need to be performed for each frame of audio data, but may be performed immediately after the start of playback of the content or at the timing when encoding parameters of a new sound source are transmitted from the encoder 11. In other words, once the transition curve parameters of an sound source are calculated, the calculated transition curve parameters will continue to be used unless the parameters κ and p of that sound source are updated.

[0199] In step S55 , the audio decoding unit 74 decodes the encoded audio data supplied from the demultiplexing unit 72 , and supplies the resulting audio data of each sound source to the directional filter superimposing unit 94 .

[0200] In step S56, the relative distance calculation unit 93 calculates the relative distance x from the sound source to the listener in the virtual space for each sound source based on the supplied sound source position / direction information and listener position / direction information for each sound source, and supplies the calculated distance x to the directional filter superimposition unit 94.

[0201] In step S57, the directional filter superimposing unit 94 determines whether or not omnidirectional approximation is possible based on the relative distance x supplied from the relative distance calculation unit 93 and the transition curve parameters supplied from the transition curve design unit 92.

[0202] Specifically, for example, the directional filter superimposition unit 94 determines that omnidirectional approximation is possible when the relative distance x is greater than the distance d1 included in the transition curve parameter, i.e., when the listener is within area A with respect to the sound source.

[0203] The process of determining whether omnidirectional approximation is possible is a process of determining whether the directional filter selection method should be the omnidirectional method. Therefore, the process of step S57 can also be said to be a process of selecting a directional filter (omnidirectional).

[0204] More specifically, the processes from step S57 to step S61 are performed for each frame of audio data for each sound source.

[0205] If it is determined in step S57 that non-directional approximation is not possible, that is, if the relative distance x is equal to or less than the distance d1, then the process proceeds to step S58.

[0206] In step S58, the directional filter superimposing unit 94 selects a directional filter to be applied to the sound source based on the supplied sound source position / direction information and listener position / direction information, the relative distance x, and the transition curve parameter.

[0207] For example, when the relative distance x is greater than the distance d3 and is equal to or less than the distance d1, the directional filter superimposing unit 94 selects the head-centered directivity method as the directional filter selection method.

[0208] Then, the directional filter superimposing unit 94 calculates the coordinates (θ, φ) of the position q shown in Fig. 8 based on the sound source position direction information and the listener position direction information, and selects the directional filter associated with the position q as the directional filter to be superimposed on the audio data of the sound source. As a result, one directional filter common to all channels, that is, the L and R channels, is selected. The directional filter selected in this case is the weighting coefficient Dir_C.

[0209] On the other hand, when the relative distance x is equal to or less than the distance d3, the directional filter superimposing unit 94 selects the ear-position-based directivity method as the directional filter selection method.

[0210] The directional filter superimposing unit 94 then calculates the coordinates (θ', φ') of position q' shown in FIG. 8 based on the sound source position direction information, the listener position direction information, and the distance b indicating the radius of the listener's head, and selects the directional filter associated with position q' as the directional filter for the L channel to be superimposed on the audio data of the sound source. The directional filter superimposing unit 94 also selects a directional filter for the R channel in the same manner as for the L channel. This results in directional filters being selected for each of multiple channels, i.e., for each of the L and R channels. The directional filters selected in this case have weighting coefficients Dir_L and Dir_R.

[0211] In step S58, a directional filter is selected using a transition curve parameter (transition curve) generated based on the parameter κ and the parameter p. Therefore, in step S58, a directional filter is selected according to the positional relationship (relationship between positions and orientations) between the sound source and the listener in the virtual space, using a selection method (selection system) determined from the relative distance x from the sound source to the listener and one or more parameters representing the characteristics of the sound source.

[0212] Furthermore, the directional filter superimposing unit 94 selects (determines) a final directional filter based on the transition curve, based on the directional filter selected by the selection method (selection system), the relative distance x, and the transition curve parameter. In other words, a weighting coefficient according to the relative distance x, which is determined by the transition curve parameter, i.e., the parameter representing the characteristics of the sound source, and the weighting coefficient of the directional filter selected by the selection method (selection system), such as the weighting coefficient Dir_C, is selected as the final directional filter.

[0213] For example, when a directional filter is selected using the head-centered directivity method, the directional filter superimposing unit 94 uses the weighting coefficient Dir_C or y1 as the final directional filter using the above-mentioned equation (3).

[0214] Specifically, when the relative distance x is greater than the distance d2 and is equal to or less than the distance d1, the weighting factor y1 is used as the final directional filter, and when the relative distance x is greater than the distance d3 and is equal to or less than the distance d2, the weighting factor Dir_C is used as the final directional filter. More specifically, the weighting factor y1 is calculated based on the relative distance x.

[0215] Thus, in the head-centered directivity method, the directional filter superimposing unit 94 selects the weighting coefficient Dir_C or y1 as the final directional filter.

[0216] Furthermore, when a directional filter is selected using the ear-position-based directivity method, the directional filter superimposing unit 94 uses Dir_L and Dir_R, or y2 and y3, as weighting factors to form the final directional filter using the above-mentioned equation (3). At this time, the weighting factor Dir_C is also used as necessary.

[0217] Specifically, when the relative distance x is greater than the distance d4 and is equal to or less than the distance d3, the weighting factors y2 and y3 are selected, and when the relative distance x is equal to or less than the distance d4, the weighting factors Dir_L and Dir_R are selected. More specifically, the weighting factors y2 and y3 are calculated based on the relative distance x.

[0218] Thus, in the ear-position-based directivity method, the directional filter superimposing unit 94 uses the weighting factor Dir_L or y2 as the final directional filter for the L channel, and the weighting factor Dir_R or y3 as the final directional filter for the R channel.

[0219] In step S59, the directional filter superimposing unit 94 superimposes the final directional filter selected in step S58 on the audio data of the sound source supplied from the audio decoding unit 74, and supplies the resulting audio data to which directivity has been added (reproduced) to the HRTF superimposing unit 95. That is, the audio data of the sound source is multiplied by a weighting coefficient as the directional filter, thereby superimposing the directional filter on the audio data. More specifically, the selection and superimposition (multiplication) of the directional filter are performed for each frequency bin.

[0220] For example, when a directional filter is selected using the head-centered directivity method, the directional filter superimposing unit 94 performs filtering processing to superimpose the finally selected directional filter, i.e., the weighting coefficient Dir_C or y1, on the audio data of the sound source. The audio data obtained by the filtering processing is then used as the L-channel and R-channel audio data for the sound source.

[0221] In this case, the audio data for the L channel and R channel will be the same, and audio data for both the L channel and the R channel can be obtained for each sound source with a single filter process. Therefore, compared to performing filter processing for each channel, it is possible to reproduce directivity with sufficient accuracy with a smaller amount of calculation.

[0222] Furthermore, for example, when a directional filter is selected using the ear-position-based directivity method, the directional filter superimposing unit 94 performs filtering for each channel.

[0223] That is, the directional filter superimposing unit 94 generates audio data for the L channel of the sound source by superimposing the weighting factor Dir_L or y2, which is the directional filter for the L channel finally selected, on the audio data of the sound source. Similarly, the directional filter superimposing unit 94 generates audio data for the R channel of the sound source by superimposing the weighting factor Dir_R or y3, which is the directional filter for the R channel finally selected, on the audio data of the sound source.

[0224] When the process of step S59 is performed and the audio data of each channel of the sound source is supplied to the HRTF superimposing unit 95, the process then proceeds to step S60.

[0225] On the other hand, if it is determined in step S57 that non-directional approximation is possible, that is, if the relative distance x is greater than the distance d1, the processes of steps S58 and S59 are not performed, and the process proceeds to step S60.

[0226] In this case, the directional filter superimposing unit 94 does not select a directional filter or superimpose the directional filter on the audio data. The directional filter superimposing unit 94 supplies the audio data of the sound source supplied from the audio decoding unit 74 to the HRTF superimposing unit 95 as L-channel and R-channel audio data for the sound source.

[0227] In this case, a directional filter with a value of 1 is essentially selected using the omnidirectional method, and the directional filter is superimposed on the audio data of the sound source. However, since the selection of a directional filter and the superimposition of a directional filter are not actually performed, the directivity (omnidirectionality) of the sound source can be reproduced with a small amount of calculation.

[0228] If it is determined in step S57 that non-directional approximation is possible or if the process of step S59 is performed, the process of step S60 is then performed. As described above, the processes of steps S57 to S59 are performed for each sound source, and therefore the method of selecting a directional filter and the selected directional filter (weighting coefficient) differ depending on the sound source.

[0229] In step S60, the HRTF superimposing unit 95 selects an HRTF for each channel for each sound source based on the supplied sound source position / direction information and listener position / direction information.

[0230] For example, in the HRTF superimposing unit 95, HRTFs are prepared for each of the L and R channels for multiple directions as seen by the listener, and the HRTF superimposing unit 95 selects, for each channel, an HRTF for the direction of the sound source as seen by the listener in the virtual space.

[0231] In step S61, the HRTF superimposing unit 95 superimposes the HRTF for each channel selected in step S60 on the audio data for each channel of the sound source for each of the L and R channels.

[0232] The HRTF superimposing unit 95 also generates output audio data by adding together the audio data of the same channel from the audio data for each channel after HRTF superimposition obtained for each of the multiple sound sources, and supplies this output audio data to the audio output unit 63.

[0233] The audio output unit 63 reproduces the sound of the content, i.e., the sound of each sound source, based on the output audio data supplied from the HRTF superimposition unit 95. In this case, the display control unit 77 may also cause the display unit 62 to display a video of the content or an image of a virtual space in which the sound sources and listeners are located. Once the content has been reproduced, the output audio data generation process ends.

[0234] In this way, the decoder 61 designs a transition curve based on parameters that represent the characteristics of the sound source, and generates output audio data of the content based on a directional filter determined from the transition curve.

[0235] In particular, the decoder 61 selects a weighting coefficient according to the transition curve as a directional filter, thereby reproducing the directivity of the sound source with fewer calculations, thereby realizing audio reproduction that is natural and has a sense of realism.

[0236] Second Embodiment Use of Multiple Transition Curves In the above, an example has been described in which one transition curve is designed for one sound source using the sharpness parameter κ that is representative of the directivity.

[0237] However, if the directivity data of one sound source has multiple sharp directivities, that is, if the directivity of the sound source has multiple peaks (hereinafter also referred to as directional peaks), a transition curve based on the parameter κ of the directional peak that is closest to the listener may be used.

[0238] In such a case, a transition curve may be designed and stored in advance for each directivity peak, or a transition curve may be designed each time as required.

[0239] For example, as shown in Fig. 10, assume that directivity data of a given sound source is defined on the surface of a virtual sphere SP11. That is, assume that the directivity of the sound source is represented by the virtual sphere SP11. Also assume that there are two directivity peaks, a directivity peak pk1 and a directivity peak pk2, on the surface of the virtual sphere SP11.

[0240] In this example, the parameter κ is obtained from the directivity data of the sound source, with a parameter κ1 indicating the sharpness of the directivity peak pk1 and a parameter κ2 indicating the sharpness of the directivity peak pk2.

[0241] When performing rendering processing, i.e., superimposing directional filters, on a sound source having such multiple directional peaks, it is possible to design a transition curve using the parameter κ of the directional peak pk1 or pk2 that is closest in distance from the extension of the listener's line of sight in the virtual space.

[0242] Here, a listener U31 is present in a virtual space, and of the directivity peaks pk1 and pk2, the directivity peak closest to the line of sight of the listener U31, i.e., the line of sight passing through the viewpoint, is the directivity peak pk1. Therefore, the transition curve is designed using the directivity peak pk1, and a directional filter (weighting coefficient) determined by the transition curve is used.

[0243] When transmitting the parameter κ for each of the multiple directional peaks of the sound source to the decoder 61, peak position information indicating the position of the directional peak on the surface of the virtual sphere (unit sphere S) is required. In other words, the peak position information can be said to be directional information indicating the direction of the position of the directional peak as viewed from the center of the sound source.

[0244] The peak position information can be, for example, polar coordinates consisting of an azimuth angle and an elevation angle indicating the position of a directional peak on the surface of a virtual sphere (unit sphere S) centered at the center position of the sound source in a polar coordinate system with the front direction of the sound source as the reference and the center of the sound source as the origin. The position of each directional peak on the virtual sphere can be identified from both the directional data after modeling and the original directional data before modeling.

[0245] When the directivity data of the sound source has a plurality of directivity peaks, the parameter generating unit 21 of the encoder 11 generates a parameter κ for each directivity peak from the directivity data, generates peak position information indicating the position of the directivity peak, and supplies the generated information to the parameter encoding unit 22. Note that peak position information may also be generated when the sound source has one directivity peak.

[0246] Therefore, the parameter encoding unit 22 encodes at least the parameter κ, the peak position information, and the parameter p, and the resulting encoded parameters are multiplexed by the multiplexing unit 24 and stored in a bit stream.

[0247] In the decoder 61, the parameter κ, peak position information, and parameter p are obtained by decoding the encoded parameters in the parameter decoding unit 91. At this time, the parameter κ and peak position information are obtained for each directivity peak of the sound source.

[0248] The transition curve design unit 92 designs a transition curve for each directional peak, for example, and supplies the resulting transition curve parameters and peak position information for each directional peak to the directional filter superimposing unit 94 .

[0249] The directional filter superimposing unit 94 identifies a directional peak that is closest to the extension of the listener's line of sight in the virtual space based on the peak position information, the sound source position / direction information, and the listener position / direction information.The directional filter superimposing unit 94 then selects a directional filter in step S58 of FIG. 9 using the transition curve of the identified directional peak, i.e., the transition curve parameter generated based on the parameter κ of the identified directional peak.

[0250] In this case, the extension of the listener's line of sight in the virtual space, i.e., the straight line of the listener's line of sight that passes through the viewpoint of the upper listener, can be obtained from the listener position / direction information. Also, the position of the directivity peak in the virtual space can be obtained from the sound source position / direction information and the peak position information.

[0251] The directivity peak closest to the extension of the listener's line of sight may be specified (selected) by the transition curve design unit 92. In such a case, the transition curve design unit 92 supplies the transition curve parameters of the selected directivity peak to the directional filter superimposition unit 94.

[0252] Furthermore, each time a directivity peak located closest to the extension of the listener's line of sight is selected, transition curve parameters for the selected directivity peak may be generated.

[0253] Furthermore, when a sound source has multiple directional peaks, the parameter generation unit 21 may select one representative directional peak from among the directional peaks, such as the directional peak with the largest value or the directional peak with the largest parameter κ. In this case, the parameter κ of the selected directional peak is coded and transmitted, as in the first embodiment.

[0254] When model parameters for obtaining modeled directivity data are transmitted to the decoder 61, the transition curve design unit 92 may generate peak position information for each directivity peak based on the modeled directivity data obtained from the model parameters. Similarly, when directivity data can be obtained on the decoder 61 side, the peak position information may be generated from the directivity data.

[0255] Additionally, if the model parameters include a parameter indicating the position of a directional peak, that parameter may be used as peak position information.

[0256] Third Embodiment Acquisition of Head Radius As described with reference to FIG. 8, in the ear-position-based directivity method, the distance b indicating the radius of the listener's head is used when selecting a directional filter.

[0257] However, since the radius of the listener's head, i.e., the distance b from the center of the head to one ear, is a parameter that varies from person to person, in order to use a more accurate distance b, it is necessary to obtain or estimate the distance b in some way.

[0258] For example, if the decoder 61 is provided in an HMD, a more accurate distance b can be obtained by utilizing the adjustment mechanism of the band part, which is a mounting device provided in the HMD for mounting the HMD on the head of the user (wearer).

[0259] Specifically, for example, it is assumed that a decoder 61 is provided in the HMD 201 shown in FIG.

[0260] 11, the portion indicated by arrow Q51 shows the HMD 201 as seen from the side, and the portions indicated by arrows Q52 and Q53 show the HMD 201 as seen from above. In particular, a simplified structure of the HMD 201 is shown here to make the drawing easier to understand.

[0261] The HMD 201 has a display unit 211 having the display unit 62 shown in FIG. 4, and a band unit 212 for attaching the HMD 201 to the head of the listener.

[0262] The band portion 212 is a band-shaped component for fixing the HMD 201 to the listener's head, and the HMD 201 is fixed to the listener's head by tightening the band portion 212, more specifically, the band of the band portion 212 that is provided along the back of the listener's head.

[0263] The band section 212 is also provided with an adjustment mechanism 213 for fitting, that is, an adjustment mechanism 213 for adjusting the length of the band section 212 (band) so that the HMD 201 fits the listener's head.

[0264] For example, the decoder 61 uses the adjustment mechanism 213 to estimate the distance b from the length of the band, that is, the extent to which the band is stretched.

[0265] In this example, as shown by arrow Q52, the initial state is the state in which the length of the band at the adjustment mechanism 213 is the shortest.

[0266] The length of the entire band portion 212 in the initial state, i.e., the circumferential length along the wearer's head passing through the left and right ears (the length of one arc around the head) which can be determined from the length of the band, is known, and this arc length will also be referred to as the initial position length.

[0267] The adjustment mechanism 213 is also provided with a memory (not shown), which allows the adjustment mechanism 213 to measure the stretch width of the band from its initial state, i.e., the length to which the band has been stretched (hereinafter also referred to as the stretch distance).

[0268] For example, when the band length is adjusted from the initial state by the adjustment mechanism 213, the band portion is stretched as shown by the arrow Q53. The length of the stretched band portion at this time is the stretched distance [m].

[0269] When the adjustment mechanism 213 is adjusted appropriately and the HMD 201 fits the listener's head, the adjustment mechanism 213 refers to the memory and acquires (measures) the extension distance [m].

[0270] Next, the adjustment mechanism 213 adds the acquired extension distance [m] to the initial position length [m] stored in advance, and as a result, the length of the arc z [m] along the listener's head passing through the left and right ears (the length of the arc once around the head) after the band length has been adjusted by the adjustment mechanism 213 is calculated.

[0271] Furthermore, the listener's head is approximated as a sphere based on the arc length z [m] of the head, and the distance b [m] indicating the radius of the listener's head is obtained by calculating the following equation (6).

[0272]

[0273] The calculation to estimate the distance b may be performed within the adjustment mechanism 213 or in a block other than the adjustment mechanism 213 of the HMD 201 and supplied to the directional filter superimposition unit 94, or may be performed in the directional filter superimposition unit 94 based on the extension distance [m], etc.

[0274] The directional filter superimposing unit 94 selects a directional filter in the ear position-based directivity method using the distance b indicating the head radius obtained as described above.

[0275] In particular, in this example, the arc length z when the HMD 201 is worn on the listener's head is obtained from the extension distance of the band portion, i.e., the amount of change in the length of the band portion, and the initial position length of the band portion, and the distance b is calculated from the arc length z. Then, the distance b is used to select a directional filter.

[0276] Therefore, it can be said that the directional filter superimposing unit 94 selects a directional filter using the ear-position-based directivity method based on the length of the band portion that is the wearing device of the HMD 201 .

[0277] <Fourth Embodiment> <Regarding GUI> Various pieces of information such as the distance b indicating the head radius used in the decoder 61 and transition curve parameters may be specified using a GUI (Graphical User Interface) displayed on the display unit 62.

[0278] For example, when a listener (user) inputs a distance b indicating the head radius using a GUI, the display control unit 77 displays an input screen 241 shown in Fig. 12 as a GUI for input on the display unit 62. The input screen 241 has an input field 242, and the listener inputs his or her own distance b by operating the input unit 76.

[0279] Then, a signal indicative of the input distance b, which corresponds to the operation of the input unit 76 by the listener, is supplied from the input unit 76 to the signal processing unit 75. In other words, the signal processing unit 75 (directional filter superimposing unit 94) acquires the signal indicative of the distance b from the input unit 76 in response to the input operation.

[0280] The directional filter superimposing unit 94 uses the distance b indicated by the signal from the input unit 76 to select a directional filter according to the ear-position-based directivity method.

[0281] Furthermore, the listener (user) may be allowed to input transition curve parameters or adjust (change) the transition curve parameters determined by the transition curve design unit 92 .

[0282] In such a case, for example, the display control unit 77 causes the display unit 62 to display an input screen 271 shown in Fig. 13 as a GUI for inputting (specifying) transition curve parameters. This input screen 271 is a screen for specifying values ​​of the distance d1 and the like, which are parameters constituting the transition curve parameters, or for changing the values ​​of the distance d1, etc. In Fig. 13, parts corresponding to those in Fig. 7 are assigned the same reference numerals, and descriptions thereof will be omitted as appropriate.

[0283] The input screen 271 displays a transition curve similar to that shown in Figure 7, and pointers PT11 to PT14 are provided on the graph of the transition curve for specifying (inputting) the transition curve parameters, more specifically, the values ​​of the distances d1 to d4 that make up the transition curve parameters.

[0284] For example, when attention is paid to the pointer PT13 for designating the distance d2, the pointer PT13 is displayed at the position of the distance d2 on the horizontal axis of the transition curve graph. For example, the initial position of the pointer PT13 may be a predetermined position, or may be the position of the distance d2 calculated by the transition curve design unit 92.

[0285] The listener (user) can specify the value of the distance d2 by operating a mouse, touch panel, keyboard, or the like as the input unit 76 to move (slide) the position of the pointer PT13 on the input screen 271 by an arbitrary distance in any direction to the left or right. In other words, the listener can input the distance d2 by moving the pointer PT13.

[0286] In this example, the listener can freely design the transition curve by operating the pointers PT11 to PT14. Note that the listener may also be allowed to freely specify the value of the vertical axis of the transition curve at each distance, i.e., the value of the weighting coefficient.

[0287] When the listener operates the input unit 76 to move the position of the pointer PT13 or the like, a signal corresponding to the operation is supplied from the input unit 76 to the initialization unit 73.

[0288] Then, the transition curve design unit 92 generates or changes transition curve parameters, more specifically, parameters such as the distance d2 that constitute the transition curve parameters, in response to the signal supplied from the input unit 76.

[0289] That is, for example, when designing a transition curve, transition curve parameters are generated based on the values ​​of parameters such as the specified distance d2, and after the transition curve has been designed, the values ​​of the parameters that make up the retained transition curve parameters are changed to the specified values.

[0290] Specifically, for example, when the position of pointer PT13 is changed on the input screen 271, the transition curve design unit 92 changes the value of the distance d2 that constitutes the transition curve parameter to the value indicated by the changed position of pointer PT13.

[0291] Fifth Embodiment <Regarding Switching of Processing Depending on Device> Based on at least one of the type of device (device type) in which the decoder 61 is provided, the remaining battery charge, and the calculation resources, it is possible to generate transition curve parameters, correct the transition curve parameters, and determine whether omnidirectional approximation is possible in step S57 of FIG. 9 .

[0292] Specifically, for example, during or after the transition curve design, the transition curve design unit 92 may determine (correct) the distance at which directional processing begins, i.e., the value of the distance d1 at which the process of superimposing a directional filter begins, based on the type of device in which the decoder 61 is installed.

[0293] For example, if the device provided with the decoder 61 is a type of device that tends to be relatively short on computing resources, such as a smartphone, the distance d1 as the transition curve parameter may be corrected to be smaller. This reduces the number of sound sources on which directional filters are superimposed, and reduces the processing load of the rendering process, i.e., the directional filter superimposition process in the directional filter superimposition unit 94.

[0294] 9, the directional filter superimposing unit 94 may determine whether or not omnidirectional approximation is possible based on the relative distance x, the transition curve parameter (distance d1), and the remaining battery charge of the device provided with the decoder 61. In other words, it may be determined whether or not to superimpose a directional filter on the audio data.

[0295] In this case, it is conceivable to determine that omnidirectional approximation is possible when the remaining battery charge of the device obtained from inside the device is equal to or less than a predetermined value. By doing so, when the remaining battery charge of the device is low, the process of superimposing a directional filter can be skipped, reducing computational costs and extending battery life.

[0296] During or after the transition curve is designed, the distance d1 may be corrected (determined) so that the distance d1 gradually decreases as the remaining battery charge of the device decreases.

[0297] Furthermore, for example, the transition curve design unit 92 may acquire computational resource information regarding the computational resources of the device in which the decoder 61 is provided when designing the transition curve, and use the computational resource information to determine the distance d1 as a transition curve parameter.

[0298] The computational resources referred to here are the computational resources (computing capacity) available for computation, such as the amount of memory and the computation time required for computation, and may be the current computational resources of the device or the maximum computational resources of the device.

[0299] 9, the directional filter superimposing unit 94 may determine whether or not omnidirectional approximation is possible by using not only the relative distance x and the transition curve parameter (distance d1) but also the remaining battery charge and calculation resource information of the device provided with the decoder 61. Specifically, for example, it may be determined that omnidirectional approximation is possible when an evaluation value calculated from a value indicating the remaining battery charge and a value indicating the calculation resource is equal to or less than a predetermined threshold.

[0300] Furthermore, in the transition curve design unit 92, the parameter p indicating the importance may be corrected (changed) based on at least one of metadata such as sound source type information, sound source position / direction information, listener position / direction information, and directional peak position information.

[0301] For example, it is conceivable to correct (make variable) the parameter p based on sound source position / direction information and listener position / direction information, i.e., depending on the relationship between the direction of the sound source in the virtual space and the direction of the listener (direction of gaze), or the relationship between the direction of the listener and the position of the sound source.

[0302] Specifically, for example, a sound source behind the listener is less important to the listener, so the parameter p is corrected (changed) to a smaller value, which reduces the distance d1 as a transition curve parameter and reduces the processing load of the directional filter superimposition process in the directional filter superimposition unit 94.

[0303] Furthermore, the parameter p may be changed according to the type (kind) of the sound source indicated by sound source type information as metadata, for example. In this case, the parameter p may be changed on the decoder 61 (transition curve design unit 92) side or on the encoder 11 (parameter generation unit 21) side.

[0304] Specifically, for example, if the content is for a first-person perspective game such as an FPS (First-Person Shooter), the importance of enemy objects (sound sources) may be set high and the importance of friendly objects may be set low. In this way, the processing load of the directional filter superimposition process in the directional filter superimposition unit 94 can be reduced without impairing the game experience.

[0305] In addition, for example, when peak position information is transmitted to the decoder 61 side, the transition curve design unit 92 may correct (change) the parameter p indicating the importance based on the peak position information, sound source position / direction information, and listener position / direction information.

[0306] In this case, the parameter p is corrected according to the relationship between the direction of the directivity peak as seen from the sound source in the virtual space and the gaze direction (face direction) of the listener. For example, if the angle between the direction of the directivity peak and the gaze direction of the listener is small, the parameter p may be corrected to a larger value.

[0307] <Example of Computer Configuration> The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the programs constituting the software are installed on a computer. Here, the computer includes a computer built into dedicated hardware, and a general-purpose personal computer, for example, that can execute various functions by installing various programs.

[0308] FIG. 14 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by a program.

[0309] In the computer, a CPU (Central Processing Unit) 501 , a ROM (Read Only Memory) 502 , and a RAM (Random Access Memory) 503 are interconnected by a bus 504 .

[0310] An input / output interface 505 is further connected to the bus 504. An input unit 506, an output unit 507, a recording unit 508, a communication unit 509, and a drive 510 are connected to the input / output interface 505.

[0311] The input unit 506 includes a keyboard, a mouse, a microphone, an image sensor, etc. The output unit 507 includes a display, a speaker, etc. The recording unit 508 includes a hard disk, a non-volatile memory, etc. The communication unit 509 includes a network interface, etc. The drive 510 drives a removable recording medium 511 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.

[0312] In a computer configured as described above, the CPU 501 loads a program recorded in the recording unit 508, for example, into the RAM 503 via the input / output interface 505 and the bus 504, and executes the program, thereby performing the above-described series of processes.

[0313] The program executed by the computer (CPU 501) can be provided by being recorded on a removable recording medium 511 such as a package medium, for example. The program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.

[0314] In a computer, a program can be installed in the recording unit 508 via the input / output interface 505 by inserting a removable recording medium 511 into the drive 510. The program can also be received by the communication unit 509 via a wired or wireless transmission medium and installed in the recording unit 508. Alternatively, the program can be installed in the ROM 502 or the recording unit 508 in advance.

[0315] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.

[0316] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present technology.

[0317] For example, the present technology can be configured as a cloud computing system in which a single function is shared and processed collaboratively by a plurality of devices via a network.

[0318] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by a plurality of devices.

[0319] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.

[0320] Furthermore, the present technology can also be configured as follows.

[0321] (1) An information processing device comprising a signal processing unit that selects, from directional filters for reproducing the directivity of the sound source, a directional filter that corresponds to a positional relationship between the sound source and the listener, using a selection method determined from one or more parameters that represent characteristics of the sound source, including at least a first parameter related to the directivity of the sound source, and a distance from the sound source to the listener in a virtual space. (2) The information processing device described in (1), in which the signal processing unit superimposes the selected directional filter on audio data of the sound source. (3) The information processing device described in (2), in which the signal processing unit superimposes, on the audio data, a weighting coefficient that corresponds to the distance, the weighting coefficient being determined by one or more parameters that represent characteristics of the sound source and the directional filter selected by the selection method, as the final directional filter. (4) The information processing device described in (3), in which the first parameter is a parameter that indicates the sharpness of the directivity of the sound source. (5) The information processing device according to (3) or (4), wherein the one or more parameters representing characteristics of the sound source include a second parameter indicating the importance of the sound source. (6) The information processing device according to any one of (2) to (5), wherein the signal processing unit does not superimpose the directional filter on the audio data when the distance is greater than a predetermined first distance. (7) The information processing device according to (6), wherein the signal processing unit selects the directional filter using a method based on the position of the center of the listener's head when the distance is greater than a second distance and equal to or less than the first distance. (8) The information processing device according to (7), wherein the signal processing unit selects one directional filter common to all channels when the distance is greater than the second distance and equal to or less than the first distance. (9) The information processing device according to (7) or (8), wherein the signal processing unit selects the directional filter using a method based on the position of the listener's ears when the distance is equal to or less than the second distance. (10) The information processing device according to (9), wherein the signal processing unit selects the directional filter for each of a plurality of channels when the distance is equal to or less than the second distance.(11) The information processing device according to (9) or (10), wherein, when the distance is greater than a third distance and equal to or less than the second distance, the signal processing unit generates the directional filter to be superimposed on the audio data based on the directional filter selected by a method based on the position of the head center and the directional filter selected by a method based on the ear position. (12) The information processing device according to any one of (3) to (5), further comprising a design unit that generates transition information indicating the selection method depending on the distance from the sound source based on parameters representing characteristics of the sound source. (13) The information processing device according to (12), wherein the design unit generates the transition information indicating the selection method and the weighting coefficient depending on the distance from the sound source. (14) The information processing device according to (12) or (13), wherein, when the directivity of the sound source has a plurality of peaks, the signal processing unit selects the directional filter using peak position information indicating the position of the peak and the transition information generated based on the first parameter of the peak selected based on sound source position direction information of the sound source and listener position direction information of the listener. (15) The information processing device according to any one of (9) to (11), wherein the signal processing unit selects the directional filter by a method based on the ear position of the listener based on a length of a wearing device for wearing the information processing device on the head of the listener. (16) The information processing device according to any one of (9) to (11), wherein the signal processing unit selects the directional filter by a method based on the ear position of the listener using a distance from the center of the head of the listener to the ear acquired in response to an input operation. (17) The information processing device according to (16), further comprising a display control unit that displays an input screen for inputting the distance from the center of the head of the listener to the ear. (18) The information processing device according to any one of (12) to (14), further comprising a display control unit that displays an input screen for specifying values ​​of parameters that constitute the transition information, wherein the design unit generates the transition information based on the values ​​of the parameters specified on the input screen, or changes the values ​​of the parameters that constitute the transition information to the specified values.(19) The information processing device according to any one of (12) to (14), wherein the design unit generates the transition information or corrects the transition information based on at least one of a computing resource, a remaining battery level, and a device type of the information processing device. (20) The information processing device according to any one of (2) to (19), wherein the signal processing unit determines whether to superimpose the directional filter on the audio data based on at least one of a computing resource, a remaining battery level, and a device type of the information processing device. (21) The information processing device according to any one of (2) to (20), wherein the signal processing unit superimposes an HRTF on the audio data on which the directional filter has been superimposed. (22) An information processing method, in which an information processing device selects, from directional filters for reproducing the directivity of the sound source, a directional filter that corresponds to a positional relationship between the sound source and the listener, by a selection method determined from one or more parameters that represent characteristics of the sound source, including at least a first parameter related to the directivity of the sound source, and a distance from the sound source to the listener in a virtual space. (23) A program that causes a computer to execute processing including a step of selecting, from directional filters for reproducing the directivity of the sound source, the directional filter that corresponds to a positional relationship between the sound source and the listener, by a selection method determined from one or more parameters that represent characteristics of the sound source, including at least a first parameter related to the directivity of the sound source, and a distance from the sound source to the listener in a virtual space. (24) An information processing device comprising: a parameter generation unit that generates one or more parameters representing characteristics of a sound source, including at least a first parameter related to the directivity of the sound source, based on directivity data that indicates the directivity of the sound source placed in a virtual space; and a communication unit that outputs a bit stream including the parameters. (25) The information processing device according to (24), wherein the first parameter is a parameter that indicates the sharpness of the directivity of the sound source. (26) The information processing device according to (24) or (25), wherein, when the directivity of the sound source has multiple peaks, the parameter generation unit generates the first parameter for each of the peaks.(27) The information processing device according to (26), wherein, when the directivity of the sound source has a plurality of peaks, the parameter generation unit generates peak position information indicating a position of the peak on a sphere surface representing the directivity of the sound source, and the communication unit outputs the bit stream including the parameter and the peak position information. (28) The information processing device according to any one of (24) to (27), wherein the parameter generation unit generates the first parameter using a mathematical model. (29) The information processing device according to any one of (24) to (28), wherein the one or more parameters representing characteristics of the sound source include a second parameter indicating importance of the sound source. (30) The information processing device according to (29), wherein the parameter generation unit generates the second parameter based on at least one of the importance acquired in response to an input operation, audio data of the sound source, and metadata of the audio data. (31) The information processing device according to (30), wherein the metadata includes at least any of position information, priority information, gain information, sound source type information, and spread information of the sound source. (32) An information processing method, in which an information processing device generates one or more parameters representing characteristics of the sound source, including at least a first parameter related to the directivity of the sound source, based on directivity data indicating the directivity of the sound source located in a virtual space, and outputs a bit stream including the parameters. (33) A program that causes a computer to execute processing including the steps of generating one or more parameters representing characteristics of the sound source, including at least a first parameter related to the directivity of the sound source, based on directivity data indicating the directivity of the sound source located in a virtual space, and outputting a bit stream including the parameters.

[0322] REFERENCE SIGNS LIST 11 Encoder, 21 Parameter generation unit, 22 Parameter encoding unit, 23 Audio encoding unit, 24 Multiplexing unit, 25 Communication unit, 61 Decoder, 71 Communication unit, 72 Demultiplexing unit, 73 Initialization unit, 74 Audio decoding unit, 75 Signal processing unit, 76 Input unit, 77 Display control unit, 91 Parameter decoding unit, 92 Transition curve design unit, 93 Relative distance calculation unit, 94 Directional filter superimposition unit, 95 HRTF superimposition unit

Claims

1. An information processing device comprising a signal processing unit that selects a directional filter corresponding to the relative positions of the sound source and the listener from among directional filters for reproducing the directivity of the sound source using a selection method determined from one or more parameters that represent the characteristics of the sound source, including at least a first parameter related to the directivity of the sound source, and the distance from the sound source to the listener in virtual space.

2. The information processing device according to claim 1, wherein the signal processing unit superimposes the selected directional filter on the audio data of the sound source.

3. The information processing device according to claim 2, wherein the signal processing unit superimposes a weighting coefficient according to the distance, determined by one or more parameters representing the characteristics of the sound source and the directional filter selected by the selection method, onto the audio data as the final directional filter.

4. The information processing device according to claim 3, wherein the first parameter is a parameter indicating the sharpness of the directionality of the sound source.

5. The information processing device according to claim 3, wherein the one or more parameters representing the characteristics of the sound source include a second parameter indicating the importance of the sound source.

6. The information processing device according to claim 2, wherein the signal processing unit does not superimpose the directional filter on the audio data when the distance is greater than a predetermined first distance.

7. The information processing device according to claim 6, wherein the signal processing unit selects the directional filter using a method based on the position of the center of the listener's head when the distance is greater than the second distance and equal to or less than the first distance.

8. The information processing device according to claim 7, wherein the signal processing unit selects one directional filter common to all channels when the distance is greater than the second distance and equal to or less than the first distance.

9. The information processing device according to claim 7, wherein the signal processing unit selects the directional filter by a method based on the ear positions of the listener when the distance is equal to or less than the second distance.

10. The information processing device according to claim 3, further comprising a design unit that generates transition information indicating the selection method according to the distance from the sound source based on parameters that represent the characteristics of the sound source.

11. The information processing device according to claim 10, wherein the design unit generates the transition information indicating the selection method and the weighting coefficient according to the distance from the sound source.

12. The information processing device according to claim 10, wherein, when the directivity of the sound source has multiple peaks, the signal processing unit selects the directional filter using peak position information indicating the position of the peak, and the transition information generated based on the first parameter of the peak selected based on sound source position direction information of the sound source and listener position direction information of the listener.

13. The information processing device according to claim 9, wherein the signal processing unit selects the directional filter using a method based on the position of the listener's ears, based on the length of a wearing device for wearing the information processing device on the listener's head.

14. The information processing device according to claim 9, wherein the signal processing unit selects the directional filter by a method based on the position of the listener's ears, using the distance from the center of the listener's head to the ears obtained in response to an input operation.

15. An information processing device as described in claim 10, further comprising a display control unit that displays an input screen for specifying the values ​​of parameters that constitute the transition information, and wherein the design unit generates the transition information based on the values ​​of the parameters specified on the input screen, or changes the values ​​of the parameters that constitute the transition information to the specified values.

16. The information processing device according to claim 10, wherein the design unit generates the transition information or corrects the transition information based on at least one of the computing resources, remaining battery power, and device type of the information processing device.

17. The information processing device according to claim 2, wherein the signal processing unit determines whether to superimpose the directional filter on the audio data based on at least one of the computing resources, remaining battery power, and device type of the information processing device.

18. An information processing method in which an information processing device selects a directional filter corresponding to the positional relationship between the sound source and the listener from among directional filters for reproducing the directionality of the sound source, using a selection method determined from one or more parameters representing the characteristics of the sound source, including at least a first parameter related to the directionality of the sound source, and the distance from the sound source to the listener in a virtual space.

19. A program that causes a computer to execute a process including a step of selecting a directional filter that corresponds to the relative positions of the sound source and the listener from among directional filters that reproduce the directivity of the sound source, using a selection method determined from one or more parameters that represent the characteristics of the sound source, including at least a first parameter related to the directivity of the sound source, and the distance from the sound source to the listener in virtual space.

20. An information processing device comprising: a parameter generation unit that generates one or more parameters representing characteristics of a sound source, including at least a first parameter related to the directivity of the sound source, based on directivity data indicating the directivity of the sound source placed in a virtual space; and a communication unit that outputs a bit stream including the parameters.

Citation Information

Patent Citations

  • Apparatus and method for processing audio signal to perform binaural rendering

    US20170325045A1

  • Signal processing device and method, and program

    WO2019116890A1

  • Method and system for controlling directivity of an audio source in a virtual reality environment

    WO2022243094A1

  • Information processing device, method, and program

    WO2023074800A1