Encoding apparatus and method, decoding apparatus and method, and program
By processing audio data, metadata, and distance sensing control information from encoding and decoding devices, the problem of flexible distance sensing control that cannot be achieved in existing technologies is solved, enabling distance sensing control based on the content creator's intent and generating realistic 3D audio reproduction.
Patent Information
- Application Number
- CN202080083336.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-01-10
- Filing Date
- 2020-12-25
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2040-12-25
AI Technical Summary
In existing technologies, content creators cannot implement distance sensing control based on their intentions, which differs from frequency characteristics and volume changes, resulting in an inability to achieve flexible distance sensing control.
By using encoding devices and methods, audio data, metadata, and distance sensing control information are multiplexed, and then demultiplexed and processed for distance sensing control in decoding devices to generate reproduced audio data, thereby achieving distance-sensing control based on the content creator's intent.
It achieves distance sensing control based on the content creator's intent, and can dynamically adjust the gain, filtering and reverberation effects of audio data according to the user's listening position and distance to generate realistic 3D audio reproduction.
Smart Images

Figure CN114762041B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present technology relates to an encoding apparatus and method, a decoding apparatus and method, and a program, and more particularly, to an encoding apparatus and method, a decoding apparatus and method, and a program capable of realizing distance sensing control based on an intention of a content creator. BACKGROUND
[0002] In recent years, object-based audio technology has attracted attention.
[0003] In object-based audio, data of an object audio is configured by a waveform signal regarding an audio object and metadata indicating positioning information of the audio object represented by a relative position to a listening position serving as a predetermined reference.
[0004] Then, the waveform signal of the audio object is rendered into a signal of a desired number of channels and reproduced by, for example, vector-based amplitude panning (VBAP) based on the metadata (see, for example, Non-Patent Literature 1 and Non-Patent Literature 2).
[0005] Furthermore, as a technology related to object-based audio, for example, a technology for realizing audio reproduction with higher degrees of freedom in which a user can specify an arbitrary listening position has also been proposed (see, for example, Patent Literature 1).
[0006] In this technology, the position information of an audio object is corrected in accordance with a listening position, and gain control or filtering processing is performed in accordance with a change in distance from the listening position to the audio object, so that a change in frequency characteristic or volume accompanying a change in listening position of a user (i.e., a sense of distance to the audio object) is reproduced.
[0007] LIST OF REFERENCES
[0008] NON-PATENT LITERATURE
[0009] Non-Patent Literature 1: ISO / IEC 23008-3 Information technology - High efficiency coding and media delivery in heterogeneous environments - Part 3: 3D audio
[0010] Non-Patent Literature 2: Ville Pulkki, "Virtual Sound Source Positioning Using Vector Base Amplitude Panning", Journal of AES, vol. 45, no. 6, pp. 456-466, 1997
[0011] PATENT LITERATURE
[0012] Patent Literature 1: WO 2015107926 A SUMMARY
[0013] PROBLEMS TO BE SOLVED BY THE INVENTION
[0014] However, in the above-described technique, the gain control and filter processing for reproducing the change in the frequency characteristic and the volume corresponding to the distance from the listening position to the audio object is predetermined.
[0015] Therefore, when the content creator desires to reproduce the sense of distance based on a different manner from the change in the frequency characteristic and the volume, such a sense of distance cannot be reproduced. That is, it is not possible to achieve the sense of distance control based on the intention of the content creator.
[0016] The present technology has been made in view of such circumstances, and it is an object to achieve the sense of distance control based on the intention of the content creator.
[0017] SOLUTION TO PROBLEM
[0018] The encoding apparatus according to a first aspect of the present technology includes an object encoding unit that encodes audio data of an object, a metadata encoding unit that encodes metadata including position information of the object, a sense-of-distance control information determination unit that determines sense-of-distance control information used for a sense-of-distance control process performed on the audio data, a sense-of-distance control information encoding unit that encodes the sense-of-distance control information, and a multiplexer that multiplexes the encoded audio data, the encoded metadata, and the encoded sense-of-distance control information to generate encoded data.
[0019] The encoding method or program according to the first aspect of the present technology includes the steps of encoding audio data of an object, encoding metadata including position information of the object, determining sense-of-distance control information used for a sense-of-distance control process performed on the audio data, encoding the sense-of-distance control information, and
[0020] multiplexing the encoded audio data, the encoded metadata, and the encoded sense-of-distance control information to generate encoded data.
[0021] In the first aspect of the present technology, the audio data of the object is encoded, the metadata including the position information of the object is encoded, the sense-of-distance control information used for the sense-of-distance control process performed on the audio data is determined, the sense-of-distance control information is encoded, the encoded audio data, the encoded metadata, and the encoded sense-of-distance control information are multiplexed to generate encoded data.
[0022] The decoding apparatus according to the second aspect of the present technology includes a demultiplexer that demultiplexes encoded data to extract encoded audio data of an object, encoded metadata including position information of the object, and encoded distance sensing control information for a distance sensing control process performed on the audio data; an object decoding unit that decodes the encoded audio data; a metadata decoding unit that decodes the encoded metadata; a distance sensing control information decoding unit that decodes the encoded distance sensing control information; a distance sensing control process unit that performs the distance sensing control process on the audio data of the object based on the distance sensing control information; and a rendering process unit that performs a reproduction process based on the audio data and the metadata obtained by the distance sensing control process to generate reproduction audio data for reproducing a sound of the object.
[0023] The decoding method or program according to the second aspect of the present technology includes the steps of demultiplexing encoded data to extract encoded audio data of an object, encoded metadata including position information of the object, and encoded distance sensing control information for a distance sensing control process performed on the audio data; decoding the encoded audio data; decoding the encoded metadata; decoding the encoded distance sensing control information; performing the distance sensing control process on the audio data of the object based on the distance sensing control information; and performing a reproduction process based on the audio data and the metadata obtained by the distance sensing control process to generate reproduction audio data for reproducing a sound of the object.
[0024] In the second aspect of the present technology, encoded data is demultiplexed to extract encoded audio data of an object, encoded metadata including position information of the object, and encoded distance sensing control information for a distance sensing control process performed on the audio data, the encoded audio data is decoded, the encoded metadata is decoded, the encoded distance sensing control information is decoded, the distance sensing control process is performed on the audio data of the object based on the distance sensing control information, and a rendering process is performed based on the audio data and the metadata obtained by the distance sensing control process to generate reproduction audio data for reproducing a sound of the object. BRIEF DESCRIPTION OF DRAWINGS
[0025] Figure 1 A diagram is shown to illustrate a configuration example of an encoding apparatus.
[0026] Figure 2 A diagram is shown to illustrate a configuration example of a decoding apparatus.
[0027] Figure 3 A diagram is shown to illustrate a configuration example of a distance sensing control process unit.
[0028] Figure 4 A diagram is shown to illustrate a configuration example of a reverberation process unit.
[0029] Figure 5 is a diagram for describing an example of a control rule for gain control processing.
[0030] Figure 6 is a diagram for describing an example of a control rule for filter processing by a high shelf filter.
[0031] Figure 7 is a diagram for describing an example of a control rule for filter processing by a low shelf filter.
[0032] Figure 8 is a diagram for describing an example of a control rule for reverb processing.
[0033] Figure 9 is a diagram for describing generation of a wet component.
[0034] Figure 10 is a diagram for describing generation of a wet component.
[0035] Figure 11 is a diagram showing an example of distance sensing control information.
[0036] Figure 12 is a diagram showing an example of parameter configuration information for gain control.
[0037] Figure 13 is a diagram showing an example of parameter configuration information for filter processing.
[0038] Figure 14 is a diagram showing an example of parameter configuration information for reverb processing.
[0039] Figure 15 is a flowchart for describing encoding processing.
[0040] Figure 16 is a flowchart for describing decoding processing.
[0041] Figure 17 is a diagram showing an example of a table and a function for obtaining a gain value.
[0042] Figure 18 is a diagram showing an example of parameter configuration information for gain control.
[0043] Figure 19 is a diagram showing an example of distance sensing control information.
[0044] Figure 20 is a diagram showing an example of distance sensing control information.
[0045] Figure 21is a diagram showing a configuration example of a distance sensing control processing unit.
[0046] Figure 22 is a diagram showing an example of distance sensing control information.
[0047] Figure 23 is a diagram showing a configuration example of a computer. DETAILED DESCRIPTION
[0048] Hereinafter, an embodiment to which the present technology is applied will be described with reference to the drawings.
[0049] <First Embodiment>
[0050] <Configuration Example of Encoding Device>
[0051] The present technology relates to audio content that reproduces object-based audio including sound of one or more audio objects.
[0052] Hereinafter, an audio object is also simply referred to as an object, and an audio content is also simply referred to as a content.
[0053] In the present technology, distance sensing control information for distance sensing control processing, which is set by a content creator and reproduces a distance sense from a listening position to an object, is transmitted to a decoding side together with audio data of the object. Therefore, distance sensing control based on the intention of the content creator can be implemented.
[0054] Here, the distance sensing control processing is processing for reproducing a distance sense from a listening position to an object when reproducing sound of the object, that is, processing for adding a distance sense to sound of the object, and is signal processing implemented by combining one or more processing steps.
[0055] Specifically, for example, in the distance sense control processing, gain control processing of the audio data, filter processing for increasing a frequency characteristic and various acoustic effects, reverberation processing, and the like are performed.
[0056] Information for causing a decoding end to reconfigure such distance sensing control processing is distance sensing control information, and the distance sensing control information includes configuration information and control rule information. That is, the distance sensing control information includes the configuration information and the control rule information.
[0057] For example, the configuration information configuring the distance sensing control information is information obtained by parameterizing a configuration of the distance sensing control processing set by the content creator, and indicates one or more signal processing steps to be combined and executed to implement the distance sensing control processing.
[0058] More specifically, the configuration information indicates the number of signal processing steps included in the distance sensing control processing, the processing performed in such signal processing, and the order of the processing.
[0059] Note that, in the case where one or more signal processing steps of the distance sensing control processing are predetermined and the order of performing these signal processing steps is predetermined, the distance sensing control information does not necessarily need to include the configuration information.
[0060] Further, the control rule information is information for obtaining a parameter that is obtained by parameterizing a control rule set by a content creator in each signal processing step of configuring the distance sensing control processing, and that is used in each signal processing step of configuring the distance sensing control processing.
[0061] More specifically, the control rule information indicates a parameter for configuring each signal processing step of the distance sensing control processing and a control rule according to which the parameter changes according to a distance from a listening position to an object.
[0062] At the encoding side, such distance sensing control information and audio data of each object are encoded and transmitted to the decoding side.
[0063] Further, at the decoding side, the distance sensing control processing is reconfigured based on the distance sensing control information, and the distance sensing control processing is performed on the audio data of each object.
[0064] At this time, a parameter corresponding to a distance from a listening position to an object is determined based on the control rule information contained in the distance sensing control information, and signal processing of configuring the distance sensing control processing is performed based on the parameter.
[0065] Then, a 3D audio rendering processing is performed based on the audio data obtained by the distance sensing control processing, and reproduction audio data of a sound (i.e., a sound of an object) for reproducing a content is generated.
[0066] Hereinafter, a more specific embodiment to which the present technology is applied will be described.
[0067] For example, a content reproduction system to which the present technology is applied includes: an encoding device that encodes audio data of each of one or more objects included in a content and distance sensing control information to generate encoded data; and a decoding device that receives supply of the encoded data to generate reproduction audio data.
[0068] For example, the encoding device configuring such a content reproduction system is configured as shown in Figure 1
[0069] Figure 1 The encoding apparatus 11 illustrated in FIG. 1 includes an object encoding unit 21, a metadata encoding unit 22, a distance sensing control information determination unit 23, a distance sensing control information encoding unit 24, and a multiplexer 25.
[0070] Audio data of each of one or more objects included in the content is supplied to the object encoding unit 21. The audio data is a waveform signal (an audio signal) for reproducing a sound of the object.
[0071] The object encoding unit 21 encodes the audio data of each of the objects supplied, and supplies the resulting encoded audio data to the multiplexer 25.
[0072] Metadata of the audio data of each of the objects is supplied to the metadata encoding unit 22.
[0073] The metadata includes at least position information indicating an absolute position of the object in the space. The position information is a coordinate indicating a position of the object in an absolute coordinate system, i.e., a three-dimensional orthogonal coordinate system based on a predetermined position in the space, for example. In addition, the metadata can include gain information for performing gain control (gain correction) on the audio data of the object, and the like.
[0074] The metadata encoding unit 22 encodes the metadata of each of the objects supplied, and supplies the resulting encoded metadata to the multiplexer 25.
[0075] The distance sensing control information determination unit 23 determines distance sensing control information in accordance with a designation operation of a user or the like, and supplies the determined distance sensing control information to the distance sensing control information encoding unit 24.
[0076] For example, the distance sensing control information determination unit 23 acquires configuration information and control rule information designated by a user in accordance with a designation operation of the user, thereby determining distance sensing control information including the configuration information and the control rule information.
[0077] In addition, for example, the distance sensing control information determination unit 23 can determine distance sensing control information based on the audio data of each of the objects of the content, information about the content such as a type of the content, information about a reproduction space of the content, and the like.
[0078] Note that in a case where each signal processing step configuring the distance sensing control processing and a processing order of the signal processing steps are known at the decoding side, the configuration information can not be included in the distance sensing control information.
[0079] The distance sensing control information encoding unit 24 encodes the distance sensing control information supplied from the distance sensing control information determination unit 23, and supplies the resulting encoded distance sensing control information to the multiplexer 25.
[0080] The multiplexer 25 multiplexes the encoded audio data supplied from the object encoding unit 21, the encoded metadata supplied from the metadata encoding unit 22, and the encoded distance sensing control information supplied from the distance sensing control information encoding unit 24 to generate encoded data (a code string). The multiplexer 25 transmits (transfers) the encoded data obtained by the multiplexing to a decoding apparatus via a communication network or the like.
[0081] <Configuration example of decoding apparatus>
[0082] Further, for example, as shown in Figure 2 A decoding apparatus included in a content reproduction system is configured.
[0083] Figure 2 The decoding apparatus 51 shown in FIG. 1 includes a demultiplexer 61, an object decoding unit 62, a metadata decoding unit 63, a distance sensing control information decoding unit 64, a user interface 65, a distance calculation unit 66, a distance sensing control processing unit 67, and a 3D audio rendering processing unit 68.
[0084] The demultiplexer 61 receives the encoded data transmitted from the encoding apparatus 11, and demultiplexes the received encoded data to extract the encoded audio data, the encoded metadata, and the encoded distance sensing control information from the encoded data.
[0085] The demultiplexer 61 supplies the encoded audio data to the object decoding unit 62, supplies the encoded metadata to the metadata decoding unit 63, and supplies the encoded distance sensing control information to the distance sensing control information decoding unit 64.
[0086] The object decoding unit 62 decodes the encoded audio data supplied from the demultiplexer 61, and supplies the generated audio data to the distance sensing control processing unit 67.
[0087] The metadata decoding unit 63 decodes the encoded metadata supplied from the demultiplexer 61, and supplies the generated metadata to the distance sensing control processing unit 67 and the distance calculation unit 66.
[0088] The distance sensing control information decoding unit 64 decodes the encoded distance sensing control information supplied from the demultiplexer 61, and supplies the generated distance sensing control information to the distance sensing control processing unit 67.
[0089] The user interface 65 supplies listening position information indicating a listening position specified by a user to the distance calculation unit 66, the distance sensing control processing unit 67, and the 3D audio rendering processing unit 68, for example, in accordance with an operation of the user or the like.
[0090] Here, the listening position indicated by the listening position information is an absolute position of a listener who listens to the sound of the content in the reproduction space. For example, the listening position information is coordinates indicating a listening position in an absolute coordinate system identical to that of the position information of the object included in the metadata.
[0091] The distance calculation unit 66 calculates the distance from the listening position to each object based on the metadata supplied from the metadata decoding unit 63 and the listening position information supplied from the user interface 65, and supplies distance information indicating the calculation result to the distance sensing control processing unit 67.
[0092] The distance sensing control processing unit 67 performs distance sensing control processing on the audio data supplied from the object decoding unit 62 based on the metadata supplied from the metadata decoding unit 63, the distance sensing control information supplied from the distance sensing control information decoding unit 64, the listening position information supplied from the user interface 65, and the distance information supplied from the distance calculation unit 66.
[0093] At this time, the distance sensing control processing unit 67 acquires parameters based on the control rule information and the distance information, and performs distance sensing control processing on the audio data based on the acquired parameters.
[0094] Through this distance sensing control processing, the audio data of the dry component of the object and the audio data of the wet component are generated.
[0095] Here, the audio data of the dry component is audio data obtained by performing one or more processing steps on the audio data of the original object, such as the direct sound component of the object.
[0096] The metadata of the original object, that is, the metadata output from the metadata decoding unit 63, is used as the metadata of the audio data of the dry component.
[0097] And, the audio data of the wet component is audio data obtained by performing one or more processing steps on the audio data of the original object, such as the reverberation component of the object sound. Therefore, it can be said that the audio data of the wet component is audio data of a new object related to the original object.
[0098] In the distance sensing control processing unit 67, the necessary data of the metadata of the original object, the control rule information, the distance information, and the listening position information are appropriately used for the metadata of the audio data of the wet component.
[0099] The metadata includes at least position information indicating the position of the object of the wet component.
[0100] For example, the position information of the object of the wet component is polar coordinates represented by an angle in a horizontal direction (horizontal angle), an angle in a vertical direction (vertical angle) indicating a position of the object viewed from the listener in the reproduction space, and a radius indicating a distance from the listening position to the object.
[0101] The distance sensing control processing unit 67 supplies the audio data and metadata of the dry component and the audio data and metadata of the wet component to the 3D audio rendering processing unit 68.
[0102] The 3D audio rendering processing unit 68 executes a 3D audio rendering process based on the audio data and metadata supplied from the distance sensing control processing unit 67 and the listening position information supplied from the user interface 65, and generates reproduction audio data.
[0103] For example, the 3D audio rendering processing unit 68 executes VBAP as the 3D audio rendering process, which is a rendering process in a polar coordinate system or the like.
[0104] In this case, for the audio data of the dry component, the 3D audio rendering processing unit 68 generates position information represented by polar coordinates based on the position information included in the metadata of the object of the dry component and the listening position information, and uses the obtained position information for the rendering process. The position information is polar coordinates represented by a horizontal angle, a vertical angle indicating a relative position of the object viewed from the listener, and a radius indicating a distance from the listening position to the object.
[0105] By such a reproduction process, for example, multi-channel reproduction audio data including audio data of channels corresponding to a plurality of speakers configuring a speaker system serving as an output destination is generated.
[0106] The 3D audio rendering processing unit 68 outputs the reproduction audio data obtained by the rendering process to a subsequent stage.
[0107] <Configuration example of distance sensing control processing unit>
[0108] Next, a specific configuration example of the distance sensing control processing unit 67 of the decoding apparatus 51 will be described.
[0109] Note that here, an example of determining in advance the configuration of the distance sensing control processing unit 67 (i.e., the configuration of one or more processing steps and the order of processing of the distance sensing control processing) will be described.
[0110] In this case, the distance sensing control processing unit 67 is configured, for example, as shown in FIG. 6. Figure 3
[0111] Figure 3 The distance sensing control processing unit 67 shown in FIG. 6 includes a gain control unit 101, an upper shelf filter processing unit 102, a lower shelf filter processing unit 103, and a reverb processing unit 104.
[0112] In this example, as the distance sensing control processing, gain control processing, filtering processing of the upper shelf filter, filtering processing of the lower shelf filter, and reverb processing are sequentially performed.
[0113] The gain control unit 101 performs gain control on the audio data of the object supplied from the object decoding unit 62 using parameters (gain values) corresponding to the control rule information and the distance information, and supplies the generated audio data to the upper shelf filter processing unit 102.
[0114] The upper shelf filter processing unit 102 performs filtering processing on the audio data supplied from the gain control unit 101 by an upper shelf filter determined by parameters corresponding to the control rule information and the distance information, and supplies the resulting audio data to the lower shelf filter processing unit 103.
[0115] In the filtering processing by the upper shelf filter, the high frequency gain of the audio data is suppressed in accordance with the distance from the listening position to the object.
[0116] The lower shelf filter processing unit 103 performs filtering processing on the audio data supplied from the upper shelf filter processing unit 102 by a lower shelf filter determined by parameters corresponding to the control rule information and the distance information.
[0117] In the filtering processing by the lower shelf filter, the low frequency of the audio data is enhanced in accordance with the distance from the listening position to the object.
[0118] The lower shelf filter processing unit 103 supplies the audio data resulting from the filtering processing to the 3D sound rendering processing unit 68 and the reverb processing unit 104.
[0119] Here, the audio data output from the lower shelf filter processing unit 103 is the above-mentioned original object's audio data, i.e., the audio data of the dry component of the object.
[0120] The reverb processing unit 104 performs reverb processing on the audio data supplied from the lower shelf filter processing unit 103 using parameters (gains) corresponding to the control rule information and the distance information, and supplies the resulting audio data to the 3D sound rendering processing unit 68.
[0121] Here, the audio data output from the reverb processing unit 104 is the above-mentioned original object's reverb component or the like, i.e., the audio data of the wet component of the object.
[0122] <Structure Example of Reverb Processing Unit>
[0123] Further, more specifically, the reverb processing unit 104 is configured, for example, as shown in Figure 4
[0124] In the example shown in Figure 4 , the reverb processing unit 104 includes a gain control unit 141, a delay generation unit 142, a comb filter bank 143, an all-pass filter bank 144, an addition unit 145, an addition unit 146, a delay generation unit 147, a comb filter bank 148, an all-pass filter bank 149, an addition unit 150, and an addition unit 151.
[0125] In this example, audio data of a stereo reverb component (i.e., two wet components located on the left and right sides of the original object) is generated by reverb processing for monaural audio data.
[0126] The gain control unit 141 performs gain control processing (gain correction processing) based on a wet gain value obtained from control rule information and distance information about the dry component audio data provided from the low shelf filter processing unit 103, and provides the generated audio data to the delay generation unit 142 and the delay generation unit 147.
[0127] The delay generation unit 142 delays the audio data provided from the gain control unit 141 by holding the audio data for a certain period of time, and provides the delayed audio data to the comb filter bank 143.
[0128] Further, the delay generation unit 142 provides two pieces of audio data obtained by delaying the audio data provided from the gain control unit 141 to the addition unit 145, the two pieces of audio data having different amounts of delay from the audio data provided to the comb filter bank 143, and having different amounts of delay from each other.
[0129] The comb filter bank 143 includes a plurality of comb filters that perform filter processing on the audio data provided from the delay generation unit 142 by the plurality of comb filters, and provides the generated audio data to the all-pass filter bank 144.
[0130] The all-pass filter bank 144 includes a plurality of all-pass filters that perform filter processing on the audio data provided from the comb filter bank 143 by the plurality of all-pass filters, and provides the generated audio data to the addition unit 146.
[0131] The addition unit 145 adds the two pieces of audio data provided from the delay generation unit 142 and provides the generated audio data to the addition unit 146.
[0132] The addition unit 146 adds the audio data supplied from the all-pass filter set 144 to the audio data supplied from the addition unit 145, and supplies the generated wet-component audio data to the 3D audio rendering processing unit 68.
[0133] The delay generation unit 147 delays the audio data supplied from the gain control unit 141 by holding the audio data for a certain period of time, and supplies the delayed audio data to the comb filter set 148.
[0134] Further, the delay generation unit 147 supplies two pieces of audio data obtained by delaying the audio data supplied from the gain control unit 141 to the addition unit 150, with different amounts of delay from the audio data supplied to the comb filter set 148, and with different amounts of delay from each other.
[0135] The comb filter set 148 includes a plurality of comb filters, performs filter processing on the audio data supplied from the delay generation unit 147 by the plurality of comb filters, and supplies the generated audio data to the all-pass filter set 149.
[0136] The all-pass filter set 149 includes a plurality of all-pass filters, performs filter processing on the audio data supplied from the comb filter set 148 by the plurality of all-pass filters, and supplies the generated audio data to the addition unit 151.
[0137] The addition unit 150 adds the two pieces of audio data supplied from the delay generation unit 147, and supplies the generated audio data to the addition unit 151.
[0138] The addition unit 151 adds the audio data supplied from the all-pass filter set 149 to the audio data supplied from the addition unit 150, and supplies the generated wet-component audio data to the 3D audio rendering processing unit 68.
[0139] Note that, although an example in which two wet components are generated for one object is described here, one wet component can be generated for one object, or three or more wet components can be generated. Further, the structure of the reverb processing unit 104 is not limited to Figure 4 the illustrated structure, and can be other structures.
[0140] <Control Rule on Parameters>
[0141] As described above, in each processing block constituting the distance-sensing control processing unit 67, the parameters used for the processing in the processing block (i.e., the characteristics of the processing) are changed according to the distance from the listening position to the object.
[0142] Here, an example of a parameter corresponding to the distance from the listening position to the object, i.e., an example of a control rule of the parameter, will be described.
[0143] For example, the gain control unit 101 determines a gain value for the gain control processing as a parameter corresponding to the distance from the listening position to the object.
[0144] In this case, for example, as shown in Figure 5 , the gain value changes according to the distance from the listening position to the object.
[0145] For example, the portion shown by an arrow Q11 indicates a change in the gain value corresponding to the distance. That is, the vertical axis indicates the gain value as the parameter, and the horizontal axis indicates the distance from the listening position to the object.
[0146] As shown by a broken line L11, when the distance d from the listening position to the object is between a predetermined minimum value Min and DO, the gain value is 0.0 dB, when the distance d is between DO and Dl, the gain value linearly decreases as the distance d increases. Further, when the distance d is between Dl and a predetermined maximum value Max, the gain value is -40.0 dB.
[0147] Thus, in the example shown in Figure 5 , it can be seen that control is performed to suppress the gain of the audio data as the distance d increases.
[0148] As a specific example, for example, in a case where the distance d is 1 m (= DO) or less, the gain value is set to 0.0 dB, and when the distance d is between 1 m and 100 m (= Dl), the gain value can linearly change to -40.0 dB as the distance d increases.
[0149] Here, when a point at which the parameter changes is referred to as a control change point, in the example of Figure 5 , the point (position) at which the distance d = DO and the point at which the distance d = Dl in the broken line L11 are control change points.
[0150] In this case, for example, as shown by an arrow Q12, when the gain value "0.0" at the distance d = DO and the gain value "-40.0" at the distance d = Dl corresponding to the control change points are transmitted to the decoding device 51, the decoding device 51 can obtain the gain value at an arbitrary distance d.
[0151] Further, in the shelving filter processing unit 102, for example, as shown by an arrow Q21 in Figure 6 , filter processing is performed to suppress the gain in the high frequency band as the distance d from the listening position to the object increases.
[0152] Note that in the portion indicated by the arrow Q21, the vertical axis indicates the gain value as a parameter, and the horizontal axis indicates the distance d from the listening position to the object.
[0153] Specifically, in this example, the shelving filter implemented by the shelving filter processing unit 102 is determined by the cutoff frequency Fc, the Q value indicating the sharpness, and the gain value at the cutoff frequency Fc.
[0154] In other words, in the shelving filter processing unit 102, the filtering processing is performed by the shelving filter determined by the cutoff frequency Fc, the Q value, and the gain value as parameters.
[0155] The broken line L21 in the portion indicated by the arrow Q21 indicates the gain value at the cutoff frequency Fc determined with respect to the distance d.
[0156] In this example, when the distance d is between the minimum value Min and DO, the gain value is 0.0 dB, and when the distance d is between DO and Dl, the gain value linearly decreases as the distance d increases.
[0157] And, when the distance d is between Dl and D2, the gain value linearly decreases as the distance d increases, similarly, when the distance d is between D2 and D3 and the distance d is between D3 and D4, the gain value linearly decreases as the distance d increases. Further, when the distance d is between D4 and the maximum value Max, the gain value is -12.0 dB.
[0158] Thus, in the example shown in Figure 6 , it can be seen that control is made in which the gain of the frequency component in the vicinity of the cutoff frequency Fc in the audio data is suppressed as the distance d increases.
[0159] As a specific example, for example, in a case where the distance d is 1 m (= DO) or less, a frequency component of 6 kHz or more as the cutoff frequency Fc can be set to pass, and in a case where the distance d is between the distance d of 1 m and the distance d of 100 m (= D4), the frequency component of 6 kHz or more can become -12.0 dB as the distance d increases.
[0160] And, in order to implement such a shelving filter in the decoding apparatus 51, for example, as indicated by the arrow Q22, it is only necessary to transmit the cutoff frequency Fc, the Q value, and the gain value as parameters for five control change points of the distances d = DO, Dl, D2, D3, and D4.
[0161] Note that here, an example in which the cutoff frequency Fc is 6 kHz and the Q value is 2.0 regardless of the distance d is described, but these cutoff frequency Fc and Q value can also be changed depending on the distance d.
[0162] Further, in the low shelf filter processing unit 103, for example, as indicated by an arrow Q31, filtering processing is performed in which a low frequency gain is amplified as the distance d from the listening position to the object decreases. Figure 7
[0163] Note that, in the portion indicated by the arrow Q31, the vertical axis indicates a gain value as a parameter, and the horizontal axis indicates the distance d from the listening position to the object.
[0164] Specifically, in this example, the low shelf filter implemented by the low shelf filter processing unit 103 is determined by a cutoff frequency Fc, a Q value indicating sharpness, and a gain value at the cutoff frequency Fc.
[0165] In other words, in the low shelf filter processing unit 103, filtering processing is performed by a low shelf filter determined by the cutoff frequency Fc, the Q value, and the gain value as parameters.
[0166] The broken line L31 in the portion indicated by the arrow Q31 indicates a gain value at the cutoff frequency Fc determined with respect to the distance d.
[0167] In this example, when the distance d is between the minimum value Min and D0, the gain value is 3.0 dB, and when the distance d is between D0 and D1, the gain value is linearly decreased as the distance d increases. Further, when the distance d is between D1 and the maximum value Max, the gain value is 0.0 dB.
[0168] Thus, in the example shown in Figure 7 It can be seen that control is performed in which, as the distance d decreases, the gain of the frequency component near the cutoff frequency Fc in the audio data is amplified.
[0169] As a specific example, for example, in a case where the distance d is 3 m (=D1) or more, a frequency component of 200 Hz or less as the cutoff frequency Fc can be set to pass, and in a case where the distance d is between 3 m and 10 cm (=D0), as the distance d decreases, the frequency component of 200 Hz or less can be changed to +3.0 dB.
[0170] Further, in order to implement such a low shelf filter in the decoding device 51, for example, as indicated by an arrow Q32, it is only necessary to transmit the cutoff frequency Fc, the Q value, and the gain value as parameters only for the two control change points of the distances d=D0 and D1.
[0171] Note that, here, an example in which the cutoff frequency Fc is 200 Hz and the Q value is 2.0 regardless of the distance d is described, but these cutoff frequency Fc and Q value can also be changed depending on the distance d.
[0172] Further, in the reverberation processing unit 104, for example, asFigure 8 The reverb processing is performed so that the gain of the wet component (wet gain value) increases as the distance d from the listening position to the object increases, as indicated by an arrow Q41.
[0173] In other words, control is performed in which the proportion of the wet component (reverberation component) generated by the reverb processing to the dry component increases as the distance d increases. Note that the wet gain value here is a gain value used in the gain control in the gain control unit 141 shown in Figure 4
[0174] In the portion indicated by the arrow Q41, the vertical axis indicates the wet gain value as a parameter, and the horizontal axis indicates the distance d from the listening position to the object. Further, the broken line L41 indicates the wet gain value determined for the distance d.
[0175] As indicated by the broken line L41, when the distance d from the listening position to the object is between the minimum value Min and DO, the wet gain value is negative infinity (-Inf dB), and when the distance d is between DO and Dl, the wet gain value linearly increases as the distance d increases. Further, when the distance d is between Dl and the maximum value Max, the wet gain value is -3.0 dB.
[0176] Thus, in the example shown in Figure 8 , it can be seen that control is performed in which the wet component increases as the distance d increases.
[0177] As a specific example, for example, in the case where the distance d is 1 m (= DO) or less, the gain of the wet component (wet gain value) is set to -Inf dB, and in the case where the distance d is between the distance d of 1 m and 50 m (= Dl), the gain can linearly change to -3.0 dB as the distance d increases.
[0178] Further, in order to implement such reverb processing in the decoding device 51, for example, as indicated by an arrow Q42, it is only necessary to transmit the wet gain value as a parameter for the two control change points of the distances d = DO and Dl.
[0179] Further, in the reverb processing, audio data of an arbitrary number of wet components (reverberation components) can be generated.
[0180] Specifically, for example, as indicated in Figure 9 , audio data of a stereo reverberation component can be generated for audio data of one object (i.e., monaural audio data).
[0181] In this example, the origin O of the XYZ coordinate system, which is a three-dimensional orthogonal coordinate system in the reproduction space, is the listening position, and one object OB11 is disposed in the reproduction space.
[0182] Now, the position of any object in the reproduction space is represented by a horizontal angle that indicates the position in the horizontal direction as viewed from the origin O and a vertical angle that indicates the position in the vertical direction as viewed from the origin O, and the position of the object OB11 is represented as (az, el) from the horizontal angle az and the vertical angle el.
[0183] Note that the horizontal angle az is the angle formed by the straight line LN' and the Z-axis when the straight line connecting the origin O and the object OB11 is LN and the straight line obtained by projecting the straight line LN on the XZ plane is LN'. Also, the vertical angle el is the angle formed by the straight line LN and the XZ plane.
[0184] In the example of Figure 9 , for the object OB11, two objects OB12 and OB13 are generated as wet component objects.
[0185] In particular, here, the objects OB12 and OB13 are arranged at positions symmetrically on both sides with respect to the object OB11 when viewed from the origin O.
[0186] That is, the objects OB12 and OB13 are arranged at positions shifted by 60 degrees to the left and right, respectively, with respect to the object OB11.
[0187] Therefore, the position of the object OB12 is the position (az + 60, el) represented by the horizontal angle (az + 60) and the vertical angle el, and the position of the object OB13 is the position (az - 60, el) represented by the horizontal angle (az - 60) and the vertical angle el.
[0188] As described above, in the case where wet components at positions symmetrically on both sides with respect to one object are generated, the positions of the wet components can be specified by offset angles with respect to the position of the one object. For example, in this example, only the offset angles of ±60 degrees of the horizontal angle need to be specified.
[0189] Note that, although an example of generating two right and left wet components located on the right and left sides with respect to one object is described here, the number of wet components generated with respect to one object can be any number, and, for example, wet components at upper, lower, left, and right positions can be generated.
[0190] Further, for example, in the case where wet components symmetrically on both sides are generated as shown in Figure 9 , the offset angles used to specify the positions of the wet components can be changed depending on the distance from the listening position to the object as shown in Figure 10 .
[0191] In the portion represented by the arrow Q51 in Figure 10 , it is shown that the object OB11 is generated as a wet component object in the case where the object OB11 is generated as a wet component object in the example of Figure 9the offset angle of the horizontal angle between the wet component object OB12 and the object OB13 shown in the middle.
[0192] That is, in the part of the arrow Q51, the vertical axis represents the offset angle of the horizontal angle, and the horizontal axis represents the distance d from the listening position to the object OB11.
[0193] Further, the broken line L51 represents the offset angle of the left wet component object OB12 determined for each distance d. In this example, as the distance d decreases, the offset angle increases, and the object OB12 is arranged at a position away from the original object OB11.
[0194] On the other hand, the broken line L52 represents the offset angle of the right wet component object OB13 determined for each distance d. In this example, as the distance d decreases, the offset angle decreases, and the object OB13 is arranged at a position away from the original object OB11.
[0195] In a case where the offset angle is changed in this way depending on the distance d, for example, as indicated by the arrow Q52, when the offset angle is transmitted to the decoding device 51 only for the control change point of the distance d=D0, the wet component can be generated at a position intended by the content creator.
[0196] As described above, by performing the distance sensing control process using the configuration and parameters corresponding to the distance d from the listening position to the object, it is possible to appropriately reproduce the distance sensing. That is, it is possible to make the listener feel the sense of the distance of the object.
[0197] At this time, when the content creator freely determines the parameters at each distance d, it is possible to realize the distance sensing control based on the intention of the content creator.
[0198] Note that the control rule of the parameters corresponding to the distance d described above is only an example, and by allowing the content creator to freely specify the control rule, it is possible to change how to feel the sense of the distance of the object.
[0199] For example, since the change of sound with respect to the distance is different between outdoors and indoors, it is necessary to change the control rule depending on whether the space to be reproduced is outdoors or indoors.
[0200] Therefore, for example, by determining (specifying) the control rule depending on the space that the content creator desires to reproduce together with the content, it is possible to realize the distance sensing control based on the intention of the content creator, and it is possible to perform the content reproduction with higher reality.
[0201] Further, in the distance sensing control processing unit 67, the parameters for the distance sensing control process can be further adjusted depending on the reproduction environment of the content (reproducing the audio data).
[0202] Specifically, for example, the gain of the wet component used in the reverb processing, that is, the above-mentioned wet gain value, can be adjusted in accordance with the reproduction environment of the content.
[0203] When the content is actually reproduced in a real space by a speaker or the like, reverberation of the sound output from the speaker or the like occurs in the real space. At this time, how much reverberation occurs depends on the real space in which the content is reproduced, that is, the reproduction environment.
[0204] For example, when the content is reproduced in an environment in which reverberation is high, reverberation is further added to the sound of the reproduced content. Therefore, in the case of actually reproducing the content, there is a case where the listener feels a sense of distance achieved by the distance sense control processing (that is, a sense of distance that is farther than the sense of distance intended by the content creator).
[0205] Therefore, in the case where the reverberation in the reproduction environment is small, the distance sense control processing is executed in accordance with the preset control rule (that is, the control rule information), but in the case where the reverberation in the reproduction environment is relatively large, a slight adjustment of the wet gain value determined in accordance with the control rule can be executed.
[0206] Specifically, for example, it is assumed that the user or the like operates the user interface 65 and inputs information on the reverberation of the reproduction environment, such as type information of the reproduction environment (such as outdoor or indoor) and information indicating whether the reproduction environment is highly reverberant. In this case, the user interface 65 supplies the information on the reverberation of the reproduction environment input by the user or the like to the distance sense control processing unit 67.
[0207] Then, the distance sense control processing unit 67 calculates the wet gain value based on the control rule information, the distance information, and the information on the reverberation of the reproduction environment supplied from the user interface 65.
[0208] Specifically, the distance sense control processing unit 67 calculates the wet gain value based on the control rule information and the distance information, and executes determination processing on whether the reproduction environment is highly reverberant based on the information on the reverberation of the reproduction environment.
[0209] Here, for example, in the case where information indicating that the reproduction environment is highly reverberant or type information of the reproduction environment indicating that the reproduction environment is highly reverberant is supplied as the information on the reverberation of the reproduction environment, it is determined that the reproduction environment is highly reverberant.
[0210] Then, the distance sense control processing unit 67 supplies the calculated wet gain value as the final wet gain value to the reverb processing unit 104 in the case where it is determined that the reproduction environment is not highly reverberant, that is, in the case where it is determined that the reproduction environment is not lowly reverberant.
[0211] On the other hand, when the distance sensing control processing unit 67 determines that the regeneration environment is a strong reverberation, it corrects (adjusts) the calculated wet path gain value using a specified correction value such as -6dB, and provides the corrected wet path gain value as the final wet path gain value to the reverberation processing unit 104.
[0212] Note that the wet gain correction value can be a predetermined value, or it can be calculated by the distance sensing control processing unit 67 based on information about the reverberation of the reproduced environment (i.e., the degree of reverberation in the reproduced environment).
[0213] By adjusting the wet gain value according to the reproduction environment in this way, the deviation from the perceived distance of the content creator caused by the content's reproduction environment can be improved.
[0214] <Transmission of Distance Sensing Control Information>
[0215] Next, the method for transmitting the distance sensing control information described above will be described.
[0216] For example, the distance sensing control information encoded by the distance sensing control information encoding unit 24 may have Figure 11 The configuration shown.
[0217] exist Figure 11 In the middle, “DistanceRender_Attn()” indicates parameter configuration information representing the control rules of the parameters used in the gain control unit 101.
[0218] In addition, “DistanceRender_Filt()” indicates parameter configuration information representing the control rules for the parameters used in the overhead filter processing unit 102 or the low-level filter processing unit 103.
[0219] Here, since the overhead filter and the low-mounted filter can be represented by the same parameter configuration, the overhead filter and the low-mounted filter are described by the same syntax of the parameter configuration information DistanceRender_Filt(). Therefore, the distance sensing control information includes the parameter configuration information DistanceRender_Filt() of the overhead filter processing unit 102 and the parameter configuration information DistanceRender_Filt() of the low-mounted filter processing unit 103.
[0220] Furthermore, "DistanceRender_Revb()" indicates parameter configuration information representing the control rules for the parameters used in the reverberation processing unit 104. The parameter configuration information DistanceRender_Attn(), DistanceRender_Filt(), and DistanceRender_Revb() included in the distance sensing control information correspond to the control rule information.
[0221] In addition, Figure 11 In the distance sensing control information shown, the parameter configuration information for the four processing steps of the distance sensing control process is arranged and stored in the order of execution of the processing steps.
[0222] Therefore, in the decoding device 51, the distance sensing control information can be used to specify... Figure 3 The distance sensing control processing unit 67 shown is configured as follows. In other words, according to... Figure 11 The distance sensing control information shown can specify how many processing steps are included in the distance sensing control process, what processing is performed in each of those steps, and in what order the processing is performed. Therefore, in this example, it can be said that the distance sensing control information essentially includes configuration information.
[0223] also, Figure 11 The parameter configuration information DistanceRender_Attn(), DistanceRender_Filt(), and DistanceRender_Revb() shown are configured as follows: Figures 12 to 14 As shown in the image.
[0224] Figure 12 This is a diagram showing a configuration instance (i.e., a syntax instance) of the parameter configuration information for the gain control processing DistanceRender_Attn().
[0225] exist Figure 12 In this context, "num_points" represents the number of control change points for the parameters processed by gain control. For example, in... Figure 5 In the example shown, the point (position) at a distance of d = D0 and the point at a distance of d = D1 are the control change points.
[0226] exist Figure 12 In this example, the "distance[i]" indicating the distance d corresponding to the control change point and the gain value "gain[i]" as a parameter at distance d are included as many times as the number of control change points. When the distance[i] and gain value gain[i] of each control change point are transmitted in this way, Figure 5 The gain control shown can be implemented in the decoding device 51.
[0227] Figure 13 is a diagram showing a configuration example (i.e., a syntax example) of the parameter configuration information DistanceRender_Filt() that shows a filter process.
[0228] In Figure 13 , "filt_type" indicates an index indicating a filter type.
[0229] For example, the index filt_type "0" indicates a low shelf filter, the index filt_type "1" indicates a high shelf filter, and the index filt_type "2" indicates a peak filter.
[0230] In addition, the index filt_type "3" indicates a low pass filter, and the index filt_type "4" indicates a high pass filter.
[0231] Therefore, for example, when the value of the index filt_type is "0", it can be seen that the parameter configuration information DistanceRender_Filt() includes information on parameters for specifying the configuration of the low shelf filter.
[0232] Note that in the example shown in Figure 3 , the high shelf filter and the low shelf filter have been described as filter examples that configure a filter process of the distance sensing control process.
[0233] On the other hand, in the example shown in Figure 13 , a peak filter, a low pass filter, a high pass filter, and the like can also be used.
[0234] Note that as filters for configuring a filter process of the distance sensing control process, only some of the low shelf filter and the high shelf filter, the peak filter, the low pass filter, and the high pass filter can be used, or other filters can be used.
[0235] In the parameter configuration information DistanceRender_Filt() shown in Figure 13 , the area after the index filt_type includes parameters and the like for specifying the configuration of the filter indicated by the index filt_type.
[0236] That is, "num_points" indicates the number of control change points of the parameters of the filter process.
[0237] Further, "distance[i]" indicating the distance d corresponding to the control change point, the frequency "freq[i]" as the parameter at the distance d, the Q value "Q[i]", and the gain value "gain[i]" are included as many as the number of control change points indicated by "num_points".
[0238] For example, when the index filt_type is "0" indicating a low shelf filter, the frequency "freq[i]" as the parameter, the Q value "Q[i]", and the gain value "gain[i]" correspond to the cutoff frequency Fc, the Q value, and the gain value shown in Figure 7
[0239] Note that the frequency freq[i] is the cutoff frequency when the filter type is a low shelf filter and a high shelf filter, a low pass filter, or a high pass filter, but is the center frequency when the filter type is a peak filter.
[0240] As described above, when the distance distance[i] of each control change point, the frequency "freq[i]", the Q value "Q[i]", and the gain value "gain[i]" are transmitted, Figure 6 Figure 7 the high shelf filter shown in
[0241] Figure 14 is a diagram showing a configuration example (i.e., a syntax example) of the parameter configuration information DistanceRender_Revb() that shows the reverb processing.
[0242] In Figure 14 In the example, "distance[i]" indicating the distance d corresponding to those control change points and the wet gain value "wet_gain[i]" as the parameter at the distance d are included as many as the number of control change points. For example, the wet gain value wet_gain[i] corresponds to the wet gain value shown in Figure 8
[0243] Further, in Figure 14 In the example, "distance[i]" indicating the distance d corresponding to those control change points and the wet gain value "wet_gain[i]" as the parameter at the distance d are included as many as the number of control change points. For example, the wet gain value wet_gain[i] corresponds to the wet gain value shown in
[0244] That is, "wet_azimuth_offset[i][j]" indicates the offset angle of the horizontal angle of the jth wet component (object) at the distance distance[i] corresponding to the ith control change point. For example, the offset angle wet_azimuth_offset[i][j] corresponds to the offset angle of the horizontal angle of the jth wet component (object) at the distance distance[i] corresponding to the ith control change point shown in Figure 10 a bias angle of a horizontal angle illustrated in FIG. 6.
[0245] Similarly, "wet_elevation_offset[i][j]" indicates a bias angle of a vertical angle of the jth wet component at a distance distance[i] corresponding to the ith control change point.
[0246] Note that the number num_wetobjs of the generated wet components is determined by the reverb processing by the decoding apparatus 51, and for example, the number num_wetobjs of the wet components is given from the outside.
[0247] As described above, in the example of Figure 14 the distance distance[i] and the wet gain value wet_gain[i] at each control change point, and the bias angles wet_azimuth_offset[i][j] and wet_elevation_offset[i][j] of each wet component are transmitted to the decoding apparatus 51.
[0248] Therefore, in the decoding apparatus 51, for example, the reverb processing unit 104 illustrated in Figure 4 can be realized, and the audio data of the dry component and the audio data and the metadata of each wet component can be obtained.
[0249] <Description of Encoding Processing>
[0250] Next, the operation of the content reproduction system will be described.
[0251] First, the encoding processing performed by the encoding apparatus 11 will be described with reference to the flowchart in Figure 15
[0252] In step S11, the object encoding unit 21 encodes the audio data of each object supplied, and supplies the obtained encoded audio data to the multiplexer 25.
[0253] In step S12, the metadata encoding unit 22 encodes the metadata of each object supplied, and supplies the obtained encoded metadata to the multiplexer 25.
[0254] In step S13, the distance sensing control information determination unit 23 determines the distance sensing control information in accordance with a designation operation of a user or the like, and supplies the determined distance sensing control information to the distance sensing control information encoding unit 24.
[0255] In step S14, the distance sensing control information encoding unit 24 encodes the distance sensing control information supplied from the distance sensing control information determination unit 23, and supplies the obtained encoded distance sensing control information to the multiplexer 25. Therefore, for example, the distance sensing control information illustrated in Figure 11 The distance sensing control information (encoded distance sensing control information) shown in the middle is provided to the multiplexer 25.
[0256] In step S15, the multiplexer 25 multiplexes the encoded audio data from the object encoding unit 21, the encoded metadata from the metadata encoding unit 22, and the encoded distance sensing control information from the distance sensing control information encoding unit 24 to generate encoded data.
[0257] In step S16, the multiplexer 25 transmits the encoded data obtained by multiplexing to the decoding apparatus 51 via a communication network or the like, and the encoding process ends.
[0258] As described above, the encoding apparatus 11 generates encoded data including distance sensing control information, and transmits the encoded data to the decoding apparatus 51.
[0259] As described above, by transmitting distance sensing control information other than audio data and metadata of each object to the decoding apparatus 51, distance sensing control based on the intention of a content creator on the decoding apparatus 51 side can be realized.
[0260] <Description of Decoding Process>
[0261] Also, in the decoding apparatus 51, the decoding process is executed when the encoding process described above is executed in the encoding apparatus 11. Figure 15 The decoding process by the decoding apparatus 51 will be described below with reference to the flowchart in Figure 16
[0262] In step S41, the demultiplexer 61 receives encoded data transmitted from the encoding apparatus 11.
[0263] In step S42, the demultiplexer 61 demultiplexes the received encoded data, and extracts encoded audio data, encoded metadata, and encoded distance sensing control information from the encoded data.
[0264] The demultiplexer 61 supplies the encoded audio data to the object decoding unit 62, supplies the encoded metadata to the metadata decoding unit 63, and supplies the encoded distance sensing control information to the distance sensing control information decoding unit 64.
[0265] In step S43, the object decoding unit 62 decodes the encoded audio data supplied from the demultiplexer 61, and supplies the obtained audio data to the distance sensing control processing unit 67.
[0266] In step S44, the metadata decoding unit 63 decodes the encoded metadata supplied from the demultiplexer 61, and supplies the obtained metadata to the distance sensing control processing unit 67 and the distance calculation unit 66.
[0267] In step S45, the distance sensing control information decoding unit 64 decodes the encoded distance sensing control information supplied from the demultiplexer 61, and supplies the obtained distance sensing control information to the distance sensing control processing unit 67.
[0268] In step S46, the distance calculation unit 66 calculates the distance from the listening position to the object based on the metadata supplied from the metadata decoding unit 63 and the listening position information supplied from the user interface 65, and supplies distance information indicating the calculation result to the distance sensing control processing unit 67. In step S46, the distance information is obtained for each object.
[0269] In step S47, the distance sensing control processing unit 67 performs the distance sensing control processing based on the audio data supplied from the object decoding unit 62, the metadata supplied from the metadata decoding unit 63, the distance sensing control information supplied from the distance sensing control information decoding unit 64, the listening position information supplied from the user interface 65, and the distance information supplied from the distance calculation unit 66.
[0270] For example, in a case where the distance sensing control processing unit 67 has the configuration shown in Figure 3 and provides the distance sensing control information shown in Figure 11 , the distance sensing control processing unit 67 calculates the parameters used in each processing step based on the distance sensing control information and the distance information.
[0271] Specifically, for example, the distance sensing control processing unit 67 obtains the gain value at the distance d indicated by the distance information based on the distance distance[i] and the gain value gain[i] of each control change point, and supplies the gain value to the gain control unit 101.
[0272] Further, based on the distance distance[i], the frequency freq[i], the Q value Q[i], and the gain value gain[i] of each control change point of the shelving filter, the distance sensing control processing unit 67 obtains the cutoff frequency, the Q value, and the gain value at the distance d indicated by the distance information, and supplies the cutoff frequency, the Q value, and the gain value to the shelving filter processing unit 102.
[0273] Accordingly, the shelving filter processing unit 102 can construct the shelving filter corresponding to the distance d indicated by the distance information.
[0274] Similarly to the case of the high shelf filter, the distance sensing control processing unit 67 obtains the cutoff frequency, the Q value, and the gain value of the low shelf filter at the distance d represented by the distance information, and supplies them to the low shelf filter processing unit 103. Thus, the low shelf filter processing unit 103 can configure the low shelf filter corresponding to the distance d represented by the distance information.
[0275] Further, the distance sensing control processing unit 67 obtains the wet gain value at the distance d represented by the distance information based on the distance distance[i] and the wet gain value wet_gain[i] of each control change point, and supplies the wet gain value to the reverb processing unit 104.
[0276] Thus, Figure 3 The distance sensing control processing unit 67 shown is configured from the distance sensing control information.
[0277] Further, the distance sensing control processing unit 67 supplies the offset angle wet_azimuth_offset[i][j] of the horizontal angle and the offset angle wet_elevation_offset[i][j] of the vertical angle, the metadata of the object, and the listening position information to the reverb processing unit 104.
[0278] The gain control unit 101 performs gain control processing on the audio data of the object based on the gain value supplied from the distance sensing control processing unit 67, and supplies the generated audio data to the high shelf filter processing unit 102.
[0279] The high shelf filter processing unit 102 performs filter processing on the audio data supplied from the gain control unit 101 by the high shelf filter determined by the cutoff frequency, the Q value, and the gain value supplied from the distance sensing control processing unit 67, and supplies the resulting audio data to the low shelf filter processing unit 103.
[0280] The low shelf filter processing unit 103 performs filter processing on the audio data supplied from the high shelf filter processing unit 102 by the low shelf filter determined by the cutoff frequency, the Q value, and the gain value supplied from the distance sensing control processing unit 67.
[0281] The distance sensing control processing unit 67 supplies the audio data obtained by the filter processing in the low shelf filter processing unit 103 together with the metadata of the object of the dry component as the audio data of the dry component to the 3D audio rendering processing unit 68. The metadata of the dry component is the metadata supplied from the metadata decoding unit 63.
[0282] In addition, the low shelf filter processing unit 103 supplies the sound data obtained by the filter processing to the reverb processing unit 104.
[0283] Then, for example, as described with reference to Figure 4 the reverb processing unit 104 performs gain control based on the wet gain value for the audio data of the dry component, delay processing on the audio data, filter processing using a comb filter and an all-pass filter, and the like, and generates the audio data of the wet component.
[0284] Further, the reverb processing unit 104 calculates the position information of the wet component based on the offset angle wet_azimuth_offset[i][j] and the offset angle wet_elevation_offset[i][j], the metadata of the object (dry component), and the listening position information, and generates the metadata of the wet component including the position information.
[0285] The reverb processing unit 104 supplies the thus generated sound data and metadata of each wet component to the 3D sound rendering processing unit 68.
[0286] In step S48, the 3D audio rendering processing unit 68 performs rendering processing based on the audio data and metadata supplied from the distance sensing control processing unit 67 and the listening position information supplied from the user interface 65, and generates reproduction audio data. For example, in step S48, VBAP or the like is executed as the rendering processing.
[0287] When the reproduction audio data is generated, the 3D audio rendering processing unit 68 outputs the generated reproduction audio data to a subsequent stage, and the decoding processing ends.
[0288] As described above, the decoding device 51 performs distance sensing control processing based on the distance sensing control information included in the encoded data, and generates reproduction audio data. In this way, distance feeling control based on the intention of the content creator can be realized.
[0289] <First Modification of the First Embodiment>
[0290] <Another Example of Parameter Configuration Information>
[0291] It is to be noted that, although the example shown in Figure 12 , Figure 13 and Figure 14 has been described above as the parameter configuration information, the parameter configuration information is not limited thereto, and any parameter configuration information can be used as long as the parameters of the distance sensing control processing are obtainable.
[0292] For example, it is also conceivable that, for each of one or more processing steps configuring the distance-sensing control processing, a table, a function (mathematical expression), or the like for obtaining a parameter of the distance d from the listening position to the object is prepared in advance, and an index indicating the table or function is included in the parameter configuration information. In this case, the index indicating the table or function is the control rule information indicating the control rule of the parameter.
[0293] In the case where the index indicating the table or function for obtaining the parameter is set as the control rule information in this way, for example, as shown in Figure 17 , a plurality of tables and functions for obtaining a gain value of the gain control processing as the parameter can be prepared.
[0294] In this example, for example, a function "20log 10 (1 / d) 2 " for obtaining a gain value of the gain control processing is prepared for the index value "1", and a gain value of the gain control processing corresponding to the distance d can be obtained by substituting the distance d into the function.
[0295] Further, for example, a table for obtaining a gain value of the gain control processing is prepared for the index value "2", and when the table is used, the gain value as the parameter decreases as the distance d increases.
[0296] The distance-sensing control processing unit 67 of the decoding apparatus 51 holds the table or function in advance in association with each index.
[0297] In this case, for example, Figure 11 the parameter configuration information DistanceRender_Attn() has Figure 18 the configuration shown in
[0298] In the example of Figure 18 , the parameter configuration information DistanceRender_Attn() includes an index "index" indicating a function or a table specified by the content creator.
[0299] Therefore, the distance-sensing control processing unit 67 reads the table or function held in association with the index "index", and obtains a gain value as the parameter based on the read table or function and the distance d from the listening position to the object.
[0300] In this way, when a plurality of patterns (i.e., a plurality of tables or functions for obtaining a parameter corresponding to the distance d) are defined in advance, the content creator can specify (select) a desired pattern from among these patterns, so that the distance-sensing control processing is performed according to his / her intention.
[0301] Note that, here, an example in which a table or a function for obtaining parameters of the gain control processing is specified by an index has been described. However, the present application is not limited to this, and in the case of filter processing of a shelving filter or the like or reverb processing, a control rule of parameters can also be similarly specified by an index.
[0302] <Second Modification of the First Embodiment>
[0303] <Another Example of Distance Sensing Control Information>
[0304] Further, in the above description, an example in which the same control rule is used to determine parameters corresponding to the distance d for all objects has been described. However, the control rule of parameters can be set (specified) for each object.
[0305] In this case, for example, as shown in Figure 19 , distance sensing control information is configured.
[0306] In the example shown in Figure 19 , "num_objs" indicates the number of objects included in the content, and for example, the number of objects num_objs is externally supplied to the distance sensing control information determination unit 23.
[0307] In the distance sensing control information, a flag "isDistanceRenderFlg" indicating whether or not an object is a target of distance sensing control is included as many as the number of objects num_objs.
[0308] For example, in the case where the value of the flag is "1" for the isDistanceRenderFlg of the i-th object, it is determined that the object is a target of distance sensing control, and distance sensing control processing is performed on the audio data of the object.
[0309] In the case where the value of the flag is "1" for the isDistanceRenderFlg of the i-th object, the distance sensing control information includes parameter configuration information DistanceRender_Attn() of the object, two parameter configuration information DistanceRender_Filt(), and parameter configuration information DistanceRender_Revb().
[0310] Therefore, in this case, as described above, the distance sensing control processing unit 67 performs distance sensing control processing on the audio data of the target object, and outputs the obtained audio data and metadata of the dry component and the wet component.
[0311] On the other hand, in a case where the value of the flag is "0" for the isDistanceRenderFlg of the i-th object, it is determined that the object is not a target of the distance sensing control, that is, is not a target, and the distance sensing control processing is not performed on the audio data of the object.
[0312] Therefore, for such an object, the audio data and the metadata of the object are supplied from the distance sensing control processing unit 67 to the 3D audio rendering processing unit 68 without change.
[0313] In a case where the value of the flag is "0" for the isDistanceRenderFlg of the i-th object, the distance sensing control information does not include the parameter configuration information DistanceRender_Attn(), the parameter configuration information DistanceRender_Filt(), and the parameter configuration information DistanceRender_Revb() of the object.
[0314] As described above, in the example shown in Figure 19 , the distance sensing control information encoding unit 24 encodes the parameter configuration information for each object. In other words, the distance sensing control information is encoded for each object. Therefore, distance sensing control based on the intention of the content creator can be implemented for each object, and content reproduction with higher realism can be performed.
[0315] Specifically, in this example, when the flag isDistanceRenderFlg is stored in the distance sensing control information, it is possible to set whether to perform distance sensing control for each object, and then perform different distance sensing control for each object.
[0316] For example, for an object of a human voice, by setting a control rule different from that of other objects than this object or not performing distance sensing control itself, it is possible to make a listener feel a smaller distance feeling, that is, to reproduce a sound that a listener always easily hears (easily heard sound).
[0317] <Third Modification of the First Embodiment>
[0318] <Another Example of Distance Sensing Control Information>
[0319] Further, the control rule of the parameter can not be set (designated) for each object, but for each object group including one or more objects.
[0320] In this case, the distance sensing control information is configured, for example, as shown in Figure 20 .
[0321] In Figure 20In the example shown, "num_obj_groups" indicates the number of target sets included in the content, and for example, the number of target sets num_obj_groups is externally supplied to the distance sensing control information determination unit 23.
[0322] In the distance sensing control information, as many flags "isDistanceRenderFlg" as the number of target sets num_obj_groups are included, which indicate whether or not the target set (more specifically, the target belonging to the target set) is a target of distance sensing control.
[0323] For example, in a case where the value of the flag is "1" for the isDistanceRenderFlg of the i-th target set, the target set is determined to be a target of distance sensing control, and distance sensing control processing is performed on the audio data of the target belonging to the target set.
[0324] In a case where the value of the flag is "1" for the isDistanceRenderFlg of the i-th target set, the distance sensing control information includes the parameter configuration information DistanceRender_Attn(), the two parameter configuration information DistanceRender_Filt(), and the parameter configuration information DistanceRender_Revb() of the target set.
[0325] Therefore, in this case, as described above, the distance sensing control processing unit 67 performs distance sensing control processing on the audio data of the object belonging to the target object group.
[0326] On the other hand, in a case where the value of the flag is "0" for the isDistanceRenderFlg of the i-th target set, the target set is determined not to be a target of distance sensing control, and distance sensing control processing is not performed on the audio data of the target of the target set.
[0327] Therefore, for the object of such an object group, the audio data and the metadata of the object are supplied from the distance sensing control processing unit 67 to the 3D audio rendering processing unit 68 without change.
[0328] In a case where the value of the flag is "0" for the isDistanceRenderFlg of the i-th target set, the distance sensing control information does not include the parameter configuration information DistanceRender_Attn(), the parameter configuration information DistanceRender_Filt(), and the parameter configuration information DistanceRender_Revb() of the target set.
[0329] As described above, in Figure 20In the example shown, the distance-sensing control information encoding unit 24 encodes the parameter configuration information for each object group. In other words, the distance-sensing control information is encoded for each target set. Thus, distance control based on the intention of the content creator can be implemented for each object group, and content reproduction with a higher sense of reality can be performed.
[0330] Specifically, in this example, when the flag isDistanceRenderFlg is stored in the distance-sensing control information, it can be set whether to perform distance-sensing control for each target set, and then different distance-sensing control is performed for each target set.
[0331] For example, in a case where the same control rule is set for a plurality of percussion instruments such as a snare drum, a bass drum, a tom-tom, cymbals, and the like that constitute a drum set, the content creator can group the plurality of percussion instruments into one object group.
[0332] In this way, the same control rule can be set for each object corresponding to each of the plurality of percussion instruments that belong to the same object group and constitute a drum set. That is, the same control rule information can be assigned to each of the plurality of objects. Moreover, as in the example shown in Figure 20 In the example shown in
[0333] <Second Embodiment>
[0334] <Configuration Example of Distance-Sensing Control Processing Unit>
[0335] Further, in the above description, an example in which the configuration of the distance-sensing control processing unit 67 provided in the decoding device 51 is determined in advance has been described. That is, an example in which one or more processing steps of the distance-sensing control processing and the processing order indicated by the configuration information configuring the distance-sensing control information are determined in advance has been described.
[0336] However, the present application is not limited thereto, and the configuration of the distance-sensing control processing unit 67 can be freely changed by the configuration information of the distance-sensing control information.
[0337] In this case, the distance-sensing control processing unit 67 is configured, for example, as shown in Figure 21
[0338] In the example shown in Figure 21 In the example shown in In the example shown in
[0339] The signal processing unit 201-1 performs signal processing on the audio data of the object supplied from the object decoding unit 62 based on the distance information supplied from the distance calculation unit 66 and the distance sensing control information supplied from the distance sensing control information decoding unit 64, and supplies the generated audio data to the signal processing unit 201-2.
[0340] At this time, in a case where the reverb processing unit 202-2 is functioning, that is, in a case where the reverb processing unit 202-2 is implemented, the signal processing unit 201-1 also supplies the sound data obtained by signal processing to the reverb processing unit 202-2.
[0341] The signal processing unit 201-2 performs signal processing on the audio data supplied from the signal processing unit 201-1 based on the distance information supplied from the distance calculation unit 66 and the distance sensing control information supplied from the distance sensing control information decoding unit 64, and supplies the generated audio data to the signal processing unit 201-3. At this time, in a case where the reverb processing unit 202-3 is functioning, the signal processing unit 201-2 also supplies the sound data obtained by signal processing to the reverb processing unit 202-3.
[0342] The signal processing unit 201-3 performs signal processing on the audio data supplied from the signal processing unit 201-2 based on the distance information supplied from the distance calculation unit 66 and the distance sensing control information supplied from the distance sensing control information decoding unit 64, and supplies the generated audio data to the 3D audio rendering processing unit 68. At this time, in a case where the reverb processing unit 202-4 is functioning, the signal processing unit 201-3 also supplies the sound data obtained by signal processing to the reverb processing unit 202-4.
[0343] Note that, hereinafter, the signal processing units 201-1 to 201-3 will also be simply referred to as the signal processing unit 201 in a case where it is not particularly necessary to distinguish the signal processing units.
[0344] The signal processing performed by the signal processing unit 201-1, the signal processing unit 201-2, and the signal processing unit 201-3 is processing indicated by the configuration information of the distance sensing control information.
[0345] Specifically, for example, the signal processing performed by the signal processing unit 201 is gain control processing and filter processing by a high shelf filter, a low shelf filter, or the like.
[0346] The reverb processing unit 202-1 performs reverb processing on the object's audio data supplied from the object decoding unit 62 based on the distance information supplied from the distance calculating unit 66 and the distance sensing control information supplied from the distance sensing control information decoding unit 64, and generates the audio data of the wet component.
[0347] Further, the reverb processing unit 202-1 generates the metadata including the position information of the wet component based on the distance sensing control information supplied from the distance sensing control information decoding unit 64, the metadata supplied from the metadata decoding unit 63, and the listening position information supplied from the user interface 65. In addition, in the reverb processing unit 202-1, the metadata of the wet component is generated using the distance information as necessary.
[0348] The reverb processing unit 202-1 supplies the metadata and the audio data of the wet component generated in this way to the 3D audio rendering processing unit 68.
[0349] The reverb processing unit 202-2 generates the metadata and the audio data of the wet component based on the distance information from the distance calculating unit 66, the distance sensing control information from the distance sensing control information decoding unit 64, the audio data from the signal processing unit 201-1, the metadata from the metadata decoding unit 63, and the listening position information from the user interface 65, and supplies the generated metadata and the audio data to the 3D audio rendering processing unit 68.
[0350] The reverb processing unit 202-3 generates the metadata and the audio data of the wet component based on the distance information from the distance calculating unit 66, the distance sensing control information from the distance sensing control information decoding unit 64, the audio data from the signal processing unit 201-2, the metadata from the metadata decoding unit 63, and the listening position information from the user interface 65, and supplies the generated metadata and the audio data to the 3D audio rendering processing unit 68.
[0351] The reverb processing unit 202-4 generates the metadata and the audio data of the wet component based on the distance information from the distance calculating unit 66, the distance sensing control information from the distance sensing control information decoding unit 64, the audio data from the signal processing unit 201-3, the metadata from the metadata decoding unit 63, and the listening position information from the user interface 65, and supplies the generated metadata and the audio data to the 3D audio rendering processing unit 68.
[0352] In the reverb processing unit 202-2, the reverb processing unit 202-3, and the reverb processing unit 202-4, the same processing as in the reverb processing unit 202-1 is performed, and the metadata and the audio data of the wet component are generated.
[0353] Further, hereinafter, the reverb processing units 202-1 to 202-4 will be simply referred to as reverb processing units 202 in a case where it is not particularly required to distinguish the reverb processing units 202-1 to 202-4.
[0354] In the distance sense control processing unit 67, none of the reverb processing units 202 can function, or one or more of the reverb processing units 202 can function.
[0355] Therefore, for example, the distance sense control processing unit 67 can include a reverb processing unit 202 that generates a wet component (dry component) located on the right and left sides of the object, and a reverb processing unit 202 that generates a wet component located on the upper and lower sides of the object.
[0356] As described above, the content creator can freely specify each signal processing step configuring the distance sense control processing and the order of executing the signal processing steps. Therefore, it is possible to realize distance sense control based on the intention of the content creator.
[0357] <Another example of distance sense control information>
[0358] Further, in a case where the configuration of the distance sense control processing unit 67 is freely changeable (specified) as illustrated in Figure 21 , for example, the distance sense control information has a configuration as illustrated in Figure 22 .
[0359] In the example illustrated in Figure 22 , "num_objs" indicates the number of objects included in the content, and in the control information in terms of distance, as many as the number of objects num_objs, a flag "isDistanceRenderFlg" indicating whether or not the object is a target of distance sense control is included.
[0360] Note that these number of objects num_objs and the flag isDistanceRenderFlg are similar to those in the example illustrated in Figure 19 , and thus the description thereof is omitted.
[0361] In a case where the value of the flag is the value of isDistanceRenderFlg of the i-th object is "1", the distance sense control information includes id information "proc_id" indicating signal processing and parameter configuration information configuring each signal processing step of the distance sense control processing to be executed on the object.
[0362] That is, for example, according to the id information "proc_id" indicating the jth (where 0 ≤ j < 4) signal processing, the parameter configuration information "DistanceRender_Attn()" of the gain control processing, the parameter configuration information "DistanceRender_Filt()" of the filter processing, the parameter configuration information "DistanceRender_Revb()" of the reverberation processing, or the parameter configuration information "DistanceRender_UserDefine()" of the user-defined processing is included in the distance sensing control information.
[0363] Specifically, for example, in a case where the id information "proc_id" is "ATTN" indicating the gain control processing, the parameter configuration information "DistanceRender_Attn()" of the gain control processing is included in the distance sensing control information.
[0364] Note that the parameter configuration information "DistanceRender_Attn()", "DistanceRender_Filt()", and "DistanceRender_Revb()" are similar to the case in Figure 11 , and thus the description thereof is omitted.
[0365] Further, the parameter configuration information "DistanceRender_UserDefine()" indicates parameter configuration information indicating a control rule of a parameter used in a user-defined processing, which is a signal processing arbitrarily defined by a user.
[0366] Thus, in this example, in addition to the gain control processing, the filter processing, and the reverberation processing, a user-defined processing individually defined by a user can be added as a signal processing configuring the distance sensing control processing.
[0367] Note that here, a case where the number of signal processing steps configuring the distance sensing control processing is four has been described as an example, but the number of signal processing steps configuring the distance sensing control processing can be any number.
[0368] In the distance sensing control information illustrated in Figure 22 , for example, when the 0th signal processing configuring the distance sensing control processing is set to the gain control processing, the first signal processing is set to the filter processing by a high shelf filter, the second signal processing is set to the filter processing by a low shelf filter, and the third signal processing is set to the reverberation processing, a distance sensing control processing unit 67 having the same configuration as that illustrated in Figure 3 is realized.
[0369] In this case, in the distance sensing control information illustrated inFigure 21 In the distance sensing control processing unit 67 shown, the signal processing units 201-1 to 201-3 and the reverberation processing unit 202-4 are implemented, and the reverberation processing units 202-1 to 202-3 are not implemented (not functioning).
[0370] Then, the signal processing units 201-1 to 201-3 and the reverberation processing unit 202-4 function as Figure 3 the gain control unit 101, the high shelf filter processing unit 102, the low shelf filter processing unit 103, and the reverberation processing unit 104 shown.
[0371] As described above, even in the case where the distance sensing control information has Figure 22 the configuration shown in FIG. 6, basically, the encoding device 11 performs the encoding processing described with reference to Figure 15 FIG. 5, and the decoding device 51 performs the decoding processing described with reference to Figure 16 FIG. 7.
[0372] However, in the encoding processing, for example, in step S13, for each object, it is determined whether the object is subjected to the distance sensing control processing, the configuration of the distance sensing control processing, and the like, and in step S14, the distance sensing control information having the configuration shown in Figure 22 FIG. 6 is encoded.
[0373] On the other hand, in the decoding processing, in step S47, the configuration of the distance sensing control processing unit 67 is determined for each object from the distance sensing control information having the configuration shown in Figure 22 FIG. 6, and the distance sensing control processing is appropriately performed.
[0374] As described above, according to the present technology, the distance sensing control information is transmitted to the decoding side together with the audio data of the object in accordance with the settings and the like of the content creator, and thus it is possible to implement the distance sensing control based on the intention of the content creator in the object-based audio.
[0375] <Configuration example of computer>
[0376] Incidentally, the series of processing described above can be executed by hardware, but can also be executed by software. In the case where the series of processing is executed by software, a program configuring the software is installed in a computer. Here, the computer includes, for example, a computer incorporating a dedicated hardware, a general-purpose personal computer capable of executing various functions by installing various programs, and the like.
[0377] Figure 23 is a block diagram showing a configuration example of hardware of a computer that executes the series of processing described above by a program.
[0378] In a computer, the central processing unit (CPU) 501, read-only memory (ROM) 502, and random access memory (RAM) 503 are interconnected via bus 504.
[0379] The input / output interface 505 is further connected to the bus 504. The input unit 506, output unit 507, recording unit 508, communication unit 509, and driver 510 are connected to the input / output interface 505.
[0380] Input unit 506 includes a keyboard, mouse, microphone, imaging element, etc. Output unit 507 includes a display, speaker, etc. Recording unit 508 includes a hard disk, non-volatile memory, etc. Communication unit 509 includes a network interface, etc. Driver 510 drives removable recording media 511 such as disk, optical disk, magneto-optical disk, or semiconductor memory.
[0381] In a computer configured as described above, for example, the above series of processes are performed in such a way that the CPU 501 loads the program recorded in the recording unit 508 into the RAM 503 via the input / output interface 505 and the bus 504, and executes the program.
[0382] For example, a program executed by a computer (CPU 501) can be recorded and provided on a removable recording medium 511, such as a packaging medium. Furthermore, the program can be provided via wired or wireless transmission media such as a local area network, the Internet, or digital satellite broadcasting.
[0383] In a computer, a program is installed in a recording unit 508 via an input / output interface 505 by installing a removable recording medium 511 into a drive 510. Alternatively, the program can be received by a communication unit 509 and installed in the recording unit 508 via a wired or wireless transmission medium. Furthermore, the program can be pre-installed in a ROM 502 or in the recording unit 508.
[0384] It should be noted that a program executed by a computer can be a program in which processing is performed sequentially in the order described in this specification, or a program in which processing is performed in parallel or at necessary time intervals, such as when a call is made.
[0385] Furthermore, the implementation of this technology is not limited to the above-described implementation, and various modifications can be made without departing from the spirit of this technology.
[0386] For example, this technology can be configured as cloud computing, where a function is shared and processed collaboratively by multiple devices via a network.
[0387] Furthermore, each step described in the flowchart above can be performed by one device or shared by multiple devices.
[0388] Further, in a case where one step includes a plurality of processes, the plurality of processes included in one step can be executed by one device or shared by a plurality of devices.
[0389] Further, the present technology can have the following configuration. (1)
[0391] An encoding device includes:
[0392] an object encoding unit that encodes audio data of an object;
[0393] a metadata encoding unit that encodes metadata including position information of the object;
[0394] a distance sensing control information determination unit that determines distance sensing control information used for distance sensing control processing performed on the audio data;
[0395] a distance sensing control information encoding unit that encodes the distance sensing control information; and
[0396] a multiplexer that multiplexes the encoded audio data, the encoded metadata, and the encoded distance sensing control information to generate encoded data. (2)
[0398] The encoding device according to (1),
[0399] wherein the distance sensing control information includes control rule information used for obtaining a parameter used in the distance sensing control processing. (3)
[0401] The encoding device according to (2),
[0402] wherein the parameter changes according to a distance from a listening position to the object. (4)
[0404] The encoding device according to (2) or (3),
[0405] wherein the control rule information is an index indicating a function or a table used for obtaining the parameter. (5)
[0407] The encoding device according to any one of (2) to (4),
[0408] wherein the distance sensing control information includes configuration information indicating one or a plurality of processing steps combined to implement the distance sensing control processing. (6)
[0410] The encoding device according to (5),
[0411] The configuration information is information indicating one or a plurality of processing steps and an order of executing the one or a plurality of processing steps. (7)
[0413] The encoding device according to any one of (5) to (6),
[0414] The processing is gain control processing, filtering processing, or reverberation processing. (8)
[0416] The encoding device according to any one of (1) to (7),
[0417] The distance sensing control information encoding unit encodes distance sensing control information for each of a plurality of objects. (9)
[0419] The encoding device according to any one of (1) to (7),
[0420] The distance sensing control information encoding unit encodes distance sensing control information for each object group including one or a plurality of objects. (10)
[0422] An encoding method performed by an encoding device, the method comprising:
[0423] encoding audio data of an object;
[0424] encoding metadata including position information of the object;
[0425] determining distance sensing control information for distance sensing control processing performed on the audio data;
[0426] encoding the distance sensing control information; and
[0427] multiplexing the encoded audio data, the encoded metadata, and the encoded distance sensing control information to generate encoded data. (11)
[0429] A program for causing a computer to execute processing including:
[0430] encoding audio data of an object;
[0431] encoding metadata including position information of the object;
[0432] determining distance sensing control information for distance sensing control processing performed on the audio data;
[0433] encoding the distance sensing control information; and
[0434] multiplexing the encoded audio data, the encoded metadata, and the encoded distance-sensing control information to generate encoded data. (12)
[0436] A decoding apparatus comprising:
[0437] a demultiplexer that demultiplexes encoded data to extract encoded audio data of an object, encoded metadata including position information of the object, and encoded distance-sensing control information for a distance-sensing control process performed on the audio data;
[0438] an object decoding unit that decodes the encoded audio data;
[0439] a metadata decoding unit that decodes the encoded metadata;
[0440] a distance-sensing control information decoding unit that decodes the encoded distance-sensing control information;
[0441] a distance-sensing control process unit that performs the distance-sensing control process on the audio data of the object based on the distance-sensing control information; and
[0442] a rendering process unit that performs a reproduction process based on the audio data and the metadata obtained by the distance-sensing control process to generate reproduction audio data for reproducing a sound of the object. (13)
[0444] The decoding apparatus according to (12),
[0445] wherein the distance-sensing control process unit performs the distance-sensing control process based on a parameter obtained from control rule information included in the distance-sensing control information and a listening position. (14)
[0447] The decoding apparatus according to (13),
[0448] wherein the parameter changes according to a distance from the listening position to the object. (15)
[0450] The decoding apparatus according to (13) or (14),
[0451] wherein the distance-sensing control process unit adjusts the parameter according to a reproduction environment of the reproduction audio data. (16)
[0453] The decoding apparatus according to any one of (13) to (15),
[0454] The distance sensing control processing unit generates the audio data of the wet component of the object by distance sensing control processing based on the distance sensing control information. (17)
[0456] The decoding device according to (16),
[0457] The processing is gain control processing, filtering processing, or reverberation processing. (18)
[0459] The decoding device according to any one of (12) to (17),
[0460] The distance sensing control processing unit generates the audio data of the wet component of the object by distance sensing control processing. (19)
[0462] A decoding method performed by a decoding device, the method comprising:
[0463] demultiplexing encoded data to extract encoded audio data of an object, encoded metadata including position information of the object, and encoded distance sensing control information for distance sensing control processing performed on the audio data;
[0464] decoding the encoded audio data;
[0465] decoding the encoded metadata;
[0466] decoding the encoded distance sensing control information;
[0467] performing the distance sensing control processing on the audio data of the object based on the distance sensing control information; and
[0468] performing rendering processing based on the audio data and the metadata obtained by the distance sensing control processing to generate reproduction audio data for reproducing a sound of the object. (20)
[0470] A program for causing a computer to execute processing including:
[0471] demultiplexing encoded data to extract encoded audio data of an object, encoded metadata including position information of the object, and encoded distance sensing control information for distance sensing control processing performed on the audio data;
[0472] decoding the encoded audio data;
[0473] decoding the encoded metadata;
[0474] decoding the encoded distance-sensing control information;
[0475] performing the distance-sensing control processing on the audio data of the object based on the distance-sensing control information; and
[0476] performing a rendering processing based on the audio data and the metadata obtained through the distance-sensing control processing to generate reproduction audio data for reproducing a sound of the object.
[0477] List of Reference Signs
[0478] 11 encoding apparatus
[0479] 21 object encoding unit
[0480] 22 metadata encoding unit
[0481] 23 distance-sensing control information determination unit
[0482] 24 distance-sensing control information encoding unit
[0483] 25 multiplexer
[0484] 51 decoding apparatus
[0485] 61 demultiplexer
[0486] 62 object decoding unit
[0487] 63 metadata decoding unit
[0488] 64 distance-sensing control information decoding unit
[0489] 66 distance calculation unit
[0490] 67 distance-sensing control processing unit
[0491] 68 3D audio rendering processing unit
[0492] 101 gain control unit
[0493] 102 high shelf filter processing unit
[0494] 103 low shelf filter processing unit
[0495] 104 reverb processing unit
Claims
1. An encoding apparatus comprising: an object encoding unit that encodes audio data of objects; a metadata encoding unit that encodes metadata including position information of the objects, wherein the metadata is metadata of audio data of each of the objects, the position information indicates an absolute position of an object in a space, and the metadata includes gain information used for performing gain control of audio data of an object; a distance-sensing control information determination unit that determines distance-sensing control information used for a distance-sensing control process performed on the audio data, wherein the distance-sensing control information includes control rule information used for obtaining a parameter used in the distance-sensing control process and the parameter changes according to a distance from a listening position to the object; a distance-sensing control information encoding unit that encodes the distance-sensing control information; and a multiplexer that multiplexes the encoded audio data, the encoded metadata, and the encoded distance-sensing control information to generate encoded data. 2.The encoding apparatus according to claim 1, wherein the control rule information is information indicating a function or a table used for obtaining the parameter. 3.The encoding apparatus according to claim 1, wherein the distance-sensing control information includes configuration information indicating one or more processing steps combined to implement the distance-sensing control process. 4.The encoding apparatus according to claim 3, wherein, the configuration information is information indicating the one or more processing steps and an order in which the one or more processing steps are performed. 5.The encoding apparatus according to claim 3, wherein the processing is gain control processing, filtering processing, or reverberation processing. 6.The encoding apparatus according to claim 1, wherein, the distance-sensing control information encoding unit encodes the distance-sensing control information for each of a plurality of the objects. 7.The encoding apparatus according to claim 1, wherein the distance-sensing control information encoding unit encodes the distance-sensing control information for each object group including one or more of the objects. 8.An encoding method performed by an encoding apparatus, the method comprising: encoding audio data of objects; encoding metadata including position information of the objects, wherein the metadata is metadata of audio data of each of the objects, the position information indicates an absolute position of an object in a space, and the metadata includes gain information used for performing gain control of audio data of an object; determining distance-sensing control information used for a distance-sensing control process performed on the audio data, wherein the distance-sensing control information includes control rule information used for obtaining a parameter used in the distance-sensing control process and the parameter changes according to a distance from a listening position to the object; encoding the distance-sensing control information; and multiplexing the encoded audio data, the encoded metadata, and the encoded distance-sensing control information to generate encoded data. 9.A computer-readable storage medium storing a program for causing a computer to execute a process, the process comprising the steps of: encode audio data of an object; encode metadata including position information of the object, wherein the metadata is metadata of the audio data of each of the objects, the position information indicates an absolute position of an object in a space, and the metadata includes gain information used for performing gain control on the audio data of an object; determine distance-sensing control information used for a distance-sensing control process performed on the audio data, wherein the distance-sensing control information includes control rule information used for obtaining a parameter used in the distance-sensing control process and the parameter changes according to a distance from a listening position to the object; encode the distance-sensing control information; and multiplex the encoded audio data, the encoded metadata, and the encoded distance-sensing control information to generate encoded data.
10. A decoding device comprising: a demultiplexer that demultiplexes encoded data to extract encoded audio data of an object, encoded metadata including position information of the object, and encoded distance-sensing control information used for a distance-sensing control process performed on the audio data; an object decoding unit that decodes the encoded audio data; a metadata decoding unit that decodes the encoded metadata; a distance-sensing control information decoding unit that decodes the encoded distance-sensing control information; a distance-sensing control process unit that performs the distance-sensing control process on the audio data of the object based on the distance-sensing control information, wherein the distance-sensing control process unit performs the distance-sensing control process based on a parameter obtained from control rule information included in the distance-sensing control information and a listening position, and the parameter changes according to a distance from the listening position to the object, and wherein the distance-sensing control process unit is capable of obtaining the parameter based on the control rule information and distance information, and performs the distance-sensing control process on the audio data based on the obtained parameter, by which dry component audio data and wet component audio data of the object are generated; and a rendering process unit that performs a rendering process based on the audio data obtained by the distance-sensing control process and the metadata to generate reproduction audio data used for reproducing a sound of the object.
11. The decoding device according to claim 10, wherein the distance-sensing control process unit adjusts the parameter according to a reproduction environment of the reproduction audio data.
12. The decoding device according to claim 10, wherein the distance-sensing control process unit performs the distance-sensing control process based on the parameter, in which one or more processing steps indicated by the distance-sensing control information are combined.
13. The decoding device according to claim 12, wherein the process is a gain control process, a filtering process, or a reverb process.
14. The decoding device according to claim 10, wherein the distance-sensing control process unit generates the wet component audio data of the object by the distance-sensing control process.
15. A decoding method performed by a decoding apparatus, the method comprising: demultiplexing encoded data to extract encoded audio data of an object, encoded metadata including position information of the object, and encoded distance sensing control information for a distance sensing control process performed on the audio data; decoding the encoded audio data; decoding the encoded metadata; decoding the encoded distance sensing control information; performing the distance sensing control process on the audio data of the object based on the distance sensing control information, wherein the distance sensing control process is performed based on a parameter obtained from control rule information included in the distance sensing control information and a listening position and the parameter changes according to a distance from the listening position to the object, and wherein the parameter is obtained based on the control rule information and distance information and the distance sensing control process is performed on the audio data based on the obtained parameter, by which dry component audio data and wet component audio data of the object are generated; and performing a rendering process based on the audio data obtained by the distance sensing control process and the metadata to generate reproduction audio data for reproducing a sound of the object.
16. A computer-readable storage medium storing a program for causing a computer to execute a process, the process comprising the steps of: demultiplexing encoded data to extract encoded audio data of an object, encoded metadata including position information of the object, and encoded distance sensing control information for a distance sensing control process performed on the audio data; decoding the encoded audio data; decoding the encoded metadata; decoding the encoded distance sensing control information; performing the distance sensing control process on the audio data of the object based on the distance sensing control information, wherein the distance sensing control process is performed based on a parameter obtained from control rule information included in the distance sensing control information and a listening position and the parameter changes according to a distance from the listening position to the object, and wherein the parameter is obtained based on the control rule information and distance information and the distance sensing control process is performed on the audio data based on the obtained parameter, by which dry component audio data and wet component audio data of the object are generated; and performing a rendering process based on the audio data obtained by the distance sensing control process and the metadata to generate reproduction audio data for reproducing a sound of the object.
Citation Information
Patent Citations
Sound processing device and method, and program
WO2015107926A1
Audio signal processing method
US20160080884A1
Concept for generating an enhanced sound-field description or a modified sound field description using a multi-layer description
WO2019012133A1
Methods, apparatus and systems for 6DOF audio rendering and data representations and bitstream structures for 6DOF audio rendering
WO2019197404A1