Information processing device and method
By employing a direction-limited spatial masking threshold to limit sound source direction range and optimize bit allocation, the method addresses the challenge of reducing coding efficiency in MPEG-I Immersive Audio while maintaining audio quality, even with changing sound source directions.
Patent Information
- Application Number
- PCT/JP2024/038341
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-02
- Filing Date
- 2024-10-28
- Publication Date
- 2025-05-08
AI Technical Summary
Existing methods for reducing coding efficiency in MPEG-I Immersive Audio while maintaining subjective quality of reproduced audio face challenges, particularly when applying spatial masking techniques in pre-prepared bitstreams.
The implementation of a direction-limited spatial masking threshold that limits the sound source direction range, allowing for bit allocation based on this threshold to generate a direction-limited bitstream, which is then used to encode audio data.
This approach effectively suppresses reductions in coding efficiency while maintaining the subjective quality of reproduced audio, even when the sound source direction changes, by optimizing bit allocation based on the direction-limited spatial masking threshold.
Smart Images

Figure JP2024038341_08052025_PF_FP_ABST
Abstract
Description
Information processing device and method
[0001] The present disclosure relates to an information processing device and method, and more particularly to an information processing device and method that are capable of suppressing a decrease in encoding efficiency while suppressing a decrease in the subjective quality of reproduced audio.
[0002] Conventionally, MPEG (Moving Picture Experts Group)-I Immersive Audio has been used as a metadata format for arranging and rendering audio objects in a 3D space (see, for example, Non-Patent Document 1). In the case of MPEG-I Immersive Audio, audio objects are encoded using the MPEG-H 3D Audio codec. MPEG-I Immersive Audio assumes that a large number of audio objects exist in a 3D space. As the number of audio objects increases, the load on transmission processing and playback processing increases.
[0003] As one method for reducing the load of these processes, a method has been proposed for encoding audio objects used in 3D space, in which, in addition to frequency masking that allocates bits based on a psychoacoustic model, spatial masking, which extends masking to the spatial domain, is applied (see, for example, Non-Patent Document 2). In this method, when encoding audio objects, a spatial masking threshold is set according to the direction of each sound source relative to the listener, and bit allocation is performed to improve encoding efficiency while suppressing degradation of subjective quality. However, because the direction of the sound source is identified based on the position and direction of the listener and the spatial masking threshold is set, it has been difficult to use this method in a model in which a pre-prepared bitstream is played back.
[0004] In order to minimize degradation of subjective quality regardless of the sound source arrangement, for example, a method of encoding by applying the minimum value of the spatial masking threshold among all directions relative to the listener can be considered. By applying such a minimum value, degradation of the subjective quality of the reproduced sound of the bitstream can be suppressed regardless of the position or direction of the listener.
[0005] "MPEG‐I (Introductory element. Main element. Part 4: Audio)", WD4 of ISO / IEC 23090-4, MPEG-I immersive audio, WG06N00211, ISO 23090‐4:202#(X), ISO / IEC JTC 1 / SC 29 / WG 6, Date: 2023-05-26; Takuhiro Kato, Masayuki Nishiguchi, Kanji Watanabe, Shoichi Takane, Koji Abe, "Basic study on coding of 3D audio signals considering spatial masking effect of auditory perception", THE INSTITUTE OF ELECTRONICS, INFORMATION AND COMMUNICATION ENGINEERS, IEICE Technical Report, EA2018-145, SIP2018-151, SP2018-107(2019-03), IEICE Technical Report, vol. 118, no. 495, EA2018-145, pp. 271-278, March 2019
[0006] However, in this method, a bitstream is generated using the minimum value of the spatial masking threshold, which limits the reduction in the code amount and may result in a decrease in coding efficiency.
[0007] The present disclosure has been made in light of such circumstances, and makes it possible to suppress a decrease in coding efficiency while suppressing a decrease in the subjective quality of reproduced audio.
[0008] An information processing device according to one aspect of the present technology is an information processing device including an encoding unit that encodes audio data using bit allocation based on a direction-limited spatial masking threshold, which is a spatial masking threshold that limits the sound source direction range, and generates a direction-limited bit stream.
[0009] An information processing method according to one aspect of the present technology is an information processing method for encoding audio data using bit allocation based on a direction-limited spatial masking threshold, which is a spatial masking threshold that limits the range of sound source directions, to generate a direction-limited bit stream.
[0010] Another aspect of the information processing device of the present technology is an information processing device including: a bitstream acquisition unit that acquires a direction-limited bitstream in which audio data is encoded with bit allocation based on a direction-limited spatial masking threshold, which is a spatial masking threshold that limits the sound source direction range; and a decoding unit that decodes the acquired direction-limited bitstream.
[0011] Another aspect of the information processing method of the present technology is an information processing method for obtaining a direction-limited bit stream in which audio data is encoded using a bit allocation based on a direction-limited spatial masking threshold, which is a spatial masking threshold that limits the range of sound source directions, and decoding the obtained direction-limited bit stream.
[0012] In an information processing device and method according to one aspect of the present technology, audio data is encoded with a bit allocation based on a direction-limited spatial masking threshold, which is a spatial masking threshold that limits the range of sound source directions, and a direction-limited bitstream is generated.
[0013] In an information processing device and method according to another aspect of the present technology, a direction-limited bitstream is obtained in which audio data is encoded using a bit allocation based on a direction-limited spatial masking threshold, which is a spatial masking threshold that limits the range of sound source directions, and the obtained direction-limited bitstream is decoded.
[0014] 1 is a diagram illustrating an example of a playback device for MPEG-I immersive audio. FIG. 1 is a diagram illustrating spatial masking. FIG. 1 is a diagram illustrating an example of encoding to which spatial masking is applied. FIG. 2 is a diagram illustrating an example of encoding to which a minimum spatial masking threshold is applied. FIG. 2 is a diagram illustrating an example of a method for providing audio data. FIG. 3 is a diagram illustrating an example of spatial masking information. FIG. 4 is a diagram illustrating sound source direction range information. FIG. 5 is a diagram illustrating sound source direction range information. FIG. 6 is a diagram illustrating an example of quality ranking. FIG. 7 is a diagram illustrating an example of spatial masking information stored in an MPD. FIG. 8 is a diagram illustrating an example of description of spatial masking information in an MPD. FIG. 9 is a diagram illustrating an example of description of spatial masking information in an MPD. FIG. 10 is a diagram illustrating an example of description of spatial masking information in an MPD. FIG. 11 is a diagram illustrating an example of description of spatial masking information in an MP4 file. FIG. 12 is a diagram illustrating an example of storage of spatial masking information in an MP4 file. FIG. 13 is a diagram illustrating an example of storage of spatial masking information in an MP4 file. FIG. 14 is a diagram illustrating an example of storage of spatial masking information in an MP4 file. FIG. 15 is a diagram illustrating an example of storage of spatial masking information in an MP4 file. Fig. 1 is a flowchart showing an example of the flow of a bitstream acquisition process; Fig. 2 is a flowchart showing an example of the flow of a playback process; Fig. 3 is a flowchart showing an example of the flow of a bitstream acquisition process; Fig. 4 is a block diagram showing an example of the main configuration of a computer;
[0015] Hereinafter, modes for carrying out the present disclosure (hereinafter referred to as embodiments) will be described. The description will be made in the following order: 1. Literature etc. supporting technical content and technical terminology 2. Audio scene playback 3. Spatial masking applying a direction-restricted spatial masking threshold 4. First embodiment (file generation device) 5. Second embodiment (playback device) 6. Supplementary notes
[0016] <1. Literature, etc. supporting technical content and technical terminology> The scope of disclosure of the present technology includes not only the content described in the embodiments, but also the content described in the following non-patent documents, etc. that were publicly known at the time of filing, and the content of other documents referenced in the following non-patent documents.
[0017] Non-patent document 1: (described above) Non-patent document 2: (described above) Non-patent document 3: "Information technology - Dynamic adaptive streaming over HTTP (DASH) - Part 1: Media presentation description and segment formats", ISO 23009-1:2021(X), ISO / IEC JTC 1 / SC 29 / WG 3, Date: 2021-06-24 Non-patent document 4: "14496-12_Ed8_DIS_potential_improvements", ISO / IEC JTC 1 / SC 29 / WG 03 N22316, 2023-02-11 Non-patent document 5: "Information technology . High efficiency coding and media delivery in heterogeneous environments . Part 3: 3D audio", ISO / IEC FDIS 23008-3:2018(E), ISO / IEC JTC 1 / SC 29 / WG 11, 2016-10-12
[0018] In other words, the contents of the above-mentioned non-patent documents and the contents of other documents referenced in the above-mentioned non-patent documents are also used as the basis for determining the support requirements. For example, even if syntax, terminology, etc. described in the above-mentioned non-patent documents are not directly defined in this disclosure, they are considered to be within the scope of this disclosure and meet the support requirements of the claims. Similarly, for example, technical terms such as parsing, syntax, and semantics are considered to be within the scope of this disclosure and meet the support requirements of the claims, even if they are not directly defined in this disclosure.
[0019] 2. Playback of Audio Scenes MPEG-I Immersive Audio Conventionally, as described in Non-Patent Document 1, for example, MPEG (Moving Picture Experts Group)-I Immersive Audio has been used as a metadata format for arranging and rendering audio objects in a three-dimensional space (also referred to as 3D space in this specification). In the case of MPEG-I immersive audio, audio objects are encoded using the MPEG-H 3D Audio codec. MPEG-I immersive audio assumes that a large number of audio objects exist in a 3D space. In this specification, these audio objects are also referred to as sound sources.
[0020] 1 is a diagram showing an example of a playback device 10 for audio data that conforms to MPEG-I Immersive Audio. As shown in Fig. 1, the playback device 10 includes an MPEG-H 3DA decoder 11, an MPEG-I Audio renderer 12, and an input unit 13.
[0021] The playback device 10 is supplied with an MPEG-H 3D audio bitstream and an MPEG-I audio bitstream. The MPEG-I audio bitstream is a bitstream of metadata, such as an audio scene description and other rendering parameters, used in processing by the MPEG-I audio renderer 12. Hereinafter, the 3D space in which sound sources are located is also referred to as an audio scene. The audio scene description is a description that defines the audio scene. The MPEG-H 3D audio bitstream is a bitstream in which all sound sources in the 3D space are coded in accordance with the MPEG-H 3D Audio standard described in Non-Patent Document 2. The MPEG-H 3D audio bitstream is converted by the MPEG-H 3DA decoder 11 into PCM data for each sound source and supplied to the MPEG-I audio renderer 12. The PCM data is sampled at, for example, 48 kHz.
[0022] The input unit 13 receives input of information such as consumption environment information, updates to audio scenes being played back, user interaction information, and user location information, and supplies the received information to the MPEG-I audio renderer 12 .
[0023] The MPEG-I audio renderer 12 generates a 3D space based on the audio scene description in the MPEG-I audio bitstream, and places in that 3D space as sound sources a group of PCM data generated by the MPEG-H 3DA decoder 11 decoding the MPEG-H 3D audio bitstream. The MPEG-I audio renderer 12 then performs rendering based on user position information, and transmits the rendered data (Audio Output) to an audio output device such as headphones.
[0024] <Spatial Masking> Generally, as the number of audio objects increases, the amount of data, such as audio data (or a bitstream obtained by encoding the audio data), increases, resulting in a greater load on the transmission and playback processes (including decoding processes). Therefore, methods for suppressing the increase in the amount of audio data (or a reduction in the coding efficiency of the bitstream obtained by encoding the audio data) and reducing the processing load have been studied. For example, conventional audio coding methods use frequency masking, which allocates bits based on a psychoacoustic model. The ease (or difficulty) of audibility of audio data to listeners can vary depending on its frequency components. Frequency masking utilizes these characteristics to mask frequency components. That is, more bits are allocated to frequency components that are more audible to listeners, and the amount of bits for less audible frequency components is reduced for encoding. This makes it possible to improve the coding efficiency of audio data while suppressing a reduction in the subjective quality of the reproduced audio perceived by listeners.
[0025] In this specification, a parameter indicating the amount of bits to be reduced in such masking (a parameter indicating the degree to which the amount of bits should be reduced) is also referred to as a masking threshold. In other words, the masking threshold can be considered as the amount of bit allocation that can be reduced while maintaining a reduction in the subjective quality of the reproduced sound perceived by the listener within an acceptable range. For example, frequency components with a smaller masking threshold are frequency components that are easier for the listener to hear, and a reduction in the amount of allocated bits has a greater impact on the subjective quality of the reproduced sound perceived by the listener. Therefore, in order to suppress a reduction in the subjective quality of the reproduced sound perceived by the listener, it is necessary to allocate more bits to such frequency components. In contrast, frequency components with a larger masking threshold are frequency components that are harder for the listener to hear, and a reduction in the amount of allocated bits has a smaller impact on the subjective quality of the reproduced sound perceived by the listener. Therefore, for such frequency components, the amount of allocated bits can be reduced more while suppressing a reduction in the subjective quality of the reproduced sound perceived by the listener.
[0026] Furthermore, as one method for reducing the load of transmission processing and playback processing, a method has been devised in which, in addition to the above-mentioned frequency masking, spatial masking, which is an extension of the masking to the spatial domain, is applied to the coding of audio objects used in 3D space, as described in Non-Patent Document 2. In other words, spatial masking is a method for allocating bits according to the direction of a sound source relative to the listener (also referred to as sound source direction in this specification).
[0027] For example, as shown in Fig. 2, in a 3D space, sound sources 21 to 24 are placed around a listener 20, and a certain reproduced sound is output from one of them. In this case, the ease with which the reproduced sound can be heard by the listener 20 may differ depending on which sound source, from sound source 21 to sound source 24, the reproduced sound is emitted from. In other words, the ease (or difficulty) of hearing the reproduced sound by the listener 20 depends on the direction of the sound source relative to the listener 20. Note that in Fig. 2, the arrow of the listener 20 indicates the orientation of the listener 20.
[0028] Furthermore, when reproduced sounds are output from multiple sound sources arranged in a 3D space, the reproduced sounds from one sound source may make it difficult for the listener 20 to hear the reproduced sounds from the other sound sources. The ease (or difficulty) of hearing such multiple reproduced sounds for the listener 20 depends on the direction of each sound source.
[0029] In spatial masking, masking is performed by utilizing such characteristics that depend on the sound source direction. In this specification, the masking threshold applied in such spatial masking is also referred to as the spatial masking threshold. In other words, the spatial masking threshold can also be said to be a parameter representing the amount of bits to be reduced in spatial masking (a parameter indicating how much the amount of bits should be reduced). In other words, the spatial masking threshold can also be said to be a parameter representing the amount of bit allocation that can be reduced while maintaining the reduction in the subjective quality of the reproduced sound to the listener within an acceptable range in spatial masking.
[0030] For example, if the spatial masking threshold is smaller than the optimal value, spatial masking may prevent a reduction in the subjective quality of the reproduced audio perceived by the listener (beyond an acceptable level), but may reduce the effect of reducing the code amount of the audio data bitstream. Also, if the spatial masking threshold is larger than the optimal value, spatial masking may sufficiently reduce the code amount of the audio data bitstream, but may reduce the subjective quality of the reproduced audio perceived by the listener (beyond an acceptable level). Therefore, the spatial masking threshold is set to an optimal value.
[0031] As described in Non-Patent Document 2, this optimal value depends on the sound source direction. In other words, the spatial masking threshold may change depending on the sound source direction. For example, a sound source direction with a smaller spatial masking threshold is a sound source direction in which the reproduced sound is easier for the listener to hear, and a reduction in the amount of allocated bits has a greater impact on the subjective quality of the reproduced sound perceived by the listener. Therefore, in order to suppress a reduction in the subjective quality of the reproduced sound perceived by the listener, it is necessary to allocate more bits to the sound data of such a sound source direction. On the other hand, a sound source direction with a larger spatial masking threshold is a sound source direction in which the reproduced sound is harder for the listener to hear, and a reduction in the amount of allocated bits has a smaller impact on the subjective quality of the reproduced sound perceived by the listener. Therefore, the amount of allocated bits can be reduced more for sound data of such a sound source direction while suppressing a reduction in the subjective quality of the reproduced sound perceived by the listener.
[0032] In spatial masking, there are sound sources that are masked and sound sources that mask. In this specification, the "masked sound source" is also referred to as the masker. The "masking sound source" is also referred to as the masker. Therefore, it can be said that the spatial masking threshold is determined by the sound source directions of the masker and the masker. In Non-Patent Document 2, the spatial masking threshold of a masker when one or two sound sources are fixed in a certain direction as a masker is derived through experiments for each direction and each frequency of the masker. The following properties of this spatial masking threshold have been reported.
[0033] For example, when there is one masker, the threshold decreases as the direction of the masker moves away from the masker. Furthermore, when the masker is in a direction symmetrical between the front and back of the masker, the threshold increases compared to the surrounding directions (thresholds are folded back in the frontal plane). Furthermore, the lower the center frequency of the masker, the more pronounced the increase in threshold when the masker is in a direction symmetrical between the front and back of the masker. For example, when there are two maskers, the threshold when these sound source signals are simultaneously present can be expressed as the sum of the spatial masking thresholds of the sound source signals present singly in different directions.
[0034] The sound source directions of such maskers and maskers are determined by the arrangement of the sound sources and the position and orientation of the listener. Audio data encoding processing when spatial masking is used is performed according to a flow as shown in Fig. 3. For example, when encoding audio data of sound source A (RAW bitstream (sound source A)), audio data of sound source B (RAW bitstream (sound source B)), and audio data of sound source C (RAW bitstream (sound source C)), encoding preprocessing 31 is performed using these data, position information of each sound source, and position and orientation information of the listener.
[0035] In this encoding preprocessing 31, a process of acquiring a spatial masking threshold and a process of calculating a masking threshold and determining bit allocation are performed. For example, when sound source A is used as a masker, in the process of acquiring the spatial masking threshold, the direction of each sound source as seen from the listener is determined, and spatial masking threshold information is acquired for each of sound source A as a masker and sound source B and sound source C as maskers. Then, in the process of calculating the masking threshold and determining bit allocation, first, a masking threshold is derived by adding spatial masking to the conventional masking threshold calculation. Then, using the sample data of sound source B (sound pressure for each frequency) and the spatial masking threshold derived from the directions of masker sound source A and masker sound source B, it is calculated how many dB are effective for each of the sample data of sound source A (sound pressure for each frequency). Then, using the sample data (sound pressure for each frequency) of sound source C and the spatial masking threshold derived from the directions of masker sound source A and masker sound source C, it is calculated how many dB will be effective for each of the sample data (sound pressure for each frequency) of sound source A. The obtained results of sound source B and sound source C are then synthesized (added), and how many dB will be effective is calculated, and the required number of bits is determined.
[0036] After the pre-processing 31 for encoding is completed, the encoding 32 is executed. In this encoding 32, the angular frequency is encoded based on the required number of bits. Then, the bit allocation is optimized to adjust the sound quality after encoding.
[0037] In this way, by determining the positional relationship between the sound source and the listener, it becomes possible to derive an optimal spatial masking threshold. However, in an audio scene in which the position of the sound source and the position and direction of the listener may change, the sound source direction changes due to these changes. In other words, in such an audio scene, it can be said that there are an infinite number of patterns of sound source direction. Therefore, in such a case, it is not possible to derive an optimal spatial masking threshold until the positional relationship between the sound source and the listener is determined. Therefore, for example, while the above-mentioned spatial masking can be applied to a model in which audio data is encoded and distributed instantly (in real time) in response to changes in the position of the sound source and the position and direction of the listener, it has been difficult to apply the above-mentioned spatial masking to a model in which a pre-prepared bitstream is played, such as MPEG-DASH (Moving Picture Experts Group Dynamic Adaptive Streaming over HTTP (Hypertext Transfer Protocol)).
[0038] For example, suppose spatial masking is performed on a sound source whose direction is in direction A, and a bitstream of the audio data is generated. Then, suppose a listener moves and the direction of the sound source changes to direction B. When audio is played back using the generated bitstream, the spatial masking threshold may also change because the sound source direction has changed from direction A to direction B. For example, if the spatial masking threshold in direction B is smaller than the spatial masking threshold in direction A, it can be said that the bit allocation amount that can be reduced in direction B while maintaining a reduction in the subjective quality of the reproduced audio to the listener within an acceptable range is smaller than in direction A. In other words, in this case, the bitstream coding amount is reduced more than necessary, which may result in a reduction in the subjective quality of the reproduced audio to the listener (below an acceptable level). Conversely, if the spatial masking threshold in direction B is larger than the spatial masking threshold in direction A, it can be said that the bit allocation amount that can be reduced in direction B while maintaining a reduction in the subjective quality of the reproduced audio to the listener within an acceptable range is larger than in direction A. That is, in this case, there is a risk that the effect of reducing the amount of code in the bitstream will not be sufficient (that is, the amount of code in the bitstream will increase compared to when the spatial masking threshold is an optimal value).
[0039] In order to minimize the degradation of subjective quality regardless of the sound source arrangement, for example, a method of encoding by applying the minimum value of the spatial masking threshold among all directions relative to the listener can be considered. In this specification, this minimum value is also referred to as the minimum spatial masking threshold. The encoding process of audio data when applying spatial masking using this minimum spatial masking threshold is performed according to the flow shown in Figure 4.
[0040] That is, in this case, in the encoding preprocessing 31, a process of calculating a masking threshold and determining bit allocation is performed. For example, when sound source A is used as the masker, first, a masking threshold is derived by adding spatial masking to the conventional masking threshold calculation. Then, using the sample data (sound pressure for each frequency) of sound source B and the minimum spatial masking threshold, it is calculated how many dB each of the sample data (sound pressure for each frequency) of sound source A will be effective. Then, using the sample data (sound pressure for each frequency) of sound source C and the minimum spatial masking threshold, it is calculated how many dB each of the sample data (sound pressure for each frequency) of sound source A will be effective. Then, the obtained results of sound source B and sound source C are combined (added), how many dB each will be effective is calculated, and the number of required bits is determined. Then, when the encoding preprocessing 31 is completed, encoding 32 is performed.
[0041] In this way, by applying a minimum spatial masking threshold, spatial masking can be applied so as to suppress a reduction in the subjective quality of the reproduced sound perceived by the listener, regardless of the sound source direction (i.e., regardless of the sound source position or the listener's position and direction).
[0042] However, in this case, there is a risk that the spatial masking threshold will be smaller than the optimal value, which may limit the effect of reducing the code amount by spatial masking and make it difficult to sufficiently reduce the code amount of the audio data bitstream. In other words, there is a risk that the code amount of the bitstream will increase compared to when the spatial masking threshold is the optimal value. Furthermore, there is a risk that the increase in the code amount of the bitstream will increase the load of the transmission process and the playback process.
[0043] In other words, in this case, more bits are allocated to information that has less impact on the subjective quality of the reproduced audio, i.e., information that is unnecessary for the listener (for the reproduced audio). Therefore, if the code amount is fixed, there is a risk that the subjective quality of the reproduced audio will be lower in a bitstream to which this minimum spatial masking threshold is applied than in a bitstream to which the optimal spatial masking threshold is applied.
[0044] As described above, applying the minimum spatial masking threshold may result in a decrease in bitstream coding efficiency compared to when the spatial masking threshold is an optimal value.
[0045] <3. Spatial masking using direction-restricted spatial masking threshold> <Method 1> Therefore, as shown in the top row of the table in Fig. 5, a direction-restricted bitstream that is coded by applying a direction-restricted spatial masking threshold that restricts the sound source direction is supplied (Method 1). In other words, a bitstream that can be used only for some sound source directions is generated, and can be used only when the sound source directions of all sound sources during playback are included in the target sound source direction.
[0046] In this specification, such a bitstream of audio data in which the target sound source direction is limited to a certain part is also referred to as a "direction-limited bitstream." The range of the target sound source direction is also referred to as a "sound source direction range." A direction-limited bitstream can also be said to be a bitstream in which the sound source direction range is limited to a certain part of the sound source directions. The spatial masking threshold for the limited sound source direction range is also referred to as a "direction-limited spatial masking threshold." The minimum value of the direction-limited spatial masking threshold, i.e., the minimum value of the spatial masking threshold within the limited sound source direction range, is also referred to as a "minimum direction-limited spatial masking threshold." In other words, a direction-limited bitstream can also be said to be a bitstream of audio data encoded by applying a direction-limited spatial masking threshold (or a minimum direction-limited spatial masking threshold).
[0047] Note that a bitstream of audio data when the sound source direction range is not limited (omnidirectional) is also referred to as an "omnidirectional bitstream." The spatial masking threshold for this unlimited sound source direction range is also referred to as an "omnidirectional spatial masking threshold." The minimum value of the omnidirectional spatial masking threshold, i.e., the minimum value of the spatial masking threshold among all sound source directions, is also referred to as a "minimum omnidirectional spatial masking threshold." In other words, an omnidirectional bitstream can be said to be a bitstream of audio data encoded by applying an omnidirectional spatial masking threshold (or minimum omnidirectional spatial masking threshold).
[0048] For example, the first information processing device may include an encoding unit that encodes audio data with bit allocation based on a direction-limited spatial masking threshold, which is a spatial masking threshold that limits the sound source direction range, and generates a direction-limited bit stream.
[0049] In addition, the first information processing method executed by the first information processing device includes encoding audio data with bit allocation based on a direction-limited spatial masking threshold, which is a spatial masking threshold that limits the sound source direction range, and generating a direction-limited bit stream.
[0050] In addition, the first program causes a computer (first information processing device) to execute a process of encoding audio data with a bit allocation based on a direction-limited spatial masking threshold, which is a spatial masking threshold that limits the sound source direction range, and generating a direction-limited bit stream.
[0051] For example, the second information processing device may include a bitstream acquisition unit that acquires a direction-limited bitstream in which audio data is encoded with a bit allocation based on a direction-limited spatial masking threshold, which is a spatial masking threshold that limits the sound source direction range, and a decoding unit that decodes the acquired direction-limited bitstream.
[0052] In addition, the second information processing method executed by the second information processing device includes obtaining a direction-limited bit stream in which audio data is encoded with a bit allocation based on a direction-limited spatial masking threshold, which is a spatial masking threshold that limits the sound source direction range, and decoding the obtained direction-limited bit stream.
[0053] In addition, the second program acquires a direction-limited bit stream in which audio data is encoded with a bit allocation based on a direction-limited spatial masking threshold, which is a spatial masking threshold that limits the sound source direction range, and causes a computer (second information processing device) to execute a process of decoding the acquired direction-limited bit stream.
[0054] Generally, by limiting the sound source direction range as described above, i.e., by limiting the application (sound source directions of all sound sources during playback), it is expected that the minimum value of the spatial masking threshold will be larger than when the sound source direction range is omnidirectional. In other words, the magnitude of the minimum direction-limited spatial masking threshold is equal to or greater than the minimum omnidirectional spatial masking threshold. Therefore, it is expected that the bitstream coding amount can be further reduced while suppressing a reduction in the subjective quality of the reproduced sound perceived by the listener. In other words, it is possible to suppress a reduction in the coding efficiency of the audio data bitstream while suppressing a reduction in the subjective quality of the reproduced sound perceived by the listener.
[0055] Note that the number of sound sources may be any number, and may be either single or multiple. In other words, the number of pieces of audio data included in the bitstream may be any number, and may be either single or multiple. Whether there is a single sound source or multiple sound sources, all of the sound sources must be included in the (limited) sound source direction range. In other words, a condition for using the bitstream is that all sound sources corresponding to the audio data included in the bitstream are included in the (limited) sound source direction range.
[0056] Therefore, a direction-restricted bitstream can be said to be a bitstream that can provide a listener with a good quality experience when certain conditions are met. This "specific condition" means that "all sound sources corresponding to the audio data contained in the bitstream are included in a (limited) sound source direction range." In other words, a direction-restricted bitstream can be said to be a bitstream that can provide a listener with a good quality experience when all sound sources are present within a limited sound source direction range. This "good quality experience" means that "deterioration of the subjective quality of the reproduced audio by the listener is suppressed."
[0057] <Method 1-1> When the above-mentioned Method 1 is applied, spatial masking information regarding spatial masking may be supplied for a direction-restricted bitstream, as shown in the second row from the top of the table in Fig. 5 (Method 1-1). In this specification, this "spatial masking information" refers to information related to spatial masking. The content of the spatial masking information may be any information related to spatial masking. For example, the spatial masking information may include direction restriction identification information, sound source direction range information, and sound source direction range quality information, which will be described later.
[0058] As described above, in order to play back a direction-limited bitstream with appropriate quality, it is necessary to satisfy "specific conditions." Therefore, by providing spatial masking information for the direction-limited bitstream, a direction-limited bitstream that can be played back so as to satisfy the specific conditions can be selected based on the spatial masking information. In other words, a direction-limited bitstream that includes all audio data (sound sources) to be played back and in which all of the sound sources are present within the sound source direction range is selected and played back based on the spatial masking information.
[0059] For example, the first information processing device may further include a spatial masking information generating unit that generates spatial masking information related to spatial masking for the direction-restricted bitstream. Also, in the second information processing device, the bitstream obtaining unit may obtain the direction-restricted bitstream based on the spatial masking information related to spatial masking.
[0060] By doing so, a playback device that plays back a direction-restricted bitstream can more easily play back the direction-restricted bitstream so as to satisfy a specific condition, i.e., it can perform playback processing so as to suppress degradation of the subjective quality of the reproduced audio as perceived by a listener.
[0061] <Method 1-1-1> When the above-mentioned Method 1-1 is applied, direction restriction identification information may be supplied as spatial masking information (Method 1-1-1), as shown in the third row from the top of the table in Fig. 5 . In other words, the spatial masking information may include direction restriction identification information. This "direction restriction identification information" is identification information that indicates whether or not the bit stream of audio data is a direction restriction bit stream.
[0062] For example, in the first information processing device and the second information processing device, the spatial masking information may include direction restriction identification information for identifying that the bit stream of audio data is a direction restricted bit stream.
[0063] For example, the direction restriction identification information may be flag information whose value is true or false. For example, when the direction restriction identification information is true, it may indicate that it is a direction-restricted bitstream, and when the direction restriction identification information is false, it may indicate that it is not a direction-restricted bitstream (i.e., it is an omnidirectional bitstream). Also, true and false may be reversed. Furthermore, when this direction restriction identification information does not exist, it may be equivalent to when the direction restriction identification information is false, or it may be equivalent to when the direction restriction identification information is true. In other words, the direction restriction identification information can be said to be a flag that indicates "this is a bitstream that can provide a good quality experience under specific conditions."
[0064] By including such direction restriction identification information in the spatial masking information, a playback device can easily identify whether the bitstream corresponding to the spatial masking information is a direction restricted bitstream. Therefore, the playback device can more easily play back the direction restricted bitstream so as to satisfy specific conditions. That is, the playback device can perform playback processing so as to suppress degradation of the subjective quality of the reproduced sound perceived by a listener.
[0065] <Method 1-1-2> When the above-mentioned Method 1-1 is applied, sound source direction range information may be supplied as spatial masking information (Method 1-1-2), as shown in the fourth row from the top of the table in Fig. 5 . That is, the sound source direction range information may be included in the spatial masking information. This "sound source direction range information" is information indicating the sound source direction range of the corresponding bitstream.
[0066] For example, in the first information processing device and the second information processing device, the spatial masking information may include sound source direction range information regarding a sound source direction range.
[0067] In this sound source direction range information, the sound source direction range may be defined using any expression. That is, the sound source direction range information may include any information. For example, the sound source direction range may be defined using the direction of the center of the sound source direction range in the horizontal direction (i.e., the azimuth angle of the center of the sound source direction range), the direction of the center of the sound source direction range in the vertical direction (i.e., the elevation angle or depression angle of the center of the sound source direction range), the angular width of the sound source direction range in the horizontal direction (i.e., the azimuth angle width of the sound source direction range), and the angular width of the sound source direction range in the vertical direction (i.e., the width of the elevation angle or depression angle of the sound source direction range). That is, the sound source direction range information may include information indicating the direction of the center of the sound source direction range in the horizontal direction (i.e., the azimuth angle of the center of the sound source direction range), information indicating the direction of the center of the sound source direction range in the vertical direction (i.e., the elevation angle or depression angle of the center of the sound source direction range), information indicating the angular width of the sound source direction range in the horizontal direction (i.e., the azimuth angle width of the sound source direction range), and information indicating the angular width of the sound source direction range in the vertical direction (i.e., the width of the elevation angle or depression angle of the sound source direction range).
[0068] The number of sound source direction ranges defined by the sound source direction range may be any number. For example, it may be one or more. Furthermore, if the spatial masking information does not include sound source direction range information, it may indicate that the sound source direction range is omnidirectional. For example, only when the direction limitation identification information is a value indicating a direction limited bitstream, the spatial masking information may include sound source direction range information.
[0069] FIG. 6 is a diagram showing an example of spatial masking information. As shown in the table in FIG. 6, sound source direction range information may be set for each MP4 file that stores an audio data bitstream. In the example of FIG. 6, 6DoFAudio1.mp4, 6DoFAudio5.mp4, 6DoFAudio2.mp4, 6DoFAudio3.mp4, and 6DoFAudio4.mp4 are prepared as MP4 files that store audio data bitstreams. FIG. 7 shows an example of a sound source direction range 102 for a listener 101. In FIG. 7, dotted lines indicate example sound source directions, and the range shown in gray indicates the sound source direction range 102.
[0070] In the table of Fig. 6, the sound source direction range information for 6DoFAudio1.mp4 and 6DoFAudio5.mp4 indicates that the sound source direction range is omnidirectional. That is, as shown in A of Fig. 7, the sound source direction range 102 covers all directions around the listener 101. Also, in the table of Fig. 6, the sound source direction range information for 6DoFAudio2.mp4 indicates that the center direction (azimuth angle) of the sound source direction range is directly in front of the listener 101, and that the angular width of the sound source direction range (azimuth angle width, elevation angle or depression angle width) is 90 degrees. That is, as shown in B of Fig. 7, the sound source direction range 102 is directly in front of the listener 101. 6, the sound source direction range information for 6DoFAudio3.mp4 indicates that the center direction (azimuth) of the sound source direction range is 90 degrees to the right of the listener 101, and the angular width of the sound source direction range (azimuth width, elevation angle or depression angle width) is 90 degrees. In other words, as shown in C of FIG. 7, the right side of the listener 101 becomes the sound source direction range 102. Also, in the table of FIG. 6, the sound source direction range information for 6DoFAudio4.mp4 indicates that the center direction (azimuth) of the sound source direction range is 90 degrees to the left of the listener 101, and the angular width of the sound source direction range (azimuth angle width, elevation angle or depression angle width) is 90 degrees. In other words, as shown in D of FIG. 7, the left side of the listener 101 becomes the sound source direction range 102.
[0071] For example, when bitstreams of multiple sound sources are stored in each MP4 file, all of those sound sources are located within the above-mentioned sound source direction range. Figure 8 is a diagram showing an example of the positions of sound sources 103-1 and 103-2 when those bitstreams are stored in an MP4 file. For example, sound sources 103-1 and 103-2 of the bitstreams included in 6DoFAudio1.mp4 and 6DoFAudio5.mp4 are located within an omnidirectional sound source direction range 102, as shown in A of Figure 8. Furthermore, sound sources 103-1 and 103-2 of the bitstreams included in 6DoFAudio2.mp4 are located within a sound source direction range 102 directly in front of the listener 101, as shown in B of Figure 8. Furthermore, sound sources 103-1 and 103-2 of the bitstreams included in 6DoFAudio3.mp4 are located within a sound source direction range 102 to the right of the listener 101, as shown in C of Figure 8. 8D, the sound source 103-1 and sound source 103-2 of the bitstream included in 6DoFAudio4.mp4 are located within the sound source direction range 102 on the left side of the listener 101.
[0072] By including such sound source direction range information in the spatial masking information, the playback device can more easily identify the sound source direction range. Therefore, the playback device can more easily play back the direction-limited bitstream so as to satisfy specific conditions. In other words, the playback device can perform playback processing so as to suppress degradation of the subjective quality of the reproduced sound perceived by the listener.
[0073] <Method 1-1-3> When the above-mentioned Method 1-1 is applied, as shown in the fifth row from the top of the table in Fig. 5, sound source direction range quality information may be supplied as spatial masking information (Method 1-1-3). In other words, the spatial masking information may include sound source direction range quality information. This "sound source direction range quality information" is information indicating the subjective quality of reproduced audio of the target bitstream when all sound source directions are within the sound source direction range.
[0074] For example, in the first information processing device and the second information processing device, the spatial masking information may include sound source direction range quality information regarding the subjective quality of the reproduced audio of the direction-limited bitstream when all sound sources are present within the sound source direction range.
[0075] The subjective quality of the reproduced audio may be expressed in any manner in this sound source direction range quality information. Furthermore, any number of pieces of sound source direction range quality information may be included in the spatial masking information. For example, the spatial masking information may include a single piece of sound source direction range quality information or multiple pieces of sound source direction range quality information. Furthermore, the spatial masking information may include a set of sound source direction range information and sound source direction range quality information. In other words, the sound source direction range quality information may indicate the subjective quality of the reproduced audio of the direction-limited bitstream when all sound sources are located within the sound source direction range defined by the sound source direction range information. Furthermore, when the sound source direction range is not specified (cannot be identified), the sound source direction range quality information may be regarded as omnidirectional quality information indicating the subjective quality of the reproduced audio regardless of the sound source direction.
[0076] As shown in the table of Figure 6, in-range quality information may be set for each MP4 file that stores an audio data bitstream. For example, the bitrate of 6DoFAudio1.mp4 is 5000000, and the in-range quality information indicates that the subjective quality of the reproduced audio is relatively high. Also, the bitrate of 6DoFAudio5.mp4 is 3000000, and the in-range quality information indicates that the subjective quality of the reproduced audio is lower than that of 6DoFAudio1.mp4.
[0077] By including such sound source direction range quality information in the spatial masking information, the playback device can more easily grasp the subjective quality of the reproduced sound of the bitstream. Therefore, the playback device can more easily reproduce the direction-limited bitstream so as to satisfy specific conditions. In other words, the playback device can perform playback processing so as to suppress a reduction in the subjective quality of the reproduced sound perceived by the listener.
[0078] <Method 1-1-3-1> When the above-mentioned Method 1-1-3 is applied, information specifying a bitstream of equivalent quality may be supplied as sound source direction range quality information, as shown in the sixth row from the top of the table in Figure 5 (Method 1-1-3-1).
[0079] For example, in the first information processing device and the second information processing device, the sound source direction range quality information may include information specifying another bitstream that provides subjective quality of reproduced audio equivalent to the direction-restricted bitstream.
[0080] In other words, another bitstream may be specified whose subjective quality of reproduced audio is equivalent to that of the bitstream corresponding to the in-source direction quality information. For example, in the table of Figure 6, the bitrate of 6DoFAudio2.mp4 is 3000000, and the in-source direction quality information indicates that the subjective quality of its reproduced audio is equivalent to that of 6DoFAudio1.mp4. The bitrate of 6DoFAudio3.mp4 is 3000000, and the in-source direction quality information indicates that the subjective quality of its reproduced audio is equivalent to that of 6DoFAudio1.mp4. The bitrate of 6DoFAudio4.mp4 is 3000000, and the in-source direction quality information indicates that the subjective quality of its reproduced audio is equivalent to that of 6DoFAudio1.mp4.
[0081] With this representation, the subjective quality of the reproduced speech of the bitstream corresponding to the sound source direction in-range quality information can be expressed using the subjective quality of the reproduced speech of another bitstream.
[0082] <Method 1-1-3-2> When the above-mentioned Method 1-1-3 is applied, information expressing quality as a level may be supplied as sound source direction range quality information, as shown in the seventh row from the top of the table in FIG. 5 (Method 1-1-3-2).
[0083] For example, in the first information processing device and the second information processing device, the sound source direction range quality information may include information representing, by level, the subjective quality of the reproduced sound of the direction-limited bitstream.
[0084] In other words, the subjective quality of the reproduced audio of the bitstream corresponding to the sound source direction in-range quality information may be indicated by a level. For example, MPEG-DASH described in Non-Patent Document 3 provides qualityRanking as shown in FIG. 9. This qualityRanking is a parameter that specifies the quality ranking of a representation compared with other representations in the same adaptation set. The lower the value of this parameter, the higher the quality of the content. Furthermore, the absence of this parameter indicates that ranking is not defined. In the sound source direction in-range quality information, such qualityRanking may be used to express the subjective quality of the reproduced audio of the corresponding bitstream.
[0085] Note that, in the sound source direction range quality information for an omnidirectional bitstream (i.e., omnidirectional quality information), such quality ranking may also be used to represent the subjective quality of the reproduced audio of the corresponding bitstream.
[0086] By applying such an expression to the sound source direction in-range quality information, it is possible to express the subjective quality of the reproduced audio of the corresponding bitstream even when there is no other bitstream with an equivalent subjective quality of the reproduced audio.
[0087] <Method 1-1-3-3> When the above-mentioned method 1-1-3 is applied, information indicating a bit rate equivalent to the quality may be supplied as sound source direction range quality information, as shown in the eighth row from the top of the table in Figure 5 (method 1-1-3-3).
[0088] For example, in the first information processing device and the second information processing device, the sound source direction in-range quality information may include information indicating a bit rate equivalent to the subjective quality of the reproduced audio of the direction-limited bitstream.
[0089] That is, in the sound source direction in-range quality information, the subjective quality of the reproduced audio of the corresponding bitstream may be indicated using a bit rate corresponding to that quality.
[0090] Note that in-sound-source-direction-range quality information for an omnidirectional bitstream (i.e., omnidirectional quality information), the subjective quality of the reproduced audio of the corresponding bitstream may also be expressed using the bitrate.
[0091] By applying such an expression to the sound source direction in-range quality information, it is possible to express the subjective quality of the reproduced audio of the corresponding bitstream even when there is no other bitstream with an equivalent subjective quality of the reproduced audio.
[0092] <Method 1-1-3-4> When the above-mentioned Method 1-1-3 is applied, as shown in the ninth row from the top of the table in Figure 5, multiple sets of sound source direction range information and sound source direction range quality information may be supplied as spatial masking information (Method 1-1-3-4).
[0093] For example, in the first information processing device and the second information processing device, the spatial masking information may include multiple sets of sound source direction range information regarding a sound source direction range and sound source direction range quality information regarding the subjective quality of the reproduced audio of the direction-limited bitstream when all sound sources are present within the sound source direction range.
[0094] In this case, the plurality of pieces of sound source direction range information indicate different sound source direction ranges. Also, the plurality of pieces of quality information within a sound source direction range correspond to different sound source direction ranges. That is, a plurality of sound source direction ranges are set for one direction-limited bitstream, and for each sound source direction range, the subjective quality of the reproduced audio when all sound sources are present within that sound source direction range is set. That is, a plurality of qualities (according to respective conditions) are set for the reproduced audio of one direction-limited bitstream.
[0095] By doing so, it is possible to express that the subjective quality of the reproduced sound differs depending on the direction of the sound source at the time of reproduction.
[0096] <Method 1-1-3-5> When the above-mentioned Method 1-1-3 is applied, guaranteed quality information may be further supplied in addition to the sound source direction range quality information, as shown in the tenth row from the top of the table in Fig. 5 (Method 1-1-3-5). The guaranteed quality information indicates the quality of audio data that is guaranteed regardless of the sound source direction, i.e., the minimum quality of reproduced audio regardless of the sound source direction. In other words, the guaranteed quality information indicates the minimum level of the range of quality that can be achieved depending on the sound source direction. In other words, the guaranteed quality information indicates the minimum subjective quality of reproduced audio when the quality experience is not good. In other words, the guaranteed quality information indicates the minimum subjective quality of reproduced audio when the above-mentioned "specific conditions" are not satisfied. In other words, the guaranteed quality information indicates the minimum subjective quality of reproduced audio when the audio source of the direction-restricted bitstream is outside the sound source direction range.
[0097] For example, in the first information processing device and the second information processing device, the spatial masking information may include, in addition to the sound source direction range quality information, guaranteed quality information indicating the subjective quality of the reproduced audio of the direction-limited bitstream that is guaranteed regardless of the sound source direction.
[0098] Note that any method for expressing this guaranteed quality information may be used. For example, it may be the same method as the sound source direction range quality information. For example, any of the above-mentioned methods 1-1-3-1 to 1-1-3-3 may be applied. For example, quality ranking may be used in the guaranteed quality information.
[0099] By providing such guaranteed quality information in addition to the sound source direction in-range quality information, the range of quality of the reproduced speech can be defined by the sound source direction in-range quality information and the guaranteed quality information. In other words, the playback device can grasp the size of the range of subjective quality of the reproduced speech based on the sound source direction in-range quality information and the guaranteed quality information. Therefore, the playback device can use this guaranteed quality information for bitstream playback control, such as selecting a direction-limited bitstream to play based on the size of the range of subjective quality of the reproduced speech.
[0100] <Method 1-1-3-6> When the above-mentioned Method 1-1-3 is applied, out-of-sound source direction range quality information may be supplied (Method 1-1-3-6), as shown in the eleventh row from the top of the table in Fig. 5. The out-of-sound source direction range quality information is information indicating the subjective quality of reproduced audio when one or more sound source directions are outside the sound source direction range. In other words, the out-of-sound source direction range quality information indicates the subjective quality of reproduced audio of a direction-restricted bitstream when the above-mentioned "specific condition" is not met (when a good quality experience cannot be provided).
[0101] For example, in the first information processing device and the second information processing device, the spatial masking information may further include out-of-sound source direction range quality information regarding the subjective quality of the reproduced audio of the direction-limited bitstream when a sound source is present outside the sound source direction range.
[0102] The number of pieces of quality information within the sound source direction range included in the spatial masking information together with the quality information outside the sound source direction range may be any number, and may be one or more.
[0103] Furthermore, any method may be used to express this out-of-sound source direction range quality information. For example, it may be the same method as the in-sound source direction range quality information. For example, any of the above-described methods 1-1-3-1 to 1-1-3-3 may be applied. For example, quality ranking may be used in the guaranteed quality information.
[0104] By providing such out-of-sound-source-direction-range quality information in addition to the in-sound-source-direction-range quality information, the range of quality of the reproduced audio can be defined by the in-sound-source-direction-range quality information and the out-of-sound-source-direction-range quality information. In other words, the playback device can grasp the size of the range of subjective quality of the reproduced audio based on the in-sound-source-direction-range quality information and the out-of-sound-source-direction-range quality information. Therefore, the playback device can use this out-of-sound-source-direction-range quality information for bitstream playback control, such as selecting a direction-limited bitstream to play based on the size of the range of subjective quality of the reproduced audio.
[0105] <Method 1-1-4> When the above-mentioned Method 1-1 is applied, spatial masking information may be supplied not only for the direction-restricted bitstream but also for the omnidirectional bitstream, as shown in the 12th row from the top of the table in Figure 5 (Method 1-1-4).
[0106] For example, in the first information processing device, the spatial masking information generation unit may further generate spatial masking information for an omnidirectional bit stream in which audio data is encoded with bit allocation based on an omnidirectional spatial masking threshold whose sound source direction range is an omnidirectional spatial masking threshold. Also, in the second information processing device, the bit stream acquisition unit may further acquire, based on the spatial masking information, an omnidirectional bit stream in which audio data is encoded with bit allocation based on an omnidirectional spatial masking threshold whose sound source direction range is an omnidirectional spatial masking threshold.
[0107] In this way, the playback device can easily compare the quality of a direction-restricted bitstream with the quality of an omnidirectional bitstream. For example, when the quality within the sound source direction range information of a direction-restricted bitstream is indicated by a level or a bitrate, it is difficult to compare the quality between the direction-restricted bitstream and the omnidirectional bitstream without spatial masking information for the omnidirectional bitstream. Even in this case, by transmitting spatial masking information for the omnidirectional bitstream, it is possible to easily compare the quality of the direction-restricted bitstream and the omnidirectional bitstream.
[0108] In this case, the sound source direction range is omnidirectional. Therefore, there is no need to specify the sound source range, and in this case, the sound source direction range information may be omitted. In other words, the spatial masking information from which the sound source direction range information is omitted may indicate that it is spatial masking information for an omnidirectional bitstream. In this way, it is possible to suppress an increase in the data amount of spatial masking information for an omnidirectional bitstream.
[0109] <Method 1-1-5> When the above-mentioned Method 1-1 is applied, a bit stream for each sound source and spatial masking information for each sound source may be supplied (Method 1-1-5), as shown in the thirteenth row from the top of the table in Fig. 5 . In other words, the bit stream of audio data may be information about a single sound source. In other words, the bit stream may include audio data of a single sound source.
[0110] For example, in the first information processing device, the encoding unit may generate a direction-limited bit stream for each sound source, and the spatial masking information generation unit may generate spatial masking information for each sound source. Also, in the second information processing device, the bit stream acquisition unit may acquire the direction-limited bit stream for each sound source based on the spatial masking information for each sound source.
[0111] In this case, a single piece of audio data (a single sound source) included in a bitstream is processed as a masker, and audio data (sound sources) included in another bitstream are processed as a masker. Therefore, the sound source direction of the masker is arbitrary. Therefore, in the sound source direction information of the spatial masking information of the direction-limited bitstream in this case, the sound source direction range for the masker is defined, and the sound source direction range of the masker is not limited (it is assumed to be omnidirectional). Note that the methods described in <Method 1-1-1> to <Method 1-1-4> may be applied to the spatial masking information in this case. For example, in the sound source direction range information, the sound source direction range of the masker may be defined using any expression, and for example, the technique defined in <Method 1-1-2> may be applied.
[0112] By doing so, it becomes possible to select the optimum masking for each sound source relative to the listener, and it is possible to reduce the bit rate while maintaining the same quality regardless of the listener's position or direction.
[0113] <Method 1-2> When the above-described method 1 is applied, listener control information related to control of the listener's position and direction according to the sound source direction may be provided (Method 1-2), as shown in the 14th row from the top of the table in FIG. 5 . For example, a playback device may have a function for moving or guiding a listener so that the listener can hear the playback audio in the appropriate sound source direction. In this case, the party providing the audio (e.g., a content creator) may want to either permit or prohibit the use of that function. By providing the above-described listener control information, it is possible to address both of these cases. In other words, the party providing the audio (e.g., a content creator) can explicitly control whether or not to permit (or prohibit) the playback device to use such a function.
[0114] For example, in the first information processing device and the second information processing device, the spatial masking information may further include listener control information related to control of the position or direction of the listener according to the direction of the sound source.
[0115] The content of this listener control information may be any. For example, the listener control information may include listener control permission information regarding permission to control the position and direction of a listener according to the direction of a sound source. For example, a listener control permission flag indicating whether or not to permit control of the position and direction of a listener according to the direction of a sound source may be applied as this listener control permission information. Furthermore, the listener control information may include listener control prohibition information regarding prohibition of control of the position and direction of a listener according to the direction of a sound source. For example, a listener control prohibition flag indicating whether or not to prohibit control of the position and direction of a listener according to the direction of a sound source may be applied as this listener control prohibition information.
[0116] <Method 2> For example, in MPEG-DASH, a control file (MPD (Media Presentation Description)) that controls the distribution of content files is used in the distribution of the content files. For example, a distribution server is provided with a plurality of content files, and the playback device acquires control files (MPD) related to the content files, selects a desired content file based on information stored in the control file (MPD), requests the distribution server to distribute the selected content file, acquires the content file distributed from the distribution server based on the request, and plays the acquired content file. The above-described method 1 may be applied to a content file distributed in such a distribution system. That is, a direction-restricted bitstream may be stored in the content file. In other words, (a content file storing) the direction-restricted bitstream may be distributed using such a control file (MPD). In this case, the above-described method 1-1 may also be applied. That is, spatial masking information may be provided.
[0117] In this case, as shown in the 15th row from the top of the table in Fig. 5 , spatial masking information for the direction-restricted bitstream may be stored in the control file (MPD) (Method 2). That is, the above-mentioned information included in the spatial masking information may be stored in the control file (MPD).
[0118] For example, the first information processing device may further include a control file generating unit that generates a control file for controlling distribution of the direction-restricted bitstream and stores spatial masking information in the control file. Also, in the second information processing device, a bitstream acquiring unit may acquire the direction-restricted bitstream based on the spatial masking information stored in the control file for controlling distribution of the direction-restricted bitstream.
[0119] By providing spatial masking information using a control file (MPD), a playback device can not only easily acquire the spatial masking information using a conventional method such as acquiring a control file, but also use the spatial masking information to select content files to be distributed. For example, the playback device can select and distribute a direction-limited bitstream that provides a good quality experience, i.e., a content file that stores a direction-limited bitstream that satisfies specific conditions, based on the listener's position, direction, etc. Therefore, it is possible to suppress a decrease in coding efficiency while suppressing a decrease in the subjective quality of the reproduced audio.
[0120] In this case, the content of the spatial masking information may be any information related to spatial masking. For example, the spatial masking information may include the above-mentioned information such as direction limitation identification information, sound source direction range information, and sound source direction range quality information.
[0121] It should be noted that the spatial masking information (each piece of information included in the spatial masking information) may be stored anywhere in the control file. Below, an example of a storage location for the spatial masking information (each piece of information included in the spatial masking information) will be described, but the location where the spatial masking information can be stored is not limited to this example. Also, below, an MPD used in MPEG-DASH will be described as an example of a control file. However, any content distribution standard may be used, and the control file is not limited to MPEG-DASH. Therefore, the control file for storing the spatial masking information may be any type as long as it controls the distribution of a content file (direction-limited bitstream), and is not limited to an MPD.
[0122] The spatial masking information may be stored in a representation that stores information on a direction-restricted bitstream. For example, when Method 1-1-1 is applied, that is, when the spatial masking information includes direction-restriction identification information, the direction-restriction identification information may be stored in the representation as an essential property. For example, an essential property may be provided in the representation, and "DirectionDependentStream" may be set as the scheme type in the essential property (<EssentialProperty schemeType="DirectionDependentStraem"> ). In other words, in this case, the essential property "schemeType="DirectionDependentStream" (that it is described) serves as direction restriction identification information. Therefore, by parsing this essential property, a playback device can determine that the bitstream corresponding to this representation is a direction-restricted bitstream.
[0123] Furthermore, when Method 1-1-2 or Method 1-1-3 is applied, that is, when the spatial masking information includes sound source direction range information or sound source direction range quality information, the information may be stored in a container element provided in the above-mentioned essential property. For example, as shown in the table of Fig. 10, a "DDS" element may be provided as a container element including an attribute or element indicating that the bitstream requires sound source direction-dependent playback, and the sound source direction range information or sound source direction range quality information may be stored therein.
[0124] For example, as shown in the table of Fig. 10, a coverage info element (DDS.coverageInfo) may be provided as sound source direction range information within the DDS element, and the sound source direction range may be defined using the coverage info element. For example, parameters such as centre_azimuth, centre_elevation, azimuth_range, and elevation_range may be set as parameters defining the sound source direction range.
[0125] As shown in FIG. 10 , centre_azimuth (DDS.coverageInfo@centre_azimuth) is a parameter that indicates the direction on the horizontal plane of the center of the sound source direction range as seen from the listener's position (the azimuth angle of the center point of the spherical region representing the sound source direction range). For example, the azimuth angle may be specified in units of (2^-16) degrees. In this case, the minimum value of centre_azimuth is -11796480 (-180 degrees) and the maximum value is 11796479 (179.99 degrees). Note that if this parameter does not exist, it may be considered that "the value of DDS.coverageInfo@centre_azimuth is 0."
[0126] The center_elevation (DDS.coverageInfo@centre_elevation) is a parameter that indicates the direction on the vertical plane of the center of the sound source direction range as seen from the listener's position (the elevation or depression angle of the center point of the spherical region representing the sound source direction range). For example, the elevation or depression angle may be specified in units of (2^-16) degrees. In this case, the minimum value of center_elevation is -5898240 (-90 degrees), and the maximum value is 5898240 (90 degrees). Note that if this parameter does not exist, it may be considered that "the value of DDS.coverageInfo@centre_elevation is 0."
[0127] azimuth_range (DDS.coverageInfo@azimuth_range) is a parameter that indicates the horizontal angular width of the sound source direction range (the azimuth angle range of the spherical region that passes through the center point of the spherical region that represents the sound source direction range). For example, the angular width may be specified in units of (2^-16) degrees. In that case, the minimum value of azimuth_range is 0 (0 degrees) and the maximum value is 23592960 (360 degrees). Note that if this parameter does not exist, it may be assumed that "the value of DDS.coverageInfo@azimuth_range is equal to 360 * 2^16."
[0128] The elevation_range (DDS.coverageInfo@elevation_range) is a parameter that indicates the vertical angle width of the sound source direction range (the width of the elevation or depression angle of the spherical region that passes through the center point of the spherical region that represents the sound source direction range). For example, the angle width may be specified in units of (2^-16) degrees. In that case, the minimum value of elevation_range is 0 (0 degrees) and the maximum value is 11796480 (180 degrees). Note that if this parameter does not exist, it may be considered that "the value of DDS.coverageInfo@elevation_range is equal to 180 * 2^16."
[0129] In this specification, "^" indicates a power. For example, in "A^B", A indicates the base and B indicates the exponent.
[0130] 10, a quality element (DDS@quality) may be provided within the DDS element as quality information within the sound source direction range. This quality element indicates the subjective quality of the reproduced sound as perceived by a listener. For example, when Method 1-1-3-1 is applied, the value of this quality element may indicate the identification information (Representation@id) of a representation that stores information on a bitstream with the same quality as the corresponding direction-limited bitstream.
[0131] A specific example of MPD description is shown in Fig. 11. In this example, direction restriction identification information (schemeType="DirectionDependentStream"), sound source direction range information (DDS.coverageInfo), sound source direction range quality information (DDS@quality), etc. are stored in the essential properties provided in the representation that stores information on each bitstream. Note that in this example, a representation that stores information on a bitstream with quality equivalent to that of the corresponding direction restricted bitstream is indicated as the sound source direction range quality information, and the "R1" in "DDS quality="R1"" indicates the identification information of the representation that stores information on 6DoFAudio1.mp4 (Representation id="R1"). In other words, this indicates that the subjective quality of the reproduced sound as perceived by the listener is equivalent to that of 6DoFAudio1.mp4.
[0132] Note that Method 1-1-3-2 may be applied, and the value of the quality element, which is quality information within the sound source direction range, may indicate a quality ranking (qualityRanking) defined in Representation@qualityRanking. Also, Method 1-1-3-3 may be applied, and the value of the quality element, which is quality information within the sound source direction range, may indicate a bit rate corresponding to the subjective quality of the reproduced sound to the listener.
[0133] Furthermore, method 1-1-3-4 may be applied, and multiple sets of sound source direction range information and sound source direction range quality information may be described in the above-mentioned essential property. An example of an MPD description in this case is shown in Fig. 12. As in the example shown in Fig. 12, by describing multiple DDS elements in the essential property, it is possible to describe different sound source direction range information and sound source direction range quality information in each of them. In other words, multiple sets of sound source direction range information and sound source direction range quality information can be described for one representation (direction-limited bitstream).
[0134] Furthermore, when Method 1-1-3-5 is applied and guaranteed quality information (guaranteed quality (i.e., the quality that results in the worst experience)) is stored in the control file (MPD), the guaranteed quality may be indicated using Representation@qualityRanking. That is, in this case, DDS@quality indicates the best quality, and Representation@qualityRanking indicates the worst quality. For example, the playback device can realize distribution control such that it selects content files for which the difference between the quality indicated by DDS@quality and the quality ranking value is small, and does not acquire content files for which the difference is large. If this difference is small, the degradation of the reproduced audio quality is small. Conversely, if this difference is large, the degradation of the reproduced audio quality is large. Therefore, by utilizing such characteristics and performing the distribution control described above, the playback device can distribute content files with higher reproduced audio quality. Note that the allowable difference may be determined by the playback device.
[0135] Furthermore, when Method 1-1-3-6 is applied and out-of-sound-direction-range quality information is stored in the control file (MPD), the out-of-sound-direction-range quality information may be stored in the above-mentioned DDS element. For example, DDS@othersQuality may be described as the out-of-sound-direction-range quality information. In other words, this DDS@othersQuality is a parameter indicating the subjective quality of the reproduced sound (the worst-experienced quality) perceived by the listener when a sound source is located outside the sound source direction range defined in the coverage info element (DDS.coverageInfo). The value of this parameter may be expressed in any manner, as in the case of the in-sound-direction-range quality information. For example, representation identification information, quality ranking, or bit rate may be used.
[0136] Furthermore, when Method 1-1-4 is applied and spatial masking information for an omnidirectional bitstream is stored in a control file (MPD), the spatial masking information may be stored in an essential property provided in the representation storing the omnidirectional bitstream, as in the case of the direction-restricted bitstream described above. The method for storing each piece of information included in this spatial masking information is the same as in the case of the direction-restricted bitstream described above. However, the sound source direction range information (coverageInfo) may be omitted. In other words, if DDS.coverageInfo is omitted, it indicates that the bitstream corresponding to that representation is an omnidirectional bitstream.
[0137] Also, when method 1-1-5 is applied, a bitstream is generated for each sound source, and spatial masking information for each sound source is stored in a control file (MPD), the method for storing each piece of information included in the spatial masking information is the same as the example described above.
[0138] Furthermore, when Method 1-2 is applied and listener control information is stored in a control file (MPD), the listener control information may be stored in the above-mentioned DDS element. For example, DDS@permitAutoViewMove may be described as the listener control information. For example, if this DDS@permitAutoViewMove is true (e.g., "1"), it may indicate that an application is permitted to change the listener's position or direction. If this DDS@permitAutoViewMove is false (e.g., "0"), it may indicate that an application is prohibited from changing the listener's position or direction.
[0139] <Method 3> For example, a content file such as an MP4 file can store multiple bitstreams by using multiple tracks, etc. In this case, the playback device selects a desired bitstream based on control information stored in the content file, obtains the selected bitstream from the content file, and plays back the obtained bitstream. The above-mentioned Method 1 may be applied to such a content file. That is, a direction-restricted bitstream may be stored in the content file. In other words, a direction-restricted bitstream stored in the content file may be obtained. In this case, the above-mentioned Method 1-1 may also be applied. That is, spatial masking information may be provided, and a direction-restricted bitstream may be obtained from the content file based on the spatial masking information.
[0140] In this case, as shown in the 16th row from the top of the table in Fig. 5, spatial masking information for the direction-restricted bitstream may be stored in a content file (e.g., an MP4 file) (Method 3). That is, the above-mentioned information included in the spatial masking information may be stored in the content file.
[0141] For example, the first information processing device may further include a content file generating unit that generates a content file and stores the direction-restricted bitstream and spatial masking information in the content file. Also, in the second information processing device, a bitstream acquiring unit may acquire the direction-restricted bitstream from the content file based on the spatial masking information stored in the content file.
[0142] By providing spatial masking information using a content file, a playback device can easily acquire the spatial masking information by parsing the content file, similar to conventional methods. For example, the playback device can select and acquire from the content file a direction-limited bitstream that provides a good quality experience, i.e., a direction-limited bitstream that satisfies specific conditions, based on the listener's position, direction, etc. Therefore, it is possible to suppress a decrease in coding efficiency while suppressing a decrease in the subjective quality of the reproduced audio.
[0143] In this case, the content of the spatial masking information may be any information related to spatial masking. For example, the spatial masking information may include the above-mentioned information such as direction limitation identification information, sound source direction range information, and sound source direction range quality information.
[0144] Note that the spatial masking information (each piece of information included in the spatial masking information) may be stored anywhere in the content file. Below, an example of a storage location for the spatial masking information (each piece of information included in the spatial masking information) will be described, but the location where the spatial masking information can be stored is not limited to this example. Also, below, an MP4 file will be used as an example of the content file. However, the content file may be of any standard and is not limited to an MP4 file.
[0145] The spatial masking information may be stored, for example, in a sample entry of a track storing a direction-restricted bitstream. For example, when Method 1-1-1 is applied, i.e., when direction restriction identification information is included in the spatial masking information, a box indicating "this is a bitstream that can provide a good quality experience under specific conditions" may be added as the direction restriction identification information to the sample entry of the track storing the direction-restricted bitstream. For example, a direction-dependent stream box (DirectionDependentStreamBox('ddst')) may be added as the box. By parsing this direction-dependent stream box, the playback device can determine that the bitstream stored in this track is a direction-restricted bitstream.
[0146] Furthermore, when Method 1-1-2 or Method 1-1-3 is applied, that is, when the spatial masking information includes sound source direction range information or sound source direction range quality information, the information may be stored in the above-mentioned direction dependent stream box. For example, as in the example of syntax shown on the left side of Fig. 14 , the sound source direction range information or sound source direction range quality information may be stored in the direction dependent stream box.
[0147] In the example of Fig. 14, parameters such as centre_azimuth, centre_elevation, azimuth_range, and elevation_range that define the sound source direction range are defined as sound source direction range information. Their semantics are as shown on the right side of Fig. 14. Centre_azimuth and centre_elevation are parameters that specify the values of the direction on the horizontal plane (azimuth angle) and the direction on the vertical plane (elevation angle or depression angle) of the center of the spherical region that indicates the sound source direction range, respectively, in units of 2^-16 degrees. The value of centre_azimuth must be within the range of -180*2^16 to 180*2^16-1. The value of centre_elevation must be within the range of -90*2^16 to 90*2^16. Furthermore, azimuth_range and elevation_range are parameters that specify the angular width of the azimuth angle and the angular width of the elevation angle, respectively, of the spherical region that indicates the sound source direction range, in units of 2^-16 degrees. The azimuth_range and elevation_range specify the angular width of the spherical region that indicates the range of sound source directions, passing through its center point. The value of azimuth_range must be in the range from 0 to 360*2^16. The value of elevation_range must be in the range from 0 to 180*2^16.
[0148] Furthermore, in the example syntax of FIG. 14 , a parameter "quality" is defined as sound source direction range quality information. This "quality" is a parameter indicating the subjective quality of the reproduced audio from the direction-restricted bitstream stored in this track, as perceived by a listener. Any method may be used to express the subjective quality in this "quality." For example, method 1-1-3-1 may be applied, and identification information of a track storing another bitstream whose subjective quality of reproduced audio is equivalent to that of the direction-restricted bitstream stored in this track may be applied as the value of this "quality." Alternatively, method 1-1-3-2 may be applied, and the quality ranking (qualityRanking) defined in "Representation@qualityRanking" may be applied as the value of this "quality." Alternatively, method 1-1-3-3 may be applied, and a bit rate corresponding to the subjective quality of the reproduced audio from the listener may be applied as the value of this "quality."
[0149] 15 and 16 are diagrams showing example configurations of MP4 files that store the above-mentioned spatial masking information (direction limitation identification information, sound source direction range information, quality information within the sound source direction range, etc.) As shown in Fig. 15 and 16, the spatial masking information is stored in the sample entry of each track.
[0150] Alternatively, method 1-1-3-4 may be applied, and multiple sets of sound source direction range information and sound source direction range quality information may be stored as spatial masking information. An example of syntax and semantics in this case is shown in Fig. 17. As shown in the example of syntax in Fig. 17, multiple sets of sound source direction range information and sound source direction range quality information can be stored by using, for example, a for statement.
[0151] Alternatively, method 1-1-3-5 may be applied, and guaranteed quality information (guaranteed quality (i.e., quality that results in the worst experience)) may be stored as spatial masking information. Alternatively, method 1-1-3-6 may be applied, and out-of-sound source direction range quality information may be stored as spatial masking information. An example of syntax and semantics in this case is shown in FIG. 18. As shown in the example of syntax in FIG. 18, otherQualityFlag and otherQuality may be stored as the out-of-sound source direction range quality information. As shown in the example of semantics in FIG. 18, otherQualityFlag is flag information indicating whether otherQuality is present. otherQuality is a parameter indicating the subjective quality of reproduced audio of a direction-restricted bitstream when the sound source exists outside the sound source direction range indicated by the sound source direction range information. In other words, by parsing this otherQuality, the playback device can grasp the subjective quality of reproduced audio of the direction-restricted bitstream stored in that track when the sound source exists outside the sound source direction range.
[0152] Alternatively, method 1-1-4 may be applied, and spatial masking information for the omnidirectional bitstream may be stored in a content file (e.g., an MP4 file). That is, spatial masking information for the omnidirectional bitstream may be stored in a track in which the omnidirectional bitstream is stored. The method for storing each piece of information included in this spatial masking information is the same as in the case of the direction-limited bitstream described above. However, sound source direction range information (centre_azimuth, centre_elevation, azimuth_range, elevation_range, etc.) may be omitted. In other words, if the definitions of these parameters are omitted, it indicates that the bitstream stored in that track is an omnidirectional bitstream.
[0153] Alternatively, Method 1-1-5 may be applied, where a bitstream is generated for each sound source and the spatial masking information for each sound source is stored in a content file (e.g., an MP4 file). In this case, the bitstream for each sound source and the spatial masking information corresponding to the bitstream may be stored in different tracks.
[0154] Alternatively, Method 1-2 may be applied, and listener control information may be stored in a content file (e.g., an MP4 file). An example of the syntax and semantics of the Direction Dependent Stream Box in this case is shown in FIG. 19. As shown in the example syntax of FIG. 19, a permitAutoViewMoveFlag may be defined as listener control information in the Direction Dependent Stream Box (DirectionDependentStreamBox('ddst')) of the sample entry of the track storing the direction-restricted bitstream. This permitAutoViewMoveFlag is flag information indicating whether or not applications, etc., are permitted to change the listener's position or direction. When this permitAutoViewMoveFlag is true (e.g., "1"), this indicates that applications, etc., are permitted to change the listener's position or direction. When this permitAutoViewMoveFlag is false (e.g., "0"), this indicates that applications, etc., are prohibited from changing the listener's position or direction.
[0155] <Method 4> Furthermore, spatial masking information for a direction-restricted bitstream may be stored in both a control file (e.g., MPD) and a content file (e.g., MP4 file), as shown in the bottom row of the table in Fig. 5. For example, some of the spatial masking information may be stored in the MPD, and the remaining information may be stored in the MP4 file.
[0156] For example, a first information processing device may include a content file generation unit that generates a content file and stores a direction-restricted bitstream and first spatial masking information related to spatial masking for the direction-restricted bitstream in the content file. The first information processing device may also include a control file generation unit that generates a control file that controls distribution of the direction-restricted bitstream and stores second spatial masking information related to spatial masking for the direction-restricted bitstream in the control file. Furthermore, in a second information processing device, a bitstream acquisition unit may acquire the direction-restricted bitstream from the content file based on the first spatial masking information related to spatial masking for the direction-restricted bitstream stored in the content file and the second spatial masking information related to spatial masking for the direction-restricted bitstream stored in the control file that controls distribution of the direction-restricted bitstream.
[0157] For example, direction limitation identification information may be stored in a control file (MPD) as second spatial masking information. Furthermore, sound source direction range information and sound source direction range quality information may be stored in a content file (e.g., an MP4 file) as first spatial masking information. The storage location of the first spatial masking information in the content file (e.g., an MP4 file) may be the same as when Method 3 is applied. The storage location of the second spatial masking information in the control file (e.g., an MPD) may be the same as when Method 2 is applied. This makes it possible to suppress an increase in the size of the control file (e.g., an MPD) due to the spatial masking information, compared to Method 2.
[0158] <Application of Each Method> Each of the above-described methods can be applied in combination with other methods as long as no contradiction occurs. For example, two or more of the methods shown in the table of FIG. 5 may be applied in appropriate combination. Of course, each of the methods shown in the table of FIG. 5 may also be applied in combination with other methods not shown.
[0159] In this specification, a description of a higher-level method may include a description of a lower-level method. For example, a description such as "applying Method 1" may also include the application of each method that directly or indirectly belongs to Method 1 (e.g., one or more of Method 1-1, Method 1-2, Method 1-1-1, Method 1-1-2, Method 1-1-3, Method 1-1-4, Method 1-1-3-1, Method 1-1-3-2, Method 1-1-3-3, Method 1-1-3-4, Method 1-1-3-5, and Method 1-1-3-6).
[0160] 4. First Embodiment File Generation Device The present technology may be applied to any device. Fig. 20 is a block diagram showing an example of the configuration of a file generation device, which is one aspect of an information processing device to which the present technology is applied. The file generation device 300 (first information processing device) shown in Fig. 20 is a device that encodes audio data to generate a bitstream and stores the bitstream in a content file (e.g., an MP4 file).
[0161] Note that Fig. 20 shows the main processing units, data flows, etc., and does not necessarily include everything shown in Fig. 20. In other words, in file generation device 300, there may be processing units that are not shown as blocks in Fig. 20, and there may be processing and data flows that are not shown as arrows, etc. in Fig. 20.
[0162] As shown in FIG. 20 , the file generation device 300 (first information processing device) has an audio data preprocessing unit 311, an encoding unit 312, a spatial masking information generation unit 313, a file generation unit 314, an MPD generation unit 315, a storage unit 316, and a supply unit 317.
[0163] The audio data preprocessing unit 311 performs preprocessing (processing before encoding) of audio data. For example, the audio data preprocessing unit 311 may acquire audio data and scene information supplied to the file generation device 300. This scene information is information indicating the situation (scene) of a 3D space in which a sound source of the audio data is located. As preprocessing, the audio data preprocessing unit 311 may set a minimum directional limited spatial masking threshold based on the scene information, or perform bit allocation of the audio data using the set minimum directional limited spatial masking threshold. The audio data preprocessing unit 311 may supply the audio data and information related to the bit allocation of the audio data to the encoding unit 312. The audio data preprocessing unit 311 may supply the scene information and information related to the audio data to the spatial masking information generation unit 313.
[0164] The encoding unit 312 executes processing related to encoding of audio data. For example, the encoding unit 312 may acquire information about the audio data supplied from the audio data preprocessing unit 311 and the bit allocation of the audio data. The encoding unit 312 may encode the audio data based on the information about the bit allocation and generate a bitstream. The encoding unit 312 may supply the generated bitstream to the file generation unit 314.
[0165] The spatial masking information generation unit 313 executes processing related to the generation of spatial masking information. For example, the spatial masking information generation unit 313 may acquire scene information and information related to audio data supplied from the audio data preprocessing unit 311. The spatial masking information generation unit 313 may use the acquired information to generate spatial masking information for the audio data to be encoded. The spatial masking information generation unit 313 may supply the generated spatial masking information to the file generation unit 314 and the MPD generation unit 315.
[0166] The file generation unit 314 executes processing related to the generation of a content file. Therefore, the file generation unit 314 can also be called a content file generation unit. For example, the file generation unit 314 may acquire a bitstream of audio data supplied from the encoding unit 312. The file generation unit 314 may acquire spatial masking information supplied from the spatial masking information generation unit 313. The file generation unit 314 may generate a content file (e.g., an MP4 file) and store the bitstream of audio data in the content file. The file generation unit 314 may further store spatial masking information in the content file. The file generation unit 314 may supply the generated content file, etc. to the MPD generation unit 315 or the storage unit 316.
[0167] The MPD generation unit 315 executes processing related to the generation of an MPD, which is a control file that controls the distribution of content files. Therefore, the MPD generation unit 315 can also be said to be a control file generation unit. For example, the MPD generation unit 315 may acquire spatial masking information supplied from the spatial masking information generation unit 313. The MPD generation unit 315 may acquire a content file or the like supplied from the file generation unit 314. The MPD generation unit 315 may generate an MPD that controls the distribution of the content file. The MPD generation unit 315 may store spatial masking information in the MPD. The MPD generation unit 315 may supply the generated MPD to the storage unit 316.
[0168] The storage unit 316 has, for example, a storage medium and executes processing related to storing information. For example, the storage unit 316 may acquire a content file supplied from the file generation unit 314 and store it in the storage medium. The storage unit 316 may acquire an MPD supplied from the MPD generation unit 315 and store it in the storage medium. The storage unit 316 may read out the content file stored in the storage medium at a predetermined timing or based on an external request or the like and supply it to the supply unit 317. The storage unit 316 may read out the MPD stored in the storage medium at a predetermined timing or based on an external request or the like and supply it to the supply unit 317.
[0169] The supply unit 317 has, for example, a communication function and executes processing related to the supply of information to the outside. For example, the supply unit 317 may acquire an MPD or content file supplied from the storage unit 316. The supply unit 317 may communicate with another device or the like using the communication function and supply the MPD or content file to the other device via that communication at a predetermined timing or based on a request from the outside. Therefore, the supply unit 317 can also be said to be a communication unit that communicates with other devices. The supply unit 317 can also be said to be a providing unit that provides the MPD or content file.
[0170] The present technology may be applied to a file generation device 300 configured as described above. For example, the above-described method 1 may be applied to this file generation device 300 as a first information processing device. In this case, the file generation device 300 includes an encoding unit 312 that encodes audio data using bit allocation based on a direction-limited spatial masking threshold, which is a spatial masking threshold that limits the sound source direction range, and generates a direction-limited bit stream.
[0171] Furthermore, the above-described method 2 may be applied to the file generation device 300 as a first information processing device. In this case, the file generation device 300 further includes an MPD generation unit 315 that generates a control file (MPD) for controlling the distribution of the direction-restricted bitstream and stores spatial masking information in the control file.
[0172] Furthermore, the file generation device 300 may be used as a first information processing device to which the above-described method 3 is applied. In this case, the file generation device 300 further includes a file generation unit 314 that generates a content file and stores the direction-limited bitstream and spatial masking information in the content file.
[0173] Furthermore, the above-described method 4 may be applied to the file generation device 300 as a first information processing device. In this case, the file generation device 300 further includes a file generation unit 314 that generates a content file and stores a direction-restricted bitstream and first spatial masking information related to spatial masking for the direction-restricted bitstream in the content file, and an MPD generation unit 315 that generates a control file (MPD) that controls distribution of the direction-restricted bitstream and stores second spatial masking information related to spatial masking for the direction-restricted bitstream in the control file.
[0174] Note that processing units of file generation device 300 that are not essential for each method may be omitted.
[0175] <File Generation Process Flow 1> An example of the flow of the file generation process executed by file generation device 300 having the above configuration will be described with reference to the flowcharts of Figures 21 to 23. First, an example of the flow of the file generation process when the above-mentioned method 2 is applied will be described with reference to the flowchart of Figure 21. That is, in this case, spatial masking information is stored in the MPD.
[0176] When the file generation process starts, in step S301, the audio data preprocessing unit 311 sets a minimum direction limited spatial masking threshold based on scene information, etc., as preprocessing, and allocates bits to the audio data. This process is basically the same as the conventional method ( FIG. 4 ), except that the sound source direction range is limited.
[0177] In step S302, the encoding unit 312 encodes the audio data in accordance with the bit allocation set in step S301, and generates a direction-limited bit stream.
[0178] In step S303, the spatial masking information generator 313 generates spatial masking information for the direction-limited bitstream.
[0179] In step S304, the file generation unit 314 generates a content file and stores the direction-restricted bitstream in the content file.
[0180] In step S305, the MPD generation unit 315 generates an MPD and stores spatial masking information in the MPD.
[0181] In step S306, the storage unit 316 may store the content file and the MPD.
[0182] In step S307, the supply unit 317 may read and supply the MPD from the storage unit 316 at a predetermined timing or based on an external request, etc. The supply unit 317 may also read and supply a content file from the storage unit 316 at a predetermined timing or based on an external request, etc.
[0183] When the process of step S307 ends, the file generation process ends.
[0184] By performing each process as described above, the file generation device 300 can provide spatial masking information using the MPD. Therefore, the file generation device 300 can suppress a decrease in the encoding efficiency of the audio data bitstream while suppressing a decrease in the subjective quality of the reproduced audio perceived by the listener.
[0185] <File Generation Process Flow 2> Next, an example of the file generation process flow when the above-described method 3 is applied will be described with reference to the flowchart in Fig. 22. That is, in this case, spatial masking information is stored in the content file.
[0186] When the file generation process is started, the processes from step S321 to step S323 are executed in the same manner as the processes from step S301 to step S303 in FIG.
[0187] In step S324, the file generation unit 314 generates a content file and stores the direction-limited bitstream and spatial masking information in the content file.
[0188] In step S325, the MPD generation unit 315 generates an MPD related to the distribution of the direction-restricted bitstream.
[0189] The processes of steps S326 and S327 are executed in the same manner as the processes of steps S306 and S307 in Fig. 21. When the process of step S327 ends, the file generation process ends.
[0190] By performing each process as described above, the file generation device 300 can provide spatial masking information using a content file. Therefore, the file generation device 300 can suppress a decrease in the encoding efficiency of the audio data bitstream while suppressing a decrease in the subjective quality of the reproduced audio as perceived by a listener.
[0191] <File Generation Process Flow 3> Next, an example of the flow of a file generation process when the above-mentioned method 4 is applied will be described with reference to the flowchart in Fig. 23. That is, in this case, spatial masking information is stored in both the control file (e.g., MPD) and the content file (e.g., MP4 file).
[0192] When the file generation process is started, the processes of steps S341 and S342 are executed in the same manner as the processes of steps S301 and S302 in FIG.
[0193] In step S343, the spatial masking information generator 313 generates first spatial masking information and second spatial masking information as spatial masking information related to the direction-limited bitstream.
[0194] In step S344, the file generation unit 314 generates a content file and stores the direction-limited bitstream and the first spatial masking information in the content file.
[0195] In step S345, the MPD generation unit 315 generates an MPD related to the distribution of the direction-restricted bitstream, and stores the second spatial masking information in the MPD.
[0196] The processes of steps S346 and S347 are executed in the same manner as the processes of steps S306 and S307 in Fig. 21. When the process of step S347 ends, the file generation process ends.
[0197] By performing each process as described above, the file generation device 300 can provide spatial masking information using the MPD and the content file. Therefore, the file generation device 300 can suppress a decrease in the encoding efficiency of the audio data bitstream while suppressing a decrease in the subjective quality of the reproduced audio perceived by the listener.
[0198] 24 is a block diagram showing an example of the configuration of a playback device, which is one aspect of an information processing device to which the present technology is applied. The playback device 400 (second information processing device) shown in Fig. 24 is a device that decodes and plays back a bitstream of audio data stored in a content file (e.g., an MP4 file).
[0199] Note that Fig. 24 shows the main processing units, data flows, etc., and does not necessarily show everything. In other words, in playback device 400, there may be processing units that are not shown as blocks in Fig. 24, and there may be processing and data flows that are not shown as arrows, etc. in Fig. 24.
[0200] As shown in FIG. 24 , the playback device 400 (second information processing device) has an MPD acquisition unit 411, a sound source direction setting unit 412, a file acquisition unit 413, a file processing unit 414, a decoding unit 415, an audio output processing unit 416, and an audio output unit 417.
[0201] The MPD acquisition unit 411 has, for example, a communication function, and executes processing related to acquisition of an MPD, which is a control file that controls distribution of content files. Therefore, the MPD acquisition unit 411 can also be called a control file acquisition unit. For example, the MPD acquisition unit 411 may communicate with an external device (another device, etc.) using the communication function, and acquire an MPD supplied from the external device (another device, etc.) via the communication. This MPD may be generated by, for example, the file generation device 300. Furthermore, this MPD may include spatial masking information. The MPD acquisition unit 411 may supply the acquired MPD to the file acquisition unit 413 or the file processing unit 414.
[0202] The sound source direction setting unit 412 executes processing related to setting of the sound source direction. For example, the sound source direction setting unit 412 may set the sound source direction relative to the listener based on scene information, changes in the audio scene, changes in the position or direction of the listener, etc. The sound source direction setting unit 412 may supply the setting of the sound source direction to the file acquisition unit 413 or the file processing unit 414.
[0203] The file acquisition unit 413, for example, has a communication function and performs processing related to acquisition of a content file storing a bitstream of audio data. For example, the file acquisition unit 413 may communicate with an external device (e.g., another device) using the communication function and acquire a content file provided from the external device (e.g., another device) via the communication. Therefore, the file acquisition unit 413 can also be referred to as a content file acquisition unit. This content file may be generated by, for example, the file generation device 300. That is, this content file may include a bitstream of audio data. For example, this content file may include a direction-limited bitstream or an omnidirectional bitstream. Furthermore, the number of bitstreams stored in the content file may be any number, and may be single or multiple. Furthermore, this content file may include spatial masking information.
[0204] At this time, for example, the file acquisition unit 413 may select and acquire a content file storing a desired bitstream from content files for distribution prepared in the distribution server based on the MPD, sound source direction settings, etc. Therefore, the file acquisition unit 413 can also be called a bitstream acquisition unit. For example, the file acquisition unit 413 may acquire an MPD supplied from the MPD acquisition unit 411. The file acquisition unit 413 may acquire a sound source direction setting supplied from the sound source direction setting unit 412. The file acquisition unit 413 may select a content file storing a desired bitstream from content files for distribution prepared in the distribution server based on the acquired MPD, sound source direction settings, etc., request the distribution server to distribute the desired content file, and acquire (the bitstream stored in) the desired content file distributed from the distribution server in response to the request. The file acquisition unit 413 may supply the acquired content file to the file processing unit 414.
[0205] The file processing unit 414 executes processing related to the content file. Therefore, the file processing unit 414 can also be called a content file processing unit. For example, the file processing unit 414 may acquire a content file supplied from the file acquisition unit 413. The file processing unit 414 may acquire a bit stream of audio data from the acquired content file.
[0206] At that time, the file processing unit 414 may acquire a desired bitstream stored in the acquired content file based on the MPD, sound source direction settings, etc. Therefore, the file processing unit 414 can also be said to be a bitstream acquisition unit. For example, the file processing unit 414 may acquire an MPD supplied from the MPD acquisition unit 411. The file processing unit 414 may acquire a sound source direction setting supplied from the sound source direction setting unit 412. The file processing unit 414 may select and acquire a desired bitstream from among the bitstreams stored in the acquired content file based on the acquired MPD, sound source direction settings, etc. The file processing unit 414 may supply the acquired bitstream to the decoding unit 415.
[0207] The decoding unit 415 executes processing related to decoding of a bitstream. For example, the decoding unit 415 may acquire a bitstream supplied from the file processing unit 414. The decoding unit 415 may decode the bitstream and generate (restore) audio data. The decoding unit 415 may supply the generated audio data to the audio output processing unit 416.
[0208] The audio output processing unit 416 executes processing related to audio data in order to output audio. For example, the audio output processing unit 416 may acquire audio data supplied from the decoding unit 415. The audio output processing unit 416 may arrange the audio data in a 3D space based on scene information or the like, and perform rendering to generate audio data for output. This audio data for output is data of audio that can be heard by a listener. The audio output processing unit 416 may supply the generated audio data for output to the audio output unit 417, which may output the audio.
[0209] The audio output unit 417 has an audio output device such as a speaker and executes processing related to audio output. For example, the audio output unit 417 may acquire audio data for output supplied from the audio output processing unit 416. The audio output unit 417 converts the acquired audio data for output into output audio by the output device and outputs the output audio.
[0210] The present technology may be applied to a playback device 400 configured as described above. For example, the above-described method 1 may be applied to this playback device 400 as a second information processing device. In this case, the playback device 400 includes a file acquisition unit 413 or a file processing unit 414 that acquires a direction-limited bitstream in which audio data is encoded using bit allocation based on a direction-limited spatial masking threshold, which is a spatial masking threshold that limits the sound source direction range, and a decoding unit 415 that decodes the acquired direction-limited bitstream.
[0211] Furthermore, the playback device 400 may be used as a second information processing device to which the above-described method 2 is applied. In this case, the file acquisition unit 413 of the playback device 400 selects and acquires a content file that stores a desired direction-limited bitstream, based on spatial masking information stored in an MPD, which is a control file that controls distribution of content files.
[0212] Furthermore, the playback device 400 may be used as a second information processing device to which the above-described method 3 is applied. In this case, the file processing unit 414 of the playback device 400 acquires a desired direction-limited bitstream from a content file based on spatial masking information stored in the content file.
[0213] Furthermore, the playback device 400 may be used as a second information processing device to which the above-described method 4 is applied. In this case, the file acquisition unit 413 of the playback device 400 selects and acquires a content file storing a desired direction-limited bitstream, based on second spatial masking information stored in an MPD, which is a control file that controls distribution of content files. Then, the file processing unit 414 of the playback device 400 acquires the desired direction-limited bitstream from the acquired content file, based on the first spatial masking information stored in the content file.
[0214] Note that processing units of the playback device 400 that are not essential for each method may be omitted.
[0215] <Playback Process Flow 1> An example of the flow of processing executed by the playback device 400 having the above configuration will be described with reference to the flowcharts of Figures 25 to 28. First, an example of the flow of playback processing when the above-mentioned method 2 is applied will be described with reference to the flowchart of Figure 25. That is, in this case, spatial masking information is stored in the MPD, and the playback device 400 uses the MPD (spatial masking information) to select a desired bitstream and distributes a content file storing the desired bitstream.
[0216] When the playback process is started, in step S401, the MPD acquisition unit 411 acquires an MPD that stores spatial masking information.
[0217] In step S402, the sound source direction setting unit 412 sets the sound source direction with respect to the listener based on, for example, scene information, changes in the audio scene, changes in the position or direction of the listener, and the like.
[0218] In step S403, the file acquisition unit 413 executes a bitstream acquisition process and acquires a desired content file based on the MPD and sound source direction.
[0219] In step S404, the file processing unit 414 obtains the desired bitstream from the obtained content file.
[0220] In step S405, the decoding unit 415 decodes the acquired bitstream to generate (reconstruct) audio data.
[0221] In step S406, the audio output processing unit 416 arranges the audio data in a 3D space based on the scene information, etc., and generates audio data for output by rendering. Then, the audio output processing unit 416 causes the audio output unit 417 to output an output audio corresponding to the generated audio data for output.
[0222] When the process of step S406 is completed, the playback process ends.
[0223] <Bitstream Acquisition Processing Flow 2> An example of the flow of the bitstream acquisition processing executed in step S403 of FIG. 25 will be described with reference to the flowchart of FIG.
[0224] When the bitstream acquisition process is started, the file acquisition unit 413 selects representations that can be transmitted based on the currently available bandwidth and the MPD, and generates a candidate list of representations in step S421. This candidate list is a list of representations that stores information about content files that are candidates for acquisition.
[0225] In step S422, the file acquisition unit 413 acquires sound source direction information indicating the sound source direction set in step S402 (FIG. 25).
[0226] In step S423, the file acquisition unit 413 acquires sound source direction range information of representations having direction limitation identification information of the spatial masking information from the candidate list generated in step S421, and leaves representations whose sound source directions are all included in the sound source direction range in the candidate list. In other words, the file acquisition unit 413 deletes from the candidate list representations whose sound source direction range does not include one or more sound source directions.
[0227] In step S424, the file acquisition unit 413 selects the representation with the maximum bit rate from the candidate list based on the quality information within the sound source direction range of the spatial masking information.
[0228] In step S425, the file acquisition unit 413 acquires a content file corresponding to the selected representation. That is, the file acquisition unit 413 requests the distribution server to distribute the content file, and acquires the content file (i.e., the content file storing the desired bitstream) distributed in response to the request.
[0229] When the process of step S425 ends, the bitstream acquisition process ends, and the process returns to FIG.
[0230] By performing each process as described above, the playback device 400 can acquire spatial masking information provided using an MPD, and acquire and play back a direction-limited bitstream using the spatial masking information. Therefore, the playback device 400 can suppress a decrease in the coding efficiency of the audio data bitstream while suppressing a decrease in the subjective quality of the reproduced audio perceived by a listener.
[0231] <Playback Process Flow 2> Next, an example of the playback process flow when the above-mentioned method 3 is applied will be described with reference to the flowchart in Fig. 27. That is, in this case, spatial masking information is stored in the content file, and the playback device 400 uses the spatial masking information to select a desired bitstream and acquire the desired bitstream from the content file.
[0232] When the playback process is started, the MPD acquisition unit 411 acquires an MPD in step S441.
[0233] In step S442, the sound source direction setting unit 412 sets the sound source direction with respect to the listener based on, for example, scene information, changes in the audio scene, changes in the position or direction of the listener, and the like.
[0234] In step S443, the file acquisition unit 413 acquires a content file that stores spatial masking information based on the MPD.
[0235] In step S444, the file processing unit 414 executes bitstream processing to obtain a desired bitstream from the obtained content file based on the sound source direction.
[0236] In step S445, the decoding unit 415 decodes the acquired bitstream to generate (reconstruct) audio data.
[0237] In step S446, the audio output processing unit 416 arranges the audio data in 3D space based on the scene information, etc., and generates audio data for output by rendering. Then, the audio output processing unit 416 causes the audio output unit 417 to output an output audio corresponding to the generated audio data for output.
[0238] When the process of step S446 is completed, the playback process ends.
[0239] <Bitstream Acquisition Processing Flow 2> An example of the flow of the bitstream acquisition processing executed in step S444 of FIG. 27 will be described with reference to the flowchart of FIG.
[0240] When the bitstream acquisition process is started, the file processing unit 414 selects processable tracks based on the currently available bitrate and MPD in step S461, and generates a candidate list of tracks. This candidate list is a list of tracks that store content files that are candidates for acquisition.
[0241] In step S462, the file processing unit 414 acquires sound source direction information indicating the sound source direction set in step S442 (FIG. 27).
[0242] In step S463, the file processing unit 414 acquires sound source direction range information stored in tracks having direction limitation identification information of the spatial masking information from the candidate list generated in step S461, and leaves tracks in which all sound source directions are included in the sound source direction range in the candidate list. In other words, the file processing unit 414 deletes from the candidate list tracks in which one or more sound source directions are not included in the sound source direction range.
[0243] In step S464, the file processing unit 414 selects the track with the highest bit rate from the candidate list based on the quality information within the sound source direction range of the spatial masking information.
[0244] In step S465, the file processing unit 414 obtains the bitstream stored in the selected track (i.e., the desired bitstream).
[0245] When the process of step S465 ends, the bitstream acquisition process ends, and the process returns to FIG.
[0246] By performing each process as described above, the playback device 400 can acquire spatial masking information provided using a content file, and acquire and play back a direction-limited bitstream using the spatial masking information. Therefore, the playback device 400 can suppress a decrease in the coding efficiency of the audio data bitstream while suppressing a decrease in the subjective quality of the reproduced audio perceived by a listener.
[0247] <When Method 4 is Applied> Note that, when the above-described Method 4 is applied, the playback process may be performed basically in the same manner as the flowcharts of Figures 25 and 27. However, acquisition of a content file to be distributed is performed based on second spatial masking information stored in the MPD, and acquisition of a bitstream from the acquired content file is performed based on first spatial masking information stored in the content file. In this case, the bitstream acquisition process performed based on the second spatial masking information may be performed basically in the same manner as the flowchart of Figure 26. Furthermore, the bitstream acquisition process performed based on the first spatial masking information may be performed basically in the same manner as the flowchart of Figure 28.
[0248] By performing each process as described above, the playback device 400 can acquire spatial masking information provided using an MPD and a content file, and acquire and play back a direction-limited bitstream using the spatial masking information. Therefore, the playback device 400 can suppress a decrease in the coding efficiency of the audio data bitstream while suppressing a decrease in the subjective quality of the reproduced audio perceived by a listener.
[0249] 6. Supplementary Notes Computer The above-described series of processes can be executed by hardware or software. When the series of processes are executed by software, the programs that make up the software are installed on a computer. Here, the term computer includes computers built into dedicated hardware, and general-purpose personal computers, for example, that can execute various functions by installing various programs.
[0250] FIG. 29 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by a program.
[0251] In a computer 900 shown in FIG. 29, a CPU (Central Processing Unit) 901, a ROM (Read Only Memory) 902, and a RAM (Random Access Memory) 903 are interconnected via a bus 904.
[0252] An input / output interface 910 is also connected to the bus 904. To the input / output interface 910, an input unit 911, an output unit 912, a storage unit 913, a communication unit 914, and a drive 915 are connected.
[0253] The input unit 911 includes, for example, a keyboard, a mouse, a microphone, a touch panel, and an input terminal. The output unit 912 includes, for example, a display, a speaker, and an output terminal. The storage unit 913 includes, for example, a hard disk, a RAM disk, and a non-volatile memory. The communication unit 914 includes, for example, a network interface. The drive 915 drives removable media 921 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0254] In a computer configured as described above, the CPU 901 performs the above-described series of processes by, for example, loading a program stored in the storage unit 913 into the RAM 903 via the input / output interface 910 and the bus 904 and executing the program. The RAM 903 also stores data necessary for the CPU 901 to execute various processes as appropriate.
[0255] The program executed by the computer can be applied by recording it on, for example, a removable medium 921 such as a package medium. In this case, the program can be installed in the storage unit 913 via the input / output interface 910 by inserting the removable medium 921 into the drive 915.
[0256] This program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, digital satellite broadcasting, etc. In this case, the program can be received by the communication unit 914 and installed in the storage unit 913.
[0257] Alternatively, this program can be installed in advance in the ROM 902 or the storage unit 913 .
[0258] <Applicable Targets of This Technology> This technology can be applied to any encoding / decoding method. Also, this technology can be applied to any distribution method or file container.
[0259] Furthermore, the present technology can be applied to any configuration, for example, various electronic devices.
[0260] Furthermore, for example, the present technology can also be implemented as part of an apparatus, such as a processor (e.g., a video processor) as a system LSI (Large Scale Integration), a module using multiple processors (e.g., a video module), a unit using multiple modules (e.g., a video unit), or a set in which other functions are added to a unit (e.g., a video set).
[0261] Furthermore, for example, the present technology can also be applied to a network system configured with multiple devices. For example, the present technology may be implemented as cloud computing in which multiple devices share and collaborate on processing via a network. For example, the present technology may be implemented in a cloud service that provides image (video)-related services to any terminal, such as a computer, an AV (Audio Visual) device, a portable information processing terminal, or an IoT (Internet of Things) device.
[0262] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are housed in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.
[0263] <Fields and uses to which this technology can be applied> Systems, devices, processing units, etc. to which this technology is applied can be used in any field, for example, transportation, medical care, crime prevention, agriculture, livestock farming, mining, beauty, factories, home appliances, weather, nature monitoring, etc. In addition, the uses thereof are also arbitrary.
[0264] For example, the present technology can be applied to systems and devices used to provide viewing content, etc. Furthermore, for example, the present technology can also be applied to systems and devices used for transportation, such as monitoring traffic conditions and controlling automatic driving. Furthermore, for example, the present technology can also be applied to systems and devices used for security. Furthermore, for example, the present technology can also be applied to systems and devices used for automatic control of machines, etc. Furthermore, for example, the present technology can also be applied to systems and devices used for agriculture and livestock farming. Furthermore, for example, the present technology can also be applied to systems and devices used to monitor natural conditions, such as volcanoes, forests, and oceans, and wildlife. Furthermore, for example, the present technology can also be applied to systems and devices used for sports.
[0265] <Others> In this specification, a "flag" refers to information for identifying multiple states, and includes not only information used to identify two states, true (1) or false (0), but also information capable of identifying three or more states. Therefore, the value that this "flag" can take may be, for example, two values, 1 / 0, or three or more values. That is, the number of bits constituting this "flag" is arbitrary, and may be one bit or multiple bits. Furthermore, identification information (including flags) can be included not only in a bitstream, but also in a bitstream that includes differential information of the identification information relative to certain reference information. Therefore, in this specification, "flag" and "identification information" encompass not only the information itself, but also differential information relative to the reference information.
[0266] Furthermore, various information (e.g., metadata) related to the coded data (bitstream) may be transmitted or recorded in any form as long as it is associated with the coded data. Here, the term "associate" means, for example, making one piece of data available (linked) when processing the other piece of data. That is, data associated with each other may be combined into one piece of data or may be stored as separate pieces of data. For example, information associated with coded data (image) may be transmitted over a transmission path separate from that of the coded data (image). Furthermore, for example, information associated with coded data (image) may be recorded on a recording medium separate from that of the coded data (image) (or on a different recording area of the same recording medium). Note that this "association" may refer not to the entire data, but to only a portion of the data. For example, an image and information corresponding to that image may be associated with each other in any unit, such as multiple frames, one frame, or a portion of a frame.
[0267] In this specification, terms such as "composite," "multiplex," "add," "integrate," "include," "store," "embed," "insert," and the like refer to combining multiple items into one, such as combining encoded data and metadata into one piece of data, and refer to one method of "associating" as described above.
[0268] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present technology.
[0269] For example, a configuration described as one device (or processing unit) may be divided and configured as multiple devices (or processing units). Conversely, configurations described above as multiple devices (or processing units) may be combined and configured as one device (or processing unit). Of course, configurations other than those described above may be added to the configuration of each device (or each processing unit). Furthermore, as long as the configuration and operation of the entire system are substantially the same, part of the configuration of one device (or processing unit) may be included in the configuration of another device (or other processing unit).
[0270] Furthermore, for example, the above-described program may be executed in any device, as long as the device has the necessary functions (functional blocks, etc.) and is able to obtain the necessary information.
[0271] Also, for example, each step of a single flowchart may be executed by a single device, or may be shared and executed by multiple devices. Furthermore, when a single step includes multiple processes, the multiple processes may be executed by a single device, or may be shared and executed by multiple devices. In other words, multiple processes included in a single step can be executed as multiple step processes. Conversely, processes described as multiple steps can be executed collectively as a single step.
[0272] For example, the steps of a program executed by a computer may be executed in chronological order in the order described herein, or may be executed in parallel or individually at the required timing, such as when a call is made. In other words, as long as no contradiction occurs, the steps may be executed in an order different from the order described above. Furthermore, the steps of this program may be executed in parallel with the processing of another program, or may be executed in combination with the processing of another program.
[0273] Furthermore, for example, multiple technologies related to the present technology can be implemented independently and independently, as long as no contradiction occurs. Of course, any multiple technologies can also be implemented in combination. For example, part or all of the present technology described in any embodiment can be implemented in combination with part or all of the present technology described in another embodiment. Furthermore, part or all of any of the above-described present technologies can be implemented in combination with other technologies not described above.
[0274] The present technology may also be configured as follows. (1) An information processing device including: an encoding unit that encodes audio data with bit allocation based on a direction-limited spatial masking threshold, which is a spatial masking threshold that limits a sound source direction range, and generates a direction-limited bitstream. (2) The information processing device according to (1), further including a spatial masking information generation unit that generates spatial masking information related to spatial masking for the direction-limited bitstream. (3) The information processing device according to (2), in which the spatial masking information includes direction limitation identification information for identifying that the bitstream of the audio data is the direction-limited bitstream. (4) The information processing device according to (2) or (3), in which the spatial masking information includes sound source direction range information related to the sound source direction range. (5) The information processing device according to any of (2) to (4), in which the spatial masking information includes sound source direction range quality information related to subjective quality of reproduced audio of the direction-limited bitstream when all sound sources are present within the sound source direction range. (6) The information processing device according to (5), wherein the quality information within sound source direction range includes information specifying another bitstream that provides subjective quality of reproduced audio equivalent to that of the direction-restricted bitstream. (7) The information processing device according to (5) or (6), wherein the quality information within sound source direction range includes information expressing the subjective quality of reproduced audio of the direction-restricted bitstream by a level. (8) The information processing device according to any of (5) to (7), wherein the quality information within sound source direction range includes information indicating a bitrate equivalent to the subjective quality of reproduced audio of the direction-restricted bitstream. (9) The information processing device according to any of (5) to (8), wherein the spatial masking information further includes guaranteed quality information indicating the subjective quality of reproduced audio of the direction-restricted bitstream that is guaranteed regardless of the sound source direction. (10) The information processing device according to any of (5) to (10), wherein the spatial masking information further includes out-of-sound source direction range quality information related to the subjective quality of reproduced audio of the direction-restricted bitstream when the sound source is present outside the sound source direction range.(11) The information processing device according to any one of (2) to (10), wherein the spatial masking information includes a plurality of sets of sound source direction range information regarding the sound source direction range and sound source direction range quality information regarding the subjective quality of reproduced sound of the direction-limited bit stream when all sound sources are present within the sound source direction range. (12) The information processing device according to any one of (2) to (11), wherein the spatial masking information generation unit further generates the spatial masking information for an omnidirectional bit stream in which the sound data is encoded using bit allocation based on an omnidirectional spatial masking threshold, the sound source direction range being an omnidirectional spatial masking threshold. (13) The information processing device according to any one of (2) to (12), wherein the encoding unit generates the direction-limited bit stream for each sound source, and the spatial masking information generation unit generates the spatial masking information for each sound source. (14) The information processing device according to any one of (2) to (13), wherein the spatial masking information further includes listener control information regarding control of a position and direction of a listener according to a sound source direction. (15) The information processing device according to any one of (2) to (14), further comprising: a control file generation unit that generates a control file that controls distribution of the direction limited bitstream and stores the spatial masking information in the control file. (16) The information processing device according to any one of (2) to (14), further comprising: a content file generation unit that generates a content file and stores the direction limited bitstream and the spatial masking information in the content file. (17) The information processing device according to any one of (2) to (14), further comprising: a content file generation unit that generates a content file and stores the direction limited bitstream and first spatial masking information related to the spatial masking of the direction limited bitstream in the content file; and a control file generation unit that generates a control file that controls distribution of the direction limited bitstream and stores second spatial masking information related to the spatial masking of the direction limited bitstream in the control file.(18) An information processing method for encoding audio data with bit allocation based on a direction-limited spatial masking threshold, which is a spatial masking threshold that limits the range of sound source directions, and generating a direction-limited bit stream. (19) A program for causing a computer to execute a process for encoding audio data with bit allocation based on a direction-limited spatial masking threshold, which is a spatial masking threshold that limits the range of sound source directions, and generating a direction-limited bit stream.
[0275] (21) An information processing device comprising: a bitstream acquisition unit that acquires a direction-limited bitstream in which audio data is coded with bit allocation based on a direction-limited spatial masking threshold that is a spatial masking threshold that limits a sound source direction range; and a decoding unit that decodes the acquired direction-limited bitstream. (22) The information processing device according to (21), in which the bitstream acquisition unit acquires the direction-limited bitstream based on spatial masking information related to spatial masking. (23) The information processing device according to (22), in which the spatial masking information includes direction limitation identification information for identifying that the bitstream of the audio data is the direction-limited bitstream. (24) The information processing device according to (22) or (23), in which the spatial masking information includes sound source direction range information related to the sound source direction range. (25) The information processing device according to any of (22) to (24), in which the spatial masking information includes sound source direction range quality information related to subjective quality of reproduced audio of the direction-limited bitstream when all sound sources are present within the sound source direction range. (26) The information processing device according to (25), wherein the quality information within the sound source direction range includes information specifying another bitstream that provides subjective quality of reproduced audio equivalent to that of the direction-restricted bitstream. (27) The information processing device according to (25) or (26), wherein the quality information within the sound source direction range includes information representing the subjective quality of reproduced audio of the direction-restricted bitstream by a level. (28) The information processing device according to any of (25) to (27), wherein the quality information within the sound source direction range includes information indicating a bitrate equivalent to the subjective quality of reproduced audio of the direction-restricted bitstream. (29) The information processing device according to any of (25) to (28), wherein the spatial masking information further includes guaranteed quality information indicating the subjective quality of reproduced audio of the direction-restricted bitstream that is guaranteed regardless of the sound source direction. (30) The information processing device according to any one of (25) to (29), wherein the spatial masking information further includes out-of-sound-source-direction-range quality information regarding subjective quality of reproduced audio of the direction-limited bitstream when the sound source is present outside the sound source direction range.(31) The information processing device according to any one of (22) to (30), wherein the spatial masking information includes a plurality of sets of sound source direction range information regarding the sound source direction range and sound source direction range quality information regarding the subjective quality of reproduced audio of the direction-limited bit stream when all sound sources are present within the sound source direction range. (32) The information processing device according to any one of (22) to (31), wherein the bit stream acquisition unit further acquires, based on the spatial masking information, an omnidirectional bit stream in which the audio data is encoded with bit allocation based on an omnidirectional spatial masking threshold, where the sound source direction range is an omnidirectional spatial masking threshold. (33) The information processing device according to any one of (22) to (32), wherein the bit stream acquisition unit acquires the direction-limited bit stream for each sound source based on the spatial masking information for each sound source. (34) The information processing device according to any one of (22) to (33), wherein the spatial masking information further includes listener control information regarding control of a listener's position and direction according to a sound source direction. (35) The information processing device according to any one of (22) to (34), wherein the bitstream acquisition unit acquires the direction-limited bitstream based on the spatial masking information stored in a control file that controls distribution of the direction-limited bitstream. (36) The information processing device according to any one of (22) to (34), wherein the bitstream acquisition unit acquires the direction-limited bitstream from a content file based on the spatial masking information stored in the content file. (37) The information processing device according to any one of (22) to (34), wherein the bitstream acquisition unit acquires the direction-limited bitstream from the content file based on first spatial masking information related to the spatial masking of the direction-limited bitstream that is stored in the content file and second spatial masking information related to the spatial masking of the direction-limited bitstream that is stored in a control file that controls distribution of the direction-limited bitstream.(38) An information processing method for acquiring a direction-limited bit stream in which audio data is encoded with bit allocation based on a direction-limited spatial masking threshold that is a spatial masking threshold that limits a sound source direction range, and decoding the acquired direction-limited bit stream. (39) A program for causing a computer to execute a process for acquiring a direction-limited bit stream in which audio data is encoded with bit allocation based on a direction-limited spatial masking threshold that is a spatial masking threshold that limits a sound source direction range, and decoding the acquired direction-limited bit stream.
[0276] 300 File generation device, 311 Audio data preprocessing unit, 312 Encoding unit, 313 Spatial masking information generation unit, 314 File generation unit, 315 MPD generation unit, 316 Storage unit, 317 Supply unit, 400 Playback device, 411 MPD acquisition unit, 412 Sound source direction setting unit, 413 File acquisition unit, 414 File processing unit, 415 Decoding unit, 416 Audio output processing unit, 417 Audio output unit, 900 Computer
Claims
1. An information processing device having an encoding unit that encodes audio data with a bit allocation based on a direction-limited spatial masking threshold, which is a spatial masking threshold that limits the sound source direction range, and generates a direction-limited bit stream.
2. The information processing device according to claim 1, further comprising a spatial masking information generating unit that generates spatial masking information regarding spatial masking for the direction limited bit stream.
3. The information processing device according to claim 2, wherein the spatial masking information includes direction restriction identification information for identifying that the bit stream of the audio data is the direction restricted bit stream.
4. The information processing device according to claim 2, wherein the spatial masking information includes sound source direction range information relating to the sound source direction range.
5. The information processing device according to claim 2, wherein the spatial masking information includes quality information within a sound source direction range regarding the subjective quality of reproduced audio of the direction-limited bitstream when all sound sources are present within the sound source direction range.
6. The information processing device according to claim 5, wherein the sound source direction in-range quality information includes information specifying another bitstream that provides a subjective quality of reproduced sound equivalent to that of the direction-limited bitstream.
7. The information processing device according to claim 5, wherein the sound source direction in-range quality information includes information expressing a subjective quality of the reproduced sound of the direction-limited bit stream at a level.
8. The information processing device according to claim 5, wherein the quality information within the sound source direction range includes information indicating a bit rate equivalent to a subjective quality of the reproduced sound of the direction-limited bit stream.
9. The information processing device according to claim 5, wherein the spatial masking information further includes guaranteed quality information indicating a subjective quality of the reproduced sound of the direction-limited bit stream that is guaranteed regardless of the sound source direction.
10. The information processing device according to claim 5, wherein the spatial masking information further includes outside sound source direction range quality information regarding the subjective quality of the reproduced sound of the direction-limited bit stream when the sound source is present outside the sound source direction range.
11. The information processing device of claim 2, wherein the spatial masking information includes multiple sets of sound source direction range information regarding the sound source direction range and sound source direction range quality information regarding the subjective quality of the reproduced sound of the direction-limited bitstream when all sound sources are present within the sound source direction range.
12. The information processing device according to claim 2, wherein the spatial masking information generation unit further generates the spatial masking information for an omnidirectional bit stream in which the audio data is encoded with a bit allocation based on an omnidirectional spatial masking threshold, the sound source direction range being an omnidirectional spatial masking threshold.
13. The information processing device according to claim 2, wherein the encoding unit generates the direction-limited bit stream for each sound source, and the spatial masking information generating unit generates the spatial masking information for each sound source.
14. The information processing device according to claim 2, wherein the spatial masking information further includes listener control information relating to control of the listener's position and direction according to the sound source direction.
15. The information processing device according to claim 2, further comprising a control file generating unit that generates a control file for controlling the distribution of said direction-limited bitstream and stores said spatial masking information in said control file.
16. The information processing device according to claim 2, further comprising a content file generating unit that generates a content file and stores the direction-limited bit stream and the spatial masking information in the content file.
17. The information processing device of claim 2, further comprising: a content file generation unit that generates a content file and stores the direction-limited bitstream and first spatial masking information regarding the spatial masking of the direction-limited bitstream in the content file; and a control file generation unit that generates a control file that controls distribution of the direction-limited bitstream and stores second spatial masking information regarding the spatial masking of the direction-limited bitstream in the control file.
18. An information processing method for encoding audio data with a bit allocation based on a direction-limited spatial masking threshold, which is a spatial masking threshold that limits the sound source direction range, and generating a direction-limited bit stream.
19. An information processing device comprising: a bit stream acquisition unit that acquires a direction-limited bit stream in which audio data is encoded with bit allocation based on a direction-limited spatial masking threshold, which is a spatial masking threshold that limits the sound source direction range; and a decoding unit that decodes the acquired direction-limited bit stream.
20. An information processing method comprising: acquiring a direction-limited bit stream in which audio data is encoded with a bit allocation based on a direction-limited spatial masking threshold, which is a spatial masking threshold that limits the sound source direction range; and decoding the acquired direction-limited bit stream.
Citation Information
Patent Citations
Acoustic signal encoding method, acoustic signal decoding method, program, encoding device, acoustic system and complexing device
WO2020171049A1
Signal processing device, acoustic output device, and signal processing method
WO2023171280A1
Encoding device and method, decoding device and method, and program
WO2023286698A1