Information processing apparatus and information processing method

By segmenting multiple types of audio data onto tracks and arranging relevant information to generate files that support easy reproduction of audio data, the problem of reproducing multiple sets of audio data is solved, and reproduction efficiency is improved.

CN113851139BActive Publication Date: 2026-01-27SONY GROUP CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111111163.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2015-06-22
Filing Date
2015-06-30
Publication Date
2026-01-27
Estimated Expiration
2035-06-30

AI Technical Summary

Technical Problem

Existing technologies have failed to effectively solve the problem of reproducing a predetermined group of audio data from multiple sets of audio data.

Method used

By segmenting audio data of multiple categories into tracks and arranging category-related information, files are generated to support easy reproduction of the audio data.

Benefits of technology

It enables easy reproduction of predetermined types of audio data from multiple categories, reduces the decoding processing load, and improves reproduction efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113851139B_ABST
    Figure CN113851139B_ABST
Patent Text Reader

Abstract

The present application discloses an information processing apparatus and an information processing method in which a predetermined kind of audio data can be easily reproduced from a plurality of kinds of audio data. A file generation device generates an audio file in which a plurality of groups of audio streams are arranged to be divided into tracks for each group or a set of more than one groups, and information related to the plurality of groups is arranged. The present application can be applied to an information processing system constituted by, for example, a file generation device for generating a file, a web server for recording a file generated by the file generation device, and a video reproduction terminal for reproducing the file, and the like.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of international application PCT / JP2015 / 068751, filed on June 30, 2015, which entered the national phase on December 23, 2016, with application number 201580034444.X and entitled "Information Processing Apparatus and Information Processing Method", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This disclosure relates to information processing apparatus and information processing method, and more particularly to information processing apparatus and information processing method that enable easy reproduction of a predetermined type of audio data from multiple types of audio data. Background Technology

[0003] In recent years, streaming services on the Internet have surpassed popular video over OTT-V as the mainstream. A technology that has become increasingly popular as a foundational technology is Moving Picture Experts Group-based HTTP-based Dynamic Adaptive Streaming (MPEG-DASH) (see, for example, Non-Patent Literature 1).

[0004] In MPEG-DASH, the allocation server prepares motion picture data sets with different screen sizes and encoding speeds for a single motion picture content, and the playback terminal requests the motion picture data set with the optimal screen size and encoding speed based on the conditions of the transmission path, thus enabling adaptive stream allocation.

[0005] Reference List

[0006] Non-patent literature

[0007] Non-patent document 1: MPEG-DASH (HTTP-based Dynamic Adaptive Streaming) (URL: http: / / mpeg.chiariglione.org / standards / mpeg-dash / media-presentation-description-and-segment-formats / text-isoiec-23009-12012-dam-1) Summary of the Invention

[0008] The problem to be solved by this invention

[0009] However, the ease of reproduction of a predetermined set of audio data in multiple sets of audio data has not yet been taken into account.

[0010] This disclosure is made in view of the above-mentioned problems, and this disclosure supports the easy reproduction of desired audio data from multiple sets of audio data.

[0011] Solution to the problem

[0012] The information processing apparatus of the first aspect of this disclosure is an information processing apparatus including a file generation unit that generates files in which multiple types of audio data are segmented into tracks and arranged for each or more of the types, and information related to the multiple types is arranged therein.

[0013] The information processing method of the first aspect of this disclosure corresponds to the information processing apparatus of the first aspect of this disclosure.

[0014] In a first aspect of this disclosure, a file is generated in which audio data of multiple categories are segmented into tracks and arranged for each or more of the categories, and information related to the multiple categories is arranged therein.

[0015] The information processing apparatus of the second aspect of this disclosure is an information processing apparatus including a reproduction unit that reproduces audio data of a predetermined track from a file, wherein multiple types of audio data in the file are segmented into tracks and arranged for each or more of the types, and information related to the multiple types is arranged.

[0016] The information processing method of the second aspect of this disclosure corresponds to the information processing apparatus of the second aspect of this disclosure.

[0017] In a second aspect of this disclosure, audio data of a predetermined track is reproduced from a file in which audio data of multiple types is segmented into tracks and arranged for each or more of the types, and information related to the multiple types is arranged.

[0018] It should be noted that the information processing apparatus of the first aspect and the information processing apparatus of the second aspect can be implemented by having a computer execute a program.

[0019] In addition, in order to realize the information processing apparatus of the first and second aspects, a program executed by a computer can be transmitted via a transmission medium or a program executed by a computer can be provided by recording it on a recording medium.

[0020] Effects of the present invention

[0021] According to the first aspect of this disclosure, a file can be generated. Furthermore, according to the first aspect of this disclosure, a file can be generated that allows for easy reproduction of a predetermined type of audio data from multiple types of frequency data.

[0022] According to the second aspect of this disclosure, audio data can be reproduced. Furthermore, according to the second aspect of this disclosure, audio data of a predetermined type among multiple types of audio data can be easily reproduced. Attached Figure Description

[0023] Figure 1 A diagram illustrating the structure of an MPD file.

[0024] Figure 2 A diagram illustrating the relationship between "Period", "Representation", and "Segment".

[0025] Figure 3 A diagram illustrating the hierarchical structure of an MPD file.

[0026] Figure 4 A diagram illustrating the relationship between the structure of an MPD file and its timeline.

[0027] Figure 5 This is a diagram illustrating the outline of the track format used to explain the 3D audio file format of MP4.

[0028] Figure 6 A diagram illustrating the structure of a moov box.

[0029] Figure 7 A diagram illustrating the hierarchical structure of 3D audio.

[0030] Figure 8 This is a diagram illustrating the general structure of the information processing system to which this disclosure is applied in the first embodiment.

[0031] Figure 9 A schematic diagram illustrating a first example of a track in a first embodiment to which this disclosure is applied.

[0032] Figure 10 A diagram illustrating an example of the syntax for sample entries of a basic track.

[0033] Figure 11 This is a diagram illustrating an example of the syntax for sample entries of tracks that form a switch group.

[0034] Figure 12 A diagram illustrating a first example of a fragment structure.

[0035] Figure 13 A diagram illustrating a second example of a fragment structure.

[0036] Figure 14 A diagram illustrating an example description of the level assignment box.

[0037] Figure 15 A diagram illustrating a first description example of an MPD file to which this disclosure applies in a first embodiment.

[0038] Figure 16 To show Figure 8 A block diagram illustrating a configuration example for the file generation device.

[0039] Figure 17 A flowchart, used to describe Figure 16 File generation processing of the file generation device.

[0040] Figure 18 The diagram illustrates the use of Figure 8 A configuration example of a streaming playback unit implemented in a motion picture playback terminal.

[0041] Figure 19 A flowchart, used to describe Figure 18 The reproduction process of the stream reproduction unit.

[0042] Figure 20 A schematic diagram illustrating a second example of a track used to describe the application of this disclosure to a first embodiment.

[0043] Figure 21 This is a diagram illustrating an example of the syntax for a sample group entry in a switch group.

[0044] Figure 22 A diagram illustrating an example of the syntax of sample entries for each group of tracks.

[0045] Figure 23 A diagram illustrating a third example of an audio file track.

[0046] Figure 24 This is a diagram illustrating a second description example of an MPD file.

[0047] Figure 25 This is a diagram illustrating another example of a second description example of an MPD file.

[0048] Figure 26 This is a schematic diagram of a fourth example used to describe the tracks in an audio file.

[0049] Figure 27 This is a diagram illustrating an example of the third description of an MPD file.

[0050] Figure 28 This is a schematic diagram of a fifth example used to describe the tracks in an audio file.

[0051] Figure 29 A diagram illustrating an example of the syntax of a sample entry where 4cc is “mha3”.

[0052] Figure 30A diagram illustrating another example of the syntax of a sample entry where 4cc is “mha3”.

[0053] Figure 31 This is a diagram illustrating an example of the fourth description of an MPD file.

[0054] Figure 32 A diagram illustrating a summary of another example of a third example used to describe the tracks in an audio file.

[0055] Figure 33 A diagram illustrating a summary of another example of a fourth example used to describe the tracks in an audio file.

[0056] Figure 34 A diagram illustrating a summary of another example of the fifth example used to describe the tracks in an audio file.

[0057] Figure 35 This is a diagram illustrating a sixth example of a track used to describe an audio file.

[0058] Figure 36 To show Figure 35 A diagram illustrating the syntax of sample entries for basic and group tracks.

[0059] Figure 37 A diagram illustrating yet another example of the syntax for a sample entry where 4cc is “mha3”.

[0060] Figure 38 This is a diagram illustrating the outline of the track in a second embodiment to which this disclosure is applied.

[0061] Figure 39 This is a diagram illustrating a first description example of an MPD file in a second embodiment to which this disclosure is applied.

[0062] Figure 40 This is a diagram illustrating the outline of an information processing system to which this disclosure is applied in a third embodiment.

[0063] Figure 41 To show Figure 40 A block diagram illustrating a configuration example for the file generation device.

[0064] Figure 42 A flowchart, used to describe Figure 41 File generation processing of the file generation device.

[0065] Figure 43 This is a block diagram showing the components... Figure 40 A configuration example of a streaming playback unit implemented in a motion picture playback terminal.

[0066] Figure 44A flowchart, used to describe Figure 43 An example of the reproduction processing of the stream reproduction unit.

[0067] Figure 45 This is a diagram illustrating a second description example of an MPD file in a second embodiment to which this disclosure is applied.

[0068] Figure 46 A diagram illustrating a third description example of an MPD file in a second embodiment to which this disclosure is applied.

[0069] Figure 47 This is a diagram illustrating a fourth description example of an MPD file in a second embodiment to which this disclosure is applied.

[0070] Figure 48 This is a diagram illustrating a fifth description example of an MPD file in a second embodiment to which this disclosure is applied.

[0071] Figure 49 A diagram illustrating a sixth descriptive example of an MPD file in a second embodiment to which this disclosure is applied.

[0072] Figure 50 This is a diagram illustrating a seventh description example of an MPD file in a second embodiment to which this disclosure is applied.

[0073] Figure 51 A diagram illustrating an example of the track structure of an audio file that includes multiple basic tracks.

[0074] Figure 52 This is a diagram illustrating another example of the track structure of an audio file that includes multiple basic tracks.

[0075] Figure 53 A block diagram illustrating an example configuration of computer hardware. Detailed Implementation

[0076] In the following text, the presuppositions of this disclosure and embodiments for implementing this disclosure (hereinafter referred to as embodiments) will be described. It should be noted that the description will be given in the following order.

[0077] 0. The presuppositions of this disclosure ( Figures 1 to 7 )

[0078] 1. First embodiment ( Figures 8 to 37 )

[0079] 2. Second embodiment ( Figures 38 to 50 )

[0080] 3. Other examples of basic tracks ( Figure 51 and Figure 52 )

[0081] 4. Third embodiment ( Figure 53 )

[0082] <Presuppositions of this disclosure>

[0083] (Explanation of the structure of an MPD file)

[0084] Figure 1 This is a diagram showing the structure of an MPEG-DASH Media Representation Description (MPD) file.

[0085] In the analysis (parsing) of MPD files, the "Representation" attribute included in the "Period" section of the MPD file is... Figure 1 (According to media reports) the best one was selected.

[0086] Then, the file is retrieved and processed by referring to the Uniform Resource Locator (URL) of the "Initialization Segment" at the beginning of the selected "Representation". Next, the file is retrieved and reproduced by referring to the URLs of subsequent "Media Segments".

[0087] It should be noted that, Figure 2 This illustrates the relationship between "Period," "Representation," and "Segment" in an MPD file. Specifically, a piece of moving image content can be managed through a "Period" with a longer time unit than a segment, and within each "Period," it can be managed in segments using "Segments." Furthermore, within each "Period," moving image content can be managed using "Representation" based on stream attributes.

[0088] Therefore, MPD files have the following in "Period" and below. Figure 3 The hierarchical structure is shown. Additionally, in Figure 4 The example shows the arrangement of the MPD file's structure along the timeline. From Figure 4 It is clear that there are multiple "Representations" for the same fragment. By adaptively selecting any of these "Representations," the flow of attributes desired by the user can be captured and reproduced.

[0089] (An overview of 3D audio file formats)

[0090] Figure 5 A diagram illustrating the outline of the track used to explain the MP4 3D audio file format.

[0091] In MP4 files, the encoding and decoding information for moving image content and its position within the file can be managed for each track. In the MP4 3D audio file format, all audio streams (elementary streams, ES) of the 3D audio (Channel audio / Object audio / SAOC Object audio / HOA audio / metadata) are recorded as a track, unit by sample (frame). Additionally, the 3D audio encoding and decoding information (Profile / level / audio configuration) is stored as sample entries.

[0092] Channel audio, which constitutes 3D audio, is audio data organized by channel, while object audio is audio data organized by object. It's important to note that the object is the sound source, and audio data is acquired object-by-object using a microphone or similar device attached to the object. Objects can be physical objects (such as a fixed microphone stand) or moving objects (such as a person).

[0093] Additionally, SAOC Object audio is audio data encoded in Spatial Audio Object (SAOC), while HOA audio is audio data in High-Order Ambient Stereo Mix (HOA), and metadata is metadata for Channel audio, Object audio, SAOC Object audio, and HOA audio.

[0094] (The structure of the Moov box)

[0095] Figure 6 This is a diagram illustrating the moov box structure of an MP4 file.

[0096] like Figure 6 As shown, in an MP4 file, image data and audio data are recorded on different tracks. Figure 6 Although details are not described, the audio data tracks are similar to the image data tracks. Sample entries are included in the sample description arranged in the stsd boxes within the moov boxes.

[0097] In this method, during the broadcast or local storage playback of MP4 files, the server typically sends the entire 3D audio stream. The client then decodes and outputs only the necessary 3D audio stream while parsing the entire stream. However, in situations with high bitrates or limitations on local memory read rates, it is desirable to reduce the decoding load by only acquiring the necessary 3D audio stream.

[0098] Furthermore, in the streaming playback of MP4 files conforming to MPEG-DASH, the server prepares multiple audio streams with different encoding speeds. Therefore, the client can select and obtain the audio stream with the optimal encoding speed for the playback environment by acquiring only the necessary 3D audio streams.

[0099] As described above, in this disclosure, by dividing the 3D audio stream into tracks according to type and arranging the audio streams in the audio file, it is possible to efficiently obtain only the audio streams of a predetermined type of 3D audio. Therefore, in broadcasting or local memory playback, the decoding processing load can be reduced. Furthermore, in streaming playback, the highest quality audio stream from the necessary 3D audio streams can be reproduced according to the frequency band.

[0100] (Description of 3D audio hierarchy)

[0101] Figure 7 A diagram illustrating the 3D audio hierarchy.

[0102] like Figure 7 As shown, the audio data of 3D audio consists of different audio elements in each audio data set. The types of audio elements include mono elements (SCE) and channel pairs elements (CPE). The audio element type for one-channel audio data is SCE, while the audio element type for two-channel audio data is CPE.

[0103] Audio elements of the same audio type (Channel / Object / SAOC Object / HOA) form a group. Therefore, instances of the group type (GroupType) include Channels, Objects, SAOC Objects, and HOA. Two or more groups can be combined to form a switch group or a group preset as needed.

[0104] A switch group is a group whose audio stream is exclusively reproduced (exclusive reproduction group). That is, such as... Figure 7As shown, in the presence of a group for object audio used in English (EN) and a group for object audio used in French (FR), only one of these groups should be reproduced. Therefore, the switch group is formed by a group for object audio used in English (group ID 2) and a group for object audio used in French (group ID 3). Thus, either the object audio used in English or the object audio used in French is reproduced exclusively.

[0105] At the same time, group Preset defines the combination of groups desired by the content creator.

[0106] In addition, the metadata for 3D audio consists of different Ext elements within various metadata sets. Ext element types include Object Metadata, SAOC 3D Metadata, HOA Metadata, DRC Metadata, SpatialFrame, and SaocFrame. The Ext element of Object Metadata contains metadata for all Object audio, and the Ext element of SAOC 3D Metadata contains metadata for all SAOC audio. Similarly, the Ext element of HOA Metadata contains metadata for all HOA audio, and the Ext element of DRC (Dynamic Range Control) Metadata contains metadata for all Object audio, SAOC audio, and HOA audio.

[0107] As mentioned above, the segmentation units of audio data in 3D audio include audio elements, group types, groups, switch groups, and group presets. Therefore, the audio stream of audio data in 3D audio can be segmented into different tracks according to various categories, where the categories are audio elements, group types, groups, switch groups, or group presets.

[0108] Furthermore, the units for segmenting metadata in 3D audio include the type of the Ext element and the corresponding audio element. Therefore, the audio stream of 3D audio metadata can be segmented into different tracks according to various categories, where the category is either an Ext element or the corresponding audio element.

[0109] In the following embodiments, the audio stream of audio data is split into tracks in one or more groups, and the audio stream of metadata is split into tracks according to the various types of Ext elements.

[0110] <First Embodiment>

[0111] (Overview of the information processing system)

[0112] Figure 8 This is a diagram illustrating an overview of an information processing system to which this disclosure is applied in a first embodiment.

[0113] Figure 8 The information processing system 140 is configured such that the network server 142, which is connected to the document generation device 141, and the motion picture reproduction terminal 144 are connected via the Internet 13.

[0114] In the information processing system 140, the network server 142 distributes the audio streams of the group of tracks to be reproduced to the motion picture reproduction terminal 144 using the MPEG-DASH method.

[0115] Specifically, file generation device 141 encodes the audio data and metadata of the 3D audio of the moving image content at various encoding speeds to generate an audio stream. File generation device 141 then creates a file from all the audio streams at each encoding speed and in time units ranging from a few seconds to ten seconds, called segments, to generate an audio file. At this time, file generation device 141 segments the audio streams according to each group and each type of Ext element, and arranges the audio streams into different tracks within the audio file. File generation device 141 then uploads (uploads) the generated audio file to network server 142.

[0116] In addition, the file generation device 141 generates MPD files (management files) for managing audio files, etc. The file generation device 141 uploads the MPD files to the network server 142.

[0117] Network server 142 stores audio files of various encoding speeds and segments, as well as MPD files, uploaded by file generation device 141. In response to a request from motion picture reproduction terminal 144, network server 142 sends the stored audio files, MPD files, etc., to motion picture reproduction terminal 144.

[0118] The motion picture replay terminal 144 includes control software (hereinafter referred to as control software) 161 for running streaming data, motion picture replay software 162, and client software (hereinafter referred to as access software) 163 for Hypertext Transfer Protocol (HTTP) access.

[0119] Control software 161 is software that controls the data streamed from network server 142. Specifically, control software 161 causes motion picture reproduction terminal 144 to retrieve MPD files from network server 142.

[0120] In addition, based on the MPD file, the control software 161 commands the access software 163 to send a transmission request for the group to be reproduced specified by the motion picture reproduction software 162, as well as an audio stream of the track corresponding to the type of the Ext element of that group.

[0121] The motion picture reconstruction software 162 is software that reconstructs an audio stream acquired from the network server 142. Specifically, the motion picture reconstruction software 162 specifies the group to be reconstructed and the type of the corresponding Ext element to the control software 161. Additionally, when a notification of reception start is received from the access software 163, the motion picture reconstruction software 162 decodes the audio stream received from the motion picture reconstruction terminal 144. The motion picture reconstruction software 162 then synthesizes and outputs the audio data obtained as a decoding result as needed.

[0122] Access software 163 is software that uses HTTP to control communication between motion picture reproduction terminal 144 and web server 142 via Internet 13. Specifically, in response to commands from control software 161, access software 163 causes motion picture reproduction terminal 144 to send a transmission request for an audio stream of a track to be reproduced, included in an audio file. Furthermore, in response to the transmission request, access software 163 causes motion picture reproduction terminal 144 to begin receiving the audio stream sent from web server 142 and provides motion picture reproduction software 162 with a notification of the commencement of reception.

[0123] It should be noted that this specification will only describe the audio files for moving image content. However, in reality, the corresponding image files are generated and reproduced together with the audio files.

[0124] (A summary of the first example of an audio file track)

[0125] Figure 9 A schematic diagram of a first example used to describe the tracks of an audio file.

[0126] Note that, in Figure 9 For ease of description, only the audio data tracks from the 3D audio are shown. This also applies to... Figure 20 , Figure 23 , Figure 26 , Figure 28 , Figure 30 , Figures 32 to 35 and Figure 38 .

[0127] like Figure 9As shown, all 3D audio streams are stored in a single audio file (3dauio.mp4). Within this file, the audio streams for each group of 3D audio are divided into different tracks and arranged accordingly. Additionally, information related to the entire 3D audio is set as the base track.

[0128] Track References are located in track boxes within each track. A Track Reference indicates the reference relationship between a given track and other tracks. Specifically, a Track Reference indicates the unique ID (hereinafter referred to as the Track ID) of each other within the reference relationship.

[0129] exist Figure 9 In the example, the track IDs of the basic track, the track in group #1 with group ID 1, the track in group #2 with group ID 2, the track in group #3 with group ID 3, and the track in group #4 with group ID 4 are 1, 2, 3, 4, and 5, respectively. Additionally, the track references of the basic track are 2, 3, 4, and 5, while the track reference for the tracks in groups #1 to #4 is 1, which is the track ID of the basic track. Therefore, the basic track and the tracks in groups #1 to #4 are in a reference relationship. That is, when reproducing the tracks in groups #1 to #4, the basic track is referenced.

[0130] Additionally, the 4cc (character code) of the sample entry for the basic track is "mha2". Within the sample entry for the basic track, there are mhaC boxes containing configuration information for all groups of 3D audio, or configuration information necessary for decoding only the basic track, and mhas boxes containing information related to all groups of 3D audio and switch groups. Group-related information is configured using the group ID, information representing the content of data categorized into groups, etc. Switch group-related information is configured using the switch group ID, the IDs of the groups forming the switch group, etc.

[0131] The 4cc of the sample entries for each group's track is "mhg1", and an mhgC box containing information related to that group can be arranged in the sample entries for each group's track. When the group forms a switch group, an mhsC box containing information related to the switch group is arranged in the sample entries for the track within that group.

[0132] Within the samples of the base track, reference information for the samples of the tracks in the group, or configuration information necessary for decoding the reference information, is arranged. By arranging the samples of the groups referenced by the reference information according to the arrangement order of the reference information, an audio stream of 3D audio before segmentation into tracks can be generated. The reference information is configured by the position and size of the samples of the tracks in the group, the group type, etc.

[0133] (An example of the syntax for sample entries of the basic track).

[0134] Figure 10 A diagram illustrating an example of the syntax for sample entries of a basic track.

[0135] like Figure 10 As shown, the sample entries for the basic track include the mhaC box (MHAC configuration box) and the mhas box (MHA AudioSceneInfo box). The mhaC box describes the configuration information for all groups of the 3D audio, or the configuration information necessary for decoding only the basic track. Additionally, the mhas box describes the audioscene information, which includes information related to all groups of the 3D audio and the switch group. Audio Scene Information Description Figure 7 The hierarchical structure.

[0136] (Examples of the syntax for sample entries of tracks in each group).

[0137] Figure 11 A diagram illustrating an example of the syntax for sample entries of the tracks in each group.

[0138] like Figure 11 As shown, each group's track sample entries contain mhaC boxes (MHAConfigurationBox), mhgC boxes (MHAGroupDefinitionBox), and mhsC boxes (MHASwitchGropuDefinitionBox).

[0139] The mhaC box describes the configuration information necessary for decoding the corresponding track. Additionally, in the mhgC box, the audio scene information associated with the corresponding group is described as a group definition. In the mhsC box, if the corresponding groups form a switch group, the audio scene information associated with the switch group is described in the switch group definition.

[0140] (First example of an audio file segment structure)

[0141] Figure 12 This is a diagram illustrating a first example of the segment structure of an audio file.

[0142] exist Figure 12 In the segment structure, the initial segment is configured by the ftyp box and the moov box. Within the moov box, trak boxes are arranged to represent each track included in the audio file. Additionally, within the moov box, mvex boxes, etc., are arranged, containing information indicating the correspondence between the track IDs of each track and the levels used in the ssix boxes within the media segment.

[0143] Furthermore, a media segment is configured with a sidx box, an ssix box, and one or more sub-segments. The sidx box contains positional information indicating the position of the sub-segment within the audio file. The ssix box contains positional information for each level of the audio stream, located within the mdat box. Note that levels correspond to tracks. Additionally, the positional information for the first track is the positional information of the data comprised of the first track's moof box and the audio stream.

[0144] Sub-segments are set to arbitrary durations and each sub-segment has a pair of moof boxes and mdat boxes, which are shared across all tracks. Within the mdat box, the audio streams of all tracks are uniformly arranged according to arbitrary durations, while the moof box contains management information for the audio streams. The audio streams of tracks arranged within the mdat box are continuous across all tracks.

[0145] exist Figure 12 In the example, track 1 with track ID 1 is the basic track, and tracks 2 to N with track IDs 2 to N are tracks in groups with group IDs 1 to N-1. This also applies to the following... Figure 13 .

[0146] (Second example of audio file segment structure)

[0147] Figure 13 This is a diagram illustrating a second example of the segment structure of an audio file.

[0148] Figure 13 fragment structure and Figure 12 The difference in the fragment structure is that the moof box and mdat box are designed for each track.

[0149] Right now, Figure 13 The initial fragment is similar to Figure 12 The initial fragment. Additionally... Figure 13 The media segments are composed of sidx boxes, ssix boxes, and one or more sub-segments, similar to... Figure 12The media clips. Within the sidx box, the positional information of the sub-segments is arranged, similar to... Figure 12 The sidx box. Within the sidx box, there is location information for the data at levels comprised of moof boxes and mdat boxes.

[0150] Sub-segments are set to arbitrary durations, and each sub-segment has a pair of moof boxes and mdat boxes for each track. That is, in the mdat boxes of each track, the audio stream of the track is uniformly arranged (interleaved) with arbitrary durations, and the management information of the audio stream is arranged in the moof boxes.

[0151] like Figure 12 and Figure 13 As shown, the audio stream of the track is uniformly arranged with arbitrary time lengths. Therefore, compared to when the audio stream is uniformly arranged with samples as units, the efficiency of acquiring audio streams via HTTP and other methods is improved.

[0152] (Example description of an Mvex box)

[0153] Figure 14 To show Figure 12 and Figure 13 A diagram illustrating an example of how hierarchical boxes are arranged within an MVEX box.

[0154] The level assignment box is a box that associates the track ID of each track with the level used in the ssix box. For example... Figure 14 As shown, the base track with track ID 1 is associated with level 0, and the channel audio track with track ID 2 is associated with level 1. Furthermore, the HOA audio track with track ID 3 is associated with level 2, and the object metadata track with track ID 4 is associated with level 3. Additionally, the object audio track with track ID 5 is associated with level 4.

[0155] (Example of the first description of an MPD file)

[0156] Figure 15 This is a diagram illustrating an example of the first description of an MPD file.

[0157] like Figure 15 As shown, the MPD file describes "Representation" which manages the segments of the audio file (3daudio.mp4) for 3D audio, and "SubRepresentation" which manages the tracks included in the segments.

[0158] "Representation" and "SubRepresentation" include "codecs" that indicate the type (profile or level) of encoding and decoding of the corresponding fragment or track as a whole in a 3D file format.

[0159] "SubRepresentation" includes the "level" value set in the level assignment box, which serves as the value indicating the level of the corresponding track. "SubRepresentation" includes "dependencyLevel", which is the value indicating the level corresponding to other tracks that have a reference relationship (dependency) (hereinafter referred to as reference tracks).

[0160] Additionally, "SubRepresentation" includes <EssentialProperty schemeIdUri=“urn:mpeg:DASH:3daudio:2014”value=“dataType,definition”。

[0161] "dataType" is a number indicating the type of content (definition) of the audio scene information described in the sample entry of the corresponding track, and this definition is its content. For example, if GroupDefinition is included in the sample entry of a track, 1 describes the track's "data type," and the group definition is described as "definition." Similarly, if SwitchGroupDefinition is included in the sample entry of a track, 2 describes the track's "data type," and SwitchGroupDefinition is described as "definition." In other words, "dataType" and "definition" are information indicating whether SwitchGroupDefinition exists in the sample entry of the corresponding track. "Definition" is binary data and is encoded using the base64 method.

[0162] It should be noted that, Figure 15 In the example, all groups form a switch group. However, there are cases where groups do not form a switch group. <EssentialProperty schemeIdUri=“urn:mpeg:DASH:3daudio:2014”value=“2,SwitchGroupDefinition”> It is not described in the "SubRepresentation" corresponding to that group. The same applies to the following: Figure 24 , Figure 25 , Figure 31 , Figure 39 , Figure 45 , Figure 47 , Figure 48 and Figure 50 .

[0163] (Configuration example of a file generation device)

[0164] Figure 16 To show Figure 8 A block diagram illustrating a configuration example of the file generation device 141.

[0165] Figure 16 The file generation device 141 is configured with an audio encoding processing unit 171, an audio file generation unit 172, an MPD generation unit 173, and a server upload processing unit 174.

[0166] The audio encoding processing unit 171 of the file generation device 141 encodes the audio data and metadata of the 3D audio of the moving image content at various encoding speeds to generate an audio stream. The audio encoding processing unit 171 provides the audio streams at various encoding speeds to the audio file generation unit 172.

[0167] The audio file generation unit 172 assigns tracks to the audio stream supplied from the audio encoding processing unit 171 for each group and each type of Ext element. The audio file generation unit 172 generates... Figure 12 or Figure 13 The audio file is a segmented audio file, wherein for each encoding speed and segment, the audio stream is arranged in tracks on a sub-segment basis. The audio file generation unit 172 supplies the generated audio file to the MPD generation unit 173.

[0168] MPD generation unit 173 determines the URL, etc., of the web server 142 where the audio file supplied from audio file generation unit 172 will be stored. Then, MPD generation unit 173 generates an MPD file in which the URL, etc., of the audio file is arranged in the "Segment" of the "Representation" of the audio file. MPD generation unit 173 supplies the generated MPD file and the audio file to server upload processing unit 174.

[0169] The server upload processing unit 174 uploads the audio files and MPD files supplied by the MPD generation unit 173 to the network server 142.

[0170] (Description of the processing of the document generation device)

[0171] Figure 17 A flowchart, used to describe Figure 16File generation processing of file generation device 141.

[0172] exist Figure 17 In step S191, the audio encoding processing unit 171 encodes the audio data and metadata of the 3D audio of the moving image content at multiple encoding speeds to generate an audio stream. The audio encoding processing unit 171 provides the audio streams at each encoding speed to the audio file generation unit 172.

[0173] In step S192, the audio file generation unit 172 assigns tracks to the audio stream supplied from the audio encoding processing unit 171 for each group and each type of Ext element.

[0174] In step S193, the audio file generation unit 172 generates... Figure 12 or Figure 13 The audio file is a segmented audio file, wherein for each encoding speed and segment, the audio stream is arranged in tracks as sub-segments. The audio file generation unit 172 supplies the generated audio file to the MPD generation unit 173.

[0175] In step S194, the MPD generation unit 173 generates an MPD file including the URL of the audio file, etc. The MPD generation unit 173 supplies the generated MPD file and audio file to the server upload processing unit 174.

[0176] In step S195, the server upload processing unit 174 uploads the audio file and MPD file supplied by the MPD generation unit 173 to the network server 142. Then, the processing is terminated.

[0177] (Example of functional configuration for a motion image reproduction terminal)

[0178] Figure 18 This is a block diagram illustrating the implementation that makes... Figure 8 Example of the configuration of the streaming playback unit of the motion picture playback terminal 144 running control software 161, motion picture playback software 162 and access software 163.

[0179] Figure 18 The streaming reproduction unit 190 is configured with an MPD acquisition unit 91, an MPD processing unit 191, an audio file acquisition unit 192, an audio decoding processing unit 194, and an audio synthesis processing unit 195.

[0180] The MPD acquisition unit 91 of the streaming reproduction unit 190 acquires the MPD file from the network server 142 and supplies the MPD file to the MPD processing unit 191.

[0181] MPD processing unit 191 extracts the URL information of the audio file of the segment to be reproduced described in the "Segment" for audio files from the MPD file supplied by MPD acquisition unit 91, and supplies the information to audio file acquisition unit 192.

[0182] The audio file acquisition unit 192 requests the web server 142 and acquires the audio stream of the track to be reproduced from the audio file identified using the URL provided by the MPD processing unit 191. The audio file acquisition unit 192 then supplies the acquired audio stream to the audio decoding processing unit 194.

[0183] The audio decoding processing unit 194 decodes the audio stream supplied by the audio file acquisition unit 192. The audio decoding processing unit 194 supplies the audio data obtained as the decoding result to the audio synthesis processing unit 195. The audio synthesis processing unit 195 synthesizes the audio data supplied by the audio decoding processing unit 194 as needed and outputs the audio data.

[0184] As described above, the audio file acquisition unit 192, the audio decoding processing unit 194, and the audio synthesis processing unit 195 serve as a reproduction unit, and acquire and reproduce the audio stream of the track to be reproduced from the audio file stored in the network server 142.

[0185] (Description of the processing of the motion picture reconstruction terminal)

[0186] Figure 19 A flowchart, used to describe Figure 18 The reproduction process of the stream reproduction unit 190.

[0187] exist Figure 19 In step S211, the MPD acquisition unit 91 of the streaming reproduction unit 190 acquires the MPD file from the network server 142 and supplies the MPD file to the MPD processing unit 191.

[0188] In step S212, the MPD processing unit 191 extracts the URL information of the audio file of the segment to be reproduced described in the “Segment” for audio files from the MPD file supplied by the MPD acquisition unit 91, and supplies the information to the audio file acquisition unit 192.

[0189] In step S213, the audio file acquisition unit 192 requests the network server 142 and acquires the audio stream of the track to be reproduced in the audio file identified by the URL based on the URL supplied from the MPD processing unit 191. The audio file acquisition unit 192 supplies the acquired audio stream to the audio decoding processing unit 194.

[0190] In step S214, the audio decoding processing unit 194 decodes the audio stream supplied by the audio file acquisition unit 192. The audio decoding processing unit 194 supplies the audio data obtained as the decoding result to the audio synthesis processing unit 195. In step S215, the audio synthesis processing unit 195 synthesizes the audio data supplied by the audio decoding processing unit 194 as needed and outputs the audio data.

[0191] (A summary of the second example of an audio file track)

[0192] It should be noted that in the above description, GroupDefinition and SwitchGroupDefinition are placed in the sample entries. However, as... Figure 20 As shown, GroupDefinition and SwitchGroupDefinition can be placed in a sample group entry, which is a sample entry for each group of subsamples in the track.

[0193] In this case, such as Figure 21 As shown, the sample group entries for tracks that form a switch group include both GroupDefinition and SwitchGroupDefinition. Although the illustration is omitted, the sample group entries for tracks that do not form a switch group only include GroupDefinition.

[0194] In addition, the sample entries for the orbits of each group become in Figure 22 One as shown. That is, as... Figure 22 As shown, in the sample entries for each group of tracks, there is an MHA group audio configuration box that describes the configuration information of the audio stream of the corresponding track, such as the configuration information of the profile (MPEGHAudioProfile) and the level (MPEGHAudioProfile).

[0195] (Summary of the third example of an audio file track)

[0196] Figure 23 A diagram illustrating a third example of a track used to describe an audio file.

[0197] Figure 23 The configuration of the audio data tracks and Figure 9 The configuration differs in that the audio streams of one or more groups of 3D audio are included in the base track, and the number of groups corresponding to the audio streams divided into tracks that do not include information related to the 3D audio as a whole (hereinafter referred to as group tracks) is 1 or more.

[0198] Right now, Figure 23The basic track sample entry is a 4cc sample entry for "mha2", which includes the basic track syntax when the audio stream of audio data in 3D audio is divided into multiple tracks and arranged, similar to... Figure 9 ( Figure 10 ).

[0199] Additionally, the sample entry for group tracks is a 4cc sample entry for "mhg1", which includes the syntax for group tracks when the audio stream of audio data in 3D audio is divided into multiple tracks and arranged, similar to... Figure 9 ( Figure 11 Therefore, basic orbitals and group orbitals are identified using 4cc of sample entries, and dependencies between orbitals can be discerned.

[0200] Additionally, similar to Figure 9 Track References are arranged in the track boxes of each track. Therefore, even when the sample entries or groups of tracks for the basic 4cc tracks "mha2" and "mhg1" are unknown, the dependencies between tracks can be identified using track references.

[0201] It should be noted that the mhgC and mhsC boxes may not be described in the sample entries for group tracks. Additionally, if the mhaC box, which includes configuration information for all groups of 3D audio, is described in the sample entries for the basic track, it may not be described in the sample entries for group tracks. However, if the mhaC box, which includes configuration information for independently reproducing the basic track, is described in the sample entries for the basic track, then the mhaC box, which includes configuration information for independently reproducing the group track, is also described in the sample entries for the group tracks. The presence / absence of configuration information in the sample entries can identify whether the previous or subsequent state is being observed. However, this can also be done by describing a flag in the sample entry or by changing the type of the sample entry. Note that although the illustration is omitted, when the previous and subsequent states are distinguishable by changing the type of the sample entry, the sample entry for the 4cc basic track is "mha2" in the previous state and "mha4" in the subsequent state.

[0202] (Example of the second description of an MPD file)

[0203] Figure 24 For illustration, it shows the configuration of the tracks in the audio file as follows. Figure 23 Here is an example description of the MPD file under the given configuration.

[0204] Figure 24 MPD files and Figure 15The difference between the MPD file and the MPD file is that it describes the "SubRepresentation" of the basic track.

[0205] The "SubRepresentation" of the base track describes the base track's "codec," "hierarchy," "dependency hierarchy," and...<EssentialProperty schemeIdUri="urn:mpeg:DASH:3daudio:2014"value="dataType,definition"> This is similar to the "SubRepresentation" of a group of orbitals.

[0206] exist Figure 24 In the example, the "codec" for the base track is "mha2.2.1", and the "level" is a value of "0" indicating the level of the base track. The "dependency level" is a value of "1" and "2" indicating the levels of the group tracks. Additionally, the "data type" is a number of "3" indicating the audio scene information of the kind described in the mhas box as a sample entry of the base track, and the "definition" is the binary metadata of the audio scene encoded by the base64 method.

[0207] Note that this is for reference only. Figure 25 In the "SubRepresentation" section of the basic track, audio scene information can be divided and described.

[0208] exist Figure 25 In the example, "1" is set as a number, indicating "Atmo" as a category, indicating "Atmo" of the content of the group with group ID "1", and the audio scene information described in the mhas box of the sample entry of the basic audio. Figure 7 ).

[0209] Additionally, "2" through "7" are set as numbers, which respectively indicate, as categories, the content of the group with group ID "2" ("Dialogue EN"), the content of the group with group ID "3" ("Dialogue FR"), the content of the group with group ID "4" ("Voiceover GE"), the content of the group with group ID "5" ("Effect"), the content of the group with group ID "6" ("Effect"), and the content of the group with group ID "7" ("Effect").

[0210] Therefore, in Figure 25The "SubRepresentation" of the basic track describes a structure where the "Data Type" is 1 and the "Definition" is "Atmo".<EssentialProperty schemeIdUri=“urn:mpeg:DASH:3daudio:2014”value=“dataType,definition”> Similarly, it describes `<"urn:mpeg:DASH:3daudio:2014"value="dataType,definition">`, where the "data type" is "2", "3", "4", "5", "6", and "7", while the values ​​are defined as "Dialogue EN", "Dialogue FR", "Voiceover GE", "Effect", "Effect", and "Effect". Figure 25 In the examples, the case where the audio scene information of the basic track is segmented and described has already been described. However, the group definition of group tracks and the switch group definition can be segmented and described similarly.

[0211] (Summary of the fourth example of an audio file track)

[0212] Figure 26 This is a schematic diagram of a fourth example used to describe the tracks in an audio file.

[0213] Figure 26 The configuration of the orbital data and the orbital configuration Figure 26 The configuration differs in that the sample entry for the group track is a "mha2" sample entry with 4cc.

[0214] exist Figure 26 In this case, the 4ccs of the sample entries for both the basic orbital and the group orbital are "mha2". Therefore, the basic orbital and the group orbital cannot be identified, and the dependencies between orbitals cannot be determined using the sample entry 4cc. Therefore, the dependencies between orbitals are identified using the orbital references in the orbital boxes arranged in each orbital.

[0215] Additionally, because the 4ccs of the sample entry is “mha2”, the corresponding track as the 3D audio track can be identified when the audio stream of the audio data is segmented and arranged in multiple tracks.

[0216] It should be noted that the mhaC box of the sample entries for the basic track describes the configuration information of all groups of 3D audio or the configuration information for independently reproducing the basic track, similar to that in Figure 9 and Figure 23 The situation is as follows. Additionally, the mhas box describes audio scene information, including information related to all groups and the switch group for 3D audio.

[0217] Meanwhile, the mhas box is not placed in the sample entries of the group track. Furthermore, if the mhaC box, which includes configuration information for all groups of 3D audio, is described in the sample entry of the basic track, the mhaC box may not be described in the sample entry of the group track. However, if the mhaC box, which includes configuration information for the basic track and can be independently reproduced, is described in the sample entry of the basic track, then the mhaC box, which also includes configuration information for the basic track, is described in the sample entry of the group track. The presence / absence of configuration information in the sample entry can identify whether it is in the previous or subsequent state. However, the previous and subsequent states can be identified by describing a flag in the sample entry or by changing the type of the sample entry. Note that although the illustration is omitted, when the previous and subsequent states are distinguishable by changing the type of the sample entry, the 4cc of the sample entry of the basic track and the 4cc of the sample entry of the group track are, for example, "mha2" in the former case and "mha4" in the latter case.

[0218] (Example of the third description of an MPD file)

[0219] Figure 27 For illustration, it shows the configuration of the tracks in the audio file as follows. Figure 26 Here is an example description of the MPD file under the given configuration.

[0220] Figure 27 MPD files and Figure 24 The MPD file differs in that the codec for the "SubRepresentation" of the group track is "mha2.2.1", and<EssentialProperty schemeIdUri=“urn:mpeg:DASH:3daudio:2014”value=“dataType,definition”> Not described in the "SubRepresentation" of the group track.

[0221] Note that although the illustrations are omitted, the audio scene information can be segmented and described in the "SubRepresentation" of the basic track, similar to... Figure 25 The situation.

[0222] (Summary of the fifth example of an audio file track)

[0223] Figure 28 This is a schematic diagram of a fifth example used to describe the tracks in an audio file.

[0224] Figure 28 The configuration of the audio data tracks and Figure 23The difference in configuration is that the sample entries for the basic track and group track are sample entries that include the syntax of both the group track and the basic track in the case where the audio stream of the audio data applicable to 3D audio is divided into multiple tracks.

[0225] exist Figure 28 In this case, the 4ccs of the sample entries for both the basic orbital and the group orbital are "mha3", which is the 4cc of the sample entries that include the syntax applicable to both the basic orbital and the group orbital.

[0226] Therefore, similar to Figure 26 In this case, track dependencies are identified using track references within each trackbox arranged in the track. Additionally, because the 4ccs of the sample entry are "mha2", the corresponding track can be identified as a track when the audio stream of the 3D audio data is segmented and arranged across multiple tracks.

[0227] (4cc is an example of the syntax for the sample entry of “mha3”).

[0228] Figure 29 A diagram illustrating an example of the syntax for a sample entry where 4cc is “mha3”.

[0229] like Figure 29 As shown, the syntax for the sample entry "mha3" with 4cc is through synthesis. Figure 10 grammar and Figure 11 The syntax obtained from the syntax.

[0230] That is, in the sample entry with 4cc as "mha3", the mhaC box (MHA configuration box), mhas box (MHA audio scene information box), mhgC box (MHA group definition box), mhsC box (MHA switch group definition box), etc. are arranged.

[0231] Within the mhaC box of the sample entries for the basic track, configuration information for all 3D audio groups or configuration information that can be independently reproduced from the basic track is described. Additionally, the mhas box describes information related to all groups and the 3D audio switchGroup, but does not include mhgC and mhsC boxes.

[0232] If the mhaC box, which includes configuration information for all groups of 3D audio, is described in the sample entry for the basic track, the mhaC box may not be described in the sample entry for the group track. However, if the mhaC box, which includes configuration information for the basic track that can be independently reproduced, is described in the sample entry for the basic track, then the mhaC box, which includes configuration information for the group track that can be independently reproduced, is also described in the sample entry for the group track. The presence or absence of configuration information in the sample entry can identify whether it is in the previous or subsequent state. However, the previous and subsequent states can be identified by describing a flag in the sample entry or by changing the type of the sample entry. Note that although the illustration is omitted, when the previous and subsequent states are distinguishable by changing the type of the sample entry, the 4ccs of the sample entries for the basic track and group track is "mha3" in the previous state and "mha5" in the subsequent state. Additionally, the mhas box is not placed in the sample entry for the group track. The mhgC box and mhsC box may or may not be placed.

[0233] Note that, if in Figure 30 As shown, in the sample entries of the basic track, mhas boxes, mhgC boxes, and mhsC boxes are arranged, along with mhaC boxes describing configuration information that can independently reproduce only the basic track, and mhaC boxes including configuration information for all groups of 3D audio. In this case, the mhaC boxes describing configuration information for all groups of 3D audio and those describing configuration information that can independently reproduce only the basic track are identified using the flags included in these mhaC boxes. Additionally, in this case, mhaC boxes may not be described in the sample entries of the group tracks. Whether a mhaC box is described in the sample entries of a group track can be identified based on the presence or absence of the mhaC box in the sample entries of the group track. However, whether a mhaC box is described in the sample entries of a group track can be identified by describing the flags in the sample entries or by changing the type of the sample entries. It should be noted that although the illustration is omitted, by changing the type of sample entry to identify whether the mhaC box is described in the sample entry of the group orbital, the 4ccs of the sample entries for the basic orbital and the group orbital are, for example, "mha3" when the mhaC box is described in the sample entry of the group orbital, and "mha5" when the mhaC box is not described in the sample entry of the group orbital. It should be noted that in Figure 30 The mhgC box and mhsC box may not be described in the sample entries for the group orbitals.

[0234] (Example of the fourth description in an MPD file)

[0235] Figure 31For illustration, it shows the configuration of the tracks in the audio file as follows. Figure 28 Example description of the MPD file in the case of configuration 30.

[0236] Figure 31 MPD files and Figure 24 The difference between the MPD files is that the codec for "Representation" is "mha3.3.1", while the codec for "SubRepresentation" is "mha3.2.1".

[0237] It should be noted that although the illustrations are omitted, audio scene information can be segmented and described in the "SubRepresentation" of the basic track, similar to... Figure 25 The situation.

[0238] Additionally, in the above description, track references are arranged in the track boxes of each track. However, track references may not be arranged. For example, Figures 32 to 34 For illustration, they respectively show those not in Figure 23 , Figure 26 and Figure 28 The arrangement of track references within the trackbox of an audio file's track. Figure 32 In this case, no orbital reference was provided, but the 4ccs of the sample entries for the basic orbit and the group orbits differed, thus allowing identification of dependencies between orbits. Figure 33 and Figure 34 In this case, because of the arrangement of the mhas box, it is possible to identify whether the track is a basic track.

[0239] The audio file's track configuration is as follows Figures 32 to 34 The configuration of the MPD file is respectively with Figure 24 , Figure 27 and Figure 31 The MPD file is the same. Note that in this case, audio scene information can be segmented and described in the "SubRepresentation" of the base track, similar to... Figure 25 The situation.

[0240] (Summary of the sixth example of an audio file track)

[0241] Figure 35 This is a diagram illustrating a sixth example of a track used to describe an audio file.

[0242] Figure 35 The configuration of the audio data tracks and Figure 33The difference in structure is that the samples of the basic track do not contain reference information for the samples of the grouped tracks and the configuration information necessary for decoding the reference information, including 0 or more groups of audio streams, while the sample entries of the basic track describe the reference information for the samples of the grouped tracks.

[0243] More specifically, the description of the group segmented mhmt boxes described in the audio scene information is arranged in a new way in the sample entry of 4cc as "mha2", which includes the syntax for the basic track when the audio stream of the 3D audio data is segmented into multiple tracks.

[0244] (4cc is another example of the syntax for the sample entry of “mha2”).

[0245] Figure 36 To show that 4cc is "mha2" Figure 35 A diagram illustrating the syntax of sample entries for basic and group tracks.

[0246] Figure 36 The configuration of the 4ccs for the sample entries of "mha2" and Figure 10 The configuration differs in that it involves arranging MHA MultiTrack Description (MHAMultiTrackDescription) boxes (mhmt boxes).

[0247] Within the MHMT box, for reference, the corresponding information between the group ID (group_ID) and track ID (track_ID) is described. It should be noted that audio elements and track IDs can be described in relation to each other within the MHMT box.

[0248] If the reference information remains unchanged in each sample, the reference information can be effectively described by arranging mhmt boxes in the sample entries.

[0249] Note that although the illustration is omitted, in Figure 9 , Figure 20 , Figure 23 , Figure 26 , Figure 28 , Figure 30 , Figure 32 and Figure 34 In this case, the MHMT boxes can be similarly arranged in the sample entries of the subsequent orbits, rather than as reference information for the samples of the orbits describing the group, similar to the samples of the basic orbits.

[0250] In this case, the syntax for the sample entry where 4cc is "mha3" becomes... Figure 37 One of them is shown. That is, Figure 36 The configuration of the 4ccs for the sample entry "mha3" and Figure 29The configuration differs in that it involves arranging MHA MultiTrack Description (MHAMultiTrackDescription) boxes (mhmt boxes).

[0251] In addition, Figure 23 , Figure 26 , Figure 28 , Figure 30 , Figures 32 to 34 and Figure 35 In this context, one or more groups of audio streams in 3D audio may not be included in the base track, similar to... Figure 9 Additionally, the number of groups corresponding to the audio streams that are divided into groups of tracks can be 1.

[0252] In addition, Figure 23 , Figure 26 , Figure 28 , Figure 30 , Figures 32 to 34 and Figure 35 In this context, group definitions and switch group definitions can be placed within the same group entry, similar to... Figure 20 The situation.

[0253] <Second Embodiment>

[0254] (Trajectory overview)

[0255] Figure 38 This is a diagram illustrating the outline of the track in the second embodiment to which this disclosure is applied.

[0256] like Figure 38 As shown, the second embodiment differs from the first embodiment in that the track records are in different files (3da_base.mp4 / 3da_group1.mp4 / 3da_group2.mp4 / 3da_group3.mp4 / 3da_group4.mp4). In this case, by obtaining the file of the desired track via HTTP, only the data of the desired track can be obtained. Therefore, the data of the desired track via HTTP can be obtained efficiently.

[0257] (Example description of an MPD file)

[0258] Figure 39 A diagram illustrating an example description of an MPD file to which this disclosure is applied in a second embodiment.

[0259] like Figure 39As shown, the MPD file contains descriptions such as "Representation" of segments of the audio files (3da_base.mp4 / 3da_group1.mp4 / 3da_group2.mp4 / 3da_group3.mp4 / 3da_group4.mp4) that manage 3D audio.

[0260] The "Representation" includes the "Codec", "id", "Association ID", and "Association Type". "id" is the ID of the "Representation" that includes the "id". "Association ID" is information indicating the reference relationship between the corresponding track and another track, and is the "id" of the reference track. "Association Type" is a code indicating the meaning of the reference relationship (dependency) with the reference track, and uses, for example, the same value as the track reference in MP4.

[0261] Additionally, the group's track "Representation" includes<EssentialProperty schemeIdUri=“urn:mpeg:DASH:3daudio:2014”value=“dataType,def inition”>.exist Figure 39 In the example, the "Representation" managing the segments of an audio file is provided under an "AdaptationSet". However, an "AdaptationSet" can be provided for each segment of the audio file, and a "Representation" managing the segment can be provided under it. In this case, within the "AdaptationSet", the "Association ID" and the meaning indicating the reference relationship with the reference track are... <EssentialProperty schemeIdUri=“urn:mpeg:DASH:3daudioAssociationData:2014”value=“dataType,id”> It can be described, similar to "association type". Additionally, the audio scene information, group definitions, and switchGroup definitions described in the "Representation" of basic tracks and group tracks can be segmented and described, similar to... Figure 25 In addition, the audio scene information, group definitions, and switch group definitions described and segmented in "Representation" can be described in "Adaptation Set".

[0262] (Overview of the information processing system)

[0263] Figure 40 This is a diagram illustrating the outline of the information processing system to which this disclosure is applied in the third embodiment.

[0264] exist Figure 40 The same configuration shown, and Figure 8 The configurations are indicated using the same reference numerals. Overlapping descriptions are omitted appropriately.

[0265] Figure 40 The information processing system 210 is configured such that a network server 212 connected to the file generation device 211 is connected to the motion picture reproduction terminal 214 via the Internet 13.

[0266] In the information processing system 210, the network server 142 distributes the audio streams of the audio files in the group to be reproduced to the motion picture reproduction terminal 144 using the MPEG-DASH method.

[0267] Specifically, file generation device 211 encodes the audio data and metadata of 3D audio of moving image content at multiple encoding speeds to generate an audio stream. File generation device 211 segments the audio stream for each group and each type of Ext element, thus placing the audio stream on different tracks. File generation device 211 creates files for the audio stream at each encoding speed for each segment and each track to generate an audio file. File generation device 211 uploads the resulting audio file to network server 212. Additionally, file generation device 211 generates an MPD file and uploads it to network server 212.

[0268] The network server 212 stores audio files for each segment and for each track at each encoding speed, as well as MPD files uploaded from the file generation device 211. In response to a request from the motion picture reproduction terminal 214, the network server 212 sends the stored audio files, stored MPD files, etc., to the motion picture reproduction terminal 214.

[0269] The motion picture reproduction terminal 214 executes control software 221, motion picture reproduction software 162, access software 223, etc.

[0270] Control software 221 is software that controls the data flowing from network server 212. Specifically, control software 221 causes motion picture reproduction terminal 214 to obtain MPD files from network server 212.

[0271] In addition, based on the MPD file, the control software 221 commands the access software 223 to transmit the sending request for the group to be reproduced specified by the motion picture reproduction software 162, as well as the audio stream of the audio file of the Ext element type corresponding to that group.

[0272] Access software 223 is software that controls communication between motion picture reproduction terminal 214 and network server 212 via the Internet 13 using HTTP. Specifically, in response to commands from control software 221, access software 223 causes motion picture reproduction terminal 144 to send a request to transmit the audio stream of the audio file to be reproduced. Furthermore, in response to the transmission request, access software 223 causes motion picture reproduction terminal 144 to begin receiving the audio stream transmitted from network server 212 and provides a notification to motion picture reproduction software 162 that reception has begun.

[0273] (Configuration example of a file generation device)

[0274] Figure 41 To show Figure 40 A block diagram illustrating a configuration example of the file generation device 211.

[0275] exist Figure 41 The same configuration shown, and Figure 16 The configurations are indicated using the same reference numerals. Overlapping descriptions are omitted appropriately.

[0276] Figure 41 The configuration of the file generation device 211 and Figure 16 The difference between the file generation device 141 and the file generation device 141 is that the audio file generation unit 241 and the MPD generation unit 242 are provided to replace the audio file generation unit 172 and the MPD generation unit 173.

[0277] Specifically, the audio file generation unit 241 of the audio file generation device 211 assigns tracks to the audio stream supplied from the audio encoding processing unit 171 for each group and each type of Ext element. The audio file generation unit 241 generates an audio file in which the audio stream is arranged at each encoding speed for each segment and for each track. The audio file generation unit 241 supplies the generated audio file to the MPD generation unit 242.

[0278] MPD generation unit 242 determines the URL, etc., of the web server 142 where the audio file supplied from audio file generation unit 172 is to be stored. MPD generation unit 242 generates an MPD file in which the URL, etc., of the audio file is arranged in a "segment" of "Representation" for the audio file. MPD generation unit 173 supplies the generated MPD file and the generated audio file to server upload processing unit 174.

[0279] (Description of the processing of the document generation device)

[0280] Figure 42 A flowchart, used to describe Figure 41 File generation processing of file generation device 211.

[0281] Figure 42 The processes in steps S301 and S302 are similar to Figure 17 The processing of steps S191 and S192 is therefore omitted from description.

[0282] In step S303, the audio file generation unit 241 generates an audio file, in which the audio stream is arranged at each encoding speed for each segment and for each track. The audio file generation unit 241 supplies the generated audio file to the MPD generation unit 242.

[0283] The processing in steps S304 and S305 is similar to Figure 17 The processing of steps S194 and S195 is therefore omitted from description.

[0284] (Example of functional configuration for a motion image reproduction terminal)

[0285] Figure 43 This is a block diagram illustrating the implementation that makes... Figure 40 Example of the configuration of the streaming playback unit of the motion picture playback terminal 214 executing control software 221, motion picture playback software 162 and access software 223.

[0286] exist Figure 43 The same configuration shown, and Figure 18 The configurations are indicated using the same reference numerals. Overlapping descriptions are omitted appropriately.

[0287] Figure 43 The configuration of the stream reproduction unit 260 and Figure 18 The difference in the configuration of the streaming reproduction unit 190 is that an audio file acquisition unit 264 is provided to replace the audio file acquisition unit 192.

[0288] The audio file acquisition unit 264 requests the network server 142 to obtain the audio stream of the audio file of the track to be reproduced based on the URL supplied by the MPD processing unit 191. The audio file acquisition unit 264 then supplies the acquired audio stream to the audio decoding processing unit 194.

[0289] That is, the audio file acquisition unit 264, the audio decoding processing unit 194 and the audio synthesis processing unit 195 are used as a reproduction unit, and the audio stream of the audio file of the track to be reproduced is obtained from the audio file stored in the network server 212, and the audio stream is reproduced.

[0290] (Description of the processing of the motion picture reconstruction terminal)

[0291] Figure 44 A flowchart, used to describe Figure 43The reproduction processing of the stream reproduction unit 260.

[0292] Figure 44 The processing of steps S321 and S322 is similar to Figure 19 The processing of steps S221 and S212 is therefore omitted from description.

[0293] In step S323, based on the URL of the audio file of the track to be reproduced, the audio file acquisition unit 192 requests the network server 142 to acquire the audio stream of the audio file provided by the MPD processing unit 191. The audio file acquisition unit 264 then supplies the acquired audio stream to the audio decoding processing unit 194.

[0294] The processing in steps S324 and S325 is similar to Figure 19 The processing of steps S214 and S215 is therefore omitted from description.

[0295] It should be noted that, in the second embodiment, similar to the first embodiment, group definitions and switch group definitions can be arranged in the sample group entries.

[0296] In addition, in the second embodiment, similar to the first embodiment, the audio data track configuration can also be... Figure 23 , Figure 26 , Figure 28 , Figure 30 , Figures 32 to 34 and Figure 35 The configuration shown.

[0297] Figures 45 to 47 For illustration, they respectively show the configuration of the audio data tracks in the second embodiment as follows: Figure 23 , Figure 26 and Figure 28 The MPD configuration shown is shown in the example. In the second embodiment, the audio data track is configured as follows: Figure 32 , Figure 33 , Figure 34 or Figure 35 The MPD file shown in the configuration is the same as in Figure 23 , Figure 26 and Figure 28 MPD in the configuration shown.

[0298] Figure 45 MPD and Figure 39The difference between MPD and other methods lies in the "encoder / decoder" and "associationId" of the basic track, and in...<EssentialProperty schemeIdUri=“urn:mpeg:DASH:3daudio:2014”value=“dataType,definition”> Included in the "Representation" of the basic track. Specifically, Figure 45 The "Representation" of the basic track of MPD has the "codec" as "mha2.2.1", while the "Association ID" is the "id" of the group track "g1" and "g2".

[0299] in addition, Figure 46 MPD and Figure 45 The difference between MPD and other methods lies in the "encoder / decoder" of the group tracks, and in...<EssentialProperty schemeIdUri=“urn:mpeg:DASH:3daudio:2014”value=“dataType,definition”> It is not included in the "Representation" section of the group track. Specifically, Figure 46 The MPD group track's "codec" is "mha2.2.1".

[0300] in addition, Figure 47 MPD and Figure 45 The difference between MPD and MPD lies in the "codec" for the basic track and the group track. Specifically, Figure 47 The MPD group track's "codec" is "mha3.2.1".

[0301] It should be noted that, Figures 45 to 47 In MPD, the "AdaptationSet" can be segmented based on the "Representation", such as... Figures 48 to 50 As shown.

[0302] <Another example of a basic track>

[0303] In the above description, only one basic track is provided. However, multiple basic tracks can be provided. In this case, the basic track is provided for each viewpoint, for example, 3D audio (details will be given below), and within the basic track, mhaC boxes containing configuration information for all groups of 3D audio for the viewpoint are arranged. Note that mhas boxes containing audio scene information for the viewpoint can also be arranged within the basic track.

[0304] The viewpoint of 3D audio is the location where 3D audio can be heard, such as the viewpoint of an image reproduced simultaneously with the 3D audio or a pre-set predetermined location.

[0305] As described above, when the basic track is segmented for each viewpoint, different audio for each viewpoint can be reproduced from the same 3D audio stream based on the position of objects on the screen, etc., included in the configuration information for each viewpoint. As a result, the amount of data in the 3D audio stream can be reduced.

[0306] That is, in cases where the viewpoint of the 3D audio is multiple viewpoints of an image of a baseball field that can be reproduced simultaneously with the 3D audio, the image with the viewpoint in the center rear screen is prepared as the main image of the image with the basic viewpoint. In addition, images with viewpoints in seats located behind the board, infield seats at first base, infield seats at third base, left field seats, right field seats, etc., are prepared as multiple images, which are images with viewpoints (which are not the basic viewpoints).

[0307] In this scenario, preparing 3D audio for all viewpoints would result in a large data volume for the 3D audio. Therefore, by describing the positions of objects on the screen within the viewpoint using basic tracks, audio streams such as Object audio and SAOC Object audio, which change based on the position of objects on the screen, can be shared across viewpoints. As a result, the data volume of the 3D audio stream can be reduced.

[0308] In the reproduction of 3D audio, for example, audio streams such as Object audio and SAOCObject audio are used, along with a base track corresponding to the main image viewpoint or multiple images reproduced using the audio stream at the same time, with different audio reproduced according to the viewpoint.

[0309] Similarly, for example, in the case where the viewpoint of the 3D audio is the pre-defined location of multiple seats in a stadium, the amount of 3D audio data becomes large if 3D audio is prepared for all viewpoints. Therefore, by describing the location of objects on the screen using a basic track, audio streams such as Object audio and SAOCObject audio can be shared between viewpoints. Thus, different audio can be reproduced based on the seat selected by the user using a seating chart, using Object audio and SAOCObject audio from a single viewpoint, and the amount of data in the 3D audio stream can be reduced.

[0310] The basic orbit is provided for use in Figure 28 In the case of 3D audio in the track structure for each viewpoint, the track structure becomes as follows: Figure 51 One shown. In Figure 51In the example shown, the number of viewpoints for the 3D audio is three. Additionally, in... Figure 51 In the example shown, channel audio is generated for each viewpoint of the 3D audio, and other audio data is shared across all viewpoints of the 3D audio. This also applies to the following... Figure 52 Examples.

[0311] In this case, three basic tracks are provided for each viewpoint of the 3D audio, such as Figure 3 As shown. Track references are arranged in the trackbox of each of the basic tracks. Additionally, the syntax of the sample entries for each basic track is the same as that of the sample entries for "mha3" with 4cc. 4cc is "mhcf," indicating that the basic track is provided for each viewpoint in the 3D audio.

[0312] The mhaC boxes, which include configuration information for all groups of 3D audio for each viewpoint, are arranged in sample entries of each base track. For example, within a viewpoint, the configuration information for all groups of 3D audio for each viewpoint is the position of objects on the screen. Additionally, mhas boxes, which include audio scene information for each viewpoint, are arranged in each base track.

[0313] The audio stream of the viewpoint's channel audio group is arranged in the sample of the basic track.

[0314] It should be noted that, in sample units, where there is ObjectMetadata describing the position of an object on the screen in each viewpoint, Object Metadata is also arranged in the sample of each basic track.

[0315] That is, when the object is a moving body (e.g., an athlete), the position of the object on the screen changes over time in each viewpoint. Therefore, this position is described as Object Metadata in the sample cell. In this case, for each viewpoint, the Object Metadata in the sample cell is arranged in a sample corresponding to the basic trajectory of the viewpoint.

[0316] Figure 51 The configuration of the group of orbits and Figure 28 The configuration is the same, except for the audio stream of the group without channel audio, so the description is omitted.

[0317] It should be noted that, Figure 51 In this track structure, the audio streams of the viewpoint's channel audio groups may not be arranged in the basic track, but rather in different group tracks. In this case, the track structure becomes... Figure 52 One of them is shown in the image.

[0318] exist Figure 52In the example shown, the audio stream of the channel audio group corresponding to the viewpoint of the basic track with track ID "1" is arranged in the group track with track ID "4". Additionally, the audio stream of the channel audio group corresponding to the viewpoint of the basic track with track ID "2" is arranged in the group track with track ID "5".

[0319] Additionally, the audio stream of the channel audio group corresponding to the viewpoint of the basic track with track ID "3" is arranged in the group track with track ID "6".

[0320] It should be noted that, Figure 51 and Figure 52 In the example, the 4cc of the sample entry for the basic orbital is "mhcf". However, 4cc can be... Figure 28 The same "mha3".

[0321] Additionally, although the illustrations are omitted, the basic orbitals are provided for use in all the orbital structures described above (except for...). Figure 28 The situation for each viewpoint of 3D audio in the orbital structure (outside of the track structure) is similar to that in... Figure 51 and 52 The situation.

[0322] <Third Embodiment>

[0323] (Description of the computer to which this disclosure applies)

[0324] The series of processes performed by the network server 142 (212) can be executed by hardware or by software. In the case where the series of processes are performed by software, a program configuring the software is installed on the computer. Here, the computer includes computers with specialized hardware and general-purpose personal computers that can perform various functions by installing various types of programs.

[0325] Figure 53 This is a block diagram illustrating an example configuration of computer hardware that utilizes a program to perform a series of processes on network server 142 (212).

[0326] In a computer, the central processing unit (CPU) 601, read-only memory (ROM) 602, and random access memory (RAM) 603 are interconnected via a bus 604.

[0327] The input / output interface 605 is also connected to the bus 604. The input unit 606, output unit 607, storage unit 608, communication unit 609, and driver 610 are connected to the input / output interface 605.

[0328] Input unit 606 consists of a keyboard, mouse, microphone, etc. Output unit 607 consists of a display, speaker, etc. Storage unit 608 consists of a hard disk, non-volatile memory, etc. Communication unit 609 consists of a network interface, etc. Driver 610 drives removable media 611 such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory.

[0329] In the computer configured as described above, the CPU 601 loads the program stored in the storage unit 608 into the RAM 603 via the input / output interface 605 and the bus 604, and executes the program to perform a series of processes.

[0330] The program executed by the computer (CPU 601) can be provided, for example, by being recorded in a removable medium 611 as a packaging medium. Alternatively, the program can be provided via wired or wireless transmission media such as a local area network, the Internet, or digital satellite broadcasting.

[0331] In a computer, a program can be installed into the storage unit 608 via the input / output interface 605 by attaching the removable medium 611 to the drive 610. Alternatively, the program can be received by the communication unit 609 via a wired or wireless transmission medium and installed into the storage unit 608. Additionally, the program can be pre-installed into the ROM 602 or the storage unit 608.

[0332] It should be noted that a program executed by a computer may be a program that is processed in a time sequence according to the order described in this specification, or it may be a program that is processed in parallel, such as when it is called, or a program that is processed at necessary time intervals.

[0333] In addition, the hardware configuration of the motion picture reproduction terminal 144 (214) can have the same as Figure 53 A computer with a similar configuration. In this case, for example, CPU 601 executes control software 161 (221), motion picture reproduction software 162, and access software 163 (223). The processing of the motion picture reproduction terminal 144 (214) can be performed by hardware.

[0334] In this specification, "system" refers to a collection of multiple configuration elements (devices, modules (components), etc.), and all configuration elements may or may not be housed in the same enclosure. Therefore, both multiple devices housed in separate enclosures and connected via a network, and a single device housing multiple modules in a single enclosure, are systems.

[0335] Note that the embodiments disclosed herein are not limited to the embodiments described above, and various changes may be made without departing from the spirit and scope of this disclosure.

[0336] Furthermore, this disclosure can be applied to information processing systems that perform broadcast or local storage playback rather than streaming playback.

[0337] In MPD implementations, information is described by basic properties defined by descriptors that can be ignored when the content described by the pattern cannot be understood. However, information can also be described by supplemental properties defined by descriptors that can be reproduced even if the content described by the pattern cannot be understood. This description method is chosen by the side that creates the intended content.

[0338] In addition, this disclosure can be configured as follows.

[0339] (1) An information processing device, comprising:

[0340] A file generation unit is configured to generate a file in which multiple types of audio data are divided into tracks and arranged according to one or more of the types, and information related to the multiple types is arranged.

[0341] (2) The information processing apparatus according to (1), wherein

[0342] Information related to the multiple categories is arranged in sample entries on a predetermined track.

[0343] (3) The information processing apparatus according to (2), wherein

[0344] The predetermined track is one of the tracks in which the various types of audio data are divided and arranged.

[0345] (4) The information processing apparatus according to any one of (1) to (3), wherein,

[0346] For each of the orbitals, information related to the type of orbital is arranged in a file.

[0347] (5) The information processing apparatus according to (4), wherein,

[0348] For each of the tracks, information related to the exclusive reproduction type is arranged in the file, which consists of the type corresponding to the track and the type of audio data that is exclusively reproduced from the audio data of the type corresponding to the track.

[0349] (6) The information processing apparatus according to (5), wherein

[0350] Information related to the type corresponding to the orbit and information related to the exclusive reproduction type are arranged in the sample entries for the corresponding orbit.

[0351] (7) The information processing apparatus according to (5) or (6), wherein

[0352] The file generation unit generates a management file, which manages information including information related to the exclusive reproduction type for the presence or absence of each track.

[0353] (8) The information processing apparatus according to any one of (1) to (7), wherein

[0354] Reference information corresponding to the various types of orbits is arranged in the document.

[0355] (9) The information processing apparatus according to (8), wherein

[0356] The reference information is arranged in a sample on a predetermined track.

[0357] (10) The information processing apparatus according to (9), wherein

[0358] The predetermined track is one of the tracks in which the various types of audio data are divided and arranged.

[0359] (11) The information processing apparatus according to any one of (1) to (10), wherein

[0360] Information indicating the reference relationships between the orbits is arranged in the document.

[0361] (12) The information processing apparatus according to any one of (1) to (11), wherein

[0362] The file generation unit generates a management file, which manages the file including information indicating the reference relationships between the tracks.

[0363] (13) The information processing apparatus according to any one of (1) to (12), wherein

[0364] The file in question is a single file.

[0365] (14) The information processing apparatus according to any one of (1) to (12), wherein

[0366] The file is the file for each of the tracks.

[0367] (15) An information processing method, comprising the following steps:

[0368] The information processing device generates a file in which multiple types of audio data are segmented into tracks and arranged for each or more of the types, and information related to the multiple types is arranged.

[0369] (16) An information processing apparatus, comprising:

[0370] The reproduction unit is configured to reproduce audio data of a predetermined track from a file, wherein multiple types of audio data in the file are segmented into tracks and arranged for each or more of the types, and information related to the multiple types is arranged.

[0371] (17) An information processing method, comprising the following steps:

[0372] Audio data of a predetermined track is reproduced from a file by an information processing device, wherein multiple types of audio data in the file are segmented into tracks and arranged for each or more of the types, and information related to the multiple types is arranged.

[0373] Reference tag list

[0374] 11 document generation devices

[0375] 192 Audio File Acquisition Unit

[0376] 194 audio decoding processing units

[0377] 195 audio synthesis processing units

[0378] 211 Document Generation Device

[0379] 264 Audio File Acquisition Unit

Claims

1. An information processing apparatus, comprising: The playback unit is configured to reproduce audio data of predetermined tracks from a file, wherein the audio data of multiple groups in the file is segmented into tracks and arranged for each or more of the multiple groups, and information associated with the multiple groups is arranged, wherein each group is associated with ID information, and the group corresponding to the ID information includes a set of audio elements of the same type. The audio element corresponds to audio of a specific language in at least one audio stream or at least one of multiple channels. The multiple tracks include a basic track, each associated with the ID information, and multiple group tracks. The basic orbit includes information associated with each of the plurality of group orbits, and The basic orbit is referenced by each of the plurality of orbit groups.

2. The information processing apparatus according to claim 1, wherein... Information related to the multiple groups is arranged in sample entries on a predetermined track.

3. The information processing apparatus according to claim 2, wherein... The predetermined track is one of the tracks in which the audio data of the plurality of groups is divided and arranged.

4. The information processing apparatus according to claim 1, wherein... For each of the tracks, information related to the group corresponding to the track is arranged in the file.

5. The information processing apparatus according to claim 4, wherein For each track, information related to an exclusive reproduction group is arranged in the file, the exclusive reproduction group consisting of a group corresponding to the track and a group of audio data that is exclusively reproduced from the audio data of the group corresponding to the track.

6. The information processing apparatus according to claim 5, wherein Information related to the group corresponding to the track and information related to the exclusive reproduction group are arranged in the sample entries of the corresponding track.

7. The information processing apparatus according to claim 1, wherein... Reference information for the orbits corresponding to the plurality of groups is arranged in the file.

8. The information processing apparatus according to claim 7, wherein The reference information is arranged in a sample on a predetermined track.

9. The information processing apparatus according to claim 8, wherein The predetermined track is one of the tracks in which the audio data of the plurality of groups is segmented and arranged.

10. The information processing apparatus according to claim 1, wherein The file in question is a single file.

11. The information processing apparatus according to claim 1, wherein The file is the file for each track in the track.

12. An information processing method, comprising the following steps: The information processing device reproduces audio data of a predetermined track from a file, wherein the audio data of multiple groups in the file is divided into tracks and arranged according to one or more of the multiple groups, and information related to the multiple groups is arranged in the file, wherein... Each group is associated with ID information, and the group corresponding to the ID information includes a set of audio elements of the same type. The audio element corresponds to audio of a specific language in at least one audio stream or at least one of multiple channels. The multiple tracks include a basic track, each associated with the ID information, and multiple group tracks. The basic orbit includes information associated with each of the plurality of group orbits, and The basic orbit is referenced by each of the plurality of orbit groups.

Citation Information

Patent Citations

  • Method and apparatus for track and track subset grouping

    CN102132562A

  • Method for creating and accessing a menu for audio content without using a display

    CN1735941A