Information processing apparatus and information processing method

The information processing apparatus and method organize audio data into tracks based on relevance, enabling easy reproduction of specific audio data groups by reducing decoding load and optimizing stream selection.

JP7708229B2Active Publication Date: 2025-07-15SONY GROUP CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024005069
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2015-06-22
Filing Date
2024-01-17
Publication Date
2025-07-15
Estimated Expiration
2035-06-30

AI Technical Summary

Technical Problem

Existing technologies do not facilitate easy reproduction of specific audio data groups among multiple audio data groups in streaming services.

Method used

An information processing apparatus and method that generates a file structure where audio data is stored in individual tracks based on relevance, with reference information linking tracks, allowing easy reproduction of desired audio data groups.

Benefits of technology

Enables efficient and easy reproduction of predetermined audio data types by reducing decoding load and optimizing audio stream selection based on playback environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007708229000001
    Figure 0007708229000001
  • Figure 0007708229000002
    Figure 0007708229000002
  • Figure 0007708229000003
    Figure 0007708229000003
Patent Text Reader

Abstract

To easily play back a prescribed type of audio data from among multiple types of audio data.SOLUTION: A file creation device creates audio files in which an audio stream of multiple groups is arranged as being divided into tracks for one or more groups, and information pertaining to multiple groups is arranged. The present invention can be applied to an information-processing system or the like configured from, e.g., a file creation device for creating files, a Web server for recording files created by the file creation device, and a video playback terminal for playing back the files.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing apparatus and an information processing method, and more particularly to an information processing apparatus and an information processing method that enable easy reproduction of predetermined types of audio data among a plurality of types of audio data.

Background Art

[0002] In recent years, the mainstream of streaming services on the Internet has become OTT-V (Over The Top Video). As a basic technology, MPEG-DASH (Moving Picture Experts Group phase - Dynamic Adaptive Streaming over HTTP) has begun to spread (see, for example, Non-Patent Document 1).

[0003] In MPEG-DASH, a delivery server prepares a group of video data with different screen sizes and encoding speeds for one video content, and a playback terminal requests a group of video data with an optimal screen size and encoding speed according to the status of the transmission path, thereby realizing adaptive streaming delivery.

Prior Art Documents

Non-Patent Documents

[0004]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, it has not been considered to easily reproduce the audio data of a predetermined group among the audio data of a plurality of groups.

[0006] The present disclosure has been made in view of such a situation, and enables easy reproduction of the audio data of a desired group among the audio data of a plurality of groups.

Means for Solving the Problem

[0007] The information processing apparatus according to the first aspect of the present disclosure includes a file generation unit that generates a file in which a plurality of types of audio data are stored in individual tracks and information regarding the plurality of types is arranged, and the audio data forms a group based on relevance. Moreover, multiple types of audio data are divided and arranged in the track for each of one or more of the types, reference information to the tracks corresponding to the multiple types is arranged in the file, the reference information is arranged in samples of a predetermined track, and the predetermined track is one of the tracks in which the multiple types of audio data are divided and arranged. It is an information processing apparatus.

[0008] The information processing method according to the first aspect of the present disclosure corresponds to the information processing apparatus according to the first aspect of the present disclosure.

[0009] In the first aspect of the present disclosure, a plurality of types of audio data are stored in individual tracks, and a file in which information regarding the plurality of types is arranged is generated, and the audio data forms a group based on relevance. Moreover, multiple types of audio data are divided and arranged in the track for each of one or more of the types, reference information to the tracks corresponding to the multiple types is arranged in the file, the reference information is arranged in samples of a predetermined track, and the predetermined track is one of the tracks in which the multiple types of audio data are divided and arranged.

[0010] The information processing apparatus according to the second aspect of the present disclosure includes a playback unit that plays back the audio data of a predetermined track from a file including a plurality of types of audio data stored in individual tracks, and the audio data forms a group based on relevance. Moreover, multiple types of audio data are divided and arranged in the track for each of one or more of the types, reference information to the tracks corresponding to the multiple types is arranged in the file, the reference information is arranged in samples of a predetermined track, and the predetermined track is one of the tracks in which the multiple types of audio data are divided and arranged. It is an information processing apparatus.

[0011] The information processing method according to the second aspect of the present disclosure corresponds to the information processing apparatus according to the second aspect of the present disclosure.

[0012] In the second aspect of the present disclosure, the audio data of a predetermined track is played back from a file including a plurality of types of audio data stored in individual tracks, and the audio data forms a group based on relevance.Moreover, multiple types of audio data are divided and arranged in the track for each of one or more of the types, reference information to the tracks corresponding to the multiple types is arranged in the file, the reference information is arranged in samples of a predetermined track, and the predetermined track is one of the tracks in which the multiple types of audio data are divided and arranged.

[0013] Note that the information processing apparatuses on the first and second sides can be realized by causing a computer to execute a program.

[0014] Also, in order to realize the information processing apparatuses on the first and second sides, the program to be executed by the computer can be transmitted via a transmission medium or recorded on a recording medium and provided.

Advantages of the Invention

[0015] According to the first aspect of the present disclosure, a file can be generated. Also, according to the first aspect of the present disclosure, a file can be generated such that predetermined types of audio data among a plurality of types of audio data can be easily reproduced.

[0016] According to the second aspect of the present disclosure, audio data can be reproduced. Also, according to the second aspect of the present disclosure, predetermined types of audio data among a plurality of types of audio data can be easily reproduced.

Brief Description of the Drawings

[0017]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Figure 35

Figure 36

Figure 37

Figure 38

Figure 39

Figure 40

Figure 41

Figure 42

Figure 43

Figure 44

Figure 45

Figure 46

Figure 47

Figure 48

Figure 49

Figure 50

Figure 51

Figure 52

Figure 53

Embodiments for Carrying Out the Invention

[0018] Hereinafter, the premise of the present disclosure and the embodiments for carrying out the present disclosure (hereinafter referred to as embodiments) will be described. The description will be made in the following order. 0. Premises of the present disclosure (Figs. 1 to 7) 1. First Embodiment (Figs. 8 to 37) 2. Second Embodiment (Figs. 38 to 50) 3. Other Examples of Base Tracks (Figs. 51 and 52) 4. Third Embodiment (Fig. 53)

[0019] <Premises of the present disclosure> (Explanation of the Structure of the MPD File) FIG. 1 is a diagram showing the structure of an MPD file (Media Presentation Description) of MPEG-DASH.

[0020] In the analysis (parsing) of the MPD file, the optimal one is selected from the attributes of "Representation" included in "Period" of the MPD file (Media Presentation in Fig. 1).

[0021] Then, the file is acquired and processed by referring to the URL (Uniform Resource Locator) etc. of the "Initialization Segment" at the head of the selected "Representation". Subsequently, the file is acquired and played back by referring to the URL etc. of the subsequent "Media Segment".

[0022] Note that the relationship among "Period", "Representation", and "Segment" in the MPD file is as shown in Fig. 2. That is, one video content can be managed in time units longer than segments by "Period", and in each "Period", it can be managed in segment units by "Segment". Also, in each "Period", the video content can be managed in units of stream attributes by "Representation".

[0023] Therefore, the MPD file has the hierarchical structure shown in FIG. 3 below the "Period". Also, when arranging the structure of this MPD file on the time axis, it becomes like the example in FIG. 4. As is clear from FIG. 4, there are multiple "Representations" for the same segment. By adaptively selecting any of these, a stream with the desired attributes of the user can be obtained and played.

[0024] (Overview of 3D audio file format) FIG. 5 is a diagram for explaining the overview of the track of the 3D audio file format of MP4.

[0025] In an MP4 file, for each track, codec information of video content and position information indicating the position within the file can be managed. In the 3D audio file format of MP4, all of the 3D audio (Channel audio / Object audio / SAOC Object audio / HOA audio / metadata) audio streams (ES (Elementary Stream)) are recorded in units of samples (frames) as one track. Also, the codec information (Profile / level / audio configuration) of the 3D audio is stored as a sample entry.

[0026] The Channel audio that constitutes the 3D audio is audio data in units of channels, and the Object audio is audio data in units of objects. Note that an object is a sound source, and the audio data in units of objects is acquired by a microphone or the like attached to the object. The object may be an object such as a fixed microphone stand, or a moving object such as a person.

[0027] Also, SAOC Object audio is the audio data of SAOC (Spatial Audio Object Coding), HOA audio is the audio data of HOA (Higher Order Ambisonics), and metadata is the metadata of Channel audio, Object audio, SAOC Object audio, and HOA audio.

[0028] (Structure of the moov box) Figure 6 is a diagram showing the structure of the moov box of an MP4 file.

[0029] As shown in Figure 6, in an MP4 file, image data and audio data are recorded as different tracks. Although the details of the audio data track are not described in Figure 6, they are the same as those of the image data track. The sample entry is included in the sample description placed in the stsd box within the moov box.

[0030] By the way, in the broadcast or local storage playback of an MP4 file, generally, the server side sends out all the audio streams of 3D audio. Then, while the client side parses all the audio streams of 3D audio, it only decodes and outputs the necessary audio streams of 3D audio. However, when the bitrate is high or there are restrictions on the reading rate of the local storage, it is desirable to reduce the load of the decoding process by only acquiring the necessary audio streams of 3D audio.

[0031] Also, in the stream playback of an MP4 file compliant with MPEG-DASH, the server side prepares audio streams with multiple encoding rates. Therefore, by the client side only acquiring the necessary audio streams of 3D audio, it can select and acquire the audio stream with the encoding rate optimal for the playback environment.

[0032] As described above, in the present disclosure, by dividing the audio stream of 3D audio into tracks according to the type and arranging them in the audio file, it is possible to efficiently obtain only the audio stream of a predetermined type of 3D audio. As a result, in broadcasting or local storage playback, the load of the decoding process can be reduced. Also, in stream playback, according to the bandwidth, it is possible to play back the highest quality audio stream among the necessary audio streams of 3D audio.

[0033] (Explanation of the hierarchical structure of 3D audio) FIG. 7 is a diagram showing the hierarchical structure of 3D audio.

[0034] As shown in FIG. 7, the audio data of 3D audio is regarded as different audio elements for each audio data. The types of audio elements are SCE (Single Channel Element) and CPE (Channel Pair Element). The type of the audio element of the audio data for one channel is SCE, and the type of the audio element corresponding to the audio data for two channels is CPE.

[0035] Audio elements form groups among the same types of audio (Channel / Object / SAOC Object / HOA). Therefore, the group types are Channels, Objects, SAOC Objects, and HOA. Two or more groups can form a switch Group or a group Preset as needed.

[0036] The switch Group is a group (exclusive playback group) in which the audio streams of the groups contained therein are played back exclusively. That is, as shown in FIG. 7, when there are a group of Object audio for English (EN) and a group of Object audio for French (FR), only one of the groups should be played back. Therefore, a switch Group is formed from the group of Object audio for English with a group ID of 2 and the group of Object audio for French with a group ID of 3. Thereby, the Object audio for English and the Object audio for French are played back exclusively.

[0037] On the other hand, the group Preset defines the combination of groups intended by the content creator.

[0038] Also, the metadata of 3D audio is different Ext elements (Ext Element) for each metadata. Types of Ext elements include Object Metadata, SAOC 3D Metadata, HOA Metadata, DRC Metadata, SpatialFrame, SaocFrame, etc. The Ext element of Object Metadata is the metadata of all Object audio, and the Ext element of SAOC 3D Metadata is the metadata of all SAOC audio. Also, the Ext element of HOA Metadata is the metadata of all HOA audio, and the Ext element of DRC (Dynamic Range Control) Metadata is the metadata of all Object audio, SAOC audio, and HOA audio.

[0039] As described above, as the division units of the audio data in 3D audio, there are audio elements, group types, groups, switch groups, and group presets. Therefore, the audio stream of the audio data in 3D audio can be divided into different tracks for each type, with audio elements, group types, groups, switch groups, or group presets as the types.

[0040] Also, as the division units of the metadata in 3D audio, there are the types of Ext elements or the audio elements corresponding to the metadata. Therefore, the audio stream of the metadata of 3D audio can be divided into different tracks for each type, with Ext elements or the audio elements corresponding to the metadata as the types.

[0041] In the following embodiments, the audio stream of the audio data is divided into tracks for each of one or more groups, and the audio stream of the metadata is divided into tracks for each type of Ext element.

[0042] <First Embodiment> (Overview of the Information Processing System) FIG. 8 is a diagram for explaining the overview of the information processing system in the first embodiment to which the present disclosure is applied.

[0043] The information processing system 140 in FIG. 8 is configured by connecting a Web server 142 connected to a file generation device 141 and a video playback terminal 144 via the Internet 13.

[0044] In the information processing system 140, in a manner conforming to MPEG-DASH, the Web server 142 distributes the audio stream of the tracks of the group to be played back to the video playback terminal 144.

[0045] Specifically, the file generation device 141 encodes each audio data and metadata of the 3D audio of the video content at a plurality of encoding rates, and generates an audio stream. The file generation device 141 files all the audio streams at the encoding rate and in time units called segments, which are about several seconds to 10 seconds, and generates an audio file. At this time, the file generation device 141 divides the audio streams into groups and by the type of Ext element, and arranges them as audio streams of different tracks in the audio file. The file generation device 141 uploads the generated audio file to the web server 142.

[0046] In addition, the file generation device 141 generates an MPD file (management file) for managing an audio file and the like. The file generation device 141 uploads the MPD file to the web server 142.

[0047] The web server 142 stores the audio files and the MPD file for each encoding rate and segment uploaded from the file generation device 141. The web server 142 transmits the stored audio files, MPD files, etc. to the video playback terminal 144 in response to a request from the video playback terminal 144.

[0048] The video playback terminal 144 executes control software for streaming data (hereinafter referred to as control software) 161, video playback software 162, client software for HTTP (HyperText Transfer Protocol) access (hereinafter referred to as access software) 163, and the like.

[0049] The control software 161 is software for controlling data streamed from the web server 142. Specifically, the control software 161 causes the video playback terminal 144 to acquire the MPD file from the web server 142.

[0050] Also, the control software 161 commands the access software 163 to send a request for the audio stream of the track of the type of the Ext element corresponding to the group of playback targets specified by the video playback software 162 based on the MPD file.

[0051] The video playback software 162 is software that plays the audio stream acquired from the web server 142. Specifically, the video playback software 162 designates the group of playback targets and the type of the Ext element corresponding to that group to the control software 161. Also, when the video playback software 162 receives a notification of the start of reception from the access software 163, it decrypts the audio stream received by the video playback terminal 144. The video playback software 162 synthesizes and outputs the audio data obtained as a result of the decryption as necessary.

[0052] The access software 163 is software that controls communication with the web server 142 via the Internet 13 using HTTP. Specifically, the access software 163 causes the video playback terminal 144 to send a request for the audio stream of the track of the playback target included in the audio file according to the command of the control software 161. Also, in response to the transmission request, the access software 163 causes the video playback terminal 144 to start receiving the audio stream transmitted from the web server 142 and supplies a notification of the start of reception to the video playback software 162.

[0053] Note that in this specification, only the audio file of the video content is described, but actually, a corresponding image file is generated and played together with the audio file.

[0054] (Outline of the first example of the track of the audio file) FIG. 9 is a diagram for explaining the outline of the first example of the track of the audio file.

[0055] In addition, in FIG. 9, for convenience of explanation, only the tracks of the audio data among the 3D audio are illustrated. The same applies to FIGS. 20, 23, 26, 28, 30, 32 to 35, and 38 described later.

[0056] As shown in FIG. 9, all the audio streams of the 3D audio are stored in one audio file (3dauio.mp4). In the audio file (3dauio.mp4), the audio streams of each group of the 3D audio are respectively divided and arranged in different tracks. Also, information regarding the entire 3D audio is arranged as a base track.

[0057] A Track Reference is arranged in the track box of each track. The Track Reference represents the reference relationship of the corresponding track with other tracks. Specifically, the Track Reference represents the ID unique to the track of the other track in the reference relationship (hereinafter referred to as the track ID).

[0058] In the example of FIG. 9, the track IDs of the base track, group #1 with group ID 1, group #2 with group ID 2, group #3 with group ID 3, and group #4 with group ID 4 are 1, 2, 3, 4, 5. Also, the Track References of the base track are 2, 3, 4, 5, and the Track References of the tracks of groups #1 to #4 are 1, which is the track ID of the base track. Therefore, the base track and the tracks of groups #1 to #4 are in a reference relationship. That is, the base track is referenced during the playback of the tracks of groups #1 to #4.

[0059] In addition, the 4cc (character code) of the sample entry of the base track is "mha2". In the sample entry of the base track, there are arranged an mhaC box containing the config information of all groups of 3D audio or the config information necessary for decoding only the base track, and an mhas box containing information regarding all groups of 3D audio and switch Group. The information regarding a group is constituted by, for example, the ID of the group, information representing the content of the data of the elements classified into the group, and the like. The information regarding a switch Group is constituted by, for example, the ID of the switch Group, the IDs of the groups forming the switch Group, and the like.

[0060] The 4cc of the sample entry of the track of each group is "mhg1". In the sample entry of the track of each group, an mhgC box containing information regarding that group may be arranged. When a group forms a switch Group, an mhsC box containing information regarding that switch Group is arranged in the sample entry of the track of that group.

[0061] In the sample of the base track, there are arranged reference information to the samples of the tracks of each group, or the config information necessary for decoding that reference information. By arranging the samples of each group referred to by the reference information in the order of arrangement of the reference information, an audio stream of 3D audio before being divided into tracks can be generated. The reference information is constituted by, for example, the position and size of the samples of the tracks of each group, the group type, and the like.

[0062] (Example of the syntax of the sample entry of the base track) FIG. 10 is a diagram showing an example of the syntax of the sample entry of the base track.

[0063] As shown in FIG. 10, in the sample entry of the base track, an mhaC box (MHA Configuration Box), an mhas box (MHA Audio Scene Info Box), etc. are arranged. In the mhaC box, config information for all groups of 3D audio or the config information necessary for decoding only the base track is described. Also, in the mhas box, AudioScene information including information regarding all groups of 3D audio and the switch Group is described. This AudioScene information describes the hierarchical structure of FIG. 7.

[0064] (Example of the syntax of the sample entry of the track of each group) FIG. 11 is a diagram showing an example of the syntax of the sample entry of the track of each group.

[0065] As shown in FIG. 11, in the sample entry of the track of each group, an mhaC box (MHA Configuration Box), an mhgC box (MHAGroupDefinitionBox), an mhsC box (MHASwitchGropuDefinition Box), etc. are arranged.

[0066] In the mhaC box, the Config information necessary for decoding the corresponding track is described. Also, in the mhgC box, the AudioScene information regarding the corresponding group is described as GroupDefinition. In the mhsC box, when the corresponding group forms a switch Group, the AudioScene information regarding that switch Group is described as SwitchGroupDefinition.

[0067] (First example of the segment structure of the audio file) FIG. 12 is a diagram showing the first example of the segment structure of the audio file.

[0068] In the segment structure of FIG. 12, the Initial segment is composed of an ftyp box and a moov box. In the moov box, a trak box is arranged for each track included in the audio file. Further, in the moov box, an mvex box is arranged which contains information representing the correspondence between the track ID of each track and the level used in the ssix box within the media segment.

[0069] In addition, the media segment is composed of an sidx box, an ssix box, and one or more subsegments. In the sidx box, position information indicating the position of each subsegment within the audio file is arranged. The ssix box contains the position information of the audio streams at each level arranged in the mdat box. Note that the level corresponds to the track. Also, the position information of the first track is the position information of the data composed of the moof box and the audio stream of the first track.

[0070] The subsegment is provided for each arbitrary time length, and in the subsegment, a pair of a moof box and an mdat box common to all tracks is provided. In the mdat box, the audio streams of all tracks are arranged together for an arbitrary time length, and in the moof box, the management information of the audio stream is arranged. The audio streams of each track arranged in the mdat box are continuous for each track.

[0071] In the example of FIG. 12, Track1 with a track ID of 1 is the base track, and Track2 to TrackN with track IDs of 2 to N are tracks of groups with group IDs of 1 to N - 1. This is the same in FIG. 13 described later.

[0072] (Second example of the segment structure of an audio file) FIG. 13 is a diagram showing a second example of the segment structure of an audio file.

[0073] The segment structure in FIG. 13 is different from the segment structure in FIG. 12 in that a moof box and an mdat box are provided for each track.

[0074] That is, the Initial segment in FIG. 13 is the same as the Initial segment in FIG. 12. Also, similar to the media segment in FIG. 12, the media segment in FIG. 13 is composed of an sidx box, an ssix box, and one or more subsegments. In the sidx box, similar to the sidx box in FIG. 12, the position information of each subsegment is arranged. The ssix box contains the position information of the data at each level consisting of a moof box and an mdat box.

[0075] The subsegment is provided for each arbitrary time length, and for each subsegment, a pair of a moof box and an mdat box is provided for each track. That is, in the mdat box of each track, the audio stream of that track is collectively arranged (interleaved and stored) for an arbitrary time length, and in the moof box, the management information of that audio stream is arranged.

[0076] As shown in FIGS. 12 and 13, since the audio stream of each track is collectively arranged for an arbitrary time length, the acquisition efficiency of the audio stream via HTTP or the like is improved as compared with the case where it is collectively arranged in sample units.

[0077] (Description example of mvex box) FIG. 14 is a diagram showing a description example of a level assignment box arranged in the mvex boxes of FIGS. 12 and 13.

[0078] The level assignment box is a box that associates the track ID of each track with the level used in the ssix box. In the example of FIG. 14, the base track with track ID 1 is associated with level 0, and the channel audio track with track ID 2 is associated with level 1. Also, the HOA audio track with track ID 3 is associated with level 2, and the object metadata track with track ID 4 is associated with level 3. Furthermore, the object audio track with track ID 5 is associated with level 4.

[0079] (First description example of MPD file) FIG. 15 is a diagram showing the first description example of the MPD file.

[0080] As shown in FIG. 15, the MPD file describes a "Representation" that manages segments of a 3D audio voice file (3daudio.mp4), a "SubRepresentation" that manages the tracks included in the segment, and the like.

[0081] The "Representation" and "SubRepresentation" include "codecs" that represent the type (profile, level) of the codec for the entire corresponding segment or track in a code defined by the 3D audio file format.

[0082] The "SubRepresentation" includes a "level" that is the value set in the level assignment box as the value representing the level of the corresponding track. The "SubRepresentation" includes a "dependencyLevel" that is the value representing the level corresponding to another track (hereinafter referred to as a reference track) having a reference relationship (dependency).

[0083] Furthermore, the "SubRepresentation" includes <essentialproperty schemeiduri=""urn:mpeg:DASH:3daudio:2014”" value=""dataType,definition”">is included.

[0084] "dataType" is a number representing the type of the content (definition) of the Audio Scene information described in the sample entry of the corresponding track, and the definition is the content thereof. For example, if the GroupDefinition is included in the sample entry of the track, 1 is described as the "dataType" of that track, and the GroupDefinition is described as the "definition". Also, if the SwitchGroupDefinition is included in the sample entry of the track, 2 is described as the "dataType" of that track, and the SwitchGroupDefinition is described as the "definition". That is, "dataType" and "definition" are information indicating whether the SwitchGroupDefinition exists in the sample entry of the corresponding track. The "definition" is binary data and is encoded in the base64 format.

[0085] In the example of FIG. 15, all groups are assumed to form a switch group. However, if there is a group that does not form a switch group, the "SubRepresentation" corresponding to that group includes <essentialproperty schemeiduri=""urn:mpeg:DASH:3daudio:2014”" value=""2,SwitchGroupDefinition”">It is not described. This also applies to FIGS. 24, 25, 31, 39, 45, 47, 48, and 50 described later.

[0086] (Configuration example of file generation device) FIG. 16 is a block diagram showing a configuration example of the file generation device 141 of FIG. 8.

[0087] The file generation device 141 in FIG. 16 is composed of an audio encoding processing unit 171, an audio file generation unit 172, an MPD generation unit 173, and a server upload processing unit 174.

[0088] The audio encoding processing unit 171 of the file generation device 141 encodes each audio data and metadata of the 3D audio of the video content at a plurality of encoding rates respectively, and generates an audio stream. The audio encoding processing unit 171 supplies the audio stream for each encoding rate to the audio file generation unit 172.

[0089] The audio file generation unit 172 assigns tracks to the audio stream supplied from the audio encoding processing unit 171 for each group and type of Ext element. The audio file generation unit 172 generates an audio file having the segment structure of FIG. 12 or FIG. 13 in which the audio stream of each track is arranged in sub-segment units for each encoding rate and segment. The audio file generation unit 172 supplies the generated audio file to the MPD generation unit 173.

[0090] The MPD generation unit 173 determines the URL of the Web server 142 that stores the audio file supplied from the audio file generation unit 172 and the like. Then, the MPD generation unit 173 generates an MPD file in which the URL of the audio file and the like are arranged in the "Segment" of the "Representation" for the audio file. The MPD generation unit 173 supplies the generated MPD file and the audio file to the server upload processing unit 174.

[0091] The server upload processing unit 174 uploads the audio file and the MPD file supplied from the MPD generation unit 173 to the Web server 142.

[0092] (Explanation of the processing of the file generation device) FIG. 17 is a flowchart for explaining the file generation process of the file generation device 141 in FIG. 16.

[0093] In step S191 of FIG. 17, the audio encoding processing unit 171 encodes each audio data and metadata of the 3D audio of the video content at a plurality of encoding rates, respectively, to generate an audio stream. The audio encoding processing unit 171 supplies the audio stream for each encoding rate to the audio file generation unit 172.

[0094] In step S192, the audio file generation unit 172 assigns tracks to the audio stream supplied from the audio encoding processing unit 171 for each group and type of Ext element.

[0095] In step S193, the audio file generation unit 172 generates an audio file having the segment structure of FIG. 12 or FIG. 13 in which the audio streams of each track are arranged in sub-segment units for each encoding rate and segment. The audio file generation unit 172 supplies the generated audio file to the MPD generation unit 173.

[0096] In step S194, the MPD generation unit 173 generates an MPD file including the URL of the audio file and the like. The MPD generation unit 173 supplies the generated MPD file and the audio file to the server upload processing unit 174.

[0097] In step S195, the server upload processing unit 174 uploads the audio file and the MPD file supplied from the MPD generation unit 173 to the Web server 142. Then, the process ends.

[0098] (Functional Configuration Example of Video Playback Terminal) FIG. 18 is a block diagram showing a configuration example of a streaming playback unit realized by the video playback terminal 144 of FIG. 8 executing control software 161, video playback software 162, and access software 163.

[0099] The streaming playback unit 190 in FIG. 18 is composed of an MPD acquisition unit 91, an MPD processing unit 191, an audio file acquisition unit 192, an audio decoding processing unit 194, and an audio synthesis processing unit 195.

[0100] The MPD acquisition unit 91 of the streaming playback unit 190 acquires an MPD file from the Web server 142 and supplies it to the MPD processing unit 191.

[0101] The MPD processing unit 191 extracts information such as the URL of the audio file of the segment to be played described in the "Segment" for the audio file from the MPD file supplied from the MPD acquisition unit 91, and supplies it to the audio file acquisition unit 192.

[0102] The audio file acquisition unit 192 requests and acquires the audio stream of the track to be played in the audio file specified by the URL supplied from the MPD processing unit 191 from the Web server 142. The audio file acquisition unit 192 supplies the acquired audio stream to the audio decoding processing unit 194.

[0103] The audio decoding processing unit 194 decodes the audio stream supplied from the audio file acquisition unit 192. The audio decoding processing unit 194 supplies the audio data obtained as a result of the decoding to the audio synthesis processing unit 195. The audio synthesis processing unit 195 synthesizes and outputs the audio data supplied from the audio decoding processing unit 194 as necessary.

[0104] As described above, the audio file acquisition unit 192, the audio decoding processing unit 194, and the audio synthesis processing unit 195 function as a playback unit, and acquire and play the audio stream of the track to be played from the audio file stored in the Web server 142.

[0105] (Explanation of the processing of the video playback terminal) FIG. 19 is a flowchart for explaining the playback processing of the streaming playback unit 190 in FIG. 18.

[0106] In step S211 of FIG. 19, the MPD acquisition unit 91 of the streaming playback unit 190 acquires the MPD file from the Web server 142 and supplies it to the MPD processing unit 191.

[0107] In step S212, the MPD processing unit 191 extracts information such as the URL of the audio file of the segment to be played described in the "Segment" for the audio file from the MPD file supplied from the MPD acquisition unit 91, and supplies it to the audio file acquisition unit 192.

[0108] In step S213, based on the URL supplied from the MPD processing unit 191, the audio file acquisition unit 192 requests and acquires the audio stream of the track to be played in the audio file specified by the URL from the Web server 142. The audio file acquisition unit 192 supplies the acquired audio stream to the audio decoding processing unit 194.

[0109] In step S214, the audio decoding processing unit 194 decodes the audio stream supplied from the audio file acquisition unit 192. The audio decoding processing unit 194 supplies the audio data obtained as a result of the decoding to the audio synthesis processing unit 195. In step S215, the audio synthesis processing unit 195 synthesizes and outputs the audio data supplied from the audio decoding processing unit 194 as necessary.

[0110] (Outline of the second example of the track of the audio file) In the above description, GroupDefinition and SwitchGroupDefinition were placed in the sample entry. However, as shown in FIG. 20, they may be placed in the sample group entry, which is the sample entry for each group of subsamples in the track.

[0111] In this case, the sample group entry of the track of the group forming the switch Group includes GroupDefinition and SwitchGroupDefinition, as shown in FIG. 21. Although not shown, the sample group entry of the track of the group not forming the switch Group includes only GroupDefinition.

[0112] Also, the sample entry of the track of each group is as shown in FIG. 22. That is, as shown in FIG. 22, in the sample entry of the track of each group, an MHAGroupAudioConfigrationBox in which Config information such as the profile (MPEGHAudioProfile) and level (MPEGHAudioLevel) of the audio stream of the corresponding track is described is arranged.

[0113] (Outline of the third example of the track of the audio file) FIG. 23 is a diagram for explaining the outline of the third example of the track of the audio file.

[0114] The configuration of the track of the audio data in FIG. 23 is different from the configuration in FIG. 9 in that the base track includes one or more groups of 3D audio audio streams, and the number of groups corresponding to the audio streams divided into each track (hereinafter referred to as the group track) that does not include information about the entire 3D audio is 1 or more.

[0115] That is, the sample entry of the base track in FIG. 23 has the syntax for the base track when the audio stream of the audio data in 3D audio is divided and arranged in a plurality of tracks, and is a sample entry (FIG. 10) with a 4cc of "mha2".

[0116] Also, the sample entry of the group track has the syntax for the group track when the audio stream of the audio data in 3D audio is divided and arranged in a plurality of tracks, similar to FIG. 9, and is a sample entry (FIG. 11) with a 4cc of "mhg1". Therefore, the 4cc of the sample entry can be used to identify the base track and the group track and recognize the dependency relationship between the tracks.

[0117] Also, similar to FIG. 9, a Track Reference is arranged in the track box of each track. Therefore, even when it is not known which of "mha2" and "mhg1" is the 4cc of the sample entry of the base track or the group track, the dependency relationship between the tracks can be recognized by the Track Reference.

[0118] Note that the sample entry of the group track may not describe the mhgC box and the mhsC box. Also, when the mhaC box containing the config information of all groups of 3D audio is described in the sample entry of the base track, the mhaC box may not be described in the sample entry of the group track. However, when the mhaC box containing the config information that allows the base track to be played independently is described in the sample entry of the base track, the mhaC box containing the config information that allows that group track to be played independently is described in the sample entry of the group track. Whether it is the former state or the latter state can be identified by the presence or absence of the config information in the sample entry, but it can also be made identifiable by describing a flag in the sample entry or changing the type of the sample entry. Although illustration is omitted, when making the former state and the latter state identifiable by changing the type of the sample entry, the 4cc of the sample entry of the base track is set to, for example, "mha2" in the case of the former state and "mha4" in the case of the latter state.

[0119] (Second description example of the MPD file) FIG. 24 is a diagram showing a description example of an MPD file when the configuration of the track of the audio file is the configuration of FIG. 23.

[0120] The MPD file of FIG. 24 is different from the MPD file of FIG. 15 in that the "SubRepresentation" of the base track is described.

[0121] In the "SubRepresentation" of the base track, similar to the "SubRepresentation" of the group track, the "codecs", "level", "dependencyLevel", and <essentialproperty schemeiduri=""urn:mpeg:DASH:3daudio:2014”" value=""dataType,definition”">is described.

[0122] In the example of FIG. 24, the “codecs” of the base track is “mha2.2.1”, and the “level” is “0” as the value representing the level of the base track. The “dependencyLevel” is “1” and “2” as the values representing the levels of the group tracks. Also, the “dataType” is “3” as the number representing the type of the AudioScene information described in the mhas box of the sample entry of the base track, and the “definition” is the binary data of the AudioScene information encoded in the base64 format.

[0123] Note that, as shown in FIG. 25, the AudioScene information may be described separately in the “SubRepresentation” of the base track.

[0124] In the example of FIG. 25, “1” is set as the number representing the type of “Atmo” representing the content of the group with the group ID “1” among the AudioScene information (FIG. 7) described in the mhas box of the sample entry of the base track.

[0125] Also, “2” to “7” are set as the numbers representing the types of “Dialog EN” representing the content of the group with the group ID “2”, “Dialog FR” representing the content of the group with the group ID “3”, “VoiceOver GE” representing the content of the group with the group ID “4”, “Effects” representing the content of the group with the group ID “5”, “Effect” representing the content of the group with the group ID “6”, and “Effect” representing the content of the group with the group ID “7”, respectively.

[0126] Therefore, in the “SubRepresentation” of the base track in FIG. 25, the “dataType” is “1” and the “definition” is “Atmo”. <essentialproperty schemeiduri=""urn:mpeg:DASH:3daudio:2014”" value=""dataType,definition”">is described. Similarly, when "dataType" is "2", "3", "4", "5", "6", "7" respectively, and "defini tion" is "Dialog EN", "Dialog FR", "VoiceOver GE", "Effects", "Effect", "Effect" respectively, "urn:mpeg:DASH:3daudio:2014” value="dataType,definition”> is described. In the example of Figure 25, the case where the AudioScene information of the base track is described separately was explained, but the GroupDefinition and SwitchGroupDefinition of the group track may also be described separately in the same way as the AudioScene information.

[0127] (Overview of the fourth example of the track of the audio file) Figure 26 is a diagram for explaining an overview of the fourth example of the track of the audio file.

[0128] The configuration of the track of the audio data in Figure 26 is different from the configuration of Figure 23 in that the sample entry of the group track is a sample entry with a 4cc of "mha2".

[0129] In the case of Figure 26, the 4cc of the sample entries of both the base track and the group track becomes "mha2". Therefore, it is not possible to identify the base track and the group track based on the 4cc of the sample entry, and recognize the dependency relationship between the tracks. Thus, the Track Reference placed in the track box of each track is used to recognize the dependency relationship between the tracks.

[0130] Also, since the 4cc of the sample entry is "mha2", it can be identified that the corresponding track is a track when the audio stream of the audio data in 3D audio is divided and arranged in a plurality of tracks.

[0131] ​Note that, in the mhaC box of the sample entry of the base track, similar to the cases of FIGS. 9 and 23, the config information of all groups of 3D audio or the config information that enables the base track to be played independently is described. Also, in the mhas box, AudioScene information including information regarding all groups of 3D audio and the switch Group is described.

[0132] On the other hand, in the sample entry of the group track, the mhas box is not arranged. Also, when the mhaC box containing the config information of all groups of 3D audio is described in the sample entry of the base track, the mhaC box may not be described in the sample entry of the group track. However, when the mhaC box containing the config information that enables the base track to be played independently is described in the sample entry of the base track, the mhaC box containing the config information that enables the group track to be played independently is described in the sample entry of the group track. Whether it is the former state or the latter state can be identified by the presence or absence of the config information in the sample entry, but it can also be made identifiable by describing a flag in the sample entry or changing the type of the sample entry. Although illustration is omitted, when making the former state and the latter state identifiable by changing the type of the sample entry, the 4cc of the sample entries of the base track and the group track is, for example, set to "mha2" in the case of the former state and "mha4" in the case of the latter state.

[0133] (The third description example of the MPD file) FIG. 27 is a diagram showing a description example of an MPD file when the configuration of the track of the audio file is the configuration of FIG. 26.

[0134] The MPD file of FIG. 27 is such that "codecs" of "SubRepresentation" of the group track is "mha2.2.1", and in "SubRepresentation" of the group track <essentialproperty schemeiduri=""urn:mpeg:DASH:3daudio:2014”" value=""dataType,definition”">The point that is not described is different from the MPD file of FIG. 24.

[0135] Although illustration is omitted, similar to the case of FIG. 25, AudioScene information may be separately described in the "SubRepresentation" of the base track.

[0136] (Outline of the fifth example of the track of the audio file) FIG. 28 is a diagram for explaining the outline of the fifth example of the track of the audio file.

[0137] The configuration of the track of the audio data in FIG. 28 is different from the configuration of FIG. 23 in that the sample entries of the base track and the group track have a syntax suitable for both the base track and the group track when the audio stream of the audio data in 3D audio is divided into a plurality of tracks.

[0138] In the case of FIG. 28, the 4cc of the sample entries of the base track and the group track both becomes "mha3", which is the 4cc of the sample entry having a syntax suitable for both the base track and the group track.

[0139] Therefore, similar to the case of FIG. 26, the dependency relationship between tracks is recognized by the Track Reference arranged in the track box of each track. Also, since the 4cc of the sample entry is "mha3", it can be identified that the corresponding track is a track when the audio stream of the audio data in 3D audio is divided and arranged in a plurality of tracks.

[0140] (Example of the syntax of the sample entry whose 4cc is "mha3") FIG. 29 is a diagram showing an example of the syntax of the sample entry whose 4cc is "mha3".

[0141] As shown in FIG. 29, the syntax of the sample entry of 4cc "mha3" is a combination of the syntax of FIG. 10 and the syntax of FIG. 11.

[0142] That is, in the sample entry where 4cc is "mha3", an mhaC box (MHA Configuration Box), an mhas box (MHA Audio Scene Info Box), an mhgC box (MHA Group Definition Box), an mhsC box (MHA Switch Group Definition Box), etc. are arranged.

[0143] In the mhaC box of the sample entry of the base track, the config information of all groups of 3D audio or the config information that can play the base track independently is described. Also, in the mhas box, AudioScene information including information on all groups and switch groups of 3D audio is described, and the mhgC box and mhsC box are not arranged.

[0144] When an mhaC box containing the config information of all groups of 3D audio is described in the sample entry of the base track, the mhaC box may not be described in the sample entry of the group track. However, when an mhaC box containing config information that allows the base track to be played independently is described in the sample entry of the base track, an mhaC box containing config information that allows the group track to be played independently is described in the sample entry of the group track. Whether it is the former state or the latter state can be identified by the presence or absence of config information in the sample entry, but it can also be made identifiable by describing a flag in the sample entry or changing the type of the sample entry. Although illustration is omitted, when the former state and the latter state can be identified by changing the type of the sample entry, the 4cc of the sample entries of the base track and the group track are, for example, set to "mha3" in the former state and "mha5" in the latter state. Also, the mhas box is not placed in the sample entry of the group track. The mhgC box and the mhsC box may or may not be placed.

[0145] Note that, as shown in FIG. 30, an mhas box, an mhgC box, and an mhsC box are arranged in the sample entry of the base track, and both an mhaC box in which config information for making only the base track reproducible independently is described and an mhaC box including config information for all groups of 3D audio may be arranged. In this case, the mhaC box in which config information for all groups of 3D audio is described and the mhaC box in which config information for making only the base track reproducible independently is described are identified by flags included in these mhaC boxes. Also, in this case, the mhaC box may not be described in the sample entry of the group track. Whether the mhaC box is described in the sample entry of the group track can be identified by the presence or absence of the mhaC box in the sample entry of the group track, but it can also be made identifiable by describing a flag in the sample entry or changing the type of the sample entry. Although illustration is omitted, when making it possible to identify whether the mhaC box is described in the sample entry of the group track by changing the type of the sample entry, the 4cc of the sample entries of the base track and the group track is, for example, set to "mha3" when the mhaC box is described in the sample entry of the group track, and set to "mha5" when the mhaC box is not described in the sample entry of the group track. Note that in FIG. 30, the mhgC box and the mhsC box may not be described in the sample entry of the base track.

[0146] (Fourth description example of the MPD file) FIG. 31 is a diagram showing a description example of an MPD file when the configuration of the track of the audio file is the configuration of FIG. 28 or FIG. 30.

[0147] The MPD file in FIG. 31 is different from the MPD file in FIG. 24 in that the "codecs" of "Representation" is "mha3.3.1" and the "codecs" of "SubRepresentation" is "mha3.2.1".

[0148] Although illustration is omitted, similar to the case of FIG. 25, the AudioScene information may be separately described in the "SubRepresentation" of the base track.

[0149] Also, in the above description, Track Reference is arranged in the track box of each track, but Track Reference may not be arranged. For example, FIGS. 32 to 34 are diagrams showing cases where Track Reference is not arranged in the track boxes of the audio files in FIGS. 23, 26, and 28 respectively. In the case of FIG. 32, although Track Reference is not arranged, since the 4cc of the sample entries of the base track and the group track are different, the dependency relationship between the tracks can be recognized. In the cases of FIGS. 33 and 34, by arranging the mhas box, it is possible to identify whether it is a base track or not.

[0150] The MPD files in the cases where the track configurations of the audio files are the configurations in FIGS. 32 to 34 are the same as the MPD files in FIGS. 24, 27, and 31 respectively. Also in this case, similar to the case of FIG. 25, the AudioScene information may be separately described in the "SubRepresentation" of the base track.

[0151] (Outline of the sixth example of the track of the audio file) FIG. 35 is a diagram for explaining the outline of the sixth example of the track of the audio file.

[0152] The configuration of the track of the audio data in FIG. 35 is different from the configuration in FIG. 33 in that reference information to the samples of each group's track and config information necessary for decoding the reference information are not arranged in the samples of the base track, and zero or more groups of audio streams are included, and in the sample entry of the base track, reference information to the samples of each group's track is described.

[0153] Specifically, in a sample entry with a 4cc of "mha2" having a syntax for the base track when the audio stream of the audio data in 3D audio is divided into a plurality of tracks, an mhmt box for describing in which track each group described in the AudioScene information is divided is newly arranged.

[0154] (Another example of the syntax of the sample entry with a 4cc of "mha2") FIG. 36 is a diagram showing an example of the syntax of the sample entry of the base track and the group track in FIG. 35 with a 4cc of "mha2".

[0155] The configuration of the sample entry with a 4cc of "mha2" in FIG. 36 is different from the configuration in FIG. 10 in that an MHAMultiTrackDescription box (mhmt box) is arranged.

[0156] In the mhmt box, as reference information, the correspondence between the group ID (group_ID) and the track ID (track_ID) is described. Note that in the mhmt box, the audio element and the track ID may be described in association with each other.

[0157] When the reference information does not change for each sample, by arranging the mhmt box in the sample entry, the reference information can be described efficiently.

[0158] Although illustration is omitted, in the cases of FIGS. 9, 20, 23, 26, 28, 30, 32, and 34 as well, similarly, instead of describing reference information from the samples of the base track to the samples of the tracks of each group, an mhmt box can be arranged in the sample entry of the base track.

[0159] In this case, the syntax of the sample entry where 4cc is "mha3" is as shown in FIG. 37. That is, the configuration of the sample entry where 4cc in FIG. 37 is "mha3" is different from the configuration in FIG. 29 in that an MHAMultiTrackDescription box (mhmt box) is arranged.

[0160] Also, in FIGS. 23, 26, 28, 30, 32 to 34, and 35, similar to FIG. 9, the base track may not include one or more groups of 3D audio audio streams. Also, the number of groups corresponding to the audio streams divided into each group track may be one.

[0161] Furthermore, in FIGS. 23, 26, 28, 30, 32 to 34, and 35, similar to the case of FIG. 20, GroupDefinition and SwitchGroupDefinition may be arranged in the sample group entry.

[0162] <Second Embodiment> (Overview of Track) FIG. 38 is a diagram for explaining an overview of a track in a second embodiment to which the present disclosure is applied.

[0163] As shown in FIG. 38, in the second embodiment, it is different from the first embodiment in that each track is recorded as a different file (3da_base.mp4 / 3da_group1.mp4 / 3da_group2.mp4 / 3da_group3.mp4 / 3da_group4.mp4). In this case, by obtaining the file of the desired track via HTTP, only the data of the desired track can be obtained. Therefore, it is possible to efficiently obtain the data of the desired track via HTTP.

[0164] (Example description of MPD file) FIG. 39 is a diagram showing an example description of an MPD file in the second embodiment to which the present disclosure is applied.

[0165] As shown in FIG. 39, in the MPD file, "Representation" and the like for managing segments of each audio file (3da_base.mp4 / 3da_group1.mp4 / 3da_group2.mp4 / 3da_group3.mp4 / 3da_group4.mp4) of 3D audio are described.

[0166] "Representation" includes "codecs", "id", "associationId", and "assciationType". "id" is the ID of the "Representation" containing it. "associationId" is information representing the reference relationship between the corresponding track and other tracks, and is the "id" of the reference track. "assciationType" is a code representing the meaning of the reference relationship (dependency relationship) with the reference track, and for example, the same value as the track reference of MP4 is used.

[0167] Also, in the "Representation" of the tracks of each group, <essentialproperty schemeiduri=""urn:mpeg:DASH:3daudio:2014”" value=""dataType,definition”">is also included. In the example of FIG. 39, under one "AdaptationSet", a "Representation" for managing segments of each audio file is provided. However, an "AdaptationSet" may be provided for each segment of each audio file, and under it, a "Representation" for managing that segment may be provided. In this case, each "AdaptationSet" includes an "associationId" and, similar to "assciationType", represents the meaning of the reference relationship with the reference track <essentialproperty schemeiduri=""urn:mpeg:DASH:3daudioAssociationData:2014”" value=""dataType,id”">may be described as such. Also, the AudioScene information, GroupDefinition, and SwitchGroupDefinition described in the "R epresentation" of the base track and group track may be described separately as in the case of FIG. 25. Further, each "AdaptationSet" may include the AudioScene information, GroupDefinition, and SwitchGroupDefinition described separately in the "Representation".

[0168] (Overview of the information processing system) FIG. 40 is a diagram for explaining an overview of an information processing system according to a second embodiment to which the present disclosure is applied.

[0169] Among the configurations shown in FIG. 40, the same configurations as those in FIG. 8 are denoted by the same reference numerals. Redundant descriptions will be omitted as appropriate.

[0170] The information processing system 210 in FIG. 40 is configured by connecting a Web server 212 and a video playback terminal 214, which are connected to a file generation device 211, via the Internet 13.

[0171] In the information processing system 210, in a manner conforming to MPEG-DASH, the Web server 142 distributes the audio stream of the audio file of the group to be played back to the video playback terminal 144.

[0172] Specifically, the file generation device 211 encodes each audio data and metadata of the 3D audio of the video content at a plurality of encoding rates to generate an audio stream. The file generation device 211 divides the audio stream for each group and type of Ext element, and makes the audio streams of different tracks. The file generation device 211 files the audio stream for each encoding rate, segment, and track to generate an audio file. The file generation device 211 uploads the resulting audio file to the Web server 212. Also, the file generation device 211 generates an MPD file and uploads it to the Web server 212.

[0173] The Web server 212 stores the audio files and MPD files for each encoding rate, segment, and track uploaded from the file generation device 211. The Web server 212 transmits the stored audio files, MPD files, etc. to the video playback terminal 214 in response to a request from the video playback terminal 214.

[0174] The video playback terminal 214 executes control software 221, video playback software 162, access software 223, etc.

[0175] The control software 221 is software that controls the data streamed from the Web server 212. Specifically, the control software 221 causes the video playback terminal 214 to acquire the MPD file from the Web server 212.

[0176] Also, the control software 221 instructs the access software 223 to send a request for the audio stream of the audio file of the type of Ext element corresponding to the group to be played back specified by the video playback software 162 based on the MPD file.

[0177] The access software 223 is software that controls communication with the web server 212 via the Internet 13 using HTTP. Specifically, the access software 223 causes the video playback terminal 144 to send a transmission request for the audio stream of the audio file to be played back in response to a command from the control software 221. Also, the access software 223 causes the video playback terminal 144 to start receiving the audio stream transmitted from the web server 212 in response to the transmission request, and supplies a notification of the start of reception to the video playback software 162.

[0178] (Configuration example of the file generation device) FIG. 41 is a block diagram showing a configuration example of the file generation device 211 of FIG. 40.

[0179] Among the configurations shown in FIG. 41, the same configurations as those in FIG. 16 are given the same reference numerals. Redundant descriptions will be omitted as appropriate.

[0180] The configuration of the file generation device 211 in FIG. 41 is different from the configuration of the file generation device 141 in FIG. 16 in that the audio file generation unit 241 and the MPD generation unit 242 are provided instead of the audio file generation unit 172 and the MPD generation unit 173.

[0181] Specifically, the audio file generation unit 241 of the file generation device 211 assigns tracks to the audio stream supplied from the audio encoding processing unit 171 for each type of group and Ext element. The audio file generation unit 241 generates an audio file in which the audio stream is arranged for each encoding rate, segment, and track. The audio file generation unit 241 supplies the generated audio file to the MPD generation unit 242.

[0182] The MPD generation unit 242 determines the URL of the Web server 142 that stores the audio file supplied from the audio file generation unit 172. The MPD generation unit 242 generates an MPD file in which the URL of the audio file and the like are arranged in the "Segment" of the "Representation" for the audio file. The MPD generation unit 173 supplies the generated MPD file and audio file to the server upload processing unit 174.

[0183] (Explanation of the processing of the file generation device) FIG. 42 is a flowchart for explaining the file generation process of the file generation device 211 in FIG. 41.

[0184] Since the processes in steps S301 and S302 in FIG. 42 are the same as the processes in steps S191 and S192 in FIG. 17, the description thereof is omitted.

[0185] In step S303, the audio file generation unit 241 generates an audio file in which the audio stream is arranged for each encoding rate, segment, and track. The audio file generation unit 241 supplies the generated audio file to the MPD generation unit 242.

[0186] Since the processes in steps S304 and S305 are the same as the processes in steps S194 and S195 in FIG. 17, the description thereof is omitted.

[0187] (Example of the functional configuration of the video playback terminal) FIG. 43 is a block diagram showing a configuration example of a streaming playback unit realized by the video playback terminal 214 in FIG. 40 executing the control software 221, the video playback software 162, and the access software 223.

[0188] Among the configurations shown in FIG. 43, the same configurations as those in FIG. 18 are denoted by the same reference numerals. Redundant descriptions are omitted as appropriate.

[0189] The configuration of the streaming playback unit 260 in FIG. 43 differs from the configuration of the streaming playback unit 190 in FIG. 18 in that an audio file acquisition unit 264 is provided instead of the audio file acquisition unit 192.

[0190] Based on the URL of the audio file of the track to be played back among the URLs supplied from the MPD processing unit 191, the audio file acquisition unit 264 requests and acquires the audio stream of that audio file from the web server 142. The audio file acquisition unit 264 supplies the acquired audio stream to the audio decoding processing unit 194.

[0191] That is, the audio file acquisition unit 264, the audio decoding processing unit 194, and the audio synthesis processing unit 195 function as a playback unit, acquire and play back the audio stream of the audio file of the track to be played back from the audio file stored in the web server 212.

[0192] (Explanation of the processing of the video playback terminal) FIG. 44 is a flowchart for explaining the playback processing of the streaming playback unit 260 in FIG. 43.

[0193] Since the processing in steps S321 and S322 in FIG. 44 is the same as the processing in steps S211 and S212 in FIG. 19, the explanation is omitted.

[0194] In step S323, based on the URL of the audio file of the track to be played back among the URLs supplied from the MPD processing unit 191, the audio file acquisition unit 192 requests and acquires the audio stream of that audio file from the web server 142. The audio file acquisition unit 264 supplies the acquired audio stream to the audio decoding processing unit 194.

[0195] Since the processing in steps S324 and S325 is the same as the processing in steps S214 and S215 in FIG. 19, the explanation is omitted.

[0196] Note that, also in the second embodiment, similar to the first embodiment, GroupDefinition and SwitchGroupDefinition may be arranged in the sample group entry.

[0197] Also, in the second embodiment, similar to the first embodiment, the configuration of the track of the audio data can be made the configuration shown in FIGS. 23, 26, 28, 30, 32 to 34, and 35.

[0198] FIGS. 45 to 47 are diagrams showing the MPD when the configuration of the track of the audio data is the configuration shown in FIGS. 23, 26, and 28 in the second embodiment. In the second embodiment, the MPD when the configuration of the track of the audio data is the configuration shown in FIGS. 32, 33 or 35, 34 is the same as the MPD when the configuration is the configuration shown in FIGS. 23, 26, and 28, respectively.

[0199] The MPD in FIG. 45 is for the "codecs" and "associationId" of the base track, and for the "Representation" of the base track <essentialproperty schemeiduri=""urn:mpeg:DASH:3daudio:2014”" value=""dataType,definition”">The points included are different from the MPD in FIG. 39. Specifically, for the "codecs" in the "Representation" of the base track of the MPD in FIG. 45, it is "mha2.2.1", and the "associationId" is "g1" and "g2" which are the "id" of the group track.

[0200] Also, the MPD in FIG. 46 is for the "codecs" of the group track and in the "Representation" of the group track <essentialproperty schemeiduri=""urn:mpeg:DASH:3daudio:2014”" value=""dataType,definition”">The point that is not included is different from the MPD in FIG. 45. Specifically, the "codecs" of the group track of the MPD in FIG. 46 is "mha2.2.1".

[0201] Also, the MPD in FIG. 47 is different from the MPD in FIG. 45 in that the "codecs" of the base track and the group track are different. Specifically, the "codecs" of the group track of the MPD in FIG. 47 is "mha3.2.1".

[0202] In the MPDs of FIGS. 45 to 47, as shown in FIGS. 48 to 50, the "AdaptationSet" can also be separated for each "Representation".

[0203] <Another example of the base track> In the above description, only one base track is provided, but a plurality of base tracks may be provided. In this case, the base tracks are provided, for example, for each viewpoint of 3D audio (details will be described later), and an mhaC box containing the config information of all groups of 3D audio for each viewpoint is arranged in the base track. Note that an mhas box containing the AudioScene information for each viewpoint may be arranged in each base track.

[0204] The viewpoint of 3D audio is the position where the 3D audio can be heard, such as the viewpoint of the image reproduced simultaneously with the 3D audio or a predetermined position set in advance.

[0205] As described above, when base tracks are provided for each viewpoint, different voices can be reproduced for each viewpoint from the audio stream of the same 3D audio based on the position on the screen of the object included in the config information of each viewpoint. As a result, the data amount of the audio stream of the 3D audio can be reduced.

[0206] That is, when the viewpoints of 3D audio are the multiple viewpoints of an image of a baseball stadium that can be played simultaneously with the 3D audio, as the main image that is the image of the basic viewpoint, for example, an image with the center back screen as the viewpoint is prepared. In addition, images with the back net, first base infield seats, third base infield seats, left support seats, right support seats, etc. as viewpoints are prepared as multi-images that are images of viewpoints other than the basic viewpoint.

[0207] In this case, if 3D audio for all viewpoints is prepared, the data volume of the 3D audio will increase. Therefore, by describing the on-screen position of an object at each viewpoint in the base track, audio streams such as Object audio and SAOC Object audio that change according to the on-screen position of the object can be shared among viewpoints. As a result, the data volume of the audio stream of the 3D audio can be reduced.

[0208] When playing 3D audio, for example, using audio streams such as Object audio and SAOC Object audio of the basic viewpoint and the base track corresponding to the viewpoint of the main image or multi-image played simultaneously, different voices are played according to the viewpoint.

[0209] Similarly, for example, when the viewpoints of 3D audio are the positions of a plurality of seats in a preset stadium, if 3D audio for all viewpoints is prepared, the data volume of the 3D audio will increase. Therefore, by describing the on-screen position of an object at each viewpoint in the base track, audio streams such as Object audio and SAOC Object audio can be shared among viewpoints. As a result, it becomes possible to play different voices according to the seats selected by the user using a seat map or the like using Object audio and SAOC Object audio of one viewpoint, and the data volume of the audio stream of the 3D audio can be reduced.

[0210] In the track structure of FIG. 28, when the base track is provided for each viewpoint of 3D audio, the track structure becomes as shown in FIG. 51. In the example of FIG. 51, there are three viewpoints of 3D audio. Also, in the example of FIG. 51, the Channel audio is generated for each viewpoint of 3D audio, and the other audio data is shared among the viewpoints of 3D audio. These are the same in the example of FIG. 52 described later.

[0211] In this case, as shown in FIG. 51, three base tracks are provided for each viewpoint of 3D audio. A Track Reference is arranged in the track box of each base track. Also, the syntax of the sample entry of each base track is the same as the syntax of the sample entry whose 4cc is "mha3", but the 4cc is "mhcf" indicating that the base track is provided for each viewpoint of 3D audio.

[0212] An mhaC box containing the config information of all groups of 3D audio for each viewpoint is arranged in the sample entry of each base track. Examples of the config information of all groups of 3D audio for each viewpoint include the position of the object on the screen in that viewpoint. Also, an mhas box containing the AudioScene information for each viewpoint is arranged in each base track.

[0213] The audio stream of the group of Channel audio for each viewpoint is arranged in the sample of each base track.

[0214] If there is Object Metadata that describes the position of the object on the screen for each sample in each viewpoint, the Object Metadata is also arranged in the sample of each base track.

[0215] That is, when the object is a moving body (for example, a sports player), since the position of the object on the screen at each viewpoint changes with time, the position is described as Object Metadata in sample units. In this case, the Object Metadata in these sample units is arranged in the samples of the base track corresponding to that viewpoint for each viewpoint.

[0216] The configuration of the group track in FIG. 51 is the same as that in FIG. 28 except that the audio streams of the group of Channel audio are not arranged, so the description is omitted.

[0217] Note that in the track structure of FIG. 51, the audio streams of the group of Channel audio for each viewpoint may not be arranged in the base track, but may be arranged in different group tracks respectively. In this case, the track structure will be as shown in FIG. 52.

[0218] In the example of FIG. 52, the audio stream of the group of Channel audio for the viewpoint corresponding to the base track with track ID "1" is arranged in the group track with track ID "4". Also, the audio stream of the group of Channel audio for the viewpoint corresponding to the base track with track ID "2" is arranged in the group track with track ID "5".

[0219] Furthermore, the audio stream of the group of Channel audio for the viewpoint corresponding to the base track with track ID "3" is arranged in the group track with track ID "6".

[0220] Note that in the examples of FIGS. 51 and 52, the 4cc of the sample entry of the base track is "mhcf", but it may also be the same "mha3" as in the case of FIG. 28.

[0221] Also, although illustration is omitted, in all of the above-described track structures other than the track structure of FIG. 28, when the base track is provided for each viewpoint of 3D audio, it is the same as in the cases of FIGS. 51 and 52.

[0222] <Third Embodiment> (Description of Computer to which the Present Disclosure is Applied) The series of processes of the above-described Web server 142(212) can be executed by hardware or by software. When the series of processes are executed by software, the program constituting the software is installed in a computer. Here, the computer includes a computer incorporated in dedicated hardware, and a general-purpose personal computer or the like that can execute various functions by installing various programs.

[0223] FIG. 53 is a block diagram showing a configuration example of the hardware of a computer that executes the series of processes of the above-described Web server 142(212) by a program.

[0224] In a computer, a CPU (Central Processing Unit) 601, a ROM (Read Only Memory) 602, and a RAM (Random Access Memory) 603 are interconnected by a bus 604.

[0225] An input / output interface 605 is further connected to the bus 604. An input unit 606, an output unit 607, a storage unit 608, a communication unit 609, and a drive 610 are connected to the input / output interface 605.

[0226] The input unit 606 consists of a keyboard, a mouse, a microphone, etc. The output unit 607 consists of a display, a speaker, etc. The storage unit 608 consists of a hard disk, a non-volatile memory, etc. The communication unit 609 consists of a network interface, etc. The drive 610 drives a removable medium 611 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.

[0227] In the computer configured as described above, for example, the CPU 601 loads and executes a program stored in the storage unit 608 via the input / output interface 605 and the bus 604 into the RAM 603, whereby the above-described series of processes are performed.

[0228] The program executed by the computer (CPU 601) can be recorded and provided, for example, on a removable medium 611 as a package medium or the like. Also, the program can be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.

[0229] In the computer, the program can be installed in the storage unit 608 via the input / output interface 605 by attaching the removable medium 611 to the drive 610. Also, the program can be received by the communication unit 609 via a wired or wireless transmission medium and installed in the storage unit 608. Additionally, the program can be installed in advance in the ROM 602 or the storage unit 608.

[0230] Note that the program executed by the computer may be a program in which processing is performed in time series in accordance with the order described in this specification, or may be a program in which processing is performed in parallel or at a necessary timing such as when a call is made.

[0231] Also, the hardware configuration of the video playback terminal 144(214) can be the same as that of the computer in FIG. 53. In this case, for example, the CPU 601 executes control software 161(221), video playback software 162, and access software 163(223). The processing of the video playback terminal 144(214) can also be executed by hardware.

[0232] In this specification, a system means a collection of a plurality of components (devices, modules (parts), etc.), and it does not matter whether all the components are in the same housing. Therefore, a plurality of devices housed in separate housings and connected via a network, and a single device in which a plurality of modules are housed in one housing are both systems.

[0233] Note that the embodiments of the present disclosure are not limited to the above-described embodiments, and various modifications are possible without departing from the gist of the present disclosure.

[0234] Also, the present disclosure can be applied to an information processing system that performs broadcast or local storage playback instead of streaming playback.

[0235] In the above-described MPD embodiments, when the content described in its schema cannot be understood, information may be described by EssentialProperty, which is a descriptor definition that can be ignored, or information may be described by SupplementalProperty, which is a descriptor definition that can be played back even when the content described in its schema cannot be understood. The selection of this description method is made according to the intention of the content creator.

[0236] Furthermore, the present disclosure can also have the following configuration.

[0237] (1) A file generation unit that generates a file in which a plurality of types of audio data are divided and arranged for each of one or more of the types, and information regarding the plurality of types is arranged An information processing apparatus including the same. (2) The information regarding the plurality of types is arranged in a sample entry of a predetermined track and is configured as follows The information processing apparatus according to (1) above. (3) The predetermined track is one of the tracks in which the plurality of types of audio data are divided and arranged and is configured as follows The information processing apparatus according to (2) above. (4) In the file, for each track, information regarding the type corresponding to that track is arranged and is configured as follows The information processing apparatus according to any one of (1) to (3) above. (5) In the file, for each track, information regarding an exclusive playback type consisting of the type corresponding to that track and the type corresponding to audio data that is exclusively played back with the audio data of that type is arranged and is configured as follows The information processing apparatus according to (4) above. (6) The information regarding the type corresponding to the track and the information regarding the exclusive playback type are arranged in a sample entry of the corresponding track and is configured as follows The information processing apparatus according to (5) above. (7) The file generation unit generates a management file that manages the file and includes information indicating whether information regarding the exclusive playback type exists for each track and is configured as follows The information processing apparatus according to (5) or (6) above. (8) In the file, reference information to tracks corresponding to the plurality of types is arranged. configured as the information processing apparatus according to any one of (1) to (7) above. (9) The reference information is arranged in samples of a predetermined track. configured as the information processing apparatus according to (8) above. (10) The predetermined track is one of the tracks in which the plurality of types of audio data are divided and arranged by type. configured as the information processing apparatus according to (9) above. (11) In the file, information representing the reference relationship between the tracks is arranged. configured as the information processing apparatus according to any one of (1) to (10) above. (12) The file generation unit generates a management file that manages the file and includes information representing the reference relationship between the tracks. configured as the information processing apparatus according to any one of (1) to (11) above. (13) The file is one file. configured as the information processing apparatus according to any one of (1) to (12) above. (14) The file is a file for each track. configured as the information processing apparatus according to any one of (1) to (12) above. (15) An information processing apparatus A file generation step of generating a file in which a plurality of types of audio data are divided and arranged in tracks for each of one or more of the types, and information regarding the plurality of types is arranged. An information processing method including. (16) A playback unit that plays the audio data of a predetermined track from a file in which a plurality of types of audio data are divided and arranged for each of one or more of the types and information regarding the plurality of types is arranged An information processing apparatus including the same. (17) The information processing apparatus A playback step of playing the audio data of a predetermined track from a file in which a plurality of types of audio data are divided and arranged for each of one or more of the types and information regarding the plurality of types is arranged An information processing method including the same.

Explanation of Signs

[0238] 11 File generation device, 192 Audio file acquisition unit, 194 Audio decoding processing unit, 195 Audio synthesis processing unit, 211 File generation device, 264 Audio file acquisition unit< / essentialproperty> < / essentialproperty> < / essentialproperty> < / essentialproperty> < / essentialproperty> < / essentialproperty> < / essentialproperty> < / essentialproperty> < / essentialproperty>

Claims

1. A file generation unit that generates a file in which a plurality of types of audio data are stored in individual tracks and information regarding the plurality of types is arranged and includes the audio data forms groups based on relevance, a plurality of types of audio data are divided and arranged in the tracks for each of the one or more types, in the file, reference information to the tracks corresponding to the plurality of types is arranged, the reference information is arranged in samples of a predetermined track, the predetermined track is one of the tracks in which the plurality of types of audio data are divided and arranged An information processing apparatus.

2. The information regarding the plurality of types is arranged in a sample entry of a predetermined track configured as The information processing apparatus according to claim 1.

3. The predetermined track is one of the tracks in which the plurality of types of audio data are divided and arranged configured as The information processing apparatus according to claim 2.

4. In the file, information regarding the type corresponding to each track is arranged for each track configured as The information processing apparatus according to claim 1.

5. In the file, for each track, information regarding an exclusive playback type including the type corresponding to that track and the type corresponding to audio data that is exclusively played back with the audio data of that type is arranged configured as The information processing apparatus according to claim 4.

6. The information regarding the type corresponding to the track and the information regarding the exclusive playback type are arranged in a sample entry of the corresponding track configured as The information processing apparatus according to claim 5.

7. The file generation unit generates a management file that manages the file and includes information indicating whether information regarding the exclusive playback type exists for each track configured as The information processing apparatus according to claim 5.

8. In the file, information representing a reference relationship between the tracks is arranged configured as The information processing apparatus according to claim 1.

9. The file generation unit generates a management file that manages the file and includes information representing a reference relationship between the tracks configured as The information processing apparatus according to claim 1.

10. The file is one file configured as The information processing apparatus according to claim 1.

11. The file is a file for each of the tracks configured as The information processing apparatus according to claim 1.

12. An information processing apparatus generates a file generation step of generating a file in which a plurality of types of audio data are stored in individual tracks and information regarding the plurality of types is arranged including the audio data forms groups based on relevance a plurality of types of audio data are divided and arranged in the tracks for each of one or more of the types in the file, reference information to the tracks corresponding to the plurality of types is arranged the reference information is arranged in samples of a predetermined track the predetermined track is one of the tracks in which the plurality of types of audio data are divided and arranged An information processing method.

13. A playback unit that plays back the audio data of a predetermined track from a file including a plurality of types of audio data stored in individual tracks comprising the audio data forms groups based on relevance a plurality of types of audio data are divided and arranged in the tracks for each of one or more of the types in the file, reference information to the tracks corresponding to the plurality of types is arranged the reference information is arranged in samples of a predetermined track the predetermined track is one of the tracks in which the plurality of types of audio data are divided and arranged An information processing apparatus.

14. An information processing apparatus includes a playback step of playing back the audio data of a predetermined track from a file including a plurality of types of audio data stored in individual tracks including the audio data forms groups based on relevance a plurality of types of audio data are divided and arranged in the tracks for each of one or more of the types in the file, reference information to the tracks corresponding to the plurality of types is arranged the reference information is arranged in samples of a predetermined track the predetermined track is one of the tracks in which the plurality of types of audio data are divided and arranged An information processing method.

15. An information processing apparatus includes a playback step of playing back the audio data of a predetermined track from a file including a plurality of types of audio data stored in individual tracks including the audio data forms groups based on relevance a plurality of types of audio data are divided and arranged in the tracks for each of one or more of the types ​ The information regarding the plurality of types is arranged in the sample entries of a predetermined track, The predetermined track is one of the tracks in which the plurality of types of audio data are divided and arranged. Information processing method. **Claim 16**: An information processing apparatus, A playback step of playing back the audio data of a predetermined track from a file including a plurality of types of audio data stored in individual tracks comprising The audio data forms groups based on relevance, The plurality of types of audio data are divided and arranged in the tracks for each of one or more of the types, In the file, for each track, information regarding the type corresponding to that track is arranged, In the file, for each track, information regarding an exclusive playback type consisting of the type corresponding to that track and the type corresponding to the audio data that is exclusively played back with the audio data of that type is arranged, A management file for managing the file is generated, the management file including information indicating whether information regarding the exclusive playback type exists for each track. Information processing method.

Citation Information

Patent Citations

  • Recording and reproducing device and group management method

    JP2002245751A

  • Information recording medium and device for reproducing the same

    JP2002313030A

  • Data generation device and data generation method, data processing device and data processing method

    JP2012033243A