Information processing apparatus and information processing method

By dividing audio data into files according to type and track, the problem of low audio data acquisition efficiency in existing technologies is solved, and efficient acquisition and generation of audio data is achieved.

CN114242081BActive Publication Date: 2025-11-07SONY GROUP CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111197608.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2014-10-01
Filing Date
2015-05-22
Publication Date
2025-11-07
Estimated Expiration
2035-05-22

AI Technical Summary

Technical Problem

Existing technologies have failed to effectively improve the efficiency of obtaining predetermined types of audio data from multiple types of audio data in video content.

Method used

By dividing various types of audio data according to type and track, generating multiple track files, and processing them through computer programs, efficient acquisition and generation of audio data can be achieved.

Benefits of technology

It improves the efficiency of acquiring specific types of audio data from multiple types of audio data, and realizes efficient acquisition of audio data and file generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114242081B_ABST
    Figure CN114242081B_ABST
Patent Text Reader

Abstract

The present disclosure discloses an information processing apparatus and an information processing method. An information processing apparatus including an acquisition unit that acquires audio data of a predetermined track in a file in which a plurality of types of audio data are divided into a plurality of tracks according to the types and the tracks are arranged. In the present invention, the audio data of the predetermined track in the file arranged by separating the plurality of types of audio data into the plurality of tracks according to the types is acquired. The present invention can be applied to, for example, a file generation apparatus for generating a file, a Web server for recording a file generated by the file generation apparatus, or an information processing system configured of a video playback terminal for playing a file.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the parent application with the application number 2015800269311 and the filing date of May 22, 2015, and the title of "Information Processing Apparatus and Information Processing Method". TECHNICAL FIELD

[0002] The present disclosure relates to an information processing apparatus and an information processing method, and more particularly, to an information processing apparatus and an information processing method capable of improving efficiency of acquiring a predetermined type of audio data among a plurality of types of audio data. BACKGROUND

[0003] One of the most popular streaming services recently is Internet video (OTT-V) via the Internet. Moving Picture Experts Group phase dynamic adaptive streaming over HTTP (MPEG-DASH) is widely used as its underlying technology (see, for example, Non-Patent Literature 1).

[0004] In MPEG-DASH, a delivery server prepares a set of video data with different screen sizes and encoding rates for one video content item, and a playback terminal requests a set of video data with the best screen size and encoding rate according to a transmission line condition, thus realizing adaptive streaming.

[0005] List of Citation Documents

[0006] Non-Patent Literature

[0007] Non-Patent Literature 1: MPEG-DASH (Dynamic Adaptive Streaming over HTTP) (URL: http: / / mpeg.chiariglione.org / standards / mpeg-dash / media-presentation-description-and-segment-formats / text-isoiec-23009-12012-dam-1) SUMMARY

[0008] Problems to be Solved by the Invention

[0009] However, there is no consideration to improve efficiency of acquiring a predetermined type of audio data among a plurality of types of audio data of a video content.

[0010] The present disclosure is made in view of the above circumstances and is capable of improving efficiency of acquiring a predetermined type of audio data among a plurality of types of audio data.

[0011] Solution to Problem

[0012] The information processing apparatus according to the first aspect of the present disclosure is an information processing apparatus including an acquisition unit that acquires audio data in a predetermined track of a file, wherein a plurality of types of audio data are divided into a plurality of tracks according to types and the tracks are arranged.

[0013] The information processing method according to the first aspect of the present disclosure corresponds to the information processing apparatus according to the first aspect of the present disclosure.

[0014] In the first aspect of the present disclosure, audio data in a predetermined track of a file is acquired, wherein a plurality of types of audio data are divided into a plurality of tracks according to types and tracks that are arranged.

[0015] The information processing apparatus according to the second aspect of the present disclosure is an information processing apparatus including a generation unit that generates a file in which a plurality of types of audio data are divided into a plurality of tracks according to types and tracks that are arranged.

[0016] The information processing method according to the second aspect of the present disclosure corresponds to the information processing apparatus according to the second aspect of the present disclosure.

[0017] In the second aspect of the present disclosure, a file in which a plurality of types of audio data are divided into a plurality of tracks according to types and tracks that are arranged is generated.

[0018] Note that the information processing apparatus according to the first aspect and the second aspect can be implemented by causing a computer to execute a program.

[0019] Further, in order to realize the information processing apparatus according to the first aspect and the second aspect, the program executed by a computer can be provided by transferring the program via a transmission medium or by recording the program in a recording medium.

[0020] Effects of the Invention

[0021] According to the first aspect of the present disclosure, audio data can be acquired. Further, according to the first aspect of the present disclosure, a specific type of audio data among a plurality of types of audio data can be efficiently acquired.

[0022] According to the second aspect of the present disclosure, a file can be generated. Further, according to the second aspect of the present disclosure, a file that improves the efficiency of acquiring a specific type of audio data among a plurality of types of audio data can be generated. BRIEF DESCRIPTION OF DRAWINGS

[0023] Figure 1 A schematic diagram showing an overview of a first example of an information processing system to which the present disclosure is applied.

[0024] Figure 2 A schematic diagram showing an example of a file.

[0025] Figure 3 A diagram for showing a subject.

[0026] Figure 4 A diagram for showing a subject position information.

[0027] Figure 5 A diagram for showing an image frame size information.

[0028] Figure 6 A diagram for showing a structure of an MPD file.

[0029] Figure 7 A diagram for showing a relationship among a "Period", a "Representation", and a "Segment".

[0030] Figure 8 A diagram for showing a hierarchical structure of an MPD file.

[0031] Figure 9 A diagram for showing a relationship between a structure of an MPD file and a time axis.

[0032] Figure 10 A diagram for showing an exemplary description of an MPD file.

[0033] Figure 11 A block diagram for showing a configuration example of a file generation apparatus.

[0034] Figure 12 A flowchart for showing a file generation process of a file generation apparatus.

[0035] Figure 13 A block diagram for showing a configuration example of a stream playback unit.

[0036] Figure 14 A flowchart for showing a stream playback process of a stream playback unit.

[0037] Figure 15 A diagram for showing an exemplary description of an MPD file.

[0038] Figure 16 A diagram for showing another exemplary description of an MPD file.

[0039] Figure 17 A diagram for showing an arrangement example of an audio stream.

[0040] Figure 18 A diagram for showing an exemplary description of gsix.

[0041] Figure 19 A diagram for showing an example of information indicating a correspondence relationship between a sample group entry and an object ID.

[0042] Figure 20 A schematic diagram showing an exemplary description of an AudioObjectSampleGroupEntry.

[0043] Figure 21 A schematic diagram showing an exemplary description of a type assignment box.

[0044] Figure 22 A schematic diagram showing an overview of a second example of an information processing system to which the present disclosure is applied.

[0045] Figure 23 A block diagram showing a configuration example of a stream playback unit of an information processing system to which the present disclosure is applied.

[0046] Figure 24 A schematic diagram showing a method of determining a position of an object.

[0047] Figure 25 A schematic diagram showing a method of determining a position of an object.

[0048] Figure 26 A schematic diagram showing a method of determining a position of an object.

[0049] Figure 27 A schematic diagram showing a relationship between a horizontal angle θ Ai and a horizontal angle θ Ai .

[0050] Figure 28 A flowchart showing a stream playback process of the stream playback unit shown in Figure 23 .

[0051] Figure 29 A flowchart showing details of the position determination process shown in Figure 28 .

[0052] Figure 30 A flowchart showing details of the horizontal angle θ Ai ' estimation process shown in Figure 29 .

[0053] Figure 31 A schematic diagram showing an overview of a track of a 3D audio file format of MP4.

[0054] Figure 32 A schematic diagram showing a structure of a moov box.

[0055] Figure 33 A schematic diagram showing an overview of a track according to a first embodiment to which the present disclosure is applied.

[0056] Figure 34 A flowchart showing a stream playback process of the stream playback unit shown in Figure 33The diagram shows an exemplary syntax of sample entries for the basic track.

[0057] Figure 35 To show in Figure 33 The diagram shows an example syntax of sample entries for the audio track.

[0058] Figure 36 To show in Figure 33 The diagram shows an example syntax of sample entries for an object audio track.

[0059] Figure 37 To show in Figure 33 The diagram shows an example syntax of sample entries for a HOA audio track.

[0060] Figure 38 To show in Figure 33 The diagram illustrates an example syntax for sample entries of the object metadata track.

[0061] Figure 39 This is a schematic diagram illustrating a first example of a fragment structure.

[0062] Figure 40 This is a schematic diagram illustrating a second example of a fragment structure.

[0063] Figure 41 A schematic diagram illustrating an exemplary description of a level assignment box.

[0064] Figure 42 This is a schematic diagram illustrating an exemplary description of an MDF file in a first embodiment of the present disclosure.

[0065] Figure 43 This is a diagram illustrating the definition of basic attributes.

[0066] Figure 44 This is a schematic diagram illustrating an overview of an information processing system in which the present disclosure is applied in a first embodiment.

[0067] Figure 45 To show in Figure 44 A block diagram showing an example configuration of a file generation device.

[0068] Figure 46 To show in Figure 45 The flowchart shown is a process for generating files using a file generation device.

[0069] Figure 47 To show by in Figure 44 The diagram shows a configuration example of a streaming playback unit implemented in a video playback terminal.

[0070] Figure 48 To show inFigure 47 a flowchart of an object specifying process of the stream playback unit shown in

[0071] Figure 49 to show an outline of a track in the second embodiment of the present disclosure. Figure 47 a flowchart of an object specifying process of the stream playback unit shown in

[0072] Figure 50 to show an outline of a track in the second embodiment of the present disclosure. Figure 47 a flowchart of a specified object audio playback process of the stream playback unit shown in

[0073] Figure 51 to show an outline of a track in the second embodiment of the present disclosure.

[0074] Figure 52 to show an example syntax of a sample entry of a track shown in Figure 51

[0075] Figure 53 to show an outline of a track in the second embodiment of the present disclosure.

[0076] Figure 54 to show an example syntax of a sample entry of a track shown in

[0077] Figure 55 to show an example of data of an extractor.

[0078] Figure 56 to show an outline of a track in the second embodiment of the present disclosure.

[0079] Figure 57 to show an outline of a track in the second embodiment of the present disclosure.

[0080] Figure 58 to show an example description of an MDF file in the fourth embodiment of the present disclosure.

[0081] Figure 59 to show an outline of a track in the second embodiment of the present disclosure.

[0082] Figure 60 to show a configuration example of a file generation apparatus shown in Figure 59

[0083] Figure 61 to show a file generation process of a file generation apparatus shown in Figure 60

[0084] to show a file generation process of a file generation apparatus shown in Figure 62 Figure 59 ​​​A block diagram showing a configuration example of a stream playback unit implemented by a video playback terminal.

[0085] Figure 63 A flowchart showing an example of a channel audio playback process of a stream playback unit shown in Figure 62

[0086] Figure 64 A flowchart showing a first example of an object audio playback process of a stream playback unit shown in Figure 62

[0087] Figure 65 A flowchart showing a second example of an object audio playback process of a stream playback unit shown in Figure 62

[0088] Figure 66 A flowchart showing a third example of an object audio playback process of a stream playback unit shown in Figure 62

[0089] Figure 67 A schematic diagram showing an example of selection of an object based on priority.

[0090] Figure 68 A schematic diagram showing an overview of a track in application of a fifth embodiment of the present disclosure.

[0091] Figure 69 A schematic diagram showing an overview of a track in application of a sixth embodiment of the present disclosure.

[0092] Figure 70 A schematic diagram showing a hierarchical structure of 3D audio.

[0093] Figure 71 A schematic diagram showing a first example of a web server process.

[0094] Figure 72 A flowchart showing a track division process of a web server.

[0095] Figure 73 A schematic diagram showing a first example of a process of an audio decoding processing unit.

[0096] Figure 74 A flowchart showing details of a first example of a decoding process of an audio decoding processing unit.

[0097] Figure 75 A schematic diagram showing a second example of a process of an audio decoding processing unit.

[0098] Figure 76 A flowchart showing details of a second example of a decoding process of an audio decoding processing unit.​​​​

[0099] Figure 77 A schematic diagram showing a second example of the process of the Web server.

[0100] Figure 78 A schematic diagram showing a third example of the process of the audio decoding processing unit.

[0101] Figure 79 A flowchart showing details of a third example of the decoding process of the audio decoding processing unit.

[0102] Figure 80 A schematic diagram showing a second example of the syntax of the configuration information set in the elementary sample.

[0103] Figure 81 An example syntax of the configuration information for the Ext element shown in Figure 80

[0104] Figure 82 A schematic diagram showing an example syntax of the configuration information for the extractor shown in Figure 81

[0105] Figure 83 A schematic diagram showing a second example of the data syntax of the frame unit set in the elementary sample.

[0106] Figure 84 A schematic diagram showing an example data syntax of the extractor shown in Figure 83

[0107] Figure 85 A schematic diagram showing a third example of the syntax of the configuration information set in the elementary sample.

[0108] Figure 86 A schematic diagram showing a third example of the data syntax of the frame unit set in the elementary sample.

[0109] Figure 87 A schematic diagram showing a configuration example of the audio stream in the seventh embodiment of the information processing system of the present disclosure.

[0110] Figure 88 A schematic diagram showing an overview of the track in the seventh embodiment.

[0111] Figure 89 A flowchart showing the file generation process in the seventh embodiment.

[0112] Figure 90 A flowchart showing the audio playback process in the seventh embodiment.

[0113] Figure 91 ​​​FIG. 8 is a diagram showing an outline of a track in the eighth embodiment of the information processing system according to the present disclosure.

[0114] Figure 92 FIG. 9 is a diagram showing a configuration example of an audio file.

[0115] Figure 93 FIG. 10 is a diagram showing another configuration example of an audio file.

[0116] Figure 94 FIG. 11 is a diagram showing yet another configuration example of an audio file.

[0117] Figure 95 FIG. 12 is a block diagram showing a configuration example of hardware of a computer. DETAILED DESCRIPTION

[0118] Modes for carrying out the present disclosure (hereinafter, referred to as embodiments) will be described below in the following order.

[0119] 0. Foretold of the present disclosure Figures 1 to 30

[0120] 1. First Embodiment Figures 31 to 50

[0121] 2. Second Embodiment Figures 51 to 55

[0122] 3. Third Embodiment Figure 56

[0123] 4. Fourth Embodiment Figures 57 to 67

[0124] 5. Fifth Embodiment Figure 68

[0125] 6. Sixth Embodiment Figure 69

[0126] 7. Explanation of the hierarchical structure of 3D audio Figure 70

[0127] 8. Explanation of the first example of the Web server process Figure 71 72

[0128] 9. Explanation of the first example of the process of the audio decoding processing unit Figure 73 74

[0129] 10. Explanation of the second example of the process of the audio decoding processing unit Figure 75 76

[0130] ​​​​​​​​​​​​​​11. Description of the second example of a web server procedure ( Figure 77 )

[0131] 12. Explanation of the third example of the audio decoding processing unit process ( Figure 78 and 79 )

[0132] 13. A second example of the syntax of the basic sample ( Figures 80 to 84 )

[0133] 14. The third example of the syntax of the basic sample ( Figure 85 and 86 )

[0134] 15. Seventh Embodiment ( Figures 87 to 90 )

[0135] 16. Eighth embodiment ( Figures 91 to 94 )

[0136] 17. Ninth Embodiment ( Figure 95 )

[0137] <Preview of this disclosure>

[0138] (An overview of the first example of an information processing system)

[0139] Figure 1 A schematic diagram illustrating an overview of a first example of an information processing system applying this disclosure.

[0140] like Figure 1 The information processing system 10 shown has a configuration in which a web server 12 (which is connected to a file generation device 11) and a video playback terminal 14 are connected via the Internet 13.

[0141] In the information processing system 10, the web server 12 transmits image data of video content in units of tiles (tile streaming) to the video playback terminal 14 using a method compatible with MPEG-DASH.

[0142] Specifically, the file generation device 11 acquires image data of the video content and encodes the image data in units of tiles to generate a video stream. The file generation device 11 processes the video stream of each tile into a file format with time intervals ranging from a few seconds to about ten seconds; this file format is referred to as a segment. The file generation device 11 uploads the resulting image file of each tile to the web server 12.

[0143] Further, the file generation apparatus 11 acquires audio data of the video content of each object (described later in detail) and encodes the image data in units of objects to generate an audio stream. The file generation apparatus 11 processes the audio stream of each object into a file format in units of segments, and uploads the resulting audio file of each object to the Web server 12.

[0144] Note that each object is a sound source. The audio data of each object is acquired by a microphone or the like attached to the object. The object can be an object such as a fixed microphone stand or can be a moving object such as a person.

[0145] The file generation apparatus 11 encodes audio metadata containing object position information (audio position information) indicating the position of each object (the position at which the audio data is acquired) and an object ID that is a unique ID of the object. The file generation apparatus 11 processes the encoded data obtained by encoding the audio metadata into a file format in units of segments, and uploads the resulting audio meta file to the Web server 12.

[0146] Further, the file generation apparatus 11 generates a Media Presentation Description (MPD) file (control information) that manages the image files and the audio files and contains image frame size information indicating the frame size of the images of the video content and position information indicating the position of each tile on the image. The file generation apparatus 11 uploads the MPD file to the Web server 12.

[0147] The Web server 12 stores the image files, the audio files, the audio meta files, and the MPD file uploaded from the file generation apparatus 11.

[0148] In the example as shown in FIG. 6, the Web server 12 stores a segment group of a plurality of segments composed of image files of the tile with the tile ID "1" and a segment group of a plurality of segments composed of image files of the tile with the tile ID "2". The Web server 12 also stores a segment group of a plurality of segments composed of audio files of the object with the object ID "1" and a segment group of a plurality of segments composed of audio files of the object with the object ID "2". Although not shown, a segment group composed of audio meta files is also similarly stored. Figure 1 Note that the file with the tile ID i is hereinafter referred to as "tile #i", and the object with the object ID i is hereinafter referred to as "object #i".

[0149] The Web server 12 functions as a sender and sends the stored image files, audio files, audio meta files, MPD file, and the like to the video playback terminal 14 in response to a request from the video playback terminal 14.

[0150] The Web server 12 functions as a sender and sends the stored image files, audio files, audio meta files, MPD file, and the like to the video playback terminal 14 in response to a request from the video playback terminal 14.

[0151] The video playback terminal 14 executes, for example, software 21 for controlling streaming data (hereinafter referred to as control software), video playback software 22, and client software 23 for Hyper Text Transfer Protocol (HTTP) access (hereinafter referred to as access software).

[0152] The control software 21 is software for controlling data delivered from the Web server 12 via streaming. Specifically, the control software 21 allows the video playback terminal 14 to acquire an MPD file from the Web server 12.

[0153] Further, the control software 21 specifies a tile in a display region based on the display region and tile position information included in the MPD file, the display region being a region in an image for displaying video content instructed by the video playback software 22. The control software 21 instructs the access software 23 to issue a request for transmitting an image file of the specified tile.

[0154] Further, the control software 21 instructs the access software 23 to issue a request for transmitting an audio element file. The control software 21 specifies an object corresponding to an image in a display region based on the display region, image frame size information included in the MPD file, and object position information included in the audio element file. The control software 21 instructs the access software 23 to issue a request for transmitting an audio file of the specified object.

[0155] The video playback software 22 is software for playing back image files and audio files acquired from the Web server 12. Specifically, when a display region is specified by a user, the video playback software 22 instructs the control software 21 of the specified display region. The video playback software 22 decodes the image files and audio files acquired from the Web server 12 in response to the instruction, and the video playback software 22 synthesizes and outputs the decoded files.

[0156] The access software 23 is software for controlling communication with the Web server 12 via the Internet 13 using HTTP. Specifically, the access software 23 allows the video playback terminal 14 to issue a request for transmitting an image file, an audio file, and an audio element file in response to an instruction of the control software 21. Further, the access software 23 allows the video playback terminal 14 to receive the image file, the audio file, and the audio element file transmitted from the Web server 12 in response to the request for transmission.

[0157] (Example of a tile)

[0158] Figure 2 A schematic view for showing an example of a tile.

[0159] As shown in Figure 2 , an image of video content is divided into a plurality of tiles. A tile ID, which is a sequential number starting from 1, is assigned to each tile. In Figure 2In the example shown, the image of the video content is divided into four tiles #1 to #4.

[0160] (Explanation of the object)

[0161] Figure 3 A schematic diagram illustrating the object.

[0162] Figure 3 The example shows the capture of eight audio objects from an image as audio for video content. Object IDs, starting from 1, are assigned to each object as sequential numbers. Objects #1 through #5 are moving objects, and objects #6 through #8 are stationary objects. Furthermore, in Figure 3 In the example, the image of the video content is divided into 7 (width) × 5 (height) tiles.

[0163] In this case, such as Figure 3 As shown, when the user specifies a display area 31 consisting of 3 (width) × 2 (height) tiles, the display area 31 only contains objects #1, #2, and #6. Therefore, the video playback terminal 14 only retrieves and plays audio files for objects #1, #2, and #6 from the web server 12.

[0164] The objects in the display area 31 can be specified based on image frame size information and object position information, as described below.

[0165] (Explanation of object location information)

[0166] Figure 4 A schematic diagram illustrating the location information of an object.

[0167] like Figure 4 As shown, the object position information includes the horizontal angle θ of object 40. A (-180°≤θ A ≤180°), vertical angle γ A (-90°≤γ A ≤90°) and distance r A (0 <r A For example, in the following settings, the horizontal angle θA is the angle in the horizontal direction formed by the line connecting object 40 and the origin O and the YZ plane: the shooting position of the image center can be set as the origin (base point) O; the horizontal direction of the image is set as the X direction; the vertical direction of the image is set as the Y direction; and the depth direction perpendicular to the XY plane is set as the Z direction. Vertical angle γ A The angle in the vertical direction formed by the straight line connecting object 40 and the origin O and the XZ plane. Distance r A This is the distance between object 40 and the origin O.

[0168] Further, in this document, the angle of rotation to the left and up is set to a positive angle, and the angle of rotation to the right and down is set to a negative angle.

[0169] (Explanation of image frame size information)

[0170] Figure 5 A diagram is shown to illustrate the image frame size information.

[0171] As Figure 5 shown, the image frame size information includes a horizontal angle θ v1 of the left end, a horizontal angle θ v2 of the right end, a vertical angle γ v1 of the upper end, a vertical angle γ v2 of the lower end, and a distance r v .

[0172] For example, when a photographing position at the center of the image is set to an origin O; a horizontal direction of the image is set to an X direction; a vertical direction of the image is set to a Y direction; and a depth direction perpendicular to the XY plane is set to a Z direction, the horizontal angle θ v1 is an angle in the horizontal direction formed by a straight line connecting the left end of the image frame and the origin O and the YZ plane. The horizontal angle θ v2 is an angle in the horizontal direction formed by a straight line connecting the right end of the image frame and the origin O and the YZ plane. Thus, an angle obtained by combining the horizontal angle θ v1 and the horizontal angle θ v2 becomes a horizontal viewing angle.

[0173] The vertical angle γ V1 is an angle formed by the XZ plane and a straight line connecting the upper end of the image frame and the origin O, and the vertical angle γ v2 is an angle formed by the XZ plane and a straight line connecting the lower end of the image frame and the origin O. An angle obtained by combining the vertical angle γ V1 and γ v2 becomes a vertical viewing angle. The distance r v is a distance between the origin O and the image plane.

[0174] As described above, the object position information indicates a positional relationship between the object 40 and the origin O, and the image frame size information indicates a positional relationship between the image frame and the origin O. Thus, it is possible to detect (recognize) a position of each object on the image based on the object position information and the image frame size information. Thus, it is possible to specify the object in the display region 31.

[0175] (Explanation of structure of MPD file)

[0176] Figure 6 A diagram is shown to illustrate the structure of the MPD file.

[0177] In the analysis (parsing) of the MPD file, the video playback terminal 14 selects the optimum attribute from among the attributes of the "Representation" contained in the "Period" of the MPD file (in Figure 6 MediaPresentation).

[0178] By referring to the uniform resource locator (URL) of the "Initialization Segment" at the head of the selected "Representation", and the like, the video playback terminal 14 acquires the file and processes the acquired file. Next, by referring to the URL of the subsequent "Media Segment", and the like, the video playback terminal 14 acquires the file and plays the acquired file.

[0179] Note that, in the MPD file, the relationship among the "Period", the "Representation", and the "Segment" becomes as shown in Figure 7 In other words, a single video content item can be managed in a longer time unit than a segment by the "Period", and can be managed in a segment unit by the "Segment" in each "Period". Furthermore, in each "Period", the video content can be managed in a stream attribute unit by the "Representation".

[0180] Thus, the MPD file has a hierarchical structure starting from the "Period" as shown in Figure 8 Furthermore, the structure of the MPD file arranged on the time axis becomes a configuration as shown in Figure 9 It is clear from Figure 9 that there are multiple "Representation" elements in the same segment. The video playback terminal 14 adaptively selects any one from among these elements, and thus can acquire the image file and the audio file in the display area selected by the user and play the acquired files.

[0181] (Explanation of the description of the MPD file)

[0182] Figure 10 A schematic diagram showing the description of the MPD file.

[0183] As described above, in the information processing system 10, the image frame size information is contained in the MPD file to allow the object in the display area to be specified by the video playback terminal 14. AsFigure 10 As shown, the scheme (urn:mpeg:DASH:viewingAngle:2013) for defining new image frame size information (viewing angle) is extended by utilizing the DescriptorType element of Viewpoint, and thus the image frame size information is arranged in the "Adaptation Set" for audio and in the "Adaptation Set" for images. The image frame size information can be arranged only in the "Adaptation Set" for images.

[0184] Further, the "Representation" for the audio metadata file is described in the "Adaptation Set" for audio of the MPD file. The URL or the like as information for specifying the audio metadata file (audio metadata.mp4) is described in the "Segment" of the "Representation". In this case, the file to be specified in the "Segment" is described as the audio metadata file (object audio metadata) utilizing the Role element.

[0185] The "Representation" for the audio metadata file of each object is also described in the "Adaptation Set" for audio of the MPD file. The URL or the like as information for specifying the audio file of each object (audioObjel.mp4, audioObjel5.mp4) is described in the "Segment" of the "Representation". In this case, the object ID (1 and 5) of the object corresponding to the audio file is also described by the extended Viewpoint.

[0186] Note that although not shown, the tile position information is arranged in the "Adaptation Set" for images.

[0187] (Configuration example of file generation apparatus)

[0188] Figure 11 A block diagram showing a configuration example of the file generation apparatus 11 shown in Figure 1 FIG. 1.

[0189] As shown in Figure 11The illustrated file generation apparatus 11 includes a screen split processing unit 51, an image encoding processing unit 52, an image file generation unit 53, an image information generation unit 54, an audio encoding processing unit 55, an audio file generation unit 56, an MPD generation unit 57, and a server upload processing unit 58.

[0190] The screen split processing unit 51 of the file generation apparatus 11 splits image data of a video content inputted from the outside into tile units. The screen split processing unit 51 provides tile position information to the image information generation unit 54. Further, the screen split processing unit 51 provides image data configured in tile units to the image encoding processing unit.

[0191] The image encoding processing unit 52 encodes image data (configured in tile units and provided from the screen split processing unit 51) for each tile to generate a video stream. The image encoding processing unit 52 provides the video stream of each tile to the image file generation unit 53.

[0192] The image file generation unit 53 processes the video stream of each tile provided from the image encoding processing unit 52 into a file format in segment units and provides the resulting image file of each tile to the MPD generation unit 57.

[0193] The image information generation unit 54 provides the tile position information provided from the screen split processing unit 51 and image frame size information inputted from the outside as image information to the MPD generation unit 57.

[0194] The audio encoding processing unit 55 encodes audio data configured in object units of a video content inputted from the outside for each object and generates an audio stream. Further, the audio encoding processing unit 55 encodes object position information of each object inputted from the outside and audio metadata containing an object ID or the like to generate encoded data. The audio encoding processing unit 55 provides the audio stream of each object and the encoded data of the audio metadata to the audio file generation unit 56.

[0195] The audio file generation unit 56 functions as an audio file generation unit, processes the audio stream of each object provided from the audio encoding processing unit 55 into a file format in segment units and provides the resulting audio file of each object to the MPD generation unit 57.

[0196] Further, the audio file generation unit 56 functions as a meta file generation unit, processes the encoded data of the audio metadata provided from the audio encoding processing unit 55 into a file format in segment units and provides the resulting audio meta file to the MPD generation unit 57.

[0197] The MPD generation unit 57 determines the URL of the web server 12 for storing image files of each tile provided by the image file generation unit 53, etc. Furthermore, the MPD generation unit 57 determines the URL of the web server 12 for storing audio files and audio metafiles of each object provided by the audio file generation unit 56, etc.

[0198] The MPD generation unit 57 arranges the image information provided by the image information generation unit 54 in the "Adaptation Set" for the image used in the MPD file. Furthermore, the MPD generation unit 57 arranges the image frame size information within the image information blocks in the "Adaptation Set" for the audio used in the MPD file. The MPD generation unit 57 also arranges the URL of the image file for each tile in the "Segment" of the "Representation" for the image file of the tile.

[0199] The MPD generation unit 57 arranges the URL of each object's audio file in the "Segment" of the "Representation" for the object's audio file. Furthermore, the MPD generation unit 57 acts as an information generation unit and arranges URLs and other information as specified audio metafiles in the "Segment" of the "Representation" for the audio metafile. The MPD generation unit 57 provides the server upload processing unit 58 with MPD files, image files, audio files, and audio metafiles, wherein various types of information are arranged in the MPD file as described above.

[0200] The server upload processing unit 58 uploads the image file of each tile, the audio file of each object, the audio metafile, and the MPD file provided by the MPD generation unit 57 to the web server 12.

[0201] (Explanation of the process of the document generation device)

[0202] Figure 12 To show in Figure 11 The flowchart of the file generation process of the file generation device 11 shown is illustrated.

[0203] exist Figure 12 In step S11, the screen splitting processing unit 51 of the file generation device 11 splits the image data of the externally input video content into tile units. The screen splitting processing unit 51 provides tile position information to the image information generation unit 54. In addition, the screen splitting processing unit 51 provides image data configured in tile units to the image encoding processing unit 52.

[0204] In step S12, the image encoding processing unit 52 encodes the image data configured in units of tiles provided from the screen split processing unit 51 for each tile to generate a video stream of each tile. The image encoding processing unit 52 provides the image file generation unit 53 with the video stream of each tile.

[0205] In step S13, the image file generation unit 53 processes the video stream of each tile provided from the image encoding processing unit 52 into a file format in units of segments to generate an image file of each tile. The image file generation unit 53 provides the MPD generation unit 57 with the image file of each tile.

[0206] In step S14, the image information generation unit 54 acquires image frame size information from the outside. In step S15, the image information generation unit 54 generates image information containing the tile position information provided from the screen split processing unit 51 and the image frame size information, and provides the MPD generation unit 57 with the image information.

[0207] In step S16, the audio encoding processing unit 55 encodes audio data configured in units of objects of the video content inputted from the outside for each object and generates an audio stream of each object. Further, the audio encoding processing unit 55 encodes object position information of each object inputted from the outside and audio metadata containing an object ID to generate encoded data. The audio encoding processing unit 55 provides the audio file generation unit 56 with the audio stream of each object and the encoded data of the audio metadata.

[0208] In step S17, the audio file generation unit 56 processes the audio stream of each object provided from the audio encoding processing unit 55 into a file format in units of segments to generate an audio file of each object. Further, the audio file generation unit 56 processes the encoded data of the audio metadata provided from the audio encoding processing unit 55 into a file format in units of segments to generate an audio meta file. The audio file generation unit 56 provides the MPD generation unit 57 with the audio file of each object and the audio meta file.

[0209] In step S18, the MPD generation unit 57 generates an MPD file containing the image information provided from the image information generation unit 54, a URL of each file, and the like. The MPD generation unit 57 provides the server upload processing unit 58 with the MPD file, the image file of each tile, the audio file of each object, and the audio meta file.

[0210] In step S19, the server upload processing unit 58 uploads the image file of each tile, the audio file of each object, the audio meta file, and the MPD file provided from the MPD generation unit 57 to the Web server 12. Then, the process is terminated.

[0211] (Functional configuration example of video playback terminal)

[0212] Figure 13 To show a block diagram of a configuration example of a stream playback unit, the stream playback unit is implemented in a manner that the video playback terminal 14 shown in Figure 1 executes the control software 21, the video playback software 22, and the access software 23.

[0213] The stream playback unit 90 shown in Figure 13 includes an MPD acquisition unit 91, an MPD processing unit 92, a metafile acquisition unit 93, an audio selection unit 94, an audio file acquisition unit 95, an audio decoding processing unit 96, an audio synthesis processing unit 97, an image selection unit 98, an image file acquisition unit 99, an image decoding processing unit 100, and an image synthesis processing unit 101.

[0214] The MPD acquisition unit 91 of the stream playback unit 90 functions as a receiver, acquires an MPD file from the Web server 12, and provides the MPD file to the MPD processing unit 92.

[0215] The MPD processing unit 92 extracts information (such as a URL described in a “Segment” for an audio metafile) from the MPD file provided from the MPD acquisition unit 91, and provides the extracted information to the metafile acquisition unit 93. Further, the MPD processing unit 92 extracts image frame size information described in an “Adaptation Set” for an image from the MPD file, and provides the extracted information to the audio selection unit 94. The MPD processing unit 92 extracts information (such as a URL described in a “Segment” for an audio file of an object requested from the audio selection unit 94) from the MPD file, and provides the extracted information to the audio selection unit 94.

[0216] The MPD processing unit 92 extracts tile position information described in an “Adaptation Set” for an image from the MPD file, and provides the extracted information to the image selection unit 98. The MPD processing unit 92 extracts information (such as a URL described in a “Segment” for an image file of a tile requested from the image selection unit 98) from the MPD file, and provides the extracted information to the image selection unit 98.

[0217] Based on information such as a URL provided from the MPD processing unit 92, the metafile acquisition unit 93 requests the Web server 12 to transmit an audio metafile specified by the URL, and acquires the audio metafile. The metafile acquisition unit 93 provides the audio selection unit 94 with object position information contained in the audio metafile.

[0218] The audio selection unit 94 functions as a position determination unit and calculates the position of each object on an image based on image frame size information provided from the MPD processing unit 92 and object position information provided from the metafile acquisition unit 93. The audio selection unit 94 selects an object in a display region designated by the user based on the position of each object on the image. The audio selection unit 94 requests the MPD processing unit 92 to transmit information such as a URL of an audio file of the selected object. The audio selection unit 94 provides the audio file acquisition unit 95 with information such as a URL provided from the MPD processing unit 92 in response to the request.

[0219] The audio file acquisition unit 95 functions as a receiver. Based on information such as a URL provided from the audio selection unit 94, the audio file acquisition unit 95 requests the Web server 12 to transmit an audio file specified by the URL and configured in units of objects, and acquires the audio file. The audio file acquisition unit 95 provides the audio decoding processing unit 96 with the acquired audio file configured in units of objects.

[0220] The audio decoding processing unit 96 decodes an audio stream contained in the audio file provided from the audio file acquisition unit 95 and configured in units of objects, to generate audio data in units of objects. The audio decoding processing unit 96 provides the audio synthesis processing unit 97 with the audio data in units of objects.

[0221] The audio synthesis processing unit 97 synthesizes the audio data provided from the audio decoding processing unit 96 and configured in units of objects, and outputs the synthesized data.

[0222] The image selection unit 98 selects a tile in a display region designated by the user based on tile position information provided from the MPD processing unit 92. The image selection unit 98 requests the MPD processing unit 92 to transmit information such as a URL of an image file of the selected tile. The image selection unit 98 provides the image file acquisition unit 99 with information such as a URL provided from the MPD processing unit 92 in response to the request.

[0223] Based on information such as a URL provided from the image selection unit 98, the image file acquisition unit 99 requests the Web server 12 to transmit an image file specified by the URL and configured in units of tiles, and acquires the image file. The image file acquisition unit 99 provides the image decoding processing unit 100 with the acquired image file configured in units of tiles.

[0224] The image decoding processing unit 100 decodes the video stream (which is contained in the image file provided from the image file acquisition unit 99 and configured in units of tiles) to generate image data in units of tiles. The image decoding processing unit 100 provides the image data in units of tiles to the image composition processing unit 101.

[0225] The image composition processing unit 101 composes the image data provided from the image decoding processing unit 100 and configured in units of tiles and outputs the composed data.

[0226] (Explanation of the process of the moving image playback terminal)

[0227] Figure 14 A flowchart showing the flow of the stream playback process of the stream playback unit (90) of the video playback terminal 14. Figure 13

[0228] In step S31 of the stream playback unit (90), the MPD acquisition unit 91 acquires the MPD file from the Web server 12 and provides the MPD file to the MPD processing unit 92. Figure 14

[0229] In step S32, the MPD processing unit 92 acquires the image frame size information and the tile position information described in the "Adaptation Set" for images from the MPD file provided from the MPD acquisition unit 91. The MPD processing unit 92 provides the image frame size information to the audio selection unit 94 and the tile position information to the image selection unit 98. Further, the MPD processing unit 92 extracts information such as the URL described in the "Segment" for the audio meta file and provides the extracted information to the meta file acquisition unit 93.

[0230] In step S33, based on the information such as the URL provided from the MPD processing unit 92, the meta file acquisition unit 93 requests the Web server 12 to transmit the audio meta file designated by the URL and acquires the audio meta file. The meta file acquisition unit 93 provides the object position information contained in the audio meta file to the audio selection unit 94.

[0231] In step S34, the audio selection unit 94 selects the object in the display area designated by the user based on the image frame size information provided from the MPD processing unit 92 and the object position information provided from the meta file acquisition unit 93. The audio selection unit 94 requests the MPD processing unit 92 to transmit information such as the URL of the audio file of the selected object.

[0232] ​​The MPD processing unit 92 extracts information such as URLs described in "Segment" for an image file of an object requested from the image selection unit 98 from the MPD file, and provides the extracted information to the image selection unit 98. The image selection unit 98 provides information such as URLs provided from the MPD processing unit 92 to the image file acquisition unit 99.

[0233] In step S37, based on information such as URLs provided from the image selection unit 98, the image file acquisition unit 99 requests the Web server 12 to transmit an image file of a selected tile designated by the URLs, and acquires the image file. The image file acquisition unit 99 provides an image file acquired in a tile unit to the image decoding processing unit 100.

[0234] In step S36, the image selection unit 98 selects a tile in a display region designated by a user based on tile position information provided from the MPD processing unit 92. The image selection unit 98 requests the MPD processing unit 92 to transmit information such as a URL of an image file of a selected tile.

[0235] The MPD processing unit 92 extracts information such as URLs described in "Segment" for an image file of an object requested from the image selection unit 98 from the MPD file, and provides the extracted information to the image selection unit 98. The image selection unit 98 provides information such as URLs provided from the MPD processing unit 92 to the image file acquisition unit 99.

[0236] In step S37, based on information such as URLs provided from the image selection unit 98, the image file acquisition unit 99 requests the Web server 12 to transmit an image file of a selected tile designated by the URLs, and acquires the image file. The image file acquisition unit 99 provides an image file acquired in a tile unit to the image decoding processing unit 100.

[0237] In step S38, the audio decoding processing unit 96 decodes an audio stream included in an audio file provided from the audio file acquisition unit 95 and configured in an object unit, to generate audio data in an object unit. The audio decoding processing unit 96 provides the audio data in an object unit to the audio synthesis processing unit 97.

[0238] In step S39, the image decoding processing unit 100 decodes a video stream included in an image file provided from the image file acquisition unit 99 and configured in a tile unit, to generate image data in a tile unit. The image decoding processing unit 100 provides the image data in a tile unit to the image synthesis processing unit 101.

[0239] In step S40, the audio synthesis processing unit 97 synthesizes the audio data provided from the audio decoding processing unit 96 and configured in units of objects and outputs the synthesized data. In step S41, the image synthesis processing unit 101 synthesizes the image data provided from the image decoding processing unit 100 and configured in units of tiles and outputs the synthesized data. Then the process is terminated.

[0240] As described above, the Web server 12 transmits the image frame size information and the object position information. Therefore, the video playback terminal 14 can specify, for example, an object in a display region to selectively acquire an audio file of the specified object so that the audio file corresponds to an image in the display region. This allows the video playback terminal 14 to acquire only necessary audio files, which makes the transmission efficient.

[0241] Note that, as shown in Figure 15 , an object ID (object of the specification information) can be described in an "Adaptation Set" for an image of an MPD file as information for specifying an object corresponding to audio to be played simultaneously with the image. The object ID can be described by defining an extension scheme (urn:mpeg:DASH:audioObj:2013) of a new object ID information (audioObj) using a DescriptorType element of a Viewpoint. In this case, the video playback terminal 14 selects an audio file of an object corresponding to the object ID described in the "Adaptation Set" for an image of an MPD file and acquires the audio file for playback.

[0242] As an alternative to generating an audio file in units of objects, encoded data of all objects can be multiplexed into a single audio stream to generate a single audio file.

[0243] In this case, as shown in Figure 16 , one "Representation" for an audio file is set in an "Adaptation Set" for audio of an MPD file, and a URL or the like for an audio file (audioObj.mp4) containing encoded data of all objects is described in a "Segment". At this time, object IDs (1, 2, 3, 4, and 5) corresponding to all objects of the audio file are described by extending a Viewpoint.

[0244] In addition, in this case, as shown in Figure 17As shown, the encoded data (audio object) of each object is arranged as a subsample in the mdat box of the audio file (hereinafter also referred to as the audio media file where appropriate) obtained by referencing the "Media Segment" of the MPD file.

[0245] Specifically, data is arranged in the audio media file in units of sub-segments, each segment being shorter than a segment at any given time. The position of the data in sub-segments is specified by the sidx box. Furthermore, the data in sub-segments consists of moof boxes and mdat boxes. The mdat box comprises multiple samples, and the encoded data for each object is arranged as each subsample of that sample.

[0246] Furthermore, the gsix box describing information about the sample is placed after the sidx box of the audio media file. The gsix box describing information about the sample is thus separated from the moof box, allowing the video playback terminal 14 to quickly obtain information about the sample.

[0247] like Figure 18 As shown, the `grouping_type`, representing the type of a sample group entry, is described in the `gsix` box. Each sample group entry contains one or more samples or subsamples managed by the `gsix` box. For example, when the sample group entry is a subsample of encoded data in object units, the type of the sample group entry is as follows: Figure 17 The "obja" grouping_type is shown. Multiple gsix boxes are arranged in the audio media file.

[0248] In addition, such as Figure 18 As shown, the index (entry_index) and the byte range (range_size) of each sample group entry, which serves as data position information indicating its position in the audio media file, are described in the gsix box. It should be noted that when the index (entry_index) is 0, the corresponding byte range indicates the byte range of the moof box (in...). Figure 17 (a1 in the example).

[0249] The information indicating which object is used to allow each sample group entry to correspond to a subsample of the encoded data is described in the audio file obtained by referring to the “Initialization Segment” of the MPD file (hereinafter also appropriately referred to as the audio initialization file).

[0250] Specifically, such as Figure 19As shown in A of FIG. 10, the information is indicated by using a type assignment box (typa) of the mvex box, which is associated with an AudioObjectSampleGroupEntry of a sample group description box (sgpd) in the sbtl box of the audio initialization file.

[0251] In other words, as shown in Figure 20 A, an object ID (audio_object_id) corresponding to encoded data contained in a sample is described in each AudioObjectSampleGroupEntry box. For example, as shown in B, object IDs 1, 2, 3, and 4 are described in each of four AudioObjectSampleGroupEntry boxes. Figure 20

[0252] On the other hand, as shown in Figure 21 , in the type assignment box, an index as a grouping type parameter (grouping_type_parameter) of a sample group entry corresponding to the AudioObjectSampleGroupEntry is described for each AudioObjectSampleGroupEntry.

[0253] The audio media file and the audio initialization file are configured as described above. Therefore, when the video playback terminal 14 acquires encoded data of an object selected as an object in a display area, an AudioObjectSampleGroupEntry in which an object ID of the selected object is described is retrieved from the stbl box of the audio initialization file. Next, an index of a sample group entry corresponding to the retrieved AudioObjectSampleGroupEntry is read from the mvex box. Next, a position of data in a subsegment is read from the sidx of the audio file, and a byte range of the index of the sample group entry is read from the gsix. Next, encoded data arranged in the mdat is acquired based on the position of data in a subsegment and the byte range. Therefore, encoded data of the selected object is acquired.

[0254] Although in the above description, an index of a sample group and an object ID of an AudioObjectSampleGroupEntry are associated with each other through the mvex box, they can be directly associated with each other. In this case, an index of a sample group entry is described in the AudioObjectSampleGroupEntry.

[0255] Further, when the audio file is composed of a plurality of tracks, the sgpd can be stored in the mvex, which allows the sgpd to be shared between the tracks.​

[0256] (An overview of the second example of an information processing system)

[0257] Figure 22 A schematic diagram illustrating a second example of an information processing system applying this disclosure.

[0258] It should be pointed out that, in Figure 22 The one shown in the middle is the same as Figure 3 The same elements shown are represented by the same reference numerals.

[0259] exist Figure 22 As shown Figure 3 In the example, the image of the video content is divided into 7 (width) × 5 (height) tiles, and the audio of objects #1 to #8 is captured just like the audio of the video content.

[0260] In this case, when the user specifies the display area 31 consisting of 3 (width) × 2 (height) tiles, the display area 31 is converted (expanded) to an area with the same size as the image of the video content, thereby obtaining a display area such as... Figure 22 The second example shown is display image 111. The audio of objects #1 to #8 is synthesized based on the positions of objects #1 to #8 in display image 111 and is output together with display image 111. In other words, in addition to the audio of objects #1, #2, and #6 within display area 31, the audio of objects #3 to #5, #7, and #8 outside display area 31 is also output.

[0261] (Configuration example of a streaming playback unit)

[0262] Configuration of a second example of the information processing system applying this disclosure Figure 1 The information processing system 10 shown has the same configuration, except for the configuration of the streaming playback unit, which will therefore be described only below.

[0263] Figure 23 A block diagram illustrating a configuration example of a streaming playback unit of an information processing system applying this disclosure.

[0264] exist Figure 23 The one shown in the middle is the same as Figure 13 The same components shown are represented by the same reference numerals, and repeated explanations are omitted where appropriate.

[0265] like Figure 23 The configuration of the streaming playback unit 120 shown is different from that of... Figure 13The configuration of the stream playback unit 90 shown is such that the MPD processing unit 121, the audio synthesis processing unit 123, and the image synthesis processing unit 124 are newly provided to replace the MPD processing unit 92, the audio synthesis processing unit 97, and the image synthesis processing unit 101, respectively, and a position determination unit 122 is additionally provided.

[0266] The MPD processing unit 121 of the stream playback unit 120 extracts information such as URLs described in "Segment" for audio meta files from the MPD file supplied from the MPD acquisition unit 91 and supplies the extracted information to the meta file acquisition unit 93. Further, the MPD processing unit 121 extracts image frame size information of images of video contents (hereinafter, referred to as content image frame size information) described in "Adaptation Set" for images from the MPD file and supplies the extracted information to the position determination unit 122. The MPD processing unit 121 extracts information such as URLs described in "Segment" for audio files of all objects from the MPD file and supplies the extracted information to the audio file acquisition unit 95.

[0267] The MPD processing unit 121 extracts tile position information described in "Adaptation Set" for images from the MPD file and supplies the extracted information to the image selection unit 98. The MPD processing unit 121 extracts information such as URLs described in "Segment" for image files of tiles requested from the image selection unit 98 from the MPD file and supplies the extracted information to the image selection unit 98.

[0268] The position determination unit 122 acquires object position information contained in the audio meta files obtained by the meta file acquisition unit 93 and the content image frame size information supplied from the MPD processing unit 121. Further, the position determination unit 122 acquires display area image frame size information as image frame size information of a display area designated by a user. The position determination unit 122 determines (recognizes) the position of each object in the display area based on the object position information, the content image frame size information, and the display area image frame size information. The position determination unit 122 supplies the determined position of each object to the audio synthesis processing unit 123.

[0269] The audio synthesis processing unit 123 synthesizes the object-unit audio data supplied from the audio decoding processing unit 96 based on the object positions supplied from the position determination unit 122. Specifically, the audio synthesis processing unit 123 determines the audio data assigned to each speaker for each object based on the object positions and the positions of each speaker of the output sound. The audio synthesis processing unit 123 synthesizes the audio data of each object for each speaker and outputs the synthesized audio data as the audio data of each speaker. The detailed description of the method of synthesizing the audio data of each object based on the object positions is disclosed in, for example, "Virtual Sound Source Positioning Using Vector Base Amplitude Panning" by Ville Pulkki, Journal of the AES, Vol. 45, No. 6, pp. 456-466, 1997.

[0270] The image synthesis processing unit 124 synthesizes the tile-unit image data supplied from the image decoding processing unit 100. The image synthesis processing unit 124 functions as a converter and converts the image size corresponding to the synthesized image data into the size of the video content to generate a display image. The image synthesis processing unit 124 outputs the display image.

[0271] (Explanation of the Object Position Determination Method)

[0272] Figures 24 to 26 Each of the above-described methods shows the object position determination method of the position determination unit 122 shown in Figure 23 .

[0273] The display region 31 is extracted from the video content and the size of the display region 31 is converted into the size of the video content so as to generate a display image 111. Therefore, the size of the display image 111 is equivalent to the size obtained by shifting the center C of the display region 31 to the center C' of the display image 111 as shown in Figure 24 , and by converting the size of the display region 31 into the size of the video content as shown in Figure 25 .

[0274] Therefore, the position determination unit 122 calculates the amount of shift θ in the horizontal direction when the center O of the display region 31 is shifted to the center O' of the display image 111 by the following equation (1) shift .

[0275] [Mathematical Equation 1]

[0276]

[0277] In equation (1), θ v1θ v2' represents a horizontal angle at the right end of the display area 31 included in the display area image frame size information. Further, θ v1 represents a horizontal angle at the left end in the content image frame size information, and θ v2 represents a horizontal angle at the right end in the content image frame size information.

[0278] Next, the position determination unit 122 calculates the displacement amount θ shift the horizontal angle θ v1_shift ' at the left end of the display area 31 after the center O of the display area 31 is displaced to the center O' of the display image 111, and the horizontal angle θ v2_shift ' at the right end thereof, by the following equation (2).

[0279] [Mathematical Equation 2]

[0280] θ v1_shift ' = mod(θ v1 + θ shift + 180°, 360°) - 180°

[0281] θ v2_shift ' = mod(θ v2 + θ shift + 180°, 360°) - 180°... (2)

[0282] According to equation (2), the horizontal angle θ v1_shift ' and the horizontal angle θ v2_shift ' are calculated so as not to exceed the range of -180° to 180°.

[0283] Note that, as described above, the display image 111 size is equivalent to the size obtained by displacing the center O of the display area 31 to the center O' of the display image 111 and by converting the size of the display area 31 to the size of the video content. Therefore, the following equation (3) satisfies the horizontal angle θ V1 and θ V2 .

[0284] [Mathematical Equation 3]

[0285]

[0286]

[0287] The position determination unit 122 calculates the displacement amount θ shift , the horizontal angle θ v1_shift ' and the horizontal angle θ v2_shiftand then calculates the horizontal angle of each object in the display image 111. Specifically, by using the displacement amount θ shift After the center C of the display region 31 is displaced to the center C' of the display image 111, the position determination unit 122 calculates the horizontal angle θ Ai_shift .

[0288] [Mathematical Formula 4]

[0289] θ Ai_shift = mod(θ Ai + θ shift + 180°, 360°) - 180°... (4)

[0290] In Equation (4), θAi denotes the horizontal angle of the object #i included in the object position information. Further, according to Equation (4), the horizontal angle θ Ai_shift is calculated so as not to exceed the range of -180° to 180°.

[0291] Next, when the object #i exists in the display region 31, that is, when the condition θ v2_shif < θ Ai_shift < θ v1_shift ' is satisfied, the position determination unit 122 calculates the horizontal angle θ A1 ' of the object #i in the display image 111 by Equation (5) below.

[0292] [Mathematical Formula 5]

[0293]

[0294] According to Equation (5), the horizontal angle θ A1 ' is calculated by expanding the distance between the position of the object #i in the display image 11 and the center C' of the display image 111 according to the ratio between the size of the display region 31 and the size of the display image 111.

[0295] On the other hand, when the object #i does not exist in the display region 31, that is, when the condition -180° ≤ θ Ai_shift ≤ θ v2_shift ' or θ v1_shift ' ≤ θ Ai_shift ≤ 180° is satisfied, the position determination unit 122 calculates the horizontal angle θ Ai ' of the object #i in the display image 111 by Equation (6) below.

[0296] [Mathematical Formula 6]

[0297]

[0298]

[0299] According to formula (6), as shown in Figure 26 , when the object #i exists at a position 151 on the right side of the display region 31 (-180° < θ Ai_shift ≤ θ v2_shift ), the horizontal angle θ Ai_shift ' is calculated by expanding the horizontal angle θ Ai according to the ratio between an angle Rl and an angle R2. Note that the angle Rl is an angle measured from the right end of the display image 111 to a position 154 just behind the viewer 153, and the angle R2 is an angle measured from the right end of the display region 31 whose center is displaced to the position 154.

[0300] Further, according to formula (6), when the object #i exists at a position 155 on the left side of the display region 31 (θ v1_shift ' < θ Ai_shift ≤ 180°), the horizontal angle θ Ai_shift ' is calculated by expanding the horizontal angle θ Ai according to the ratio between an angle R3 and an angle R4. Note that the angle R3 is an angle measured from the left end of the display image 111 to the position 154, and the angle R4 is an angle measured from the left end of the display region 31 whose center is displaced to the position 154.

[0301] In addition, the position determination unit 122 calculates the vertical angle γ Ai ' in a manner similar to the horizontal angle θ Ai '. Specifically, the position determination unit 122 calculates the displacement amount γ shift in the vertical direction when the center C of the display region 31 is displaced to the center C' of the display image 111 by the following formula (7).

[0302]

Mathematical Formula 7

[0303]

[0304] In formula (7), γ v1 ' indicates the vertical angle of the upper end of the display region 31 included in the display region image frame size information, and γ v2 ' indicates the vertical angle of the lower end. Further, γ v1 indicates the vertical angle of the upper end in the content image frame size information, and γ v2 indicates the vertical angle of the lower end in the content image frame size information.

[0305] Next, the position determination unit 122 uses the displacement amount γ shift, the vertical angle γ v1_shift ' at the lower end of the display region 31 v2_shift

[0306] [Equation 8]

[0307] γ v1_shift ' = mod(γ v1 + γ shift + 90°, 180°) - 90°

[0308] γ v2_shift ' = mod(γ v2 + γ shift + 90°, 180°) - 90° (8)

[0309] According to Equation (8), the vertical angle γ v1_shift ' and the vertical angle γ v2_shift ' are calculated so as not to exceed the range of -90° to 90°.

[0310] The position determination unit 122 calculates the displacement amount γ shift , the vertical angle γ v1_shift ', and the vertical angle γ v2_shift ' in the above-described manner, and then calculates the position of each object in the display image 111. Specifically, the position determination unit 122 calculates the vertical angle γ Ai_shift of the object #i after the center C of the display region 31 is displaced to the center C' of the display image 111 by using the displacement amount γ shift through Equation (9) below.

[0311] [Equation 9]

[0312] γ Ai_shift = mod(γ Ai + γ shift + 90°, 180°) - 90° (9)

[0313] In Equation (9), γAi denotes the vertical angle of the object #i included in the object position information. Further, according to Equation (9), the vertical angle γ Ai_shift is calculated so as not to exceed the range of -90° to 90°.

[0314] Next, the position determination unit 122 calculates the vertical angle γ A1 ' of the object #i in the display image 111 through Equation (10) below.

[0315] [Equation 10]

[0316]

[0317] Further, the position determining unit 122 determines the distance r A1 of the object #i in the display image 111 A1 . The position determining unit 122 provides the audio synthesis processing unit 123 with the horizontal angle θ Ai , the vertical angle γ A1 , and the distance r A1 of the object #i as the position of the object #i, as described above.

[0318] Figure 27 is a schematic diagram showing the relationship between the horizontal angle θ Ai and the horizontal angle θ Ai '.

[0319] In the graph of Figure 27 , the horizontal axis represents the horizontal angle θ Ai , and the vertical axis represents the horizontal angle θ Ai '.

[0320] As shown in Figure 27 , when the condition θ V2 ' < θ Ai < θ V1 ' is satisfied, the horizontal angle θ Ai is shifted by the shift amount θ shift and is expanded, and then the horizontal angle θ Ai becomes equal to the horizontal angle θ Ai '. Further, when the condition -180° ≤ θ Ai ≤ θ v2 ' or θ v1 ' ≤ θ Ai ≤ 180° is satisfied, the horizontal angle θ Ai is shifted by the shift amount θ shift and is reduced, and then the horizontal angle θ Ai becomes equal to the horizontal angle θ Ai '.

[0321] (Explanation of the process of the stream playback unit)

[0322] Figure 28 is a flowchart showing the stream playback process of the stream playback unit 120 shown in Figure 23 .

[0323] In step S131 of Figure 28 , the MPD acquiring unit 91 of the stream playback unit 120 acquires the MPD file from the Web server 12 and provides the MPD to the MPD processing unit 121.

[0324] In step S132, the MPD processing unit 121 acquires content image frame size information and tile position information described in "Adaptation Set" for an image from the MPD file supplied from the MPD acquisition unit 91. The MPD processing unit 121 supplies the image frame size information to the position determination unit 122 and supplies the tile position information to the image selection unit 98. Further, the MPD processing unit 121 extracts information such as a URL described in "Segment" for an audio meta file and supplies the extracted information to the meta file acquisition unit 93.

[0325] In step S133, the meta file acquisition unit 93 requests the Web server 12 to transmit an audio meta file designated by the URL based on the information such as the URL supplied from the MPD processing unit 121 and acquires the audio meta file. The meta file acquisition unit 93 supplies object position information contained in the audio meta file to the position determination unit 122.

[0326] In step S134, the position determination unit 122 performs a position determination process for determining the position of each object in a display image based on the object position information, the content image frame size information, and the display region image frame size information. The position determination process will be described in detail with reference to Figure 29 described later.

[0327] In step S135, the MPD processing unit 121 extracts information such as a URL described in "Segment" for an audio file for all objects from the MPD file and supplies the extracted information to the audio file acquisition unit 95.

[0328] In step S136, the audio file acquisition unit 95 requests the Web server 12 to transmit an audio file for all objects designated by the URL based on the information such as the URL supplied from the MPD processing unit 121 and acquires the audio file. The audio file acquisition unit 95 supplies the acquired audio file in units of objects to the audio decoding processing unit 96.

[0329] The processes of steps S137 to S140 are similar to the processes of steps S36 to S39 as Figure 14 shown in FIG. 4 and thus the description thereof will be omitted.

[0330] In step S141, the audio synthesis processing unit 123 synthesizes the audio data in units of objects supplied from the audio decoding processing unit 96 based on the position of each object supplied from the position determination unit 122 and outputs the audio data.

[0331] In step S142, the image synthesis processing unit 124 synthesizes the image data in tile units supplied from the image decoding processing unit 100.

[0332] In step S143, the image synthesis processing unit 124 converts the image size corresponding to the synthesized image data into the size of the video content and generates a display image. Next, the image synthesis processing unit 124 outputs the display image, and the process terminates.

[0333] Figure 29 A flowchart showing details of the position determination process in step S134 of Figure 28 will be described. The position determination process is executed, for example, for each object.

[0334] In step S151 of Figure 29 , the position determination unit 122 executes a horizontal angle θ Ai ' estimation process for estimating a horizontal angle θ Ai ' in the display image. Details of the horizontal angle θ Ai ' estimation process will be described with reference to Figure 30 described later.

[0335] In step S152, the position determination unit 122 executes a vertical angle γ Ai ' estimation process for estimating a vertical angle γ Ai ' in the display image. Details of the vertical angle γ Ai ' estimation process are similar to those of the horizontal angle θ Ai ' estimation process in step S151 except that the vertical direction is used instead of the horizontal direction, and thus a detailed description thereof will be omitted.

[0336] In step S153, the process determination unit 122 determines a distance r Ai ' in the display image as the distance r Ai contained in the object position information supplied from the meta file acquisition unit 93.

[0337] In step S154, the position determination unit 122 outputs the horizontal angle θ Ai ', the vertical angle γ A1 ', and the distance r A1 as the position of the object #i to the audio synthesis processing unit 123. Next, the process returns to step S135 of Figure 28 and proceeds to step S136.

[0338] Figure 30 A flowchart showing details of the horizontal angle θ Ai ' estimation process in step S151 of Figure 29 will be described.

[0339] In such Figure 30 In step S171 shown, the position determination unit 122 obtains the horizontal angle θ contained in the object position information provided by the metafile acquisition unit 93. Ai .

[0340] In step S172, the position determination unit 122 acquires the content image frame size information provided by the MPD processing unit 121 and the display area image frame size information specified by the user.

[0341] In step S173, the position determination unit 122 calculates the displacement θ based on the content image frame size information and the display area image frame size information using the formula (1) above. shift .

[0342] In step S174, the position determination unit 122 uses the displacement θ shift The horizontal angle θ is calculated using the above formula (2) to determine the size of the image frame in the display area. v1_shift ' and θ v2_shift '.

[0343] In step S175, the position determination unit 122 uses the horizontal angle θ Ai and displacement θ shift The horizontal angle θ is calculated using the formula (4) above. Ai_shift .

[0344] In step S176, the position determination unit 122 determines whether object #i exists in the display area 31 (the horizontal angle of object #i is between the horizontal angles at both ends of the display area 31), that is, whether θ is satisfied. v2_shift '<θ Ai_shift <θ v1_shift 'Conditions.

[0345] When it is determined in step S176 that object #i exists in display area 31, that is, when condition θ is satisfied... v2_shift '<θ Ai_shift <θ v1_shift At this point, the process proceeds to step S177. In step S177, the position determination unit 122 determines the position based on the content image frame size information and the horizontal angle θ. v1_shift ' and θ v2_shift 'and horizontal angle θ Ai_shift The horizontal angle θ is calculated using the formula (5) above. A1 '.

[0346] On the other hand, when it is determined in step S176 that object #i does not exist in display area 31, that is, when the condition -180°≤θ is satisfied... Ai_shift ≤θ v2_shift 'or θv1_shift '≤θ Ai_shift ≤180°, the process proceeds to step S178. In step S178, the position determination unit 122 determines the horizontal angle θ v1_shift 'orθ v2_shift 'and the horizontal angle θ Ai_shift The horizontal angle θ Ai 'is calculated by the above-described equation (6).

[0347] After the process of step S177 or step S178, the process returns to step S151 of FIG. 15 and proceeds to step S152. Figure 29

[0348] Note that in the second example, the size of the display image is the same as the size of the video content, but instead, the size of the display image can be different from the size of the video content.

[0349] Further, in the second example, the audio data of all objects is not synthesized and output, but instead, only the audio data of some objects (e.g., objects in the display region, objects within a predetermined range of the display region, etc.) is synthesized and output. The method for selecting the objects of the audio data to be output can be determined in advance or can be designated by the user.

[0350] Further, in the above description, only the audio data of unit objects is used, but the audio data can include the audio data of channel audio, the audio data of higher order ambisonics (HOA) audio, the audio data of spatial audio object coding (SAOC), and the metadata (scene information, dynamic or static metadata) of the audio data. In this case, for example, not only the encoded data of each object but also the encoded data of these data blocks are arranged as sub-samples.

[0351] <First Embodiment>

[0352] (Outline of 3D Audio File Format)

[0353] Before describing the first embodiment of the present disclosure, an outline of the channels of the 3D audio file format of MP4 will be described with reference to Figure 31

[0354] In the MP4 file, the codec information of the video content and the position information indicating the position in the file can be managed for each track. In the 3D audio file format of MP4, all audio streams (elementary streams (ES)) of 3D audio (channel audio / object audio / HOA audio / metadata) are recorded as one track in units of samples (frames). Further, the codec information (profile / level / audio configuration) of 3D is stored as a sample entry.​​

[0355] The channel audio constituting the 3D audio is audio data in units of channels; the object audio is audio data in units of objects; the HOA audio is spherical audio data; and the metadata is metadata of the channel audio / object audio / HOA audio. In this case, the audio data in units of objects is used as the object audio, but alternatively audio data of SAOC can be used instead.

[0356] Structure of moov box (structure of moov box)

[0357] Figure 32 The structure of the moov box of the MP4 file is shown.

[0358] As shown in Figure 32 In the MP4 file, the image data and the audio data are recorded in different tracks. Figure 32 The details of the track of the audio data are not shown, but a track of the audio data similar to the track of the image data is shown. The sample entry is contained in the sample description in the stsd box arranged in the moov box.

[0359] Incidentally, in broadcast or local storage playback, when all the audio streams are parsed and output (reproduced), the Web server transmits all the audio streams, and the video playback terminal (client) decodes the audio streams of the necessary 3D audio. When the Bitrate is high or there is a limit to the reading rate of the local storage, there is a demand to reduce the load of the decoding process by acquiring only the audio streams of the necessary 3D audio.

[0360] Further, in streaming playback, there is a demand that the video playback terminal (client) acquires only the encoded data of the necessary 3D audio, thereby acquiring the audio stream of the encoding rate optimum for the playback environment.

[0361] Therefore, in the present disclosure, the encoded data of the 3D audio is divided into tracks for each type of data and the tracks are arranged in the audio file, which makes it possible to efficiently acquire only the encoded data of a predetermined type. Therefore, the load on the system is reduced in broadcast and local storage playback. Further, in streaming playback, the highest quality encoded data of the necessary 3D audio can be played back according to the frequency band. Further, since only the position information of the audio stream of the 3D file in the audio file needs to be recorded in units of tracks of subfragments, the amount of position information can be reduced compared to the case where the encoded data in units of objects is arranged in sub samples.

[0362] Overview of track

[0363] Figure 33 A schematic diagram showing an overview of the track in the first embodiment of the present disclosure is shown.

[0364] As Figure 33 shown in the first embodiment, the channel audio / object audio / HOA audio / metadata constituting the 3D audio are respectively set as audio streams of different tracks (channel audio track / object audio track / HOA audio track / object metadata track). The audio stream of the audio metadata is arranged in the object metadata track.

[0365] Further, a base track (base track) is provided as a track for arranging information about the entire 3D audio. In the base track as Figure 33 shown, information about the entire 3D audio is arranged in a sample entry when no sample is arranged in the sample entry. Further, the base track, the channel audio track, the object audio track, the HOA audio track, and the object metadata are recorded as the same audio file (3dauio.mp4).

[0366] A track reference number (Track Reference) is arranged in, for example, a track box, and indicates a reference relationship between a corresponding track and another track. Specifically, the track reference number indicates an ID (hereinafter, referred to as a track ID) that is unique to a track in other referenced tracks. In Figure 33 the example shown, the track IDs of the base track, the channel audio track, the HOA audio track, the object metadata track, and the object audio track are 1, 2, 3, 4, 10,... respectively. The track reference numbers of the base track are 2, 3, 4, 10,..., and the track reference numbers of the channel audio track / HOA audio track / object metadata track / object audio track are 1, which corresponds to the track ID of the base track.

[0367] Accordingly, the base track and the channel audio track / HOA audio track / object metadata track / object audio track have a reference relationship. Specifically, the base track is referenced in the process of playing the channel audio track / HOA audio track / object metadata track / object audio track.

[0368] (Example syntax of a sample entry of a base track)

[0369] Figure 34 To show a schematic diagram of the example syntax of the sample entry of the base track shown in Figure 33

[0370] As information about the entire 3D audio, as Figure 34 ​The configurationVersion, MPEGHAudioProfile, and MPEGHAudioLevel shown represent the configuration information, profile information, and level information (for normal 3D audio streams) of the entire 3D audio stream, respectively. Additionally, as information about the entire 3D audio stream, such as... Figure 34 The width and height shown represent the number of pixels in the horizontal direction and the number of pixels in the vertical direction of the video content, respectively. As information about the entire 3D audio, θ1, θ2, γ1, and γ2 represent the horizontal angle θv1 at the left end of the image frame, the horizontal angle θv2 at the right end of the image frame, the vertical angle γv1 at the top end of the image frame, and the vertical angle γv2 at the bottom end of the image frame, respectively, in the image frame size information of the video content.

[0371] (Example syntax for sample entries of audio track channels)

[0372] Figure 35 To show in Figure 33 The diagram shows an example syntax of sample entries for the channel audio track (channel audio track).

[0373] Figure 35 The configurationVersion, MPEGHAudioProfile, and MPEGHAudioLevel are displayed, representing the configuration information, profile information, and level information of the audio channel, respectively.

[0374] (Example syntax for sample entries of an object audio track)

[0375] Figure 36 To show in Figure 33 The diagram illustrates an exemplary syntax for sample entries of an object audio track (object audio track).

[0376] In one or more object audio tracks contained within an object audio track, such as Figure 36 The ConfigurationVersion, MPEGHAudioProfile, and MPEGHAudioLevel shown represent configuration information, profile information, and level information, respectively. object_is_fixed indicates whether one or more object audio objects included in the object audio track are fixed. When object_is_fixed is 1, it indicates that the object is fixed; when object_is_fixed is 0, it indicates that the object is displaced. mpegh3daConfig represents the configuration of the identification information for one or more object audio objects included in the object audio track.

[0377] Further, objectTheta1 / objectTheta2 / objectGamma1 / objectGamma2 / objectRength represent object information of one or more object audios included in the object audio track. This object information is information that is valid when Object_is_fixed = 1.

[0378] maxobjectTheta1, maxobjectTheta2, maxobjectGamma1, maxobjectGamma2, and maxobjectRength represent maximum values of object information when one or more object audios included in the object audio track are displaced.

[0379] (Exemplary syntax of sample entry of HOA audio track)

[0380] Figure 37 A schematic diagram showing the exemplary syntax of the sample entry of the HOA audio track shown in Figure 33

[0381] ConfigurationVersion, MPEGHAudioProfile, and MPEGHAudioLevel shown in Figure 37

[0382] (Exemplary syntax of sample entry of object metadata track)

[0383] Figure 38 A schematic diagram showing the exemplary syntax of the sample entry of the object metadata track (object metadata track) shown in Figure 33

[0384] ConfigurationVersion shown in Figure 38

[0385] (First example of segment structure of audio file of 3D audio)

[0386] Figure 39 A schematic diagram showing the first example of the segment structure of the audio file of the 3D audio in which the first embodiment of the present disclosure is applied.

[0387] In the first example of the segment structure of the audio file of the 3D audio shown in Figure 39 ​​​​In the illustrated segment structure, an initial segment is composed of an ftyp box and a moov box. A trak box for each track included in the audio file is arranged in the moov box. An mvex box is arranged in the moov box, where the mvex box contains information indicating a correspondence between a track ID of each track and a level used in an ssix box within a media segment.

[0388] Further, a media segment is composed of an sidx box, an ssix box, and one or more subsegments. Position information indicating a position in the audio file of each subsegment is arranged in the sidx box. The ssix box contains position information of an audio stream at each level arranged in an mdat box. Note that each level corresponds to each track. Further, the position information of the first track is position information of data composed of an audio stream of the first track and a moof box.

[0389] A subsegment is set for any time length. A pair of a moof box and an mdat box common to all tracks is set in the subsegment. In the mdat box, audio streams of all tracks are arranged in concentration with respect to any time length. In the moof box, management information of the audio streams is arranged. The audio stream of each track arranged in the mdat box is continuous for each track.

[0390] In Figure 39 the example, track 1 with a track ID of 1 is a base track, and tracks 2 to N with track IDs of 2 to N are a channel audio track, an object audio track, an HOA audio track, and an object metadata track, respectively. The case described later Figure 40 is the same.

[0391] (Second Example of Segment Structure of Audio File of 3D Audio)

[0392] Figure 40 A schematic diagram showing a second example of a segment structure of an audio file of 3D audio in which the first embodiment of the present disclosure is applied.

[0393] The segment structure as Figure 40 shown is different from the segment structure as Figure 39 shown in that a moof box and an mdat box are set for each track.

[0394] Specifically, an initial segment as Figure 40 shown is similar to an initial segment as Figure 39 shown. Like a media segment as Figure 39 shown, a media segment as Figure 40The illustrated media segment is composed of an sidx box, an ssix box, and one or more subsegments. Further, like the illustrated sidx box, the position information of each subsegment is arranged in the sidx box. The ssix box contains the position information of the data of each level composed of a moof box and an mdat box. Figure 39

[0395] Subsegments are set for any time length. A pair of a moof box and an mdat box is set for each track in the subsegment. Specifically, the audio stream of each track is centrally arranged (interleaved and stored) in the mdat box of each track in any time length, and the management information of the audio stream is arranged in the moof box.

[0396] As illustrated in Figure 39 and 40 , the audio stream for each track is centrally arranged in any time length, so that, compared to the case where the audio stream is centrally arranged in samples, the efficiency of acquiring the audio stream via HTTP or the like can be improved.

[0397] (Exemplary description of mvex box)

[0398] Figure 41 A schematic diagram for an exemplary description of a level assignment box arranged in the mvex box as illustrated in Figure 39 and 40 .

[0399] The level assignment box is a box for associating the track ID of each track with the level used in the ssix box. In the example of Figure 41 , the base track of track ID 1 is associated with level 0, and the channel audio track of track ID 2 is associated with level 1. Further, the HOA audio track of track ID 3 is associated with level 2, and the object metadata track of track ID 4 is associated with level 3. Further, the object audio track of track ID 10 is associated with level 4.

[0400] (Exemplary description of MPD file)

[0401] Figure 42 A schematic diagram for an exemplary description of an MDF file in which the first embodiment of the present disclosure is applied.

[0402] As illustrated in Figure 42 , the "Representation" ("representation") of a segment of an audio file (3daudio.mp4) for managing 3D audio, the "SubRepresentation" ("subrepresentation") for managing the tracks contained in the segment, and the like are described in the MPD file.

[0403] ​In the "Representation" and the "SubRepresentation", a "codecs" which indicates a type of a codec of a corresponding fragment or track in a code defined in the 3D audio file format is contained. Further, an "id", an "associationId", and an "assciationType" are contained in the "Representation".

[0404] The "id" indicates an ID of the "Representation" containing the "id". The "associationId" indicates information indicating a reference relationship between a corresponding track and another track and indicates an "id" of a reference track. The "assciationType" indicates a code indicating a meaning of the reference relationship (correlation relationship) with respect to the reference track. For example, a value identical to a value of a track reference serial number of the MP4 is used.

[0405] Further, a "level" is contained in the "SubRepresentation", which is a value set in a level assignment box as a value indicating a corresponding track and a corresponding level. A "dependencyLevel" is contained in the "SubRepresentation", which is a value indicating a level corresponding to another track (hereinafter, referred to as a reference track) having a reference relationship (correlation).

[0406] Further, the "SubRepresentation" contains <EssentialProperty schemeIdUri="urn:mpeg:DASH:3daudio:2014" value="audioType, contentkind, priority"> as information required for selecting 3D audio.

[0407] Further, the "SubRepresentation" in the object audio track contains <EssentialProperty schemeIdUri="urn:mpeg:DASH:viewingAngle:2014" value="Θ, γ, r">. When an object corresponding to the "SubRepresentation" is fixed, Θ, γ, and r indicate a horizontal angle, a vertical angle, and a distance in object position information, respectively. On the other hand, when the object is displaced, the values Θ, γ, and r indicate a maximum value of a horizontal angle, a maximum value of a vertical angle, and a maximum value of a distance among maximum values of the object position information, respectively.

[0408] Figure 43 A diagram for illustrating a definition of the basic attribute shown in Figure 42

[0409] In the upper left side of Figure 43 , audioType (audio type) of <EssentialProperty schemeIdUri="urn:mpeg:DASH:3daudio:2014" value="audioType, contentkind, priority"> is defined. The audioType indicates a type of 3D audio of the corresponding track.

[0410] In the example of Figure 43 , when the audioType indicates 1, it indicates that the audio data of the corresponding track is a channel audio of 3D audio, and when the audioType indicates 2, it indicates that the audio data of the corresponding track is HOA audio. Further, when the audioType indicates 3, it indicates that the audio data of the corresponding track is object audio, and when the audioType is 4, it indicates that the audio data of the corresponding track is metadata.

[0411] Further, in the right side of Figure 43 , contentkind (content kind) of <EssentialProperty schemeIdUri="urn:mpeg:DASH:3daudio:2014" value="audioType, contentkind, priority"> is defined. The contentkind indicates a content of the corresponding audio. For example, in the example of Figure 43 , when the contentkind indicates 3, the corresponding audio is music.

[0412] As shown in the lower left side of Figure 43 , priority (priority) is defined by 23008-3 and indicates a processing priority of the corresponding object. Only when the value does not change in the course of the audio stream, a value indicating the processing priority of the object is described, and when the value changes in the course of the audio stream, a value of "0" is described.

[0413] (Outline of information processing system)

[0414] Figure 44 A diagram for illustrating an outline of an information processing system according to a first embodiment of the present disclosure.

[0415] In the example shown in Figure 44 , the outline of the information processing system according to the first embodiment of the present disclosure is shown. Figure 1 ​Components that are the same as those shown in the drawing are denoted by the same reference numerals. Repetitive explanation is omitted as appropriate.

[0416] As Figure 44 shown, the information processing system 140 has a configuration in which a Web server 142 (connected to a file generation apparatus 141) is connected to a video playback terminal 144 via the Internet 13.

[0417] In the information processing system 140, the Web server 142 transmits a video stream of a video content to the video playback terminal 144 in units of tiles by a method compatible with MPEG-DASH (tile streaming). Further, in the information processing system 140, the Web server 142 transmits an audio stream of object audio, channel audio, or HOA audio corresponding to a tile to be played to the video playback terminal 144.

[0418] The file generation apparatus 141 of the information processing system 140 is similar to the file generation apparatus 11 as Figure 11 shown, except that, for example, the audio file generation unit 56 generates an audio file in the first embodiment and the MPD generation unit 57 generates an MPD file in the first embodiment.

[0419] Specifically, the file generation apparatus 141 acquires image data of a video content and encodes the image data in units of tiles to generate a video stream. The file generation apparatus 141 processes the video stream of each tile into a file format. The file generation apparatus 141 uploads an image file of each tile obtained as a result of the processing to the Web server 142.

[0420] Further, the file generation apparatus 141 acquires 3D audio of a video content and encodes the 3D audio for each type (channel audio / object audio / HOA audio / metadata) of the 3D audio to generate an audio stream. The file generation apparatus 141 assigns a track to the audio stream for each type of the 3D audio. The file generation apparatus 141 generates an audio file of a segment structure (in which the audio stream of each track is arranged in sub-segments) as Figure 39 or 40 and uploads the audio file to the Web server 142.

[0421] The file generation apparatus 141 generates an MPD file containing image frame size information, tile position information, and object position information. The file generation apparatus 141 uploads the MPD file to the Web server 142.

[0422] The Web server 142 stores the image file, the audio file, and the MPD file uploaded from the file generation apparatus 141.

[0423] In Figure 44In the example of FIG. 6, the Web server 142 stores a segment group formed of image files of a plurality of segments of tile #1 and a segment group formed of image files of a plurality of segments of tile #2. The Web server 142 also stores a segment group formed of an audio file of 3D audio.

[0424] The Web server 142 transmits the image files, the audio file, the MPD file, and the like stored in the Web server to the video playback terminal 144 in response to a request from the video playback terminal 144.

[0425] The video playback terminal 144 executes the control software 161, the video playback software 162, the access software 163, and the like.

[0426] The control software 161 is software for controlling data streamed from the Web server 142. Specifically, the control software 161 causes the video playback terminal 144 to acquire the MPD file from the Web server 142.

[0427] Further, the control software 161 specifies a tile in the display region based on the display region instructed from the video playback software 162 and tile position information included in the MPD file. Then, the control software 161 instructs the access software 163 to transmit a request for the image file of the tile.

[0428] When object audio is to be played, the control software 161 instructs the access software 163 to transmit a request for image frame size information in the audio file. Further, the control software 161 instructs the access software 163 to transmit a request for an audio stream of the metadata. The control software 161 specifies an object corresponding to an image in the display region based on the image frame size information and object position information included in the audio stream of the metadata, which is transmitted from the Web server 142 according to the instruction and the display region. Then, the control software 161 instructs the access software 163 to transmit a request for an audio stream of the object.

[0429] Further, when channel audio or HOA audio is to be played, the control software 161 instructs the access software 163 to transmit a request for an audio stream of the channel audio or the HOA audio.

[0430] Video playback software 162 is used to play image and audio files obtained from web server 142. Specifically, when the display area is specified by the user, video playback software 162 commands control software 161 to send the display area. Furthermore, video playback software 162 decodes the image and audio files obtained from web server 142 according to instructions. Video playback software 162 synthesizes the tile-based image data obtained as a result of decoding and outputs the image data. Additionally, when needed, video playback software 162 synthesizes object audio, channel audio, or HOA audio obtained as a result of decoding and outputs the audio.

[0431] Access software 163 is software used to control communication between the video playback terminal 142 and the web server 142 via the Internet 13 using HTTP. Specifically, access software 163 causes video playback terminal 144 to send requests regarding image frame size information or predetermined audio streams in image and audio files in response to instructions from control software 161. Furthermore, access software 163 causes video playback terminal 144 to receive image frame size information or predetermined audio streams in image and audio files sent from the web server 12 in response to these sending requests.

[0432] (Example of document generation device configuration)

[0433] Figure 45 To show in Figure 44 A block diagram showing an example configuration of the file generation device 141.

[0434] exist Figure 45 The one shown in the middle is the same as Figure 11 Components that are identical in the diagram are indicated by the same reference numeral. Where appropriate, repeated explanations are omitted.

[0435] like Figure 45 The configuration of the document generation device 141 shown is different from that of... Figure 11 The file generation apparatus 11 shown is configured to provide an audio encoding processing unit 171, an audio file generation unit 172, an MPD generation unit 173, and a server upload processing unit 174 in place of the audio encoding processing unit 55, the audio file generation unit 56, the MPD generation unit 57, and the server upload processing unit 58.

[0436] Specifically, the audio encoding processing unit 171 of the file generation device 141 encodes the 3D audio of the video content input from the outside for each type (channel audio / object audio / HOA audio / metadata) to generate an audio stream. The audio encoding processing unit 171 provides the audio stream for each type of 3D audio to the audio file generation unit 172.

[0437] The audio file generation unit 172 allocates a track to the audio stream supplied from the audio encoding processing unit 171 for each type of 3D audio. The audio file generation unit 172 generates an audio file in which the audio stream of each track is arranged in units of subsegments, as shown in the segment structure of FIG. 40. At this time, the audio file generation unit 172 stores the image frame size information inputted from the outside in the sample entry. The audio file generation unit 172 supplies the generated audio file to the MPD generation unit 173. Figure 39

[0438] The MPD generation unit 173 determines the URL or the like of the Web server 142 that stores the image file of each tile supplied from the image file generation unit 53. Further, the MPD generation unit 173 determines the URL or the like of the Web server 142 that stores the audio file supplied from the audio file generation unit 172.

[0439] The MPD generation unit 173 arranges the image information supplied from the image information generation unit 54 in the "Adaptation Set" for the image of the MPD file. Further, the MPD generation unit 173 arranges the URL or the like of the image file of each tile in the "Segment" for the "Representation" of the tile.

[0440] The MPD generation unit 173 arranges the URL or the like of the audio file in the "Segment" for the "Representation" of the audio file. Further, the MPD generation unit 173 arranges the object position information or the like of each object inputted from the outside in the "SubRepresentation" for the object metadata track of the object. The MPD generation unit 173 supplies the MPD file in which various pieces of information are arranged as described above, and the image file and the audio file to the server upload processing unit 174.

[0441] The server upload processing unit 174 uploads the image file, the audio file, and the MPD file of each tile supplied from the MPD generation unit 173 to the Web server 142.

[0442] (Explanation of the procedure of the file generation apparatus)

[0443] Figure 46 A flowchart showing the file generation procedure of the file generation apparatus 141 shown in Figure 45

[0444] The procedure of steps S191 to S195 shown in Figure 46 Figure 12 ​​​The process of steps S11 to S15 is shown, and therefore its description is omitted.

[0445] In step S196, the audio encoding processing unit 171 encodes the 3D audio of the externally input video content for each type (channel audio / object audio / HOA audio / metadata) to generate an audio stream. The audio encoding processing unit 171 provides the audio stream to the audio file generation unit 172 for each type of 3D audio.

[0446] In step S197, the audio file generation unit 172 assigns tracks to the audio stream provided from the audio encoding processing unit 171 for each type of 3D audio.

[0447] In step S198, the audio file generation unit 172 generates an audio file as follows: Figure 39 Or, as shown in section 40, an audio file with a segment structure, where the audio stream of each track is arranged in units of sub-segments. In this case, the audio file generation unit 172 stores the image frame size information input from the outside in a sample entry. The audio file generation unit 172 provides the generated audio file to the MPD generation unit 173.

[0448] In step S199, the MPD generation unit 173 generates an MPD file containing image information provided by the image information generation unit 54, the URL of each file, and object location information. The MPD generation unit 173 then uploads the image file, audio file, and MPD file to the server upload processing unit 174.

[0449] In step S200, the server upload processing unit 174 uploads the image file, audio file, and MPD file provided by the MPD generation unit 173 to the web server 142. The process then terminates.

[0450] (Example of video playback terminal function configuration)

[0451] Figure 47 A block diagram illustrating an example configuration of a streaming playback unit, which is configured as follows: Figure 44 The video playback terminal 144 shown implements control software 161, video playback software 162, and access software 163.

[0452] exist Figure 47 The one shown in the middle is the same as Figure 13 Components that are identical in the diagram are indicated by the same reference numeral. Where appropriate, repeated explanations are omitted.

[0453] like Figure 47 The configuration of the streaming unit 190 shown is different from that of... Figure 13The configuration of the stream playback unit 90 shown is such that an MPD processing unit 191, an audio selection unit 193, an audio file acquisition unit 192, an audio decoding processing unit 194, and an audio synthesis processing unit 195 are provided in place of the MPD processing unit 92, the audio selection unit 94, the audio file acquisition unit 95, the audio decoding processing unit 96, and the audio synthesis processing unit 97, and the meta file acquisition unit 93 is not provided.

[0454] The stream playback unit 190 is similar to the stream playback unit 90 shown, except for the method of acquiring, for example, the audio data to be played of the selected object. Figure 13

[0455] Specifically, the MPD processing unit 191 of the stream playback unit 190 extracts information such as the URL of the audio file of the segment to be played described in the "Segment" for the audio meta file from the MPD file provided from the MPD acquisition unit 91, and provides the extracted information to the audio file acquisition unit 192.

[0456] The MPD processing unit 191 extracts tile position information described in the "Adaptation Set" for the image from the MPD file, and provides the extracted information to the image selection unit 98. The MPD processing unit 191 extracts information such as the URL described in the "Segment" for the image file of the tile requested from the image selection unit 98 from the MPD file, and provides the extracted information to the image selection unit 98.

[0457] When the object audio is to be played, the audio file acquisition unit 192 requests the Web server 142 to send the initial segment of the elementary track in the audio file specified by the URL based on information such as the URL provided from the MPD processing unit 191, and acquires the initial segment of the elementary track.

[0458] Further, the audio file acquisition unit 192 requests the Web server 142 to send the audio stream of the object metadata track in the audio file specified by the URL based on information such as the URL of the audio file, and acquires the audio stream of the object metadata track. The audio file acquisition unit 192 provides the object position information contained in the audio stream of the object metadata track, the image frame size information contained in the initial segment of the elementary track, and information such as the URL of the audio file to the audio selection unit 193.

[0459] ​Further, when the channel audio is to be played, the audio file acquisition unit 192 requests the Web server 142 to send an audio stream of the object audio track in the audio file specified by the URL of the audio file based on information such as the URL of the audio file and acquires the audio stream of the object audio track. The audio file acquisition unit 192 provides the audio decoding processing unit 194 with the acquired audio stream of the object audio track.

[0460] When the HOA audio is to be played, the audio file acquisition unit 192 performs a process similar to that performed when the channel audio is to be played. Thus, the audio stream of the HOA audio track is provided to the audio decoding processing unit 194.

[0461] It should be noted that which one of the object audio, the channel audio and the HOA audio is determined to be played, for example, in accordance with an instruction of the user.

[0462] The audio selection unit 193 calculates the position of each object on the image based on the image frame size information and the object position information provided from the audio file acquisition unit 192. The audio selection unit 193 selects the object in the display region specified by the user based on the position of each object on the image. Based on information such as the URL of the audio file provided from the audio file acquisition unit 192, the audio selection unit 193 requests the Web server 142 to send an audio stream of the object audio track of the selected object in the audio file specified by the URL and acquires the audio stream of the object audio track. The audio selection unit 193 provides the audio decoding processing unit 194 with the acquired audio stream of the object audio track.

[0463] The audio decoding processing unit 194 decodes the audio stream of the channel audio track or the HOA audio track provided from the audio file acquisition unit 192, or decodes the audio stream of the object audio track provided from the audio selection unit 193. The audio decoding processing unit 194 provides the audio synthesis processing unit 195 with one of the channel audio, the HOA audio and the object audio obtained as a result of the decoding.

[0464] The audio synthesis processing unit 195 synthesizes the object audio, the channel audio or the HOA audio provided from the audio decoding processing unit 194 and outputs the audio, as necessary.

[0465] (Explanation of the process of the video playback terminal)

[0466] Figure 48 A flowchart showing a channel audio playback process of the stream playback unit 190 shown in Figure 47 is executed, for example, when the user selects the channel audio as an object to be played.

[0467] In Figure 48In step S221, the MPD processing unit 191 analyzes the MPD file supplied from the MPD acquisition unit 91, and specifies "SubRepresentation" ("subrepresentation") of the channel audio of the segment to be played based on the basic attribute and the codec described in "SubRepresentation" ("subrepresentation"). Further, the MPD processing unit 191 extracts information such as the URL described in "Segment" ("segment") of the audio file for the segment to be played from the MPD file, and supplies the extracted information to the audio file acquisition unit 192.

[0468] In step S222, the MPD processing unit 191 specifies the level of the basic track as the reference track based on the dependencyLevel of the "SubRepresentation" ("subrepresentation") specified in step S221, and supplies the specified level of the basic track to the audio file acquisition unit 192.

[0469] In step S223, the audio file acquisition unit 192 requests the Web server 142 to transmit the initial segment of the segment to be played and acquires the initial segment based on information such as the URL supplied from the MPD processing unit 191.

[0470] In step S224, the audio file acquisition unit 192 acquires the track ID corresponding to the level of the channel audio track and the basic track as the reference track from the level assignment box in the initial segment.

[0471] In step S225, the audio file acquisition unit 192 acquires the sample entry of the initial segment in the track box corresponding to the track ID of the initial segment based on the track ID of the channel audio track and the basic track as the reference track. The audio file acquisition unit 192 supplies the codec information contained in the acquired sample entry to the audio decoding processing unit 194.

[0472] In step S226, the audio file acquisition unit 192 transmits a request to the Web server 142 and acquires the sidx box and the ssix box from the header of the audio file of the segment to be played based on information such as the URL supplied from the MPD processing unit 191.

[0473] In step S227, the audio file acquisition unit 192 acquires the position information of the reference track and the channel audio track of the segment to be played from the sidx box and the ssix box acquired in step S223. In this case, since the basic track as the reference track does not contain any audio stream, there is no position information of the reference track.

[0474] In step S228, the audio file acquisition unit 192 requests the Web server 142 to transmit the audio stream of the channel audio track arranged in the mdat box based on the position information of the channel audio track and information such as the URL of the audio file of the segment to be played, and acquires the audio stream of the channel audio track. The audio file acquisition unit 192 provides the audio stream of the acquired channel audio track to the audio decoding processing unit 194.

[0475] In step S229, the audio decoding processing unit 194 decodes the audio stream of the channel audio track based on the codec information provided from the audio file acquisition unit 192. The audio file acquisition unit 192 provides the channel audio obtained as a result of the decoding to the audio synthesis processing unit 195.

[0476] In step S230, the audio synthesis processing unit 195 outputs the channel audio. Then, the process is terminated.

[0477] It should be noted that although not shown, the HOA audio playback process for playing the HOA audio by the stream playback unit 190 is executed in a manner similar to the channel audio playback process as shown in FIG. 8. Figure 48

[0478] Figure 49 A flowchart showing the object designation process of the stream playback unit 190 shown in FIG. 6 is shown. The object designation process is executed, for example, when the user selects the object audio as the object to be played and the playback area is changed. Figure 47

[0479] In step S251 of FIG. 6, the audio selection unit 193 acquires the display area designated by the user through the user's operation or the like. Figure 49

[0480] In step S252, the MPD processing unit 191 analyzes the MPD file provided from the MPD acquisition unit 91, and designates the "SubRepresentation" ("sub representation") of the metadata of the segment to be played based on the basic attribute and the codec described in the "SubRepresentation" ("sub representation"). Further, the MPD processing unit 191 extracts information such as the URL of the audio file of the segment to be played described in the "Segment" ("segment") for the audio meta file from the MPD file, and provides the extracted information to the audio file acquisition unit 192.

[0481] ​​​In step S253, the MPD processing unit 191 specifies the level of the base track that is the reference track based on the dependencyLevel of the "SubRepresentation" specified in step S252, and supplies the audio file acquisition unit 192 with the specified level of the base track.

[0482] In step S254, the audio file acquisition unit 192 requests the Web server 142 to send the initial segment of the segment to be played and acquires the initial segment based on information such as the URL supplied from the MPD processing unit 191.

[0483] In step S255, the audio file acquisition unit 192 acquires the track ID corresponding to the level of the object metadata track and the base track that is the reference track from the level assignment box in the initial segment.

[0484] In step S256, the audio file acquisition unit 192 acquires the sample entry of the initial segment in the track box corresponding to the track ID of the initial segment based on the track ID of the object metadata track and the base track that is the reference track. The audio file acquisition unit 192 supplies the audio selection unit 193 with the image frame size information contained in the sample entry of the base track that is the reference track. Further, the audio file acquisition unit 192 supplies the audio selection unit 193 with the initial segment.

[0485] In step S257, the audio file acquisition unit 192 sends a request to the Web server 142 and acquires the sidx box and the ssix box from the header of the audio file of the segment to be played based on information such as the URL supplied from the MPD processing unit 191.

[0486] In step S258, the audio file acquisition unit 192 acquires the position information of the reference track and the object metadata track of the sub-segment to be played from the sidx box and the ssix box acquired in step S257. In this case, since the base track that is the reference track does not contain any audio stream, there is no position information of the reference track. The audio file acquisition unit 192 supplies the audio selection unit 193 with the sidx box and the ssix box.

[0487] In step S259, the audio file acquisition unit 192 requests the Web server 142 to send the audio stream of the object metadata track arranged in the mdat box based on the position information of the object metadata track and information such as the URL of the audio file of the segment to be played, and acquires the audio stream of the object metadata track.

[0488] In step S260, the audio file acquisition unit 192 decodes the audio stream of the object metadata track acquired in step S259 on the basis of the codec information contained in the sample entry acquired in step S256. The audio file acquisition unit 192 provides the object position information contained in the metadata obtained as a result of the decoding to the audio selection unit 193. Further, the audio file acquisition unit 192 provides information such as the URL of the audio file provided from the MPD processing unit 191 to the audio selection unit 193.

[0489] In step S261, the audio selection unit 193 selects the object in the display region on the basis of the image frame size information and the object position information provided from the audio file acquisition unit 192 and on the basis of the display region specified by the user. The process is then terminated.

[0490] Figure 50 A flowchart showing the specified object audio playback process performed by the stream playback unit 190 after the object specifying process shown in Figure 49

[0491] In step S281 of Figure 50 , the MPD processing unit 191 analyzes the MPD file provided from the MPD acquisition unit 91 and specifies the "SubRepresentation" ("subrepresentation") of the object audio of the selected object on the basis of the basic attribute and the codec described in "SubRepresentation" ("subrepresentation").

[0492] In step S282, the MPD processing unit 191 specifies the level of the base track that is the reference track on the basis of the dependencyLevel of the "SubRepresentation" ("subrepresentation") specified in step S281 and provides the specified level of the base track to the audio file acquisition unit 192.

[0493] In step S283, the audio file acquisition unit 192 acquires the track ID corresponding to the level of the object audio track and the base track that is the reference track from the level assignment box in the initial segment and provides the track ID to the audio selection unit 193.

[0494] In step S284, the audio selection unit 193 acquires the sample entry of the initial segment in the track box corresponding to the track ID of the initial segment on the basis of the track IDs of the object audio track and the base track that is the reference track. The initial segment is obtained as described in Figure 49 ​The audio file acquisition unit 192 shown in step S256 provides. The audio selection unit 193 provides the audio decoding processing unit 194 with codec information contained in the acquired sample entry.

[0495] In step S285, the audio selection unit 193 acquires position information of the object audio track of the selected object of the reference track and the sub-fragment to be played from the sidx box and the ssix box provided from the audio file acquisition unit 192 in step S258. In this case, since the base track as the reference track does not contain any audio stream, there is no position information of the reference track.

[0496] In step S286, the audio selection unit 193 requests the Web server 142 to send the audio stream of the object audio track of the selected object arranged in the mdat box based on the position information of the object audio track and information such as the URL of the audio file of the fragment to be played, and acquires the audio stream of the object audio track. The audio selection unit 193 provides the audio decoding processing unit 194 with the acquired audio stream of the object audio track.

[0497] In step S287, the audio decoding processing unit 194 decodes the audio stream of the object audio track based on the codec information provided from the audio selection unit 193. The audio selection unit 193 provides the audio synthesis processing unit 195 with the object audio obtained as a result of the decoding.

[0498] In step S288, the audio synthesis processing unit 195 synthesizes and outputs the object audio provided from the audio decoding processing unit 194. Then the process is terminated.

[0499] As described above, in the information processing system 140, the file generation apparatus 141 generates an audio file in which 3D audio is divided into a plurality of tracks and the tracks are arranged according to the type of the 3D audio. The video playback terminal 144 acquires the audio stream of the predetermined type of 3D audio in the audio file. Therefore, the video playback terminal 144 can efficiently acquire the audio stream of the predetermined type of 3D audio. Therefore, it can be said that the file generation apparatus 141 generates an audio file capable of improving the efficiency of acquiring the audio stream of the predetermined type of 3D audio.

[0500] <Second Embodiment>

[0501] (Outline of Track)

[0502] Figure 51 A schematic diagram showing an outline of a track in the second embodiment of the present disclosure.

[0503] As Figure 51The second embodiment differs from the first embodiment in that the basic sample is recorded as a sample of the basic track. The basic sample is formed of information referred to by the samples of the channel audio / object audio / HOA audio / metadata. The samples of the channel audio / object audio / HOA audio / metadata that refer to the reference information contained in the basic sample are arranged in the order of arrangement of the reference information, making it possible to generate an audio stream of the 3D audio before the 3D audio is divided into tracks.

[0504] (Exemplary syntax of a sample entry of a basic track)

[0505] Figure 52 A schematic diagram showing the exemplary syntax of the sample entry of the basic track shown in Figure 51

[0506] The syntax shown in Figure 52 is the same as the syntax shown in Figure 34 except that "mha2" is described to indicate that the sample entry is a sample entry of the basic track shown in Figure 51 instead of "mhal" described to indicate that the sample entry is a sample entry of the basic track shown in Figure 33

[0507] (Exemplary structure of a basic entry)

[0508] Figure 53 A schematic diagram showing the exemplary structure of the basic sample.

[0509] As shown in Figure 53 , the basic sample is configured using an extractor of the channel audio / object audio / HOA audio / metadata in units of samples that are sub-samples. The extractor of the channel audio / object audio / HOA audio / metadata consists of the type of the extractor and the offset and size of the sub-samples of the corresponding channel audio track / object audio track / HOA audio track / object metadata track. The offset is the difference between the position of the basic sample in the file of the sub-sample of the basic sample and the position of the channel audio track / object audio track / HOA audio track / object metadata track in the file of the sample. In other words, the offset is information indicating the position within the file of the sample of another track corresponding to the sub-sample of the basic sample containing the offset.

[0510] Figure 54 A schematic diagram showing the exemplary syntax of the basic sample.

[0511] As shown in Figure 54 , in the basic sample, the SCE element used to store the object audio in the sample of the object audio track is replaced with the EXT element storing the extractor.

[0512] ​​Figure 55 A diagram for showing an example of extractor data.

[0513] As Figure 55 shown, the type of extractor and the offset and size of the sub-sample of the corresponding channel audio track / object audio track / HOA audio track / object metadata track are described in the extractor.

[0514] Note that the extractor can utilize a network abstraction layer (NAL) structure extension defined in advanced video coding (AVC) / high efficiency video coding (HEVC) so that the audio elements and configuration information can be stored.

[0515] The information processing system in the second embodiment and the process executed by the information processing system are similar to those of the first embodiment, and thus the description thereof is omitted.

[0516] <Third Embodiment>

[0517] (Outline of tracks)

[0518] Figure 56 A diagram for showing an outline of tracks in the third embodiment of the present disclosure is shown.

[0519] As Figure 56 shown, the third embodiment differs from the first embodiment in that the basic samples and the samples of metadata are recorded as samples of the basic track and the object metadata track is not provided.

[0520] The information processing system in the third embodiment and the process executed by the information processing system are similar to those of the first embodiment except that the audio stream of the basic track instead of the object metadata track is acquired so as to acquire the object position information. Thus, the description thereof is omitted.

[0521] <Fourth Embodiment>

[0522] (Outline of tracks)

[0523] Figure 57 A diagram for showing an outline of tracks in the fourth embodiment of the present disclosure is shown.

[0524] As Figure 57 shown, the fourth embodiment differs from the first embodiment in that the tracks are recorded as different files (3da_base.mp4 / 3da_channel.mp4 / 3da_object_1.mp4 / 3da_hoa.mp4 / 3da_meta.mp4). In this case, only the audio data of the desired track can be acquired via HTTP by acquiring the file of the desired track. Thus, the audio data of the desired track can be effectively acquired via HTTP.

[0525] (Example description of an MPD file)

[0526] Figure 58 This is a schematic diagram illustrating an exemplary description of an MDF file according to a fourth embodiment of the present disclosure.

[0527] like Figure 58 As shown, the "Representation" ("Representation") of each segment of the 3D audio file (3da_base.mp4 / 3da_channel.mp4 / 3da_object_1.mp4 / 3da_hoa.mp4 / 3da_meta.mp4) is described in the MPD file.

[0528] The "Representation" field contains "codecs", "id", "associationId", and "assciationType". Additionally, the "Representation" field for channel audio tracks / object audio tracks / HOA audio tracks / object metadata tracks also contains... <EssentialProperty schemeIdUri="urn:mpeg:DASH:3daudio:2014"value="audioType,contentkind,priority"> Furthermore, the "Representation" of the object's audio track contains...<EssentialProperty schemeIdUri="urn:mpeg:DASH:viewingAngle:2014"value="θ,γ,r"> .

[0529] (Overview of Information Processing Systems)

[0530] Figure 59 This is a schematic diagram illustrating an overview of an information processing system in which the present disclosure is applied in a fourth embodiment.

[0531] exist Figure 59 The one shown in the middle is the same as Figure 1 Components that are identical in the diagram are indicated by the same reference numeral. Where appropriate, repeated descriptions are omitted.

[0532] like Figure 59 The information processing system 210 shown has the following configuration: a web server 212 connected to a file generation device 211 and a video playback terminal 214 are connected via the Internet 13.

[0533] In the information processing system 210, the web server 212 transmits video streams of video content to the video playback terminal 214 in units of tiles (tile streaming) using an MPEG-DASH compatible method. Furthermore, in the information processing system 210, the web server 212 transmits audio files corresponding to the object audio, channel audio, or HOA audio of the file to be played to the video playback terminal 214.

[0534] Specifically, the file generation device 211 acquires image data of the video content and encodes the image data in units of tiles to generate a video stream. The file generation device 211 processes the video stream of each tile into a file format for each segment. The file generation device 211 uploads the image file of each file obtained as a result of the above processing to the web server 212.

[0535] In addition, the file generation device 211 acquires 3D audio of the video content and encodes the 3D audio for each type (channel audio / object audio / HOA audio / metadata) to generate an audio stream. The file generation device 211 assigns tracks to the audio stream of each type of 3D audio. The file generation device 211 generates an audio file (with the audio stream arranged in the audio file) for each track and uploads the generated audio file to the web server 212.

[0536] The file generation device 211 generates an MPD file, which includes image frame size information, tile position information, and object position information. The file generation device 211 uploads the MPD file to the web server 212.

[0537] Web server 212 stores image files uploaded from file generation device 211, audio files for each type of 3D audio, and MPD files.

[0538] exist Figure 59 In the example, Web server 212 stores a group of segments formed by image files of multiple fragments of tile #1 and a group of segments formed by image files of multiple fragments of tile #2. Web server 212 also stores a group of segments formed by audio files of channel audio and a group of segments of audio files of object #1.

[0539] In response to a request from the video playback terminal 214, the web server 212 transmits image files, audio files of a predetermined type of 3D audio, MPD files, etc., stored in the web server to the video playback terminal 214.

[0540] The video playback terminal 214 executes control software 221, video playback software 222, access software 223, etc.

[0541] The control software 221 is software for controlling data streamed from the Web server 212. Specifically, the control software 221 causes the video playback terminal 214 to acquire the MPD file from the Web server 212.

[0542] Further, the control software 221 specifies a tile in the MPD file based on the display region instructed from the video playback software 222 and the tile position information contained in the MPD file. Then, the control software 221 instructs the access software 223 to send a request for transmitting an image file of the tile.

[0543] When object audio is to be played, the control software 221 instructs the access software 223 to send a request for transmitting an audio file of the basic track. Then, the control software 221 instructs the access software 223 to send a request for transmitting an audio file of the object metadata track. The control software 221 acquires image frame size information in the audio file of the basic track and object position information contained in the audio file of the metadata, which are transmitted from the Web server 142 according to the instruction. The control software 221 specifies an object corresponding to an image in the display region based on the image frame size information, the object position information, and the display region. Further, the control software 221 instructs the access software 223 to send a request for transmitting an audio file of the object.

[0544] Further, when channel audio or HOA audio is to be played, the control software 221 instructs the access software 223 to send a request for transmitting an audio file of the channel audio or the HOA audio.

[0545] The video playback software 222 is software for playing back image files and audio files acquired from the Web server 212. Specifically, when a display region is specified by a user, the video playback software 222 gives an instruction on the display region to the control software 221. Further, the video playback software 222 decodes image files and audio files acquired from the Web server 212 according to the instruction. The video playback software 222 synthesizes image data obtained in units of tiles as a result of the decoding and outputs the image data. Further, when necessary, the video playback software 222 synthesizes object audio, channel audio, or HOA audio obtained as a result of the decoding and outputs the audio.

[0546] The access software 223 is software for controlling communication with the Web server 212 via the Internet 13 using HTTP. Specifically, the access software 223 causes the video playback terminal 214 to send a request for transmitting an image file and a predetermined audio file in response to an instruction from the control software 221. Further, the access software 223 causes the video playback terminal 214 to receive an image file and a predetermined audio file transmitted from the Web server 212 according to the transmission request.

[0547] (Example of document generation device configuration)

[0548] Figure 60 In order to be in Figure 59 The block diagram of the document generation device 211 shown is shown in the figure.

[0549] exist Figure 60 The one shown in the middle is the same as Figure 45 Components that are identical in the diagram are indicated by the same reference numeral. Where appropriate, repeated explanations are omitted.

[0550] like Figure 60 The configuration of the document generation device 211 shown is different from that of... Figure 45 The file generation apparatus 141 shown is configured such that an audio file generation unit 241, an MPD generation unit 242, and a server upload processing unit 243 are provided to replace the audio file generation unit 172, the MPD generation unit 173, and the server upload processing unit 174, respectively.

[0551] Specifically, the audio file generation unit 241 of the file generation device 211 allocates tracks to the audio stream for each type of 3D audio, which is provided by the audio encoding processing unit 171. The audio file generation unit 241 generates an audio file (containing the audio stream) for each track. At this time, the audio file generation unit 241 stores externally input image frame size information in sample entries of the basic track. The audio file generation unit 241 provides the audio file for each type of 3D audio to the MPD generation unit 242.

[0552] MPD generation unit 242 determines the URL of the web server 212 that stores the image files of each tile provided by image file generation unit 53, etc. Furthermore, for each type of 3D audio, MPD generation unit 242 determines the URL of the web server 212 that stores the audio files provided by audio file generation unit 241, etc.

[0553] The MPD generation unit 242 arranges the image information provided by the image information generation unit 54 in the "Adaptation Set" of the image for the MPD file. Furthermore, the MPD generation unit 242 arranges the URL of the image file for each tile in the "Segment" of the "Representation" of the image file for the tile.

[0554] The MPD generation unit 242 arranges, in a "Segment" of a "Representation" for an audio file, a URL or the like of the audio file for each type of 3D audio. Further, the MPD generation unit 242 arranges, in a "Representation" for an object metadata track of an object, object position information or the like of each object inputted from the outside. The MPD generation unit 242 supplies the server upload processing unit 243 with an MPD file in which various pieces of information are arranged as described above, image files, and audio files for each type of 3D audio.

[0555] The server upload processing unit 243 uploads, to the Web server 212, the image files of each tile, the audio files of each type of 3D audio, and the MPD file supplied from the MPD generation unit 242.

[0556] (Explanation of the process of the file generation apparatus)

[0557] Figure 61 A flowchart showing the file generation process of the file generation apparatus 211 shown in Figure 60 will be described.

[0558] The processes of steps S301 to S307 shown in Figure 61 are similar to those of steps S191 to S197 shown in Figure 46 , and thus the description thereof is omitted.

[0559] In step S308, the audio file generation unit 241 generates an audio file (in which an audio stream is arranged) for each track. At this time, the audio file generation unit 241 stores the image frame size information inputted from the outside in a sample entry in the audio file of the base track. The audio file generation unit 241 supplies the MPD generation unit 242 with the generated audio files for each type of 3D audio.

[0560] In step S309, the MPD generation unit 242 generates an MPD file containing the image information supplied from the image information generation unit 54, a URL of each file, and object position information. The MPD generation unit 242 supplies the server upload processing unit 243 with the image files, the audio files for each type of 3D audio, and the MPD file.

[0561] In step S310, the server upload processing unit 243 uploads, to the Web server 212, the image files, the audio files of each type of 3D audio, and the MPD file supplied from the MPD generation unit 242. Then the process is terminated.

[0562] (Functional configuration example of the video playback terminal)

[0563] Figure 62 A block diagram showing a configuration example of the stream playback unit is shown in FIG. 26. The stream playback unit is implemented in such a manner that the video playback terminal 214 executes the control software 221, the video playback software 222, and the access software 223 as shown in FIG. 26. Figure 59

[0564] In Figure 62 , the same components as those shown in FIG. 25 are denoted by the same reference numerals. Repetitive explanation is omitted as appropriate. Figure 13 47 The configuration of the stream playback unit 260 as shown in FIG. 26 is different from the configuration of the stream playback unit 90 as shown in FIG. 25 in that the MPD processing unit 261, the meta file acquisition unit 262, the audio selection unit 263, the audio file acquisition unit 264, the audio decoding processing unit 194, and the audio synthesis processing unit 195 are provided to replace the MPD processing unit 92, the meta file acquisition unit 93, the audio selection unit 94, the audio file acquisition unit 95, the audio decoding processing unit 96, and the audio synthesis processing unit 97, respectively.

[0565] Specifically, when the object audio is to be played back, the MPD processing unit 261 of the stream playback unit 260 extracts information such as a URL described in the "Segment" of the audio file of the object metadata track of the segment to be played back from the MPD file supplied from the MPD acquisition unit 91, and supplies the extracted information to the meta file acquisition unit 262. Further, the MPD processing unit 261 extracts information such as a URL described in the "Segment" of the audio file of the object audio track of the object requested from the audio selection unit 263 from the MPD file, and supplies the extracted information to the audio selection unit 263. Further, the MPD processing unit 261 extracts information such as a URL described in the "Segment" of the audio file of the basic track of the segment to be played back from the MPD file, and supplies the extracted information to the meta file acquisition unit 262. Figure 62 Figure 13 Further, when the channel audio or the HOA audio is to be played back, the MPD processing unit 261 extracts information such as a URL described in the "Segment" of the audio file of the channel audio track or the HOA audio track of the segment to be played back from the MPD file. The MPD processing unit 261 supplies the information such as the URL to the audio file acquisition unit 264 via the audio selection unit 263.

[0566]

[0567]

[0568] ​​​​​It should be noted that which one of the object audio, the channel audio, and the HOA audio is to be played is determined, for example, in accordance with an instruction of the user.

[0569] The MPD processing unit 261 extracts tile position information described in an "Adaptation Set" for an image from the MPD file and supplies the extracted tile position information to the image selection unit 98. The MPD processing unit 261 extracts information such as a URL described in a "Segment" for an image file of a tile requested from the image selection unit 98 from the MPD file and supplies the extracted information to the image selection unit 98.

[0570] Based on information such as a URL supplied from the MPD processing unit 261, the meta file acquisition unit 262 requests the Web server 212 to transmit an audio file of an object metadata track designated by the URL and acquires the audio file of the object metadata track. The meta file acquisition unit 93 supplies object position information contained in an audio meta file of the object metadata track to the audio selection unit 263.

[0571] Further, based on information such as a URL of an audio file, the meta file acquisition unit 262 requests the Web server 142 to transmit an initial segment of an audio file of a base track designated by the URL and acquires the initial segment. The meta file acquisition unit 262 supplies image frame size information contained in a sample entry of the initial segment to the audio selection unit 263.

[0572] The audio selection unit 263 calculates a position of each object on an image based on the image frame size information and the object position information supplied from the meta file acquisition unit 262. The audio selection unit 263 selects an object in a display region designated by the user based on the position of each object on the image. The audio selection unit 263 requests the MPD processing unit 261 to transmit information such as a URL of an audio file of an object audio track of the selected object. The audio selection unit 263 supplies information such as a URL supplied from the MPD processing unit 261 to the audio file acquisition unit 264 in accordance with the request.

[0573] Based on information such as a URL of an audio file of an object audio track, a channel audio track, or a HOA audio track supplied from the audio selection unit 263, the audio file acquisition unit 264 requests the Web server 12 to transmit an audio stream of an audio file designated by the URL and acquires the audio stream of the audio file. The audio file acquisition unit 95 supplies the acquired object-unit audio file to the audio decoding processing unit 194.

[0574] (Explanation of the process of the video playback terminal)

[0575] Figure 63 To show in Figure 62 The flowchart shown is a process for playing audio channels in the streaming unit 260. For example, the audio channel playback process is executed when the user selects an audio channel as the object to be played.

[0576] exist Figure 63 In step S331, the MPD processing unit 261 analyzes the MPD file provided by the MPD acquisition unit 91 and specifies the "Representation" of the channel audio of the segment to be played based on the basic attributes and the codec described in the "Representation". Furthermore, the MPD processing unit 261 extracts information (such as the URL of the audio file for the channel audio track of the segment to be played, described in the "Segment" contained in the "Representation") and provides the extracted information to the audio file acquisition unit 264 via the audio selection unit 263.

[0577] In step S332, based on the associationId of the “Representation” specified in step S331, the MPD processing unit 261 specifies the “Representation” of the base track as the reference track. The MPD processing unit 261 extracts information (such as the URL of the audio file of the reference track described in the “Segment” contained in the “Representation”) and provides the extracted file to the audio file acquisition unit 264 via the audio selection unit 263.

[0578] In step S333, the audio file acquisition unit 264 requests the Web server 212 to send the initial segment of the audio file of the channel audio track and the reference track of the segment to be played, based on information such as the URL provided by the audio selection unit 263, and acquires the initial segment.

[0579] In step S334, the audio file acquisition unit 264 acquires sample entries in the trak box of the acquired initial segment. The audio file acquisition unit 264 provides the encoding and decoding information contained in the acquired sample entries to the audio decoding processing unit 194.

[0580] In step S335, the audio file acquisition unit 264 sends a request to the web server 142 based on information such as the URL provided by the audio selection unit 263, and acquires the sidx box and ssix box from the header of the audio file of the channel audio track of the segment to be played.

[0581] In step S336, the audio file acquisition unit 264 obtains the position information of the sub-segment to be played from the sidx box and ssix box obtained in step S333.

[0582] In step S337, the audio selection unit 263, based on the location information obtained in step S337 and information such as the URL of the audio file of the channel audio track of the segment to be played, requests the web server 142 to send the audio stream of the channel audio track arranged in the mdat box of the audio file, and obtains the audio stream of the channel audio track. The audio selection unit 263 provides the obtained audio stream of the channel audio track to the audio decoding processing unit 194.

[0583] In step S338, the audio decoding processing unit 194 decodes the audio stream of the channel audio track provided by the audio selection unit 263 based on the encoding and decoding information provided by the audio file acquisition unit 264. The audio selection unit 263 provides the channel audio obtained as the result of decoding to the audio synthesis processing unit 195.

[0584] In step S339, the audio synthesis processing unit 195 outputs channel audio. Then the process terminates.

[0585] Although not shown, the HOA audio playback process for playing HOA audio via the streaming playback unit 260 is similar to that shown in... Figure 63 The audio playback process shown is executed in a manner that allows for channel-based audio playback.

[0586] Figure 64 To show in Figure 62 The flowchart shown illustrates the object audio playback process of the streaming playback unit 260. For example, this object audio playback process is executed when the user selects object audio as the object to be played and the playback area is changed.

[0587] exist Figure 64 In step S351, the audio selection unit 263 obtains the display area specified by the user through user operations, etc.

[0588] In step S352, the MPD processing unit 261 analyzes the MPD file supplied from the MPD acquisition unit 91, and specifies the "Representation" ("representation") of the metadata of the segment to be played based on the basic attribute and the codec described in the "Representation" ("representation"). Further, the MPD processing unit 261 extracts information such as the URL of the audio file of the object metadata track of the segment to be played described in the "Segment" ("segment") included in the "Representation" ("representation"), and supplies the extracted information to the meta file acquisition unit 262.

[0589] In step S353, the MPD processing unit 261 specifies the "Representation" ("representation") of the base track that is the reference track based on the associationId of the "Representation" ("representation") specified in step S352. The MPD processing unit 261 extracts information such as the URL of the audio file of the reference track described in the "Segment" ("segment") included in the "Representation" ("representation"), and supplies the extracted information to the meta file acquisition unit 262.

[0590] In step S354, the meta file acquisition unit 262 requests the Web server 212 to transmit the initial segment of the audio file of the object metadata track and the reference track of the segment to be played based on information such as the URL supplied from the MPD processing unit 261, and acquires the initial segment.

[0591] In step S355, the meta file acquisition unit 262 acquires the sample entry in the trak box of the acquired initial segment. The meta file acquisition unit 262 supplies the image frame size information included in the sample entry of the base track that is the reference track to the audio file acquisition unit 264.

[0592] In step S356, the meta file acquisition unit 262 transmits a request to the Web server 142 based on information such as the URL supplied from the MPD processing unit 261, and acquires the sidx box and the ssix box from the header of the audio file of the object metadata track of the segment to be played.

[0593] In step S357, the meta file acquisition unit 262 acquires the position information of the sub-segment to be played from the sidx box and the ssix box acquired in step S356.

[0594] In step S358, the metafile acquisition unit 262, based on the location information obtained in step S357 and information such as the URL of the audio file of the object metadata track of the segment to be played, requests the Web server 142 to transmit the audio stream of the object metadata track arranged in the mdat box of the audio file, and acquires the audio stream of the object metadata track.

[0595] In step S359, the metadata acquisition unit 262 decodes the audio stream of the object metadata track acquired in step S358 based on the encoding / decoding information contained in the sample entries acquired in step S355. The metadata acquisition unit 262 provides the audio selection unit 263 with the object location information contained in the metadata as a result of the decoding.

[0596] In step S360, the audio selection unit 263 selects an object in the display area based on the image frame size information and the object location information provided by the metafile acquisition unit 262, and based on the display area specified by the user. The audio selection unit 263 requests the MPD processing unit 261 to send information such as the URL of the audio file of the object's audio track.

[0597] In step S361, MPD processing unit 261 analyzes the MPD file provided by MPD acquisition unit 91 and specifies the "Representation" of the selected object's audio based on basic attributes and the codec described in the "Representation". Furthermore, MPD processing unit 261 extracts information (such as the URL of the audio file of the selected object's audio track, described in the "Segment" contained in the "Representation"), and provides the extracted information to audio file acquisition unit 264 via audio selection unit 263.

[0598] In step S362, based on the associationId of the “Representation” specified in step S361, the MPD processing unit 261 specifies the “Representation” of the base track as the reference track. The MPD processing unit 261 extracts information (such as the URL of the audio file of the reference track described in the “Segment” contained in the “Representation”) and provides the extracted information to the audio file acquisition unit 264 via the audio selection unit 263.

[0599] In step S363, the audio file acquisition unit 264 requests the Web server 212 to transmit the initial fragment of the audio file of the object audio track and the reference track of the segment to be played based on information such as the URL supplied from the audio selection unit 263, and acquires the initial fragment.

[0600] In step S364, the audio file acquisition unit 264 acquires the sample entry in the trak box of the acquired initial fragment. The audio file acquisition unit 264 supplies the codec information contained in the sample entry to the audio decoding processing unit 194.

[0601] In step S365, the audio file acquisition unit 264 transmits a request to the Web server 142 based on information such as the URL supplied from the audio selection unit 263, and acquires the sidx box and the ssix box from the head of the audio file of the object audio track of the segment to be played.

[0602] In step S366, the audio file acquisition unit 264 acquires the position information of the sub-segment to be played from the sidx box and the ssix box acquired in step S365.

[0603] In step S367, the audio file acquisition unit 264 requests the Web server 142 to transmit the audio stream of the object audio track arranged in the mdat box within the audio file based on the position information acquired in step S366 and information such as the URL of the audio file of the object audio track of the segment to be played, and acquires the audio stream of the object audio track. The audio file acquisition unit 264 supplies the acquired audio stream of the object audio track to the audio decoding processing unit 194.

[0604] The processes of steps S368 and S369 are similar to those of steps S287 and S288 as Figure 50 shown in FIG. 14, and thus the description thereof is omitted.

[0605] Note that in the above description, the audio selection unit 263 selects all the objects in the display region. However, the audio selection unit 263 can select only the objects having a high processing priority in the display region, or can select only the audio objects of predetermined content.

[0606] Figure 65 A flowchart of the object audio playback process when the audio selection unit 263 selects only the objects having a high processing priority among the objects in the display region.

[0607] The object audio playback process as Figure 65 shown in FIG. 15 is similar to the object audio playback process as Figure 64 shown in FIG. 14, except that the process of step S390 as Figure 65 shown in FIG. 15 is performed instead of the process of step S365 asFigure 64 The process of steps S360 and S361 to S369 shown in FIG. 12 is similar to the process of steps S351 to S359 and steps S361 to S369 shown in FIG. 11. Thus, only the process of step S360 will be described below. Figure 65 The process of steps S381 to S389 and steps S391 to S399 shown in FIG. 13 is similar to the process of steps S381 to S389 and steps S391 to S399 shown in FIG. 12. Thus, only the process of step S390 will be described below. Figure 64 The process of steps S351 to S359 and steps S361 to S369 shown in FIG. 12 is similar to the process of steps S351 to S359 and steps S361 to S369 shown in FIG. 11. Thus, only the process of step S360 will be described below.

[0608] The process of steps S351 to S359 and steps S361 to S369 shown in FIG. 12 is similar to the process of steps S351 to S359 and steps S361 to S369 shown in FIG. 11. Thus, only the process of step S360 will be described below. Figure 65 In step S390 shown in FIG. 12, the audio file acquisition unit 264 selects an object having a high processing priority among the objects in the display area based on the image frame size information, the object position information, the display area, and the priority of each object. Specifically, the audio file acquisition unit 264 specifies each object of the display area based on the image frame size information, the object position information, and the display area. The audio file acquisition unit 264 selects an object having a priority equal to or higher than a predetermined value from among the specified objects. Note that, for example, the MPD processing unit 261 analyzes the MPD file to thereby acquire the priority from the "Representation" of the object audio of the specified object. The audio selection unit 263 requests the MPD processing unit 261 to transmit information such as the URL of the audio file of the object audio track of the selected object.

[0609] Figure 66 A flowchart showing the process of the object audio playback when the audio selection unit 263 selects only the audio object of the predetermined content having a high processing priority among the objects in the display area.

[0610] The process of the object audio playback shown in FIG. 13 is similar to the process of the object audio playback shown in FIG. 12, except that the process of step S420 shown in FIG. 13 is performed instead of the process of step S360 shown in FIG. 12. Specifically, the process of steps S381 to S389 and steps S391 to S399 shown in FIG. 13 is similar to the process of steps S381 to S389 and steps S391 to S399 shown in FIG. 12. Figure 66 The process of steps S351 to S359 and steps S361 to S369 shown in FIG. 12 is similar to the process of steps S351 to S359 and steps S361 to S369 shown in FIG. 11. Thus, only the process of step S360 will be described below. Figure 64 The process of steps S351 to S359 and steps S361 to S369 shown in FIG. 12 is similar to the process of steps S351 to S359 and steps S361 to S369 shown in FIG. 11. Thus, only the process of step S360 will be described below. Figure 66 The process of steps S351 to S359 and steps S361 to S369 shown in FIG. 12 is similar to the process of steps S351 to S359 and steps S361 to S369 shown in FIG. 11. Thus, only the process of step S360 will be described below. Figure 64 The process of steps S351 to S359 and steps S361 to S369 shown in FIG. 12 is similar to the process of steps S351 to S359 and steps S361 to S369 shown in FIG. 11. Thus, only the process of step S360 will be described below. Figure 66 The process of steps S351 to S359 and steps S361 to S369 shown in FIG. 12 is similar to the process of steps S351 to S359 and steps S361 to S369 shown in FIG. 11. Thus, only the process of step S360 will be described below. Figure 64 The process of steps S351 to S359 and steps S361 to S369 shown in FIG. 12 is similar to the process of steps S351 to S359 and steps S361 to S369 shown in FIG. 11. Thus, only the process of step S360 will be described below.

[0611] The process of steps S351 to S359 and steps S361 to S369 shown in FIG. 12 is similar to the process of steps S351 to S359 and steps S361 to S369 shown in FIG. 11. Thus, only the process of step S360 will be described below. Figure 66In step S420 shown, the audio file acquisition unit 264 selects an audio object of predetermined content having a high processing priority in the display area based on the image frame size information, the object position information, the display area, the priority of each object, and the content category of each object. Specifically, the audio file acquisition unit 264 specifies each object in the display area based on the image frame size information, the object position information, and the display area. The audio file acquisition unit 264 selects an object having a priority equal to or higher than a predetermined value and having a content category indicated by the predetermined value from among the specified objects.

[0612] It should be noted that, for example, the MPD processing unit 261 analyzes the MPD file to thereby acquire the priority and the content category from the "Representation" of the object audio of the specified object. The audio selection unit 263 requests the MPD processing unit 261 to transmit information such as the URL of the audio file of the object audio track of the selected object.

[0613] Figure 67 A diagram for illustrating an example of the selection of the object based on the priority.

[0614] In Figure 67 the example, the objects #1 (object 1) to #4 (object 4) are objects in the display area, and an object having a priority equal to or lower than 2 is selected from among the objects in the display area. It is assumed that the smaller the value, the higher the processing priority. Further, in Figure 67 , the value in the circle indicates the value of the priority of the corresponding object.

[0615] In the example as shown in Figure 67 , when the priorities of the objects #1 to #4 are 1, 2, 3, and 4, respectively, the object #1 and the object #2 are selected. Further, when the priorities of the objects #1 to #4 are changed to 3, 2, 1, and 4, respectively, the object #2 and the object #3 are selected. Further, when the priorities of the objects #1 to #4 are changed to 3, 4, 1, and 2, respectively, the object #3 and the object #4 are selected.

[0616] As described above, only the audio stream of the object audio of the object having a high processing priority is selectively acquired from among the objects in the display area, and the frequency band between the Web server 142 (212) and the video playback terminal 144 (214) is efficiently utilized. The same applies to the case where the object is selected based on the content category of the object.

[0617] <5th Embodiment>

[0618] (Outline of Track)

[0619] Figure 68 A diagram for illustrating an outline of the track in which the 5th embodiment of the present disclosure is applied.

[0620] As Figure 68 shown in FIG. 6, the fifth embodiment differs from the second embodiment in that the tracks are recorded as different files (3da_base.mp4 / 3da_channel.mp4 / 3da_object_1.mp4 / 3da_hoa.mp4 / 3da_meta.mp4).

[0621] The information processing system according to the fifth embodiment and the process executed by the information processing system are similar to the fourth embodiment, and thus the description thereof is omitted.

[0622] <Sixth Embodiment>

[0623] (Outline of Tracks)

[0624] Figure 69 A schematic diagram showing an outline of tracks in the sixth embodiment of the present disclosure.

[0625] As Figure 69 shown in FIG. 6, the fifth embodiment differs from the second embodiment in that the tracks are recorded as different files (3da_base.mp4 / 3da_channel.mp4 / 3da_object_1.mp4 / 3da_hoa.mp4 / 3da_meta.mp4).

[0626] The information processing system according to the sixth embodiment and the process executed by the information processing system are similar to the fourth embodiment, except that the audio stream of the base track instead of the object metadata track is acquired in order to acquire the object position information. Thus, the description thereof is omitted.

[0627] Note that in the first to third embodiments, the fifth embodiment, and the sixth embodiment, the object in the display area can also be selected based on the priority or the content category of the object.

[0628] Further, in the first to sixth embodiments, the stream playback unit can acquire the audio stream of the object outside the display area and synthesize the object audio of the object and output the object audio, as with the stream playback unit 120 shown in Figure 23 .

[0629] Further, in the first to sixth embodiments, the object position information is acquired from the metadata, but instead, the object position information can be acquired from the MPD file.

[0630] <Explanation of Hierarchical Structure of 3D Audio>

[0631] Figure 70 A schematic diagram showing a hierarchical structure of 3D audio.

[0632] As Figure 70As shown, an audio element (element) different for each audio data is used as the audio data of the 3D audio. As the type of the audio element, there are a single channel element (SCE) and a channel pair element (CPE). The type of the audio element of the audio data for one channel is SCE, and the type of the audio element of the audio data corresponding to two channels is CPE.

[0633] Audio elements of the same audio type (channel / object / SAOC object / HOA) form a group. Examples of the group type (GroupType) include channel, object, SAOC object, and HOA. When necessary, two or more groups can form a switch group or a group preset.

[0634] A switch group defines a group of audio elements to be played separately. Specifically, as shown in Figure 70 when there are an object audio group for English (EN) and an object audio group for French (FR), one of the groups is to be played. Thus, the switch group is formed of the object audio group for English with a group ID of 2 and the object audio group for French with a group ID of 3. Thus, the object audio for English and the object audio for French are played separately.

[0635] On the other hand, a group preset defines a combination of groups predetermined by a content producer.

[0636] Ext elements (Ext Elements) different for each metadata are used as the metadata of the 3D audio. Examples of the type of the Ext element include object metadata, SAOC 3D metadata, HOA metadata, DRC metadata, SpatialFrame, and SaocFrame. The Ext element of the object metadata is all metadata of the object audio, and the Ext element of the SAOC 3D metadata is all metadata of the SAOC audio. Further, the Ext element of the HOA metadata is all metadata of the HOA audio and the Ext element of the dynamic range control (DRC) metadata is all metadata of the object audio, the SAOC audio, and the HOA audio.

[0637] As described above, the audio data of the 3D audio is divided in units of the audio element, the group type, the group, the switch group, and the group preset. Thus, the audio data can be divided into the audio element, the group, the switch group, or the group preset, instead of dividing the audio data into the track for each group type as described in the first to sixth embodiments (however, in this case, the object audio is divided for each object).

[0638] Further, the metadata of the 3D audio is divided in units of Ext element types (ExtElementType) or in units of audio elements corresponding to the metadata. Thus, the metadata can be divided for each audio element corresponding to the metadata, instead of being divided for each type of Ext element as described in the first to sixth embodiments.

[0639] It is assumed in the following description that the audio data is divided for each audio element; the metadata is divided for each type of Ext element; and the data of different tracks is arranged. The same applies also when using other units of division.

[0640] <Explanation of the first example of the Web server process>

[0641] Figure 71 A schematic diagram showing the first example of the process of the Web server 142 (212).

[0642] In the example of Figure 71 , the 3D audio corresponding to the audio file uploaded from the file generation apparatus 141 (211) is composed of the channel audio of five channels, the object audio of three objects, and the metadata of the object audio (object metadata).

[0643] The channel audio of the five channels is divided into the channel audio of the front center (FC) channel, the channel audio of the front left / front right (FL, FR) channels, and the channel audio of the rear left / rear right (RL, RR) channels, which are arranged as data of different tracks. Further, the object audio of each object is arranged as data of different tracks. Further, the object metadata is arranged as data of one track.

[0644] Further, as shown in Figure 71 , each audio stream of the 3D audio is composed of configuration information and data in units of frames (samples). In the example of Figure 71 , in the audio stream of the audio file, the configuration information of the channel audio of the five channels, the object audio of the three objects, and the object metadata is arranged collectively, and the data items of each frame are arranged collectively.

[0645] In this case, as shown in Figure 71 , the Web server 142 (212) divides the audio stream of the audio file uploaded from the file generation apparatus 141 (211) for each track and generates the audio stream of seven tracks. Specifically, the Web server 142 (212) extracts the configuration information and the audio data of each track from the audio stream of the audio file according to information such as the ssix box, and generates the audio stream of each track. The audio stream of each track is composed of the configuration information of the track and the audio data of each frame.

[0646] Figure 72 A flowchart showing the track division process of the Web server 142 (212) is shown. This track division process is started, for example, when an audio file is uploaded from the file generation apparatus 141 (211).

[0647] In the case shown in FIG. 44, the Web server 142 (212) holds the audio stream of each track as shown in FIG. 45. The track to be played is the track of the channel audio of the front left / front right channels, the track of the channel audio of the rear left / rear right channels, the track of the object audio of the first object, and the track of the object metadata. The case described later is the same. Figure 72 In step S441, the Web server 142 (212) stores the audio file uploaded from the file generation apparatus 141.

[0648] In step S442, the Web server 142 (212) divides the audio stream constituting the audio file for each track based on the information such as the ssix box of the audio file.

[0649] In step S443, the Web server 142 (212) holds the audio stream of each track. Then the process is terminated. When the audio stream is requested from the audio file acquisition unit 192 (264) of the video playback terminal 144 (214), the audio stream is transmitted from the Web server 142 (212) to the video playback terminal 144 (214).

[0650] <Explanation of the first example of the process of the audio decoding processing unit>

[0651] Figure 73 A schematic diagram showing the first example of the process of the audio decoding processing unit 194 when the process described above with reference to FIG. 44 is executed by the Web server 142 (212) is shown. Figure 71 and 72

[0652] In the example shown in FIG. 44, the Web server 142 (212) holds the audio stream of each track as shown in FIG. 45. The track to be played is the track of the channel audio of the front left / front right channels, the track of the channel audio of the rear left / rear right channels, the track of the object audio of the first object, and the track of the object metadata. The case described later is the same. Figure 73 Figure 71 In this case, the audio file acquisition unit 192 (264) acquires the track of the channel audio of the front left / front right channels, the track of the channel audio of the rear left / right channels, the track of the object audio of the first object, and the track of the object metadata. Figure 75 The audio decoding processing unit 194 first extracts the audio stream of the metadata of the object audio of the first object from the audio stream of the track of the object metadata acquired by the audio file acquisition unit 192 (264).

[0653] Next, the audio decoding processing unit 194 extracts the audio stream of the channel audio of the front left / front right channels from the audio stream of the track of the channel audio of the front left / front right channels acquired by the audio file acquisition unit 192 (264).

[0654] Next, the audio decoding processing unit 194 extracts the audio stream of the channel audio of the rear left / rear right channels from the audio stream of the track of the channel audio of the rear left / rear right channels acquired by the audio file acquisition unit 192 (264).

[0655] Figure 73 ​​​As shown, the audio decoding processing unit 194 synthesizes the audio stream of the audio track to be played and the audio stream of the extracted metadata. Specifically, the audio decoding processing unit 194 generates an audio stream in which configuration information items contained in all audio streams are centrally arranged, and data items for each frame are centrally arranged. Furthermore, the audio decoding processing unit 194 decodes the generated audio stream.

[0656] As mentioned above, when the audio stream to be played includes an audio stream in addition to the audio stream of one channel audio track, the audio streams of two or more tracks are to be played. Therefore, the audio streams are synthesized before decoding.

[0657] On the other hand, when an audio stream with only one channel audio track is to be played, it is not necessary to synthesize the audio stream. Therefore, the audio decoding processing unit 194 directly decodes the audio stream acquired by the audio file acquisition unit 192 (264).

[0658] Figure 74 To illustrate the execution of the above reference on Web server 142 (212) Figure 71 and 72 The flowchart shows a detailed example of the decoding process of the audio decoding processing unit 194 during the process. This decoding process is performed when the track to be played includes other tracks besides a single-channel audio track. Figure 48 The steps S229 and as shown are as follows Figure 50 At least one of the processes in step S287 shown.

[0659] exist Figure 74 In step S461, the audio decoding processing unit 194 sets the number of all elements representing the number of elements included in the generated audio stream to "0". In step S462, the audio decoding processing unit 194 resets (clears) all element type information indicating the type of elements included in the generated audio stream.

[0660] In step S463, the audio decoding processing unit 194 sets tracks that were not previously identified as tracks to be processed as tracks to be processed. In step S464, the audio decoding processing unit 194 obtains, for example, the number and type of elements contained in the track to be processed from the audio stream of the track to be processed.

[0661] In step S465, the audio decoding processing unit 194 adds the number of acquired elements to the total number of elements. In step S466, the audio decoding processing unit 194 adds the type of the acquired elements to all element type information.

[0662] In step S467, the audio decoding processing unit 194 determines whether all the tracks to be played are set as the tracks to be processed. When it is determined in step S467 that not all the tracks to be played are set as the tracks to be processed, the process returns to step S463, and the process of steps S463 to S467 is repeated until all the tracks to be played are set as the tracks to be processed.

[0663] On the other hand, when it is determined in step S467 that all the tracks to be played are set as the tracks to be processed, the process proceeds to step S468. In step S468, the audio decoding processing unit 194 arranges the total number of elements and all the element type information at a predetermined position on the generated audio stream.

[0664] In step S469, the audio decoding processing unit 194 sets the tracks, among the tracks to be played, which are not determined as the tracks to be processed, as the tracks to be processed. In step S470, the audio decoding processing unit 194 sets the elements, which are contained in the tracks to be processed and are not determined as the elements to be processed, as the elements to be processed, when the elements are to be processed.

[0665] In step S471, the audio decoding processing unit 194 acquires the configuration information of the elements to be processed from the audio stream of the tracks to be processed and arranges the configuration information on the generated audio stream. At this time, the configuration information items of all the elements of all the tracks to be played are arranged continuously.

[0666] In step S472, the audio decoding processing unit 194 determines whether all the elements contained in the tracks to be processed are set as the elements to be processed. When it is determined in step S472 that not all the elements are set as the elements to be processed, the process returns to step S470, and the process of steps S470 to S472 is repeated until all the elements are set as the elements to be processed.

[0667] On the other hand, when it is determined in step S472 that all the elements are set as the elements to be processed, the process proceeds to step S473. In step S473, the audio decoding processing unit 194 determines whether all the tracks to be played are set as the tracks to be processed. When it is determined in step S473 that not all the tracks to be played are set as the tracks to be processed, the process returns to step S469, and the process of steps S469 to S473 is repeated until all the tracks to be played are set as the tracks to be processed.

[0668] On the other hand, when it is determined in step S473 that all the to-be-played tracks are set as the to-be-processed tracks, the process proceeds to step S474. In step S474, the audio decoding processing unit 194 determines the to-be-processed frame. In the process of step S474 at the first time, the head frame is determined as the to-be-processed frame. In the process of step S474 at the second and subsequent times, the frame immediately following the current frame to-be-processed is determined as the new frame to-be-processed.

[0669] In step S475, the audio decoding processing unit 194 sets the track that is not determined as the to-be-processed track among the to-be-played tracks as the to-be-processed track. In step S476, the audio decoding processing unit 194 sets the element that is not determined as the to-be-processed element among the elements included in the to-be-processed track as the to-be-processed element.

[0670] In step S477, the audio decoding processing unit 194 determines whether the to-be-processed element is an EXT element. When it is determined in step S477 that the to-be-processed element is not an EXT element, the process proceeds to step S478.

[0671] In step S478, the audio decoding processing unit 194 acquires the audio data of the to-be-processed frame of the to-be-processed element from the audio stream of the to-be-processed track and arranges the audio data on the generated audio stream. At this time, the data in the same frame of all the elements of all the to-be-played tracks are arranged continuously. After the process of step S478, the process proceeds to step S481.

[0672] On the other hand, when it is determined in step S477 that the to-be-processed element is an EXT element, the process proceeds to step S479. In step S479, the audio decoding processing unit 194 acquires the metadata of all the objects in the to-be-processed frame of the to-be-processed element from the audio stream of the to-be-processed track.

[0673] In step S480, the audio decoding processing unit 194 arranges the metadata of the to-be-played object among the acquired metadata of all the objects on the generated audio stream. At this time, the data items in the same frame of all the elements of all the to-be-played tracks are arranged continuously. After the process of step S480, the process proceeds to step S481.

[0674] In step S481, the audio decoding processing unit 194 determines whether all the elements included in the to-be-processed track are set as the to-be-processed elements. When it is determined in step S481 that not all the elements are set as the to-be-processed elements, the process returns to step S476, and the processes of steps S476 to S481 are repeated until all the elements are set as the to-be-processed elements.

[0675] On the other hand, when it is determined in step S481 that all the elements are set as elements to be processed, the process proceeds to step S482. In step S482, the audio decoding processing unit 194 determines whether all the tracks to be played are set as tracks to be processed. When it is determined in step S482 that not all the tracks to be played are set as tracks to be processed, the process returns to step S475, and the process of steps S475 to S482 is repeated until all the tracks to be played are set as tracks to be processed.

[0676] On the other hand, when it is determined in step S482 that all the tracks to be played are set as tracks to be processed, the process proceeds to step S483.

[0677] In step S483, the audio decoding processing unit 194 determines whether all the frames are set as frames to be processed. When it is determined in step S483 that not all the frames are set as frames to be processed, the process returns to step S474, and the process of steps S474 to S483 is repeated until all the frames are set as frames to be processed.

[0678] On the other hand, when it is determined in step S483 that all the frames are set as frames to be processed, the process proceeds to step S484. In step S484, the audio decoding processing unit 194 decodes the generated audio stream. Specifically, the audio decoding processing unit 194 decodes the audio stream in which the total number of elements, all the element type information, the configuration information, the audio data, and the metadata of the object to be played are arranged. The audio decoding processing unit 194 provides the audio data (object audio, channel audio, HOA audio) obtained as a result of the decoding to the audio synthesis processing unit 195. Then the process is terminated.

[0679] <Explanation of the second example of the process of the audio decoding processing unit>

[0680] Figure 75 To show the process of the audio decoding processing unit 194 when the Web server 142 (212) executes the above-described process with reference to Figure 71 and 72 , a schematic diagram of the second example of the process of the audio decoding processing unit 194.

[0681] As shown in Figure 75 , the second example of the process of the audio decoding processing unit 194 differs from the first example in that the audio stream of all the tracks is arranged on the generated audio stream and the indication of zero decoding result stream or the flag (hereinafter, referred to as zero stream) is arranged as the audio stream of the track not to be played.

[0682] Specifically, the audio file acquisition unit 192 (264) acquires the configuration information included in the audio stream of all the tracks held in the Web server 142 (212) and the data of each frame included in the audio stream of the track to be played back.

[0683] As shown in Figure 75 , the audio decoding processing unit 194 arranges the configuration information items of all the tracks on the generated audio stream. Further, the audio decoding processing unit 194 arranges the data of each frame of the track to be played back and the zero stream as the data of each frame of the track not to be played back on the generated audio stream.

[0684] As described above, since the audio decoding processing unit 194 arranges the zero stream as the audio stream of the track not to be played back on the generated audio stream, there is also the audio stream of the object not to be played back. Therefore, the metadata of the object not to be played back can be included in the generated audio stream. This eliminates the need for the audio decoding processing unit 194 to extract the audio stream of the metadata of the object to be played back from the audio stream of the track of the object metadata.

[0685] Note that the zero stream can be arranged as the configuration information of the track not to be played back.

[0686] Figure 76 To show the details of the second example of the decoding process of the audio decoding processing unit 194 when the above-described process with reference to Figure 71 and 72 is executed by the Web server 142 (212), a flowchart of the decoding process is shown in FIG. 27. The decoding process is executed when the track to be played back includes a track in addition to one channel audio track, as at least one of the process of step S229 shown in Figure 48 and the process of step S287 shown in Figure 50 .

[0687] The processes of steps S501 and S502 shown in Figure 76 are similar to the processes of steps S461 to S462 shown in Figure 74 , and thus the description thereof is omitted.

[0688] In step S503, the audio decoding processing unit 194 sets the track corresponding to the track not determined as the track to be processed among the tracks of the audio stream held in the Web server 142 (212) as the track to be processed.

[0689] The processes of steps S504 to S506 are similar to the processes of steps S464 to S466, and thus the description thereof will be omitted.

[0690] In step S507, the audio decoding processing unit 194 determines whether all the tracks corresponding to the audio stream held in the Web server 142 (212) are set as the tracks to be processed. When it is determined in step S507 that not all the tracks are set as the tracks to be processed, the process returns to step S503, and the process of steps S503 to S507 is repeated until all the tracks are set as the tracks to be processed.

[0691] On the other hand, when it is determined in step S507 that all the tracks are set as the tracks to be processed, the process proceeds to step S508. In step S508, the audio decoding processing unit 194 arranges the total number of elements and all the element type information at a predetermined position on the generated audio stream.

[0692] In step S509, the audio decoding processing unit 194 sets the tracks corresponding to the tracks of the audio stream held in the Web server 142 (212) which are not determined as the tracks to be processed, as the tracks to be processed. In step S510, the audio decoding processing unit 194 sets the elements included in the tracks to be processed which are not determined as the elements to be processed, as the elements to be processed.

[0693] In step S511, the audio decoding processing unit 194 acquires the configuration information of the elements to be processed from the audio stream of the tracks to be processed and generates the configuration information on the generated audio stream. At this time, the configuration information items of all the elements corresponding to all the tracks of the audio stream held in the Web server 142 (212) are arranged continuously.

[0694] In step S512, the audio decoding processing unit 194 determines whether all the elements included in the tracks to be processed are set as the elements to be processed. When it is determined in step S512 that not all the elements are set as the elements to be processed, the process returns to step S510, and the process of steps S510 to S512 is repeated until all the elements are set as the elements to be processed.

[0695] On the other hand, when it is determined in step S512 that all the elements are set as the elements to be processed, the process proceeds to step S513. In step S513, the audio decoding processing unit 194 determines whether all the tracks corresponding to the audio stream held in the Web server 142 (212) are set as the tracks to be processed. When it is determined in step S513 that not all the tracks are set as the tracks to be processed, the process returns to step S509, and the process of steps S509 to S513 is repeated until all the tracks are set as the tracks to be processed.

[0696] On the other hand, when it is determined in step S513 that all the tracks are set as the tracks to be processed, the process proceeds to step S514. In step S514, the audio decoding processing unit 194 determines the frame to be processed. In the process of step S514 at the first time, the head frame is determined as the frame to be processed. In the process of step S514 at the second and subsequent times, the frame immediately following the current frame to be processed is determined as the new frame to be processed.

[0697] In step S515, the audio decoding processing unit 194 sets the track, which does not correspond to the track to be processed, among the tracks of the audio stream held in the Web server 142 (212), as the track to be processed.

[0698] In step S516, the audio decoding processing unit 194 determines whether the track to be processed is the track to be played. When it is determined in step S516 that the track to be processed is the track to be played, the process proceeds to step S517.

[0699] In step S517, the audio decoding processing unit 194 sets the element, which does not correspond to the element to be processed, among the elements contained in the track to be processed, as the element to be processed.

[0700] In step S518, the audio decoding processing unit 194 acquires the audio data of the frame to be processed of the element to be processed from the audio stream of the track to be processed and arranges the audio stream on the generated audio stream. At this time, the data items in the same frame of all the elements of all the tracks corresponding to the audio stream held in the Web server 142 (212) are arranged continuously.

[0701] In step S519, the audio decoding processing unit 194 determines whether all the elements contained in the track to be processed are set as the elements to be processed. When it is determined in step S519 that not all the elements are set as the elements to be processed, the process returns to step S517, and the process of steps S517 to S519 is repeated until all the elements are set as the elements to be processed.

[0702] On the other hand, when it is determined in step S519 that all the elements are set as the elements to be processed, the process proceeds to step S523.

[0703] Further, when it is determined in step S516 that the track to be processed is not the track to be played, the process proceeds to step S520. In step S520, the audio decoding processing unit 194 sets the element, which does not correspond to the element to be processed, among the elements contained in the track to be processed, as the element to be processed.

[0704] In step S521, the audio decoding processing unit 194 arranges the zero stream of the data of the frame to be processed, which is an element to be processed, on the generated audio stream. At this time, the data items in the same frame of all elements corresponding to all tracks of the audio stream held in the Web server 142 (212) are arranged continuously.

[0705] In step S522, the audio decoding processing unit 194 determines whether all elements contained in the track to be processed are set as elements to be processed. Upon determining in step S522 that not all elements are set as elements to be processed, the process returns to step S520, and the process of steps S520 to S522 is repeated until all elements are set as elements to be processed.

[0706] On the other hand, upon determining in step S522 that all elements are set as elements to be processed, the process proceeds to step S523.

[0707] In step S523, the audio decoding processing unit 194 determines whether all tracks corresponding to the audio stream held in the Web server 142 (212) are set as tracks to be processed. Upon determining in step S522 that not all tracks are set as tracks to be processed, the process returns to step S515, and the process of steps S515 to S523 is repeated until all tracks to be played are set as tracks to be processed.

[0708] On the other hand, upon determining in step S523 that all tracks are set as tracks to be processed, the process proceeds to step S524.

[0709] In step S524, the audio decoding processing unit 194 determines whether all frames are set as frames to be processed. Upon determining in step S524 that not all frames are set as frames to be processed, the process returns to step S514, and the process of steps S514 to S524 is repeated until all frames are set as frames to be processed.

[0710] On the other hand, upon determining in step S524 that all frames are set as frames to be processed, the process proceeds to step S525. In step S525, the audio decoding processing unit 194 decodes the generated audio stream. Specifically, the audio decoding processing unit 194 decodes the audio stream in which the total number of elements, all element type information and configuration information, and data corresponding to all tracks of the audio stream held in the Web server 142 (212) are arranged. The audio decoding processing unit 194 provides the audio data (object audio, channel audio, HOA audio) obtained as a result of the decoding to the audio synthesis processing unit 195. The process is then terminated.

[0711] Explanation of a second example of the process of the Web server

[0712] Figure 77 A schematic diagram showing a second example of the process of the Web server 142 (212) is shown.

[0713] The second example of the process of the Web server 142 (212) shown in Figure 77 is the same as the first example shown in Figure 71 except that the object metadata of each object is arranged as data of different tracks in the audio file.

[0714] Therefore, as shown in Figure 77 , the Web server 142 (212) divides the audio stream of the audio file uploaded from the file generation apparatus 141 (211) for each track and generates the audio stream of nine tracks.

[0715] In this case, the track division process of the Web server 142 (212) is similar to the track division process shown in Figure 72 , and thus the description thereof is omitted.

[0716] Explanation of a third example of the audio decoding processing unit

[0717] Figure 78 A schematic diagram showing the process of the audio decoding processing unit 194 when the process described above with reference to Figure 77 is performed by the Web server 142 (212) is shown.

[0718] In the example shown in Figure 78 , the Web server 142 (212) holds the audio stream of each track shown in Figure 77 . The tracks to be played are the channel audio of the front left / front right channels, the channel audio of the rear left / rear right channels, the object audio of the first object, and the object metadata of the first object.

[0719] In this case, the audio file acquisition unit 192 (264) acquires the audio stream of the tracks of the channel audio of the front left / front right channels, the channel audio of the rear left / rear right channels, the object audio of the first object, and the object metadata of the first object. The audio decoding processing unit 194 synthesizes the acquired audio streams of the tracks to be played and decodes the generated audio stream.

[0720] As described above, when the object metadata is arranged as data of different tracks for each object, the audio decoding processing unit 194 does not need to extract the audio stream of the object metadata of the object to be played. Therefore, the audio decoding processing unit 194 can easily generate the audio stream to be decoded.

[0721] Figure 79To show details of the decoding process of the audio decoding processing unit 194 when the above-described process is executed by the Web server 142 (212), a flowchart of the decoding process is shown in FIG. 23. The decoding process is one of the processes of step S229 shown in FIG. 22 and step S287 shown in FIG. 27, which are executed when the track to be played contains a track other than a one-channel audio track. Figure 77 Figure 48 Figure 50

[0722] Figure 79 Figure 74 Figure 79 Figure 74 Figure 79 Figure 74

[0723] It should be noted that in the above-described process, the video playback terminal 144 (214) generates the audio stream to be decoded, but instead, the Web server 142 (212) can generate a combination of audio streams of combinations of tracks assumed to be played. In this case, the video playback terminal 144 (214) can play the audio of the tracks to be played only by acquiring the audio stream having the combination of the tracks to be played from the Web server 142 (212) and decoding the audio stream.

[0724] Further, the audio decoding processing unit 194 can decode the audio stream of the tracks to be played acquired from the Web server 142 (212) for each track. In this case, the audio decoding processing unit 194 needs to synthesize the audio data and metadata obtained as a result of decoding.

[0725] <Second Example of Syntax of Configuration Information Arranged in a Basic Sample>

[0726] <Second Example of Syntax of Configuration Information Arranged in a Basic Sample>

[0727] Figure 80 A schematic diagram showing a second example of the syntax of configuration information arranged in a basic sample.

[0728] In Figure 80 ​​​​​​​​​​In the example of FIG. 6, the number of elements (numElements) arranged in the basic sample is described as the configuration information. Further, as the type of each element (usacElementType) arranged in the basic sample, "ID_USAC_EXT" indicating an Ext element is described and the configuration information (mpegh3daExtElementCongfig) of the Ext element for each element is also described.

[0729] Figure 81 An example of the configuration information (mpegh3daExtElementCongfig) of the Ext element shown in FIG. 5 is shown as a schematic diagram. Figure 80

[0730] As shown in FIG. 7, "ID_EXT_ELE_EXTRACTOR" indicating an extractor as the type of the Ext element is described as the configuration information (mpegh3daExtElementCongfig) of the Ext element shown in FIG. 6. Further, the configuration information (ExtractorConfig) for the extractor is described. Figure 81 Figure 80 An example of the configuration information (ExtractorConfig) of the extractor shown in FIG. 5 is shown as a schematic diagram.

[0731] Figure 82 As shown in FIG. 8, as the configuration information (ExtractorConfig) of the extractor shown in FIG. 6, the type of the element to be referred to by the extractor (usac ElementType Extractor) is described. Further, the type of the Ext element (usacExtElementTypeExtractor) when the element type (usac ElementType Extractor) is "ID_USAC_EXT" indicating an Ext element is described. Further, the size (configLength) and the position (configOffset) of the configuration information of the element (subsample) to be referred to are described. Figure 81

[0732] Figure 82 Figure 81

[0733] (Second example of data syntax of frame unit arranged in basic sample)

[0734] Figure 83 A schematic diagram showing a second example of the data syntax in the frame unit arranged in the basic sample is shown.

[0735] As shown in FIG. 9, the type of the element (usacElementType) arranged in the basic sample is described as the configuration information. Further, as the type of each element (usacElementType) arranged in the basic sample, "ID_USAC_EXT" indicating an Ext element is described and the configuration information (mpegh3daExtElementCongfig) of the Ext element for each element is also described. Figure 83 ​​​​​​As shown, as data arranged in a frame unit in the basic sample, "ID_EXT_ELE_EXTRACTOR" indicating an extractor as a type of an Ext element, which is a data element, is described. Extractor data (Extractor Metadata) is also described.

[0736] Figure 84 A schematic diagram showing an exemplary syntax of the extractor data (Extractor Metadata) shown in Figure 83

[0737] As shown in Figure 84 , the size (elementLength) and the position (elementOffset) of the data of the element to be referenced by the extractor are described as the extractor data (Extractor Metadata) as shown in Figure 83

[0738] <Third example of syntax of a basic sample>

[0739] (Third example of syntax of configuration information arranged in a basic sample)

[0740] Figure 85 A schematic diagram showing a third example of the syntax of the configuration information arranged in the basic sample.

[0741] In the example of Figure 85 , the number of elements (numElements) arranged in the basic sample is described as the configuration information. Further, "1" indicating the extractor is described as a flag Extractor, which indicates whether the sample in which the configuration information is arranged is the extractor. Further, "1" is described as elementLengthPresent.

[0742] Further, the element type to be referenced by the element is described as a type (usacElementType) of each element arranged in the basic sample. When the element type (usacElementType) is "ID_USAC_EXT" indicating an Ext element, a type (usacExtElementType) of the Ext element is described. Further, the size (configLength) and the position (configOffset) of the configuration information of the element to be referenced are described.

[0743] (Third example of syntax of data arranged in a frame unit in a basic sample)

[0744] Figure 86 ​​A diagram for showing a third example of data syntax arranged in a frame unit in a basic sample.

[0745] As Figure 86 shown, as data arranged in a frame unit in a basic sample, the size (elementLength) and position (elementOffset) of data of an element to be referenced by data are described.

[0746] <Seventh Embodiment>

[0747] (Configuration Example of Audio Stream)

[0748] Figure 87 A diagram for showing a configuration example of an audio stream stored in an audio file in the seventh embodiment of the information processing system to which the present disclosure is applied.

[0749] As Figure 87 shown, in the seventh embodiment, the audio file stores, in units of samples of 3D audio for each group type, encoded data (however, in this case, object audio is stored for each object) and an audio stream (3D audio stream) arranged as a sub-sample.

[0750] Further, the audio file stores a cue stream (3D audio cue stream) in which an extractor containing the size, position, and group type of encoded data in units of samples of 3D audio for each group type is arranged as a sub-sample. The configuration of this extractor is similar to the above-described configuration, and the group type is described as the type of the extractor.

[0751] (Outline of Track)

[0752] Figure 88 A diagram for showing an outline of a track in the seventh embodiment.

[0753] As Figure 88 shown, in the seventh embodiment, different tracks are respectively assigned to the audio stream and the cue stream. The track ID "2" of the track corresponding to the cue stream is described as the track reference number of the track of the audio stream. Further, the track ID "1" of the track of the corresponding audio stream is described as the track reference number of the track of the cue stream.

[0754] The syntax of the sample entry of the track of the audio stream is the syntax as Figure 34 shown, and the syntax of the sample entry of the track of the cue stream contains the syntax as Figures 35 to 38 shown.

[0755] (Explanation of Process of File Generation Apparatus)

[0756] Figure 89A flowchart showing a file generation process of the file generation apparatus in the seventh embodiment.

[0757] It should be noted that the file generation apparatus according to the seventh embodiment is the same as the file generation apparatus 141 as shown in FIG. 1 except for the processes of the audio encoding processing unit 171 and the audio file generation unit 172. Therefore, the file generation apparatus, the audio encoding processing unit, and the audio file generation unit according to the seventh embodiment are hereinafter referred to as a file generation apparatus 301, an audio encoding processing unit 341, and an audio file generation unit 342, respectively. Figure 45

[0758] The processes of steps S601 to S605 as shown in FIG. 6 are similar to the processes of steps S191 to S195 as shown in FIG. 2, and therefore the description thereof is omitted. Figure 89 Figure 46

[0759] In step S606, the audio encoding processing unit 341 encodes the 3D audio of the video content inputted from the outside for each group type and generates an audio stream as shown in FIG. 4. The audio encoding processing unit 341 provides the generated audio stream to the audio file generation unit 342. Figure 87

[0760] In step S607, the audio file generation unit 342 acquires sub-sample information from the audio stream provided from the audio encoding processing unit 341. The sub-sample information indicates the size, position, and group type of the encoded data in units of samples of the 3D audio of each group type.

[0761] In step S608, the audio file generation unit 342 generates a cue stream as shown in FIG. 5 based on the sub-sample information. In step S609, the audio file generation unit 342 multiplexes the audio stream and the cue stream into different tracks and generates an audio file. At this time, the audio file generation unit 342 stores the image frame size information inputted from the outside in a sample entry. The audio file generation unit 342 provides the generated audio file to the MPD generation unit 173. Figure 87

[0762] The processes of steps S610 and S611 as shown in FIG. 6 are similar to the processes of steps S199 and S200 as shown in FIG. 2, and therefore the description thereof is omitted. Figure 46

[0763] (Explanation of the process of the video playback terminal)

[0764] Figure 90 A flowchart showing an audio playback process of the stream playback unit of the video playback terminal in the seventh embodiment.

[0765] It should be noted that the stream playback unit according to the seventh embodiment is the same as the stream playback unit 151 as shown in FIG. 1 except for the processes of the audio decoding processing unit 351 and the audio file playback unit 352. Therefore, the stream playback unit, the audio decoding processing unit, and the audio file playback unit according to the seventh embodiment are hereinafter referred to as a stream playback unit 451, an audio decoding processing unit 451, and an audio file playback unit 452, respectively. Figure 47 ​​​​​​The stream playback unit 190 shown is the same except that the processes of the MPD processing unit 191, the audio file acquisition unit 192, and the audio decoding processing unit 194 are different and the audio selection unit 193 is not provided. Therefore, the stream playback unit, the MPD processing unit, the audio file acquisition unit, and the audio decoding processing unit according to the seventh embodiment are hereinafter referred to as a stream playback unit 360, an MPD processing unit 381, an audio file acquisition unit 382, and an audio decoding processing unit 383, respectively.

[0766] In the step S621 shown, the MPD processing unit 381 of the stream playback unit 360 analyzes the MPD file provided from the MPD acquisition unit 91, acquires information such as the URL of the audio file of the segment to be played, and provides the acquired information to the audio file acquisition unit 382. Figure 90

[0767] In the step S622, the audio file acquisition unit 382 requests the Web server to send and acquire the initial segment of the segment to be played based on information such as the URL provided from the MPD processing unit 381.

[0768] In the step S623, the audio file acquisition unit 382 acquires the track ID of the track of the audio stream that is the reference track from the sample entry of the track of the hint stream of the moov box in the initial segment (hereinafter referred to as the hint track).

[0769] In the step S624, the audio file acquisition unit 382 requests the Web server to send and acquire the sidx box and the ssix box from the header of the media segment of the segment to be played based on information such as the URL provided from the MPD processing unit 381.

[0770] In the step S625, the audio file acquisition unit 382 acquires the position information of the hint track from the sidx box and the ssix box acquired in the step S624.

[0771] In the step S626, the audio file acquisition unit 382 requests the Web server to send and acquire the hint stream based on the position information of the hint track acquired in the step S625. Further, the audio file acquisition unit 382 acquires the extractor of the group type of the 3D audio to be played from the hint stream. Note that when the 3D audio to be played is object audio, the object to be played is selected based on the image frame size information and the object position information.

[0772] ​In step S627, the audio file acquisition unit 382 acquires the position information of the reference track from the sidx box and the ssix box acquired in step S624. In step S628, the audio file acquisition unit 382 determines the position information of the audio stream of the group type of the 3D audio to be played based on the position information of the reference track acquired in step S627 and the subsample information contained in the acquired extractor.

[0773] In step S629, the audio file acquisition unit 382 requests the Web server to send the audio stream of the group type of the 3D audio to be played based on the position information determined in step S627 and acquires the audio stream. The audio file acquisition unit 382 provides the acquired audio stream to the audio decoding processing unit 383.

[0774] In step S630, the audio decoding processing unit 383 decodes the audio stream provided from the audio file acquisition unit 382 and provides the audio data obtained as a result of the decoding to the audio synthesis processing unit 195.

[0775] In step S631, the audio synthesis processing unit 195 outputs the audio data. Then the process is terminated.

[0776] Note that in the seventh embodiment, the track of the audio stream and the cue track are stored in the same audio file, but can be stored in different files.

[0777] <Embodiment 8>

[0778] (Outline of Track)

[0779] Figure 91 A schematic diagram showing an outline of a track in the eighth embodiment of the information processing system to which the present disclosure is applied.

[0780] The audio file of the eighth embodiment differs from the audio file of the seventh embodiment in that the cue stream stored is a stream for each group type. Specifically, the cue stream of the eighth embodiment is generated for each group type, and an extractor containing the size, position, and group type of the encoded data in units of samples of the 3D audio of each group type is arranged in each cue stream. Note that when the 3D audio contains object audio of a plurality of objects, the extractor is arranged as a subsample for each object.

[0781] Further, as Figure 91 shown in the eighth embodiment, a different track is assigned to the audio stream and each cue stream. The track of this audio stream is the same as the track of the audio stream as Figure 88 shown, and thus the description thereof is omitted.

[0782] The track ID "1" of the track of the corresponding audio stream is described as the track reference number of the cue track of the group type "channel", "object", "HOA", and "metadata".

[0783] The syntax of the sample entry of the cue track of each of the group types "channel", "object", "HOA", and "metadata" is the same as the syntax shown in Figures 35 to 38 except for the information indicating the type of the sample entry. The information indicating the type of the sample entry of the cue track of each of the group types "channel", "object", "HOA", and "metadata" is the same as the information shown in Figures 35 to 38 except that the number "1" of the information is replaced with "2". The number "2" indicates the sample entry of the cue track.

[0784] (Configuration example of audio file)

[0785] Figure 92 A schematic diagram showing a configuration example of an audio file.

[0786] As shown in Figure 92 , the audio file stores all the tracks shown in Figure 91 . Specifically, the audio file stores the audio stream and the cue stream of each group type.

[0787] The file generation process of the file generation apparatus according to the eighth embodiment is similar to the file generation process shown in Figure 89 except that the cue stream is generated for each group type, contrary to the cue stream shown in Figure 87 .

[0788] Further, the audio playback process of the stream playback unit of the video playback terminal according to the eighth embodiment is similar to the audio playback process shown in Figure 90 except that the track ID of the cue track of the group type to be played and the track ID of the reference track acquired in step S623; the position information of the cue track of the group type to be played is acquired in step S625; and the cue stream of the group type to be played is acquired in step S626.

[0789] It should be noted that in the eighth embodiment, the tracks of the audio stream and the cue track are stored in the same audio file, but can be stored in different files.

[0790] For example, as shown in Figure 93 , the tracks of the audio stream can be stored in one audio file (3D audio stream MP4 file), and the cue track can be stored in one audio file (3D audio cue stream MP4 file). Further, as shown in Figure 94 , the cue track can be divided into a plurality of audio files to be stored. In Figure 94In the example of the seventh embodiment, the cue track is stored in a different audio file.

[0791] Further, in the eighth embodiment, a cue stream is generated for each group type even when the group type indicates an object. However, when the group type indicates an object, a cue stream can be generated for each object. In this case, a different track is assigned to the cue stream of each object.

[0792] As described above, in the audio files of the seventh and eighth embodiments, the audio streams of 3D audio are stored in one track. Therefore, the video playback terminal can play all the audio streams of 3D audio by acquiring the track.

[0793] Further, the cue stream is stored in the audio files of the seventh and eighth embodiments. Therefore, the video playback terminal can acquire the audio stream of the desired group type among all the audio streams of 3D audio without referring to the moof box in which the table associating a subsample with the size or position of the subsample is described, thus making it possible to play the audio stream.

[0794] Further, in the audio files of the seventh and eighth embodiments, the video playback terminal can be caused to acquire the audio stream of each group type only by storing all the audio streams of 3D audio and the cue stream. Therefore, it is not necessary to prepare the audio stream of 3D audio separately for each group type from all the generated audio streams of 3D audio for the purpose of broadcasting or local storage in order to be able to acquire the audio stream for each group type.

[0795] Note that in the seventh and eighth embodiments, the extractor is generated for each group type, but can be generated in units of audio elements, groups, switch groups, or group presets.

[0796] When the extractor is generated in units of groups, the sample entry of each cue track of the eighth embodiment contains information on the corresponding group. The information on the group is composed of, for example, information indicating the ID of the group and the contents of the data classified as the elements of the group. When the group forms a switch group, the sample entry of the cue track of the group also contains information on the switch group. The information on the switch group is composed of, for example, the ID of the switch group and the IDs of the groups forming the switch group. The sample entry of the cue track of the seventh embodiment contains the information contained in the sample entries of all the cue tracks of the eighth embodiment.

[0797] Further, the segment structure in the seventh and eighth embodiments is the same as the segment structure as shown in Figure 39 and 40 .

[0798] <Embodiment 9>

[0799] (Explanation of a computer to which the present disclosure is applied)

[0800] The series of processes of the Web server described above can also be executed by hardware or software. When the series of processes is executed by software, a program constituting the software is installed in a computer. Examples of the computer include a computer incorporating a dedicated hardware and a general personal computer capable of executing various functions by installing various programs.

[0801] Figure 95 A block diagram showing an example of the configuration of the hardware of the computer that executes the series of processes of the Web server by using a program.

[0802] In the computer, a central processing unit (CPU) 601, a read only memory (ROM) 602, and a random access memory (RAM) 603 are interconnected via a bus 604.

[0803] The bus 604 is also connected to an input / output interface 605. The input / output interface 605 is connected to each of an input unit 606, an output unit 607, a storage unit 608, a communication unit 609, and a drive 610.

[0804] The input unit 606 is formed of a keyboard, a mouse, a microphone, and the like. The output unit 607 is formed of a display, a speaker, and the like. The storage unit 608 is formed of a hardware, a non-volatile memory, and the like. The communication unit 609 is formed of a network interface and the like. The drive 610 drives a removable medium 611 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.

[0805] In the computer configured as described above, the CPU 601 loads a program stored in the storage unit 608 in the RAM 603, for example, via the input / output interface 605 and the bus 604 and executes the program, thereby executing the series of processes described above.

[0806] The program executed by the computer (CPU 601) can be provided recorded in the removable medium 611 serving as a package medium or the like. Further, the program can be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.

[0807] The program can be installed in the storage unit 608 via the input / output interface 605 by loading the removable medium 611 in the drive 610. Further, the program can be received via the wired or wireless transmission medium by the communication unit 609 and installed in the storage unit 608. Further, the program can be installed in the ROM 602 or the storage unit 608 in advance.

[0808] Note that the program executed by the computer can be a program that executes the processes in a time series in the order described in this description, or can be a program that executes the processes in parallel or at necessary times, for example, when solicited.

[0809] The video playback terminal described above can have a hardware configuration similar to that of the computer as shown in FIG. 6. In this case, for example, the CPU 601 can execute the control software 161 (221), the video playback software 162 (222), and the access software 163 (223). The processes of the video playback terminal 144 (214) can be executed by hardware. Figure 95

[0810] In the present description, a system has a plurality of components (such as devices or modules (parts)) and it is not considered whether all the components are in the same housing. Therefore, the system can be a plurality of devices that can be stored in separate housings and connected through a network and a plurality of modules within a single housing.

[0811] Note that the embodiments of the present disclosure are not limited to the above-described embodiments and various changes can be made without departing from the gist of the present disclosure.

[0812] For example, the file generation device 141 (211) can generate a video stream by multiplexing the encoded data of all tiles to generate one image file instead of generating image files in tile units.

[0813] The present disclosure can be applied not only to MPEG-H 3D audio but also to a general audio codec capable of forming a stream of each object.

[0814] Further, the present disclosure can also be applied to an information processing system that performs broadcasting and local storage playback as well as stream playback.

[0815] Further, the present disclosure can have the following configurations. (1)

[0817] An information processing device including an acquisition unit that acquires audio data of a predetermined track in a file in which a plurality of types of audio data are divided into a plurality of tracks according to the types and the tracks are arranged. (2)

[0819] The information processing device according to the above item (1), in which the type is configured as an element of the audio data, a type of the element, or a group into which the element is classified. (3)

[0821] The information processing device according to the above item (1) or (2), further including a decoding unit that decodes the audio data of the predetermined track acquired by the acquisition unit. (4)

[0823] ​The information processing apparatus according to the above item (3), wherein, when there are a plurality of predetermined tracks, the decoding unit synthesizes audio data of the predetermined tracks acquired by the acquisition unit and decodes the synthesized audio data. (5)

[0825] The information processing apparatus according to the above item (4), wherein

[0826] The file is configured in such a manner that audio data in units of a plurality of objects is divided into the tracks different for each object and the tracks are arranged, and metadata items of all the audio data in units of objects are collectively arranged in a track different from the tracks,

[0827] The acquisition unit is configured to acquire audio data of the tracks of the object to be played as the audio data of the predetermined tracks and acquire the metadata, and

[0828] The decoding unit is configured to extract metadata of the object to be played from the metadata acquired by the acquisition unit and synthesize the metadata and the audio data acquired by the acquisition unit. (6)

[0830] The information processing apparatus according to the above item (4), wherein

[0831] The file is configured in such a manner that audio data in units of a plurality of objects is divided into the tracks different for each object and the tracks are arranged, and metadata items of all the audio data in units of objects are collectively arranged in a track different from the tracks,

[0832] The acquisition unit is configured to acquire audio data of the tracks of the object to be played as the audio data of the predetermined tracks and acquire the metadata, and

[0833] The decoding unit is configured to synthesize zero data and audio data and the metadata acquired by the acquisition unit, the zero data indicating a decoding result of zero as audio data of a track not played. (7)

[0835] The information processing apparatus according to the above item (4), wherein

[0836] The file is configured in such a manner that audio data in units of a plurality of objects is divided into the tracks different for each object and the tracks are arranged, metadata items of the audio data in units of objects are arranged in the tracks different for each object,

[0837] The acquisition unit is configured to acquire audio data of a track of an object to be played as audio data of a predetermined track and acquire metadata of the object to be played, and

[0838] The decoding unit is configured to synthesize the audio data and the metadata acquired by the acquisition unit. (8)

[0840] The information processing apparatus according to any one of the above (1) to (7), wherein the audio data items of the plurality of tracks are configured to be arranged in one file. (9)

[0842] The information processing apparatus according to any one of the above (1) to (7), wherein the audio data items of the plurality of tracks are configured to be arranged in files different for each track. (10)

[0844] The information processing apparatus according to any one of the above (1) to (9), wherein the file is configured in such a manner that information on the plurality of types of the audio data is arranged as a track different from the plurality of tracks. (11)

[0846] The information processing apparatus according to the above (10), wherein the information on the plurality of types of the audio data is configured to contain image frame size information indicating a size of an image frame of image data corresponding to the audio data. (12)

[0848] The information processing apparatus according to any one of the above (1) to (9), wherein the file is configured in such a manner that, as the audio data different from the plurality of tracks, information indicating a position of the audio data of another track corresponding to the audio data is arranged. (13)

[0850] The information processing apparatus according to any one of the above (1) to (9), wherein the file is configured in such a manner that, as the data different from the plurality of tracks, information indicating a position of the audio data of another track corresponding to the data and the metadata of the other track is arranged. (14)

[0852] The information processing apparatus according to the above (13), wherein the metadata of the audio data is configured to contain information indicating a position at which the audio data is acquired. (15)

[0854] The information processing apparatus according to any one of (1) to (14) above, wherein the file is configured to contain information indicating a reference relationship between the track and another track. (16)

[0856] The information processing apparatus according to any one of (1) to (15) above, wherein the file is configured to contain codec information of audio data of each track. (17)

[0858] The information processing apparatus according to any one of (1) to (16) above, wherein the predetermined type of audio data is information indicating a position at which another type of audio data is acquired. (18)

[0860] An information processing method including an acquisition step of acquiring, by an information processing apparatus, audio data of a predetermined track in a file in which a plurality of types of audio data are divided into a plurality of tracks according to the types and the tracks are arranged. (19)

[0862] An information processing apparatus including a generation unit that generates a file in which a plurality of types of audio data are divided into a plurality of tracks according to the types and the tracks are arranged. (20)

[0864] An information processing method including a generation step of generating, by an information processing apparatus, a file in which a plurality of types of audio data are divided into a plurality of tracks according to the types and the tracks are arranged.

[0865] List of Reference Signs

[0866] 141 file generation apparatus

[0867] 144 moving image playback terminal

[0868] 172 audio file generation unit

[0869] 192 audio file acquisition unit

[0870] 193 audio selection unit

[0871] 211 file generation apparatus

[0872] 214 moving image playback terminal

[0873] 241 audio file generation unit

[0874] 264 audio file acquisition unit

Claims

1. An information processing apparatus comprising: circuitry configured to: acquire an audio file composed of a plurality of audio streams respectively assigned to tracks by type, wherein a plurality of types of audio streams are divided into a plurality of tracks according to the type, and the tracks are arranged, wherein each audio stream contains object position information and information about a priority of each object, and wherein an object is selected based on the information about the priority; decode at least one audio stream assigned to a track of the plurality of tracks from the audio file, perform rendering based on the audio data, the object position information, and the priority of each object.

2. The information processing apparatus according to claim 1, wherein the type is configured as an element of audio data, a type of the element, or a group to which the element is classified.

3. The information processing apparatus according to claim 1, wherein the circuitry is configured to decode audio data of a predetermined track.

4. The information processing apparatus according to claim 3, wherein when there are a plurality of predetermined tracks, the circuitry synthesizes the audio data of the predetermined tracks and decodes the synthesized audio data. 5.The information processing apparatus according to claim 4, wherein the file is configured in such a manner that audio data in units of a plurality of objects is divided into the tracks that differ with respect to each object and the tracks are arranged, and metadata items of all the audio data in units of objects are arranged collectively in a track that is different from the tracks, the circuitry is configured to acquire audio data of the track of the object to be played as the audio data of the predetermined track and acquire the metadata, and the circuitry is configured to extract metadata of the object to be played from the metadata and synthesize the metadata and the audio data. 6.The information processing apparatus according to claim 4, wherein the file is configured in such a manner that audio data in units of a plurality of objects is divided into the tracks that differ with respect to each object and the tracks are arranged, and metadata of all the audio data in units of objects are arranged collectively in a track that is different from the tracks, the circuitry is configured to acquire audio data of the track of the object to be played as the audio data of the predetermined track and acquire the metadata, and the circuitry is configured to synthesize the audio data and the metadata with zero data indicating a decoding result of zero as audio data of a track that is not played. 7.The information processing apparatus according to claim 4, wherein the file is configured in such a manner that audio data in units of a plurality of objects is divided into the tracks that differ with respect to each object and the tracks are arranged, metadata of the audio data in units of objects are arranged in the tracks that differ with respect to each object, the circuitry is configured to acquire audio data in the track of the object to be played as the audio data of the predetermined track and acquire metadata of the object to be played, and the circuitry synthesizes the audio data and the metadata.

8. The information processing apparatus according to claim 1, wherein the audio data of the plurality of tracks is arranged in one file.

9. The information processing apparatus according to claim 1, wherein the audio data of the plurality of tracks is arranged in the files that differ with respect to each track.

10. The information processing apparatus according to claim 1, wherein The file is configured in such a way that information about the plurality of types of the audio data is arranged as a track different from the plurality of tracks.

Citation Information

Patent Citations

  • Method and apparatus for track and track subset grouping

    CN102132562A

  • Method for creating and accessing a menu for audio content without using a display

    CN1735941A