Information processing apparatus and information processing method
By dividing audio data into multiple files according to type and track, the problem of low audio data acquisition efficiency in existing technologies is solved, and efficient acquisition of audio data and file generation are achieved.
Patent Information
- Application Number
- CN202111197667.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2014-10-01
- Filing Date
- 2015-05-22
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2035-05-22
AI Technical Summary
Existing technologies have failed to effectively improve the efficiency of obtaining predetermined types of audio data from multiple types of audio data in video content.
By dividing various types of audio data according to type and track, generating multiple track files, and processing and transmitting them through computer programs, the efficiency of audio data acquisition can be improved.
It enables efficient file generation from specific types of audio data, thus improving the efficiency of audio data acquisition.
Smart Images

Figure CN114242082B_ABST
Abstract
Description
[0001] This application is a divisional application of the parent application with the application number 2015800269311 and the filing date of May 22, 2015, and the invention title of "Information processing apparatus and information processing method". TECHNICAL FIELD
[0002] The present disclosure relates to an information processing apparatus and an information processing method, and more particularly, to an information processing apparatus and an information processing method capable of improving efficiency of acquiring a predetermined type of audio data among a plurality of types of audio data. BACKGROUND
[0003] One of the most popular streaming services recently is Internet video (OTT-V) via the Internet. Moving Picture Experts Group phase dynamic adaptive streaming over HTTP (MPEG-DASH) is widely used as its underlying technology (see, for example, Non-Patent Literature 1).
[0004] In MPEG-DASH, a delivery server prepares a set of video data having different screen sizes and encoding rates for one video content item, and a playback terminal requests a set of video data having the best screen size and encoding rate according to a transmission line condition, thus realizing adaptive streaming.
[0005] List of cited documents
[0006] Non-Patent Literature
[0007] Non-Patent Literature 1: MPEG-DASH (Dynamic Adaptive Streaming over HTTP) (URL: http: / / mpeg.chiariglione.org / standards / mpeg-dash / media-presentation-description-and-segment-formats / text-isoiec-23009-12012-dam-1) SUMMARY
[0008] Problems to be Solved by the Invention
[0009] However, there is no consideration to improve efficiency of acquiring a predetermined type of audio data among a plurality of types of audio data of a video content.
[0010] The present disclosure is made in view of the above circumstances and is capable of improving efficiency of acquiring a predetermined type of audio data among a plurality of types of audio data.
[0011] Solution to Problem
[0012] The information processing apparatus according to the first aspect of the present disclosure is an information processing apparatus including an acquisition unit that acquires audio data in a predetermined track of a file, wherein a plurality of types of audio data are divided into a plurality of tracks according to types and the tracks are arranged.
[0013] The information processing method according to the first aspect of the present disclosure corresponds to the information processing apparatus according to the first aspect of the present disclosure.
[0014] In the first aspect of the present disclosure, audio data in a predetermined track of a file is acquired, wherein a plurality of types of audio data are divided into a plurality of tracks according to types and tracks that are arranged.
[0015] The information processing apparatus according to the second aspect of the present disclosure is an information processing apparatus including a generation unit that generates a file in which a plurality of types of audio data are divided into a plurality of tracks according to types and tracks that are arranged.
[0016] The information processing method according to the second aspect of the present disclosure corresponds to the information processing apparatus according to the second aspect of the present disclosure.
[0017] In the second aspect of the present disclosure, a file in which a plurality of types of audio data are divided into a plurality of tracks according to types and tracks that are arranged is generated.
[0018] Note that the information processing apparatus according to the first aspect and the second aspect can be implemented by causing a computer to execute a program.
[0019] Further, in order to realize the information processing apparatus according to the first aspect and the second aspect, the program executed by the computer can be provided by transmitting the program via a transmission medium or by recording the program in a recording medium.
[0020] Effects of the Invention
[0021] According to the first aspect of the present disclosure, audio data can be acquired. Further, according to the first aspect of the present disclosure, a specific type of audio data among a plurality of types of audio data can be efficiently acquired.
[0022] According to the second aspect of the present disclosure, a file can be generated. Further, according to the second aspect of the present disclosure, a file that improves the efficiency of acquiring a specific type of audio data among a plurality of types of audio data can be generated. BRIEF DESCRIPTION OF DRAWINGS
[0023] Figure 1 A schematic diagram showing an overview of a first example of an information processing system to which the present disclosure is applied.
[0024] Figure 2 A schematic diagram showing an example of a file.
[0025] Figure 3 A diagram for showing a schematic view of an object.
[0026] Figure 4 A diagram for showing a schematic view of object position information.
[0027] Figure 5 A diagram for showing a schematic view of image frame size information.
[0028] Figure 6 A diagram for showing a schematic view of a structure of an MPD file.
[0029] Figure 7 A diagram for showing a relationship among a "Period", a "Representation", and a "Segment".
[0030] Figure 8 A diagram for showing a schematic view of a hierarchical structure of an MPD file.
[0031] Figure 9 A diagram for showing a relationship between a structure of an MPD file and a time axis.
[0032] Figure 10 A diagram for showing an exemplary description of an MPD file.
[0033] Figure 11 A block diagram for showing a configuration example of a file generation apparatus.
[0034] Figure 12 A flowchart for showing a file generation process of a file generation apparatus.
[0035] Figure 13 A block diagram for showing a configuration example of a stream playback unit.
[0036] Figure 14 A flowchart for showing a stream playback process of a stream playback unit.
[0037] Figure 15 A diagram for showing an exemplary description of an MPD file.
[0038] Figure 16 A diagram for showing another exemplary description of an MPD file.
[0039] Figure 17 A diagram for showing an arrangement example of an audio stream.
[0040] Figure 18 A diagram for showing an exemplary description of gsix.
[0041] Figure 19 A diagram for showing an example of information indicating a correspondence relationship between a sample group entry and an object ID.
[0042] Figure 20 A schematic diagram for illustrating an exemplary description of an AudioObjectSampleGroupEntry.
[0043] Figure 21 A schematic diagram for illustrating an exemplary description of a type assignment box.
[0044] Figure 22 A schematic diagram for illustrating an overview of a second example of an information processing system to which the present disclosure is applied.
[0045] Figure 23 A block diagram for illustrating a configuration example of a stream playback unit of an information processing system to which the present disclosure is applied.
[0046] Figure 24 A schematic diagram for illustrating a method of determining a position of an object.
[0047] Figure 25 A schematic diagram for illustrating a method of determining a position of an object.
[0048] Figure 26 A schematic diagram for illustrating a method of determining a position of an object.
[0049] Figure 27 A schematic diagram for illustrating a relationship between a horizontal angle θ Ai and a horizontal angle θ Ai .
[0050] Figure 28 A flowchart for illustrating a stream playback process of a stream playback unit illustrated in Figure 23 .
[0051] Figure 29 A flowchart for illustrating details of a position determination process illustrated in Figure 28 .
[0052] Figure 30 A flowchart for illustrating details of a horizontal angle θ Ai ' estimation process illustrated in Figure 29 .
[0053] Figure 31 A schematic diagram for illustrating an overview of a track of a 3D audio file format of MP4.
[0054] Figure 32 A schematic diagram for illustrating a structure of a moov box.
[0055] Figure 33 A schematic diagram for illustrating an overview of a track according to a first embodiment to which the present disclosure is applied.
[0056] Figure 34To show in Figure 33 The diagram shows an exemplary syntax of sample entries for the basic track.
[0057] Figure 35 To show in Figure 33 The diagram shows an example syntax of sample entries for the audio track.
[0058] Figure 36 To show in Figure 33 The diagram shows an example syntax of sample entries for an object audio track.
[0059] Figure 37 To show in Figure 33 The diagram shows an example syntax of sample entries for a HOA audio track.
[0060] Figure 38 To show in Figure 33 The diagram illustrates an example syntax for sample entries of the object metadata track.
[0061] Figure 39 This is a schematic diagram illustrating a first example of a fragment structure.
[0062] Figure 40 This is a schematic diagram illustrating a second example of a fragment structure.
[0063] Figure 41 A schematic diagram illustrating an exemplary description of a level assignment box.
[0064] Figure 42 This is a schematic diagram illustrating an exemplary description of an MDF file in a first embodiment of the present disclosure.
[0065] Figure 43 This is a diagram illustrating the definition of basic attributes.
[0066] Figure 44 This is a schematic diagram illustrating an overview of an information processing system in which the present disclosure is applied in a first embodiment.
[0067] Figure 45 To show in Figure 44 A block diagram showing an example configuration of a file generation device.
[0068] Figure 46 To show in Figure 45 The flowchart shown is a process for generating files using a file generation device.
[0069] Figure 47 To show by in Figure 44 The diagram shows a configuration example of a streaming playback unit implemented in a video playback terminal.
[0070] Figure 48 A flowchart for illustrating a channel audio playback process of the stream playback unit shown in Figure 47
[0071] Figure 49 A flowchart for illustrating an object designation process of the stream playback unit shown in Figure 47
[0072] Figure 50 A flowchart for illustrating a designated object audio playback process of the stream playback unit shown in Figure 47
[0073] Figure 51 A schematic diagram for illustrating an overview of a track in the second embodiment of the present disclosure.
[0074] Figure 52 A schematic diagram for illustrating an example syntax of a sample entry of a basic track shown in Figure 51
[0075] A schematic diagram for illustrating a structure of a basic sample. Figure 53
[0076] A schematic diagram for illustrating an example syntax of a basic sample. Figure 54
[0077] A schematic diagram for illustrating an example of data of an extractor. Figure 55
[0078] A schematic diagram for illustrating an overview of a track in the third embodiment of the present disclosure. Figure 56
[0079] A schematic diagram for illustrating an overview of a track in the fourth embodiment of the present disclosure. Figure 57
[0080] A schematic diagram for illustrating an example description of an MDF file in the fourth embodiment of the present disclosure. Figure 58
[0081] A schematic diagram for illustrating an overview of an information processing system in the fourth embodiment of the present disclosure. Figure 59
[0082] A block diagram for illustrating a configuration example of the file generation apparatus shown in Figure 60 Figure 59 A flowchart for illustrating a file generation process of the file generation apparatus shown in
[0083] Figure 61 Figure 60 A flowchart for illustrating a file generation process of the file generation apparatus shown in
[0084] Figure 62 FIG. 1 is a block diagram showing a configuration example of a stream playback unit implemented by a video playback terminal shown in Figure 59 FIG. 2 is a block diagram showing a configuration example of a stream playback unit implemented by a video playback terminal shown in
[0085] Figure 63 FIG. 3 is a flowchart showing an example of a channel audio playback process of the stream playback unit shown in Figure 62 FIG. 4 is a flowchart showing an example of an object audio playback process of the stream playback unit shown in
[0086] Figure 64 FIG. 5 is a flowchart showing a first example of an object audio playback process of the stream playback unit shown in Figure 62 FIG. 6 is a flowchart showing a second example of an object audio playback process of the stream playback unit shown in
[0087] Figure 65 FIG. 7 is a flowchart showing a third example of an object audio playback process of the stream playback unit shown in Figure 62 FIG. 8 is a schematic diagram showing an example of selection of an object based on priority.
[0088] Figure 66 FIG. 9 is a schematic diagram showing an overview of a track in application of a fifth embodiment of the present disclosure. Figure 62 FIG. 10 is a schematic diagram showing an overview of a track in application of a sixth embodiment of the present disclosure.
[0089] Figure 67 FIG. 11 is a schematic diagram showing a hierarchical structure of 3D audio.
[0090] Figure 68 FIG. 12 is a schematic diagram showing a first example of a Web server process.
[0091] Figure 69 FIG. 13 is a flowchart showing a track division process of a Web server.
[0092] Figure 70 FIG. 14 is a schematic diagram showing a first example of a process of an audio decoding processing unit.
[0093] Figure 71 FIG. 15 is a flowchart showing a first example of details of a decoding process of an audio decoding processing unit.
[0094] Figure 72 FIG. 16 is a schematic diagram showing a second example of a process of an audio decoding processing unit.
[0095] Figure 73 FIG. 17 is a flowchart showing a second example of details of a decoding process of an audio decoding processing unit.
[0096] Figure 74 FIG. 18 is a flowchart showing a third example of details of a decoding process of an audio decoding processing unit.
[0097] Figure 75 FIG. 19 is a flowchart showing a fourth example of details of a decoding process of an audio decoding processing unit.
[0098] Figure 76A flowchart showing details of a second example of the decoding process of the audio decoding processing unit.
[0099] Figure 77 A diagram showing a second example of the Web server process.
[0100] Figure 78 A diagram showing a third example of the process of the audio decoding processing unit.
[0101] Figure 79 A flowchart showing details of a third example of the decoding process of the audio decoding processing unit.
[0102] Figure 80 A diagram showing a second example of the syntax of the configuration information set in the elementary sample.
[0103] Figure 81 An example syntax of the configuration information for the Ext element shown in Figure 80
[0104] An example syntax of the configuration information for the extractor shown in Figure 82 Figure 81 An example syntax of the data syntax of the frame unit set in the elementary sample shown in
[0105] Figure 83 An example data syntax of the extractor shown in
[0106] Figure 84 Figure 83 An example data syntax of the extractor shown in
[0107] Figure 85 A diagram showing a third example of the syntax of the configuration information set in the elementary sample.
[0108] Figure 86 A diagram showing a third example of the data syntax of the frame unit set in the elementary sample.
[0109] Figure 87 A diagram showing a configuration example of the audio stream in the seventh embodiment of the information processing system of the present disclosure.
[0110] Figure 88 A diagram showing an overview of the track in the seventh embodiment.
[0111] Figure 89 A flowchart showing the file generation process in the seventh embodiment.
[0112] Figure 90 A flowchart showing the audio playback process in the seventh embodiment.
[0113] Figure 91 FIG. 8 is a schematic diagram showing an outline of a track in the eighth embodiment of the information processing system to which the present disclosure is applied.
[0114] Figure 92 FIG. 9 is a schematic diagram showing a configuration example of an audio file.
[0115] Figure 93 FIG. 10 is a schematic diagram showing another configuration example of an audio file.
[0116] Figure 94 FIG. 11 is a schematic diagram showing yet another configuration example of an audio file.
[0117] Figure 95 FIG. 12 is a block diagram showing a configuration example of hardware of a computer. DETAILED DESCRIPTION
[0118] Modes for carrying out the present disclosure (hereinafter, referred to as embodiments) will be described below in the following order.
[0119] 0. Foretold of the present disclosure Figures 1 to 30
[0120] 1. First Embodiment Figures 31 to 50
[0121] 2. Second Embodiment Figures 51 to 55
[0122] 3. Third Embodiment Figure 56
[0123] 4. Fourth Embodiment Figures 57 to 67
[0124] 5. Fifth Embodiment Figure 68
[0125] 6. Sixth Embodiment Figure 69
[0126] 7. Explanation of hierarchical structure of 3D audio Figure 70
[0127] 8. Explanation of first example of Web server process Figure 71 and 72
[0128] 9. Explanation of first example of process of audio decoding processing unit Figure 73 and 74
[0129] 10. Explanation of second example of process of audio decoding processing unit Figure 75 and 76
[0130] 11. Explanation of a second example of a process of a Web server Figure 77 )
[0131] 12. Explanation of a third example of a process of an audio decoding processing unit Figure 78 and 79 )
[0132] 13. Second example of syntax of a base sample Figures 80 to 84 )
[0133] 14. Third example of syntax of a base sample Figure 85 and 86 )
[0134] 15. Seventh embodiment Figures 87 to 90 )
[0135] 16. Eighth embodiment Figures 91 to 94 )
[0136] 17. Ninth embodiment Figure 95 )
[0137] <Preview of the present disclosure>
[0138] (Outline of a first example of an information processing system)
[0139] Figure 1 An explanatory diagram showing an outline of a first example of an information processing system to which the present disclosure is applied.
[0140] As shown in FIG. 1, an information processing system 10 has a configuration in which a Web server 12 (which is connected to a file generation device 11) and a video playback terminal 14 are connected via the Internet 13. Figure 1 In the information processing system 10, the Web server 12 transmits image data of a video content in units of tiles (tile streaming) to the video playback terminal 14 by a method compatible with MPEG-DASH.
[0141] Specifically, the file generation device 11 acquires image data of a video content and encodes the image data in units of tiles to generate a video stream. The file generation device 11 processes the video stream of each tile into a file format of a time interval from several seconds to about ten seconds, which is called a segment. The file generation device 11 uploads the resulting image file of each tile to the Web server 12.
[0142]
[0143] In addition, the file generation apparatus 11 acquires the audio data of the video content of each object (described in detail below) and encodes the image data on an object-by-object basis to generate an audio stream. The file generation apparatus 11 processes the audio stream of each object into a file format on a segment-by-segment basis and uploads the resulting audio file of each object to the web server 12.
[0144] It should be noted that each object is a sound source. The audio data of each object is acquired by a microphone or similar device attached to that object. The object may be an object such as a fixed microphone stand or a moving body such as a person.
[0145] The file generation device 11 encodes audio metadata, which includes object location information (audio location information) indicating the location of each object (the location where audio data is acquired) and an object ID serving as a unique ID for the object. The file generation device 11 processes the encoded data obtained by encoding the audio metadata into a file format based on segments and uploads the resulting audio metadata file to the web server 12.
[0146] In addition, the file generation apparatus 11 generates a Media Representation Description (MPD) file (control information), which manages image and audio files and contains image frame size information indicating the frame size of the image indicating the video content and position information indicating the position of each tile on the image. The file generation apparatus 11 uploads the MPD file to the web server 12.
[0147] Web server 12 stores image files, audio files, audio metafiles, and MPD files uploaded from file generation device 11.
[0148] In such Figure 1 In the example shown, Web server 12 stores a group of segments consisting of image files of tiles with tile ID "1" and a group of segments consisting of image files of tiles with tile ID "2". Web server 12 also stores a group of segments consisting of audio files of objects with object ID "1" and a group of segments consisting of audio files of objects with object ID "2". Although not shown, a similar group of segments consisting of audio metafiles is also stored.
[0149] It should be noted that the file with tile ID i is referred to as "tile #i" below, and the object with object ID i is referred to as "object #i" below.
[0150] Web server 12 acts as a transmitter and responds to requests from video playback terminal 14 by sending stored image files, audio files, audio metafiles, MPD files, etc. to video playback terminal 14.
[0151] The video playback terminal 14 executes, for example, software 21 for controlling streaming data (hereinafter referred to as control software), video playback software 22, and client software 23 for Hyper Text Transfer Protocol (HTTP) access (hereinafter referred to as access software).
[0152] The control software 21 is software for controlling data delivered from the Web server 12 via streaming. Specifically, the control software 21 allows the video playback terminal 14 to acquire an MPD file from the Web server 12.
[0153] Further, the control software 21 specifies a tile in a display region based on the display region and tile position information included in the MPD file, the display region being a region in an image for displaying video content instructed by the video playback software 22. The control software 21 instructs the access software 23 to issue a request for transmitting an image file of the specified tile.
[0154] Further, the control software 21 instructs the access software 23 to issue a request for transmitting an audio meta file. The control software 21 specifies an object corresponding to an image in a display region based on the display region, image frame size information included in the MPD file, and object position information included in the audio meta file. The control software 21 instructs the access software 23 to issue a request for transmitting an audio file of the specified object.
[0155] The video playback software 22 is software for playing back image files and audio files acquired from the Web server 12. Specifically, when a display region is specified by a user, the video playback software 22 instructs the control software 21 of the specified display region. The video playback software 22 decodes the image files and audio files acquired from the Web server 12 in response to the instruction, and the video playback software 22 synthesizes and outputs the decoded files.
[0156] The access software 23 is software for controlling communication with the Web server 12 via the Internet 13 using HTTP. Specifically, the access software 23 allows the video playback terminal 14 to transmit a request for transmitting an image file, an audio file, and an audio meta file in response to an instruction of the control software 21. Further, the access software 23 allows the video playback terminal 14 to receive the image file, the audio file, and the audio meta file transmitted from the Web server 12 in response to the transmission request.
[0157] (Example of a tile)
[0158] Figure 2 A schematic view for illustrating an example of a tile.
[0159] As Figure 2As shown, the video content image is divided into multiple tiles. Tile IDs, starting from 1, are assigned to each tile as sequential numbers. Figure 2 In the example shown, the image of the video content is divided into four tiles #1 to #4.
[0160] (Explanation of the object)
[0161] Figure 3 A schematic diagram illustrating the object.
[0162] Figure 3 The example shows the capture of eight audio objects from an image as audio for video content. Object IDs, starting from 1, are assigned to each object as sequential numbers. Objects #1 through #5 are moving objects, and objects #6 through #8 are stationary objects. Furthermore, in Figure 3 In the example, the image of the video content is divided into 7 (width) × 5 (height) tiles.
[0163] In this case, such as Figure 3 As shown, when the user specifies a display area 31 consisting of 3 (width) × 2 (height) tiles, the display area 31 only contains objects #1, #2, and #6. Therefore, the video playback terminal 14 only retrieves and plays audio files for objects #1, #2, and #6 from the web server 12.
[0164] The objects in the display area 31 can be specified based on image frame size information and object position information, as described below.
[0165] (Explanation of object location information)
[0166] Figure 4 A schematic diagram illustrating the location information of an object.
[0167] like Figure 4 As shown, the object position information includes the horizontal angle θ of object 40. A (-180°≤θ A ≤180°), vertical angle γ A (-90°≤γ A ≤90°) and distance r A (0 <r A For example, in the following settings, the horizontal angle θA is the angle in the horizontal direction formed by the line connecting object 40 and the origin O and the YZ plane: The center of the image can be set as the origin (base point) O; the horizontal direction of the image is set as the X direction; the vertical direction of the image is set as the Y direction; and the depth direction perpendicular to the XY plane is set as the Z direction. Vertical angle γ A The angle in the vertical direction formed by the straight line connecting object 40 and the origin O and the XZ plane. Distance rA is the distance between the object 40 and the origin O.
[0168] Further, in this document, the angle of rotation to the left and up is set to a positive angle, and the angle of rotation to the right and down is set to a negative angle.
[0169] (Explanation of image frame size information)
[0170] Figure 5 is a schematic view showing the image frame size information.
[0171] As shown in Figure 5 , the image frame size information includes a horizontal angle θ v1 of the left end, a horizontal angle θ v2 of the right end, a vertical angle γ v1 of the upper end, a vertical angle γ v2 of the lower end, and a distance r v .
[0172] For example, when a photographing position at the center of the image is set to the origin O; a horizontal direction of the image is set to the X direction; a vertical direction of the image is set to the Y direction; and a depth direction perpendicular to the XY plane is set to the Z direction, the horizontal angle θ v1 is an angle in the horizontal direction formed by a straight line connecting the left end of the image frame and the origin O and the YZ plane. The horizontal angle θ v2 is an angle in the horizontal direction formed by a straight line connecting the right end of the image frame and the origin O and the YZ plane. Thus, an angle obtained by combining the horizontal angle θ v1 and the horizontal angle θ v2 becomes a horizontal viewing angle.
[0173] The vertical angle γ V1 is an angle formed by the XZ plane and a straight line connecting the upper end of the image frame and the origin O, and the vertical angle γ v2 is an angle formed by the XZ plane and a straight line connecting the lower end of the image frame and the origin O. An angle obtained by combining the vertical angle γ V1 and the vertical angle γ v2 becomes a vertical viewing angle. The distance r v is a distance between the origin O and the image plane.
[0174] As described above, the object position information indicates a positional relationship between the object 40 and the origin O, and the image frame size information indicates a positional relationship between the image frame and the origin O. Thus, it is possible to detect (recognize) a position of each object on the image based on the object position information and the image frame size information. Thus, it is possible to specify the object in the display region 31.
[0175] (Explanation of structure of MPD file)
[0176] Figure 6 A diagram showing the structure of the MPD file.
[0177] In the analysis (parsing) of the MPD file, the video playback terminal 14 selects the optimum attribute from among the attributes of the "Representation" contained in the "Period" of the MPD file (in Figure 6 MediaPresentation).
[0178] By referring to the uniform resource locator (URL) of the "Initialization Segment" at the head of the selected "Representation", and the like, the video playback terminal 14 acquires the file and processes the acquired file. Next, by referring to the URL of the subsequent "Media Segment", and the like, the video playback terminal 14 acquires the file and plays the acquired file.
[0179] Note that, in the MPD file, the relationship between the "Period", the "Representation", and the "Segment" becomes as shown in Figure 7 In other words, a single video content item can be managed in a longer time unit than the segment by the "Period", and can be managed in the unit of the segment by the "Segment" in each "Period". Furthermore, in each "Period", the video content can be managed in the unit of the stream attribute by the "Representation".
[0180] Therefore, the MPD file has a hierarchical structure starting from the "Period" as shown in Figure 8 Furthermore, the structure of the MPD file arranged on the time axis becomes a configuration as shown in Figure 9 It is clear from Figure 9 that there are a plurality of "Representation" elements in the same segment. The video playback terminal 14 adaptively selects any one from among these elements, and thus can acquire the image file and the audio file in the display region selected by the user and play the acquired file.
[0181] (Explanation of the description of the MPD file)
[0182] Figure 10 A diagram showing the description of the MPD file.
[0183] As described above, in the information processing system 10, the image frame size information is contained in the MPD file to allow the object in the display area to be specified by the video playback terminal 14. As shown in Figure 10 The scheme (urn:mpeg:DASH:viewingAngle:2013) for defining the new image frame size information (viewing angle) is extended by utilizing the DescriptorType element of the Viewpoint, and thus the image frame size information is arranged in the "Adaptation Set" for audio and in the "Adaptation Set" for images. The image frame size information can be arranged only in the "Adaptation Set" for images.
[0184] Further, the "Representation" for the audio metadata file is described in the "Adaptation Set" for audio of the MPD file. The URL or the like as information for specifying the audio metadata file (audio metadata.mp4) is described in the "Segment" of the "Representation". In this case, it is described that the file to be specified in the "Segment" is the audio metadata file (object audio metadata) by the Role element.
[0185] The "Representation" for the audio metadata file of each object is also described in the "Adaptation Set" for audio of the MPD file. The URL or the like as information for specifying the audio file (audioObjel.mp4, audioObjel5.mp4) of each object is described in the "Segment" of the "Representation". In this case, the object IDs (1 and 5) of the objects corresponding to the audio files are also described by the extended Viewpoint.
[0186] It should be noted that although not shown, the tile position information is arranged in the "Adaptation Set" for images.
[0187] (Configuration example of file generation apparatus)
[0188] Figure 11 To show the above-mentioned Figure 1A block diagram of a configuration example of the file generation apparatus 11 shown in FIG. 1 is shown in FIG. 2.
[0189] As shown in FIG. 1, the file generation apparatus 11 includes a screen split processing unit 51, an image encoding processing unit 52, an image file generation unit 53, an image information generation unit 54, an audio encoding processing unit 55, an audio file generation unit 56, an MPD generation unit 57, and a server upload processing unit 58. Figure 11
[0190] The screen split processing unit 51 of the file generation apparatus 11 splits image data of a video content inputted from the outside into tile units. The screen split processing unit 51 provides tile position information to the image information generation unit 54. In addition, the screen split processing unit 51 provides image data configured in tile units to the image encoding processing unit 52.
[0191] The image encoding processing unit 52 encodes image data (configured in tile units and provided from the screen split processing unit 51) for each tile to generate a video stream. The image encoding processing unit 52 provides the video stream of each tile to the image file generation unit 53.
[0192] The image file generation unit 53 processes the video stream of each tile provided from the image encoding processing unit 52 into a file format in segment units and provides the resulting image file of each tile to the MPD generation unit 57.
[0193] The image information generation unit 54 provides the tile position information provided from the screen split processing unit 51 and image frame size information inputted from the outside as image information to the MPD generation unit 57.
[0194] The audio encoding processing unit 55 encodes audio data configured in object units of a video content inputted from the outside for each object and generates an audio stream. In addition, the audio encoding processing unit 55 encodes object position information of each object inputted from the outside and audio metadata containing an object ID or the like to generate encoded data. The audio encoding processing unit 55 provides the audio stream of each object and the encoded data of the audio metadata to the audio file generation unit 56.
[0195] The audio file generation unit 56 functions as an audio file generation unit, processes the audio stream of each object provided from the audio encoding processing unit 55 into a file format in segment units, and provides the resulting audio file of each object to the MPD generation unit 57.
[0196] Further, the audio file generation unit 56 functions as a meta file generation unit, processes encoded data of the audio meta data provided from the audio encoding processing unit 55 into a file format in units of segments, and provides the resultant audio meta file to the MPD generation unit 57.
[0197] The MPD generation unit 57 determines a URL or the like of the Web server 12 for storing the image file of each tile provided from the image file generation unit 53. Further, the MPD generation unit 57 determines a URL or the like of the Web server 12 for storing the audio file and the audio meta file of each object provided from the audio file generation unit 56.
[0198] The MPD generation unit 57 arranges the image information provided from the image information generation unit 54 in an "Adaptation Set" for images of the MPD file. Further, the MPD generation unit 57 arranges the image frame size information among the image information blocks in an "Adaptation Set" for audio of the MPD file. The MPD generation unit 57 arranges the URL or the like of the image file of each tile in a "Segment" of a "Representation" for the image file of the tile.
[0199] The MPD generation unit 57 arranges the URL or the like of the audio file of each object in a "Segment" of a "Representation" for the audio file of the object. Further, the MPD generation unit 57 functions as an information generation unit and arranges the URL or the like in a "Segment" of a "Representation" for the audio meta file as information for specifying the audio meta file. The MPD generation unit 57 provides the MPD file in which various types of information are arranged as described above, the image file, the audio file, and the audio meta file to the server upload processing unit 58.
[0200] The server upload processing unit 58 uploads the image file of each tile, the audio file of each object, the audio meta file, and the MPD file provided from the MPD generation unit 57 to the Web server 12.
[0201] (Explanation of the procedure of the file generation apparatus)
[0202] Figure 12 To show the procedure of the file generation process of the file generation apparatus 11 shown in Figure 11 To show the procedure of the file generation process of the file generation apparatus 11 shown in
[0203] To show the procedure of the file generation process of the file generation apparatus 11 shown in Figure 12In step Sll, the screen split processing unit 51 of the file generation apparatus 11 splits the image data of the video content inputted from the outside into tile units. The screen split processing unit 51 supplies tile position information to the image information generation unit 54. Further, the screen split processing unit 51 supplies the image data configured in tile units to the image encoding processing unit 52.
[0204] In step S12, the image encoding processing unit 52 encodes the image data configured in tile units supplied from the screen split processing unit 51 for each tile to generate a video stream of each tile. The image encoding processing unit 52 supplies the video stream of each tile to the image file generation unit 53.
[0205] In step S13, the image file generation unit 53 processes the video stream of each tile supplied from the image encoding processing unit 52 into a file format in segment units to generate an image file of each tile. The image file generation unit 53 supplies the image file of each tile to the MPD generation unit 57.
[0206] In step S14, the image information generation unit 54 acquires image frame size information from the outside. In step S15, the image information generation unit 54 generates image information containing the tile position information supplied from the screen split processing unit 51 and the image frame size information, and supplies the image information to the MPD generation unit 57.
[0207] In step S16, the audio encoding processing unit 55 encodes the audio data configured in object units of the video content inputted from the outside for each object to generate an audio stream of each object. Further, the audio encoding processing unit 55 encodes object position information of each object inputted from the outside and audio metadata containing an object ID to generate encoded data. The audio encoding processing unit 55 supplies the audio stream of each object and the encoded data of the audio metadata to the audio file generation unit 56.
[0208] In step S17, the audio file generation unit 56 processes the audio stream of each object supplied from the audio encoding processing unit 55 into a file format in segment units to generate an audio file of each object. Further, the audio file generation unit 56 processes the encoded data of the audio metadata supplied from the audio encoding processing unit 55 into a file format in segment units to generate an audio metadata file. The audio file generation unit 56 supplies the audio file of each object and the audio metadata file to the MPD generation unit 57.
[0209] In step S18, the MPD generation unit 57 generates an MPD file containing the image information supplied from the image information generation unit 54, the URL of each file, and the like. The MPD generation unit 57 supplies the MPD file, the image file of each tile, the audio file of each object, and the audio meta file to the server upload processing unit 58.
[0210] In step S19, the server upload processing unit 58 uploads the image file of each tile, the audio file of each object, the audio meta file, and the MPD file supplied from the MPD generation unit 57 to the Web server 12. Then the process is terminated.
[0211] (Functional configuration example of video playback terminal)
[0212] Figure 13 To show a block diagram of a configuration example of the stream playback unit, the stream playback unit is implemented in such a way that the video playback terminal 14 shown in Fig. 1 executes the control software 21, the video playback software 22, and the access software 23. Figure 1
[0213] As shown in Fig. 9, the stream playback unit 90 includes an MPD acquisition unit 91, an MPD processing unit 92, a meta file acquisition unit 93, an audio selection unit 94, an audio file acquisition unit 95, an audio decoding processing unit 96, an audio synthesis processing unit 97, an image selection unit 98, an image file acquisition unit 99, an image decoding processing unit 100, and an image synthesis processing unit 101. Figure 13 The MPD acquisition unit 91 of the stream playback unit 90 functions as a receiver, acquires the MPD file from the Web server 12, and supplies the MPD file to the MPD processing unit 92.
[0214] The MPD processing unit 92 extracts information (such as the URL described in the "Segment" for the audio meta file) from the MPD file supplied from the MPD acquisition unit 91, and supplies the extracted information to the meta file acquisition unit 93. Further, the MPD processing unit 92 extracts the image frame size information described in the "Adaptation Set" for the image from the MPD file, and supplies the extracted information to the audio selection unit 94. The MPD processing unit 92 extracts information (such as the URL described in the "Segment" for the audio file of the object requested from the audio selection unit 94) from the MPD file, and supplies the extracted information to the audio selection unit 94.
[0215]
[0216] The MPD processing unit 92 extracts tile position information described in an "Adaptation Set" for images from the MPD file and supplies the extracted information to the image selection unit 98. The MPD processing unit 92 extracts information such as a URL described in a "Segment" for an image file of a tile requested from the image selection unit 98 from the MPD file and supplies the extracted information to the image selection unit 98.
[0217] Based on information such as the URL supplied from the MPD processing unit 92, the meta file acquisition unit 93 requests the Web server 12 to transmit an audio meta file designated by the URL and acquires the audio meta file. The meta file acquisition unit 93 supplies object position information contained in the audio meta file to the audio selection unit 94.
[0218] The audio selection unit 94 functions as a position determination unit and calculates the position of each object on an image based on image frame size information supplied from the MPD processing unit 92 and object position information supplied from the meta file acquisition unit 93. The audio selection unit 94 selects an object in a display region designated by a user based on the position of each object on an image. The audio selection unit 94 requests the MPD processing unit 92 to transmit information such as a URL of an audio file of the selected object. The audio selection unit 94 supplies information such as the URL supplied from the MPD processing unit 92 to the audio file acquisition unit 95 in response to the request.
[0219] The audio file acquisition unit 95 functions as a receiver. Based on information such as a URL supplied from the audio selection unit 94, the audio file acquisition unit 95 requests the Web server 12 to transmit an audio file designated by the URL and configured in units of objects and acquires the audio file. The audio file acquisition unit 95 supplies the acquired audio file configured in units of objects to the audio decoding processing unit 96.
[0220] The audio decoding processing unit 96 decodes an audio stream contained in the audio file supplied from the audio file acquisition unit 95 and configured in units of objects to generate audio data in units of objects. The audio decoding processing unit 96 supplies the audio data in units of objects to the audio synthesis processing unit 97.
[0221] The audio synthesis processing unit 97 synthesizes the audio data supplied from the audio decoding processing unit 96 and configured in units of objects and outputs the synthesized data.
[0222] The image selection unit 98 selects a tile from the display area specified by the user based on the tile location information provided by the MPD processing unit 92. The image selection unit 98 requests the MPD processing unit 92 to send information such as the URL of the image file of the selected tile. In response to this request, the image selection unit 98 provides the image file acquisition unit 99 with information such as the URL provided by the MPD processing unit 92.
[0223] Based on information such as the URL provided by the image selection unit 98, the image file acquisition unit 99 requests the web server 12 to send an image file specified by the URL and configured in tile units, and acquires the image file. The image file acquisition unit 99 provides the acquired tile-based image file to the image decoding processing unit 100.
[0224] The image decoding processing unit 100 decodes the video stream (which is contained in an image file provided by the image file acquisition unit 99 and configured in tile units) to generate tile-based image data. The image decoding processing unit 100 provides the tile-based image data to the image compositing processing unit 101.
[0225] The image synthesis processing unit 101 synthesizes image data provided by the image decoding processing unit 100 and configured in units of tiles, and outputs the synthesized data.
[0226] (Explanation of the process of a motion picture playback terminal)
[0227] Figure 14 To illustrate the streaming unit of the video playback terminal 14 ( Figure 13 The flowchart of the streaming playback process.
[0228] exist Figure 14 In step S31, the MPD acquisition unit 91 of the streaming playback unit 90 acquires the MPD file from the Web server 12 and provides the MPD file to the MPD processing unit 92.
[0229] In step S32, the MPD processing unit 92 obtains image frame size information and tile position information described in the "Adaptation Set" for images from the MPD file provided by the MPD acquisition unit 91. The MPD processing unit 92 provides the image frame size information to the audio selection unit 94 and the tile position information to the image selection unit 98. Furthermore, the MPD processing unit 92 extracts information such as URLs described in the "Segment" for audio metadata and provides the extracted information to the metadata acquisition unit 93.
[0230] In step S33, based on information such as the URL supplied from the MPD processing unit 92, the metafile acquisition unit 93 requests the Web server 12 to transmit an audio metafile designated by the URL, and acquires the audio metafile. The metafile acquisition unit 93 supplies the object position information contained in the audio metafile to the audio selection unit 94.
[0231] In step S34, the audio selection unit 94 selects an object in the display area designated by the user based on the image frame size information supplied from the MPD processing unit 92 and the object position information supplied from the metafile acquisition unit 93. The audio selection unit 94 requests the MPD processing unit 92 to transmit information such as the URL of the audio file of the selected object.
[0232] The MPD processing unit 92 extracts information such as the URL described in the "Segment" of the audio file for the object requested from the audio selection unit 94 from the MPD file, and supplies the extracted information to the audio selection unit 94. The audio selection unit 94 supplies information such as the URL supplied from the MPD processing unit 92 to the audio file acquisition unit 95.
[0233] In step S35, based on information such as the URL supplied from the audio selection unit 94, the audio file acquisition unit 95 requests the Web server 12 to transmit the audio file of the selected object designated by the URL, and acquires the audio file. The audio file acquisition unit 95 supplies the acquired audio file in units of objects to the audio decoding processing unit 96.
[0234] In step S36, the image selection unit 98 selects a tile in the display area designated by the user based on the tile position information supplied from the MPD processing unit 92. The image selection unit 98 requests the MPD processing unit 92 to transmit information such as the URL of the image file of the selected tile.
[0235] The MPD processing unit 92 extracts information such as the URL described in the "Segment" of the image file for the object requested from the image selection unit 98 from the MPD file, and supplies the extracted information to the image selection unit 98. The image selection unit 98 supplies information such as the URL supplied from the MPD processing unit 92 to the image file acquisition unit 99.
[0236] In step S37, based on information such as the URL supplied from the image selection unit 98, the image file acquisition unit 99 requests the Web server 12 to transmit the image file of the selected tile specified by the URL, and acquires the image file. The image file acquisition unit 99 supplies the acquired image file in units of tiles to the image decoding processing unit 100.
[0237] In step S38, the audio decoding processing unit 96 decodes the audio stream contained in the audio file supplied from the audio file acquisition unit 95 and configured in units of objects, to generate audio data in units of objects. The audio decoding processing unit 96 supplies the audio data in units of objects to the audio synthesis processing unit 97.
[0238] In step S39, the image decoding processing unit 100 decodes the video stream contained in the image file supplied from the image file acquisition unit 99 and configured in units of tiles, to generate image data in units of tiles. The image decoding processing unit 100 supplies the image data in units of tiles to the image synthesis processing unit 101.
[0239] In step S40, the audio synthesis processing unit 97 synthesizes the audio data supplied from the audio decoding processing unit 96 and configured in units of objects, and outputs the synthesized data. In step S41, the image synthesis processing unit 101 synthesizes the image data supplied from the image decoding processing unit 100 and configured in units of tiles, and outputs the synthesized data. The process then terminates.
[0240] As described above, the Web server 12 transmits the image frame size information and the object position information. Therefore, the video playback terminal 14 can specify, for example, an object in a display region to selectively acquire the audio file of the specified object so that the audio file corresponds to the image in the display region. This allows the video playback terminal 14 to acquire only the necessary audio file, which results in improved transmission efficiency.
[0241] Note that, as Figure 15As shown, an object ID (object specifying information) can be described in an "Adaptation Set" for an image of an MPD file as information for specifying an object corresponding to audio to be played simultaneously with the image. The object ID can be described by defining an extension scheme (urn:mpeg:DASH:audioObj:2013) of a new object ID information (audioObj) using a DescriptorType element of a Viewpoint. In this case, the video playback terminal 14 selects an audio file of an object corresponding to the object ID described in the "Adaptation Set" for an image of an MPD file, and acquires the audio file for playback.
[0242] As an alternative to generating an audio file in units of objects, the encoded data of all objects can be multiplexed into a single audio stream to generate a single audio file.
[0243] In this case, as shown in FIG. 6, one "Representation" for an audio file is set in an "Adaptation Set" for audio of an MPD file, and a URL of an audio file (audioObje.mp4) containing encoded data of all objects, and the like are described in a "Segment". Figure 16
[0244] In addition, in this case, as shown in FIG. 7, the encoded data of each object (audio object) is arranged as a sub-sample in an mdat box of an audio file (hereinafter, also referred to as an audio media file) acquired by referring to a "Media Segment" of an MPD file. Figure 17
[0245] Specifically, data is arranged in an audio media file in units of sub-segments, which are shorter than segments at any time. The position of data in units of sub-segments is specified by an sidx box. Furthermore, data in units of sub-segments is composed of a moof box and an mdat box. The mdat box is composed of a plurality of samples, and the encoded data of each object is arranged as each sub-sample of the sample.
[0246] Further, a gsix box describing information on samples is arranged after the sidx box of the audio media file. The gsix box describing information on samples is set apart from the moof box in this way, and thus the video playback terminal 14 can quickly acquire information on samples.
[0247] As shown in Figure 18 , a grouping_type indicating a type of a sample group entry is described in the gsix box, where each sample group entry contains one or more samples or sub-samples managed by the gsix box. For example, when the sample group entry is a sub-sample of encoded data in units of objects, the type of the sample group entry is "obja" as shown in Figure 17 . Multiple gsix boxes of the grouping_type are arranged in the audio media file.
[0248] Further, as shown in Figure 18 , an entry_index of each sample group entry and a range_size of data position information indicating a position in the audio media file are described in the gsix box. Note that when the entry_index is 0, the corresponding range_size indicates a byte range of a moof box (a1) in the example of Figure 17 .
[0249] Information indicating which object is used to allow each sample group entry to correspond to a sub-sample of encoded data is described in an audio file acquired by referring to an "Initialization Segment" of an MPD file (hereinafter also referred to as an audio initialization file as appropriate).
[0250] Specifically, as shown in Figure 19 , this information is indicated by using a type assignment box (typa) of the mvex box, which is associated with an AudioObjectSampleGroupEntry of a sample group description box (sgpd) in an sbtl box of the audio initialization file.
[0251] In other words, as shown in A of Figure 20 , an object ID (audio_object_id) corresponding to encoded data contained in a sample is described in each AudioObjectSampleGroupEntry box. For example, as shown in B of Figure 20 , object IDs 1, 2, 3, and 4 are described in each of four AudioObjectSampleGroupEntry boxes.
[0252] On the other hand, such as Figure 21 As shown, in the type allocation box, the index of the parameter (grouping_type_parameter) corresponding to the sample group entry of AudioObjectSampleGroupEntry is described for each AudioObjectSampleGroupEntry.
[0253] The audio media file and audio initialization file are configured as described above. Therefore, when the video playback terminal 14 acquires the encoded data of the object selected as the object in the display area, the AudioObjectSampleGroupEntry, which describes the object ID of the selected object, is retrieved from the stbl box of the audio initialization file. Next, the index of the sample group entry corresponding to the retrieved AudioObjectSampleGroupEntry is read from the mvex box. Then, the position of the data in sub-segments is read from the sidx of the audio file, and the byte range of the sample group entry whose index is read is read from gsix. Next, the encoded data arranged in mdat is acquired based on the position and byte range of the data in sub-segments. Thus, the encoded data of the selected object is acquired.
[0254] Although in the above description, the index of the sample group and the object ID of the AudioObjectSampleGroupEntry are associated with each other through the mvex box, they can be directly associated with each other. In this case, the index of the sample group entry is described in the AudioObjectSampleGroupEntry.
[0255] Furthermore, when an audio file consists of multiple tracks, the sgpd can be stored in mvex, which allows the sgpd to be shared between tracks.
[0256] (An overview of the second example of an information processing system)
[0257] Figure 22 A schematic diagram illustrating a second example of an information processing system applying this disclosure.
[0258] It should be pointed out that, in Figure 22 The one shown in the middle is the same as Figure 3 The same elements shown are represented by the same reference numerals.
[0259] exist Figure 22 As shown Figure 3 In the example, the image of the video content is divided into 7 (width) × 5 (height) tiles, and the audio of objects #1 to #8 is captured just like the audio of the video content.
[0260] In this case, when the user instructs the display region 31 composed of 3 (width) x 2 (height) tiles, the display region 31 is converted (expanded) to a region having the same size as the size of the image of the video content, and thus the display image 111 in the second example as shown in FIG. 11 is obtained. The audio of the objects #1 to #8 is synthesized based on the positions of the objects #1 to #8 in the display image 111 and output together with the display image 111. In other words, the audio of the objects #3 to #5, #7 and #8 outside the display region 31 is also output in addition to the audio of the objects #1, #2 and #6 inside the display region 31. Figure 22
[0261] (Configuration example of stream playback unit)
[0262] The configuration of the second example of the information processing system to which the present disclosure is applied is the same as the configuration of the information processing system 10 as shown in FIG. 1 except for the configuration of the stream playback unit, and thus only the stream playback unit is described below. Figure 1
[0263] A block diagram showing a configuration example of the stream playback unit of the information processing system to which the present disclosure is applied. Figure 23 The same components as shown in FIG. 1 are denoted by the same reference numerals in
[0264] Figure 23 Figure 13 The configuration of the stream playback unit 120 as shown in FIG. 2 is different from the configuration of the stream playback unit 90 as shown in FIG. 1 in that an MPD processing unit 121, an audio synthesis processing unit 123 and an image synthesis processing unit 124 are newly provided to replace the MPD processing unit 92, the audio synthesis processing unit 97 and the image synthesis processing unit 101, respectively, and a position determination unit 122 is further provided.
[0265] The configuration of the stream playback unit 120 as shown in FIG. 2 is different from the configuration of the stream playback unit 90 as shown in FIG. 1 in that an MPD processing unit 121, an audio synthesis processing unit 123 and an image synthesis processing unit 124 are newly provided to replace the MPD processing unit 92, the audio synthesis processing unit 97 and the image synthesis processing unit 101, respectively, and a position determination unit 122 is further provided. Figure 23 Figure 13
[0266] The MPD processing unit 121 of the streaming playback unit 120 extracts information such as URLs described in "Segment" for audio meta files from the MPD file supplied from the MPD acquisition unit 91, and supplies the extracted information to the meta file acquisition unit 93. Further, the MPD processing unit 121 extracts image frame size information (hereinafter, referred to as content image frame size information) of images of video contents described in "Adaptation Set" for images from the MPD file, and supplies the extracted information to the position determination unit 122. The MPD processing unit 121 extracts information such as URLs described in "Segment" for audio files of all objects from the MPD file, and supplies the extracted information to the audio file acquisition unit 95.
[0267] The MPD processing unit 121 extracts tile position information described in "Adaptation Set" for images from the MPD file, and supplies the extracted information to the image selection unit 98. The MPD processing unit 121 extracts information such as URLs described in "Segment" for image files of tiles requested from the image selection unit 98 from the MPD file, and supplies the extracted information to the image selection unit 98.
[0268] The position determination unit 122 acquires object position information contained in the audio meta file obtained by the meta file acquisition unit 93 and the content image frame size information supplied from the MPD processing unit 121. Further, the position determination unit 122 acquires display area image frame size information as image frame size information of a display area designated by a user. The position determination unit 122 determines (recognizes) the position of each object in the display area on the basis of the object position information, the content image frame size information, and the display area image frame size information. The position determination unit 122 supplies the determined position of each object to the audio synthesis processing unit 123.
[0269] The audio synthesis processing unit 123 synthesizes object-based audio data provided by the audio decoding processing unit 96 based on the object positions provided by the position determination unit 122. Specifically, the audio synthesis processing unit 123 determines the audio data assigned to each speaker for each object based on the object positions and the positions of each speaker outputting the sound. The audio synthesis processing unit 123 synthesizes the audio data for each object for each speaker and outputs the synthesized audio data as the audio data for each speaker. A detailed description of the method for synthesizing the audio data for each object based on the object position is disclosed, for example, in Ville Pulkki's "Virtual Sound Source Positioning Using Vector Base Amplitude Panning" published in the AES Journal, Vol. 45, No. 6, 1997, pp. 456-466.
[0270] The image compositing processing unit 124 composes image data in tile units provided by the image decoding processing unit 100. The image compositing processing unit 124 acts as a converter, transforming the image size corresponding to the compositing image data into the size of the video content to generate a display image. The image compositing processing unit 124 outputs this display image.
[0271] (Explanation of methods for determining object location)
[0272] Figures 24 to 26 Each of them is shown as follows Figure 23 The method for determining the object position of the position determination unit 122 shown.
[0273] Display area 31 is extracted from the video content, and the size of display area 31 is converted to the size of the video content to generate display image 111. Therefore, the size of display image 111 is equivalent to, for example, Figure 24 As shown, by shifting the center C of display area 31 to the center C′ of display image 111, and as... Figure 25 The dimensions shown are obtained by converting the size of the display area 31 to the size of the video content.
[0274] Therefore, the position determination unit 122 calculates the horizontal displacement θ when the center O of the display area 31 is displaced to the center O′ of the display image 111 using the following formula (1). shift .
[0275]
Mathematical Formula 1
[0276]
[0277] In formula (1), θv1 denotes a horizontal angle at the left end of the display area 31 included in the display area image frame size information, and θ v2' denotes a horizontal angle at the right end of the display area 31 included in the display area image frame size information. Further, θ v1 denotes a horizontal angle at the left end in the content image frame size information, and θ v2 denotes a horizontal angle at the right end in the content image frame size information.
[0278] Next, the position determination unit 122 calculates the displacement amount θ shift the horizontal angle θ v1_shift ' at the right end of the display area 31 after the center O of the display area 31 is displaced to the center O' of the display image 111 by using the following equation (2). v2_shift
[0279]
Mathematical Equation 2
[0280] θ v1_shift ' = mod(θ v1 + θ shift + 180°, 360°) - 180°
[0281] θ v2_shift ' = mod(θ v2 + θ shift + 180°, 360°) - 180°... (2)
[0282] According to equation (2), the horizontal angle θ v1_shift ' and the horizontal angle θ v2_shift ' are calculated so as not to exceed the range of -180° to 180°.
[0283] Note that, as described above, the display image 111 size is equivalent to the size obtained by displacing the center O of the display area 31 to the center O' of the display image 111 and by converting the size of the display area 31 to the size of the video content. Therefore, the following equation (3) satisfies the horizontal angle θ V1 and θ V2 .
[0284]
Mathematical Equation 3
[0285]
[0286] The position determination unit 122 calculates the displacement amount θ shift , the horizontal angle θ v1_shift ' and the horizontal angle θ v2_shift and then calculates the horizontal angle of each object in the display image 111. Specifically, the position determining unit 122 calculates the horizontal angle θ shift After the center C of the display region 31 is shifted to the center C' of the display image 111, the position determining unit 122 calculates the horizontal angle θ Ai_shift .
[0287] [Mathematical Formula 4]
[0288] θ Ai_shift = mod(θ Ai + θ shift + 180°, 360°) - 180°... (4)
[0289] In the formula (4), θAi denotes the horizontal angle of the object #i included in the object position information. Further, according to the formula (4), the horizontal angle θ Ai_shift is calculated so as not to exceed the range of -180° to 180°.
[0290] Next, when the object #i exists in the display region 31, that is, when the condition θ v2_shif < θ Ai_shift < θ v1_shift ' is satisfied, the position determining unit 122 calculates the horizontal angle θ A1 ' of the object #i in the display image 111 by the following formula (5).
[0291] [Mathematical Formula 5]
[0292]
[0293] According to the formula (5), the horizontal angle θ A1 ' is calculated by expanding the distance between the position of the object #i in the display image 11 and the center C' of the display image 111 according to the ratio between the size of the display region 31 and the size of the display image 111.
[0294] On the other hand, when the object #i does not exist in the display region 31, that is, when the condition -180° ≤ θ Ai_shift ≤ θ v2_shift ' or θ v1_shift ' ≤ θAi_shift ≤ 180° is satisfied, the position determining unit 122 calculates the horizontal angle θ Ai ' of the object #i in the display image 111 by the following formula (6).
[0295] [Mathematical Formula 6]
[0296]
[0297] According to formula (6), as shown in Figure 26 Fig. 16, when the object #i exists at a position 151 on the right side of the display area 31 (-180° < θ Ai_shift ≤ θ v2_shift '), the horizontal angle θ Ai_shift is calculated by expanding the horizontal angle θ Ai ' according to the ratio between an angle Rl and an angle R2. Note that the angle Rl is an angle measured from the right end of the display image 111 to a position 154 just behind the viewer 153, and the angle R2 is an angle measured from the right end of the display area 31 whose center is displaced to the position 154.
[0298] Further, according to formula (6), when the object #i exists at a position 155 on the left side of the display area 31 (θ v1_shift ' < θ Ai_shift ≤ 180°), the horizontal angle θ Ai_shift is calculated by expanding the horizontal angle θ Ai ' according to the ratio between an angle R3 and an angle R4. Note that the angle R3 is an angle measured from the left end of the display image 111 to the position 154, and the angle R4 is an angle measured from the left end of the display area 31 whose center is displaced to the position 154.
[0299] In addition, the position determination unit 122 calculates the vertical angle γ Ai ' in a manner similar to the horizontal angle θ Ai '. Specifically, the position determination unit 122 calculates the displacement amount γ shift in the vertical direction when the center C of the display area 31 is displaced to the center C' of the display image 111 by the following formula (7).
[0300] [mathematical formula 7]
[0301]
[0302] In formula (7), γ v1 ' indicates the vertical angle of the upper end of the display area 31 included in the display area image frame size information, and γ v2 ' indicates the vertical angle of the lower end thereof. Further, γ v1 indicates the vertical angle of the upper end in the content image frame size information, and γ v2 indicates the vertical angle of the lower end in the content image frame size information.
[0303] Next, the position determination unit 122 calculates the displacement amount γ shift, the vertical angle γ v1_shift ' at the lower end of the display region 31 v2_shift .
[0304] [Equation 8]
[0305] γ v1_shift ' = mod(γ v1 + γ shift + 90°, 180°) - 90°
[0306] γ v2_shift ' = mod(γ v2 + γ shift + 90°, 180°) - 90°...(8)
[0307] According to Equation (8), the vertical angle γ v1_shift ' and the vertical angle γ v2_shift ' are calculated so as not to exceed the range of -90° to 90°.
[0308] The position determination unit 122 calculates the displacement amount γ shift , the vertical angle γ v1_shift ' and the vertical angle γ v2_shift ' in the above-described manner, and then calculates the position of each object in the display image 111. Specifically, the position determination unit 122 calculates the vertical angle γ shift of the object #i after the center C of the display region 31 is displaced to the center C' of the display image 111 by using the displacement amount γ Ai_shift in Equation (9) below.
[0309] [Equation 9]
[0310] γ Ai_shift = mod(γ Ai + γ shift + 90°, 180°) - 90°...(9)
[0311] In Equation (9), γAi denotes the vertical angle of the object #i included in the object position information. Further, according to Equation (9), the vertical angle γ Ai_shift is calculated so as not to exceed the range of -90° to 90°.
[0312] Next, the position determination unit 122 calculates the vertical angle γ A1 ' of the object #i in the display image 111 by Equation (10) below.
[0313] [Equation 10]
[0314]
[0315] Further, the position determining unit 122 determines the distance r A1 of the object #i in the display image 111 A1 . The position determining unit 122 provides the audio synthesis processing unit 123 with the horizontal angle θ Ai , the vertical angle γ A1 , and the distance r A1 of the object #i as the position of the object #i, as described above.
[0316] Figure 27 is a schematic diagram showing the relationship between the horizontal angle θ Ai and the horizontal angle θ Ai '.
[0317] In the graph of Figure 27 , the horizontal axis represents the horizontal angle θ Ai , and the vertical axis represents the horizontal angle θ Ai '.
[0318] As shown in Figure 27 , when the condition θ V2 < θ Ai < θ V1 ' is satisfied, the horizontal angle θ Ai is shifted by the shift amount θ shift and is expanded, and then the horizontal angle θ Ai becomes equal to the horizontal angle θ Ai '. Further, when the condition -180°≤ θ Ai ≤ θ v2 ' or θ v1 '≤ θ Ai ≤ 180° is satisfied, the horizontal angle θ Ai is shifted by the shift amount θ shift and is reduced, and then the horizontal angle θ Ai becomes equal to the horizontal angle θ Ai '.
[0319] (Explanation of the process of the stream playback unit)
[0320] Figure 28 is a flowchart showing the stream playback process of the stream playback unit 120 shown in Figure 23 .
[0321] In Figure 28In step S131, the MPD acquiring unit 91 of the stream playback unit 120 acquires the MPD file from the Web server 12 and supplies the MPD file to the MPD processing unit 121.
[0322] In step S132, the MPD processing unit 121 acquires the content image frame size information and tile position information described in the "Adaptation Set" for the image from the MPD file supplied from the MPD acquiring unit 91. The MPD processing unit 121 supplies the image frame size information to the position determining unit 122 and supplies the tile position information to the image selection unit 98. Further, the MPD processing unit 121 extracts information such as the URL described in the "Segment" for the audio meta file and supplies the extracted information to the meta file acquiring unit 93.
[0323] In step S133, the meta file acquiring unit 93 requests the Web server 12 to transmit the audio meta file designated by the URL based on the information such as the URL supplied from the MPD processing unit 121 and acquires the audio meta file. The meta file acquiring unit 93 supplies the object position information included in the audio meta file to the position determining unit 122.
[0324] In step S134, the position determining unit 122 performs a position determining process for determining the position of each object in the display image based on the object position information, the content image frame size information, and the display region image frame size information. The position determining process will be described in detail later with reference to FIG. 6. Figure 29
[0325] In step S135, the MPD processing unit 121 extracts information such as the URL described in the "Segment" for the audio file of all the objects from the MPD file and supplies the extracted information to the audio file acquiring unit 95.
[0326] In step S136, the audio file acquiring unit 95 requests the Web server 12 to transmit the audio file of all the objects designated by the URL based on the information such as the URL supplied from the MPD processing unit 121 and acquires the audio file. The audio file acquiring unit 95 supplies the acquired audio file in units of objects to the audio decoding processing unit 96.
[0327] The processes of steps S137 to S140 are similar to the processes of steps S36 to S39 as shown in FIG. 4, and thus the description thereof will be omitted. Figure 14
[0328] In step S141, the audio synthesis processing unit 123 synthesizes the object-unit audio data supplied from the audio decoding processing unit 96 based on the position of each object supplied from the position determination unit 122 and outputs the audio data.
[0329] In step S142, the image synthesis processing unit 124 synthesizes the tile-unit image data supplied from the image decoding processing unit 100.
[0330] In step S143, the image synthesis processing unit 124 converts the image size corresponding to the synthesized image data into the size of the video content and generates a display image. Next, the image synthesis processing unit 124 outputs the display image, and the process terminates.
[0331] Figure 29 A flowchart showing details of the position determination process in step S134 of Figure 28 will be described. The position determination process is executed, for example, for each object.
[0332] In step S151 of Figure 29 , the position determination unit 122 executes a horizontal angle θ Ai ' estimation process for estimating a horizontal angle θ Ai ' in the display image. Details of the horizontal angle θ Ai ' estimation process will be described with reference to Figure 30 described later.
[0333] In step S152, the position determination unit 122 executes a vertical angle γ Ai ' estimation process for estimating a vertical angle γ Ai ' in the display image. Details of the vertical angle γ Ai ' estimation process are similar to those of the horizontal angle θ Ai ' estimation process in step S151 except that the vertical direction is used instead of the horizontal direction, and thus a detailed description thereof will be omitted.
[0334] In step S153, the process determination unit 122 determines a distance r Ai ' in the display image as the distance r Ai included in the object position information supplied from the meta file acquisition unit 93.
[0335] In step S154, the position determination unit 122 outputs the horizontal angle θ Ai ', the vertical angle γ A1 ', and the distance r A1 as the position of the object #i to the audio synthesis processing unit 123. Next, the process returns to Figure 28Proceed to step S134 and then proceed to step S135.
[0336] Figure 30 To show in Figure 29 The horizontal angle θ in step S151 Ai A flowchart detailing the estimation process.
[0337] In such Figure 30 In step S171 shown, the position determination unit 122 obtains the horizontal angle θ contained in the object position information provided by the metafile acquisition unit 93. Ai .
[0338] In step S172, the position determination unit 122 obtains the content image frame size information provided by the MPD processing unit 121 and the display area image frame size information specified by the user.
[0339] In step S173, the position determination unit 122 calculates the displacement θ based on the content image frame size information and the display area image frame size information using the formula (1) above. shift .
[0340] In step S174, the position determination unit 122 uses the displacement θ shift The horizontal angle θ is calculated using the above formula (2) to determine the size of the image frame in the display area. v1_shift ' and θ v2_shift '.
[0341] In step S175, the position determination unit 122 uses the horizontal angle θ Ai and displacement θ shift The horizontal angle θ is calculated using the formula (4) above. Ai_shift .
[0342] In step S176, the position determination unit 122 determines whether object #i exists in the display area 31 (the horizontal angle of object #i is between the horizontal angles at both ends of the display area 31), that is, whether θ is satisfied. v2_shift '<θ Ai_shift <θ v1_shift 'Conditions.
[0343] When it is determined in step S176 that object #i exists in display area 31, that is, when condition θ is satisfied... v2_shift '<θ Ai_shift <θ v1_shift At this point, the process proceeds to step S177. In step S177, the position determination unit 122 determines the position based on the content image frame size information and the horizontal angle θ. v1_shift ' and θ v2_shift 'and horizontal angle θ Ai_shiftThe horizontal angle θ A1 .
[0344] On the other hand, when it is determined in step S176 that the object #i does not exist in the display region 31, i.e., when the condition -180°≤ θ Ai_shift ≤ θ v2_shift ' or θ v1_shift '≤ θ Ai_shift ≤ 180° is satisfied, the process proceeds to step S178. In step S178, the position determination unit 122 determines the position of the object #i based on the content image frame size information, the horizontal angle θ v1_shift ' or θ v2_shift ' and the horizontal angle θ Ai_shift The horizontal angle θ Ai ' is calculated by the above formula (6).
[0345] After the process of step S177 or step S178, the process returns to step S151 of Figure 29 and proceeds to step S152.
[0346] Note that in the second example, the size of the display image is the same as the size of the video content, but alternatively, the size of the display image can be different from the size of the video content.
[0347] Further, in the second example, the audio data of all the objects is not synthesized and output, but instead, the audio data of only some of the objects (e.g., the objects in the display region, the objects within a predetermined range of the display region, etc.) is synthesized and output. The method for selecting the objects of the audio data to be output can be determined in advance or can be designated by the user.
[0348] Further, in the above description, only the audio data of unit objects is used, but the audio data can include the audio data of channel audio, the audio data of higher order high fidelity (HOA), the audio data of spatial audio object coding (SAOC), and the metadata (scene information, dynamic or static metadata) of the audio data. In this case, for example, not only the encoded data of each object but also the encoded data of these data blocks are arranged as sub-samples.
[0349] <First Embodiment>
[0350] (Overview of 3D Audio File Format)
[0351] Before describing the first embodiment of the present disclosure, an overview of the channels of the 3D audio file format of MP4 will be described with reference to Figure 31
[0352] In the MP4 file, codec information of video content and position information indicating a position in the file can be managed for each track. In the 3D audio file format of MP4, all audio streams (ES) of 3D audio (channel audio / object audio / HOA audio / metadata) are recorded as one track in units of samples (frames). Further, the codec information (profile / level / audio configuration) of 3D is stored as a sample entry.
[0353] Channel audio constituting 3D audio is audio data in units of channels; object audio is audio data in units of objects; HOA audio is spherical audio data; and metadata is metadata of channel audio / object audio / HOA audio. In this case, audio data in units of objects is used as object audio, but alternatively audio data of SAOC can be used instead.
[0354] Structure of moov box
[0355] Figure 32 The structure of the moov box of the MP4 file is shown.
[0356] As shown in Figure 32 In the MP4 file, image data and audio data are recorded in different tracks. Figure 32 Details of the track of audio data are not shown, but a track of audio data similar to the track of image data is shown. A sample entry is contained in a sample description arranged in a stsd box within the moov box.
[0357] Incidentally, in broadcast or local storage playback, when all audio streams are parsed and output (rendered), a Web server transmits all audio streams, and a video playback terminal (client) decodes the audio stream of necessary 3D audio. When the Bitrate is high or there is a limitation in the reading rate of local storage, there is a demand to reduce the load of the decoding process by acquiring only the audio stream of necessary 3D audio.
[0358] Further, in streaming playback, there is a demand that a video playback terminal (client) acquires only the encoded data of necessary 3D audio, thereby acquiring an audio stream of the encoding rate optimal for the playback environment.
[0359] Accordingly, in the present disclosure, encoded data of 3D audio is divided into tracks for each type of data and the tracks are arranged in an audio file, which makes it possible to efficiently acquire only a predetermined type of encoded data. Accordingly, load on a system at the time of broadcasting and local storage playback is reduced. Furthermore, at the time of streaming playback, the highest quality encoded data of 3D audio necessary can be played back according to a frequency band. Furthermore, since it is only necessary to record position information of an audio stream of a 3D file within an audio file in units of tracks of subsegments, the amount of position information can be reduced compared to a case where encoded data in units of objects is arranged in subsamples.
[0360] (Overview of tracks)
[0361] Figure 33 A schematic diagram showing an overview of tracks in the first embodiment of the present disclosure.
[0362] As Figure 33 shown, in the first embodiment, channel audio / object audio / HOA audio / metadata constituting 3D audio are respectively set as audio streams of different tracks (channel audio track / object audio track / HOA audio track / object metadata track). The audio stream of audio metadata is arranged in the object metadata track.
[0363] Furthermore, a base track (base track) is provided as a track for arranging information about the entire 3D audio. In the base track as Figure 33 shown, information about the entire 3D audio is arranged in a sample entry when no sample is arranged in the sample entry. Furthermore, the base track, the channel audio track, the object audio track, the HOA audio track, and the object metadata are recorded as the same audio file (3dauio.mp4).
[0364] A track reference number (Track Reference) is arranged in, for example, a track box, and indicates a reference relationship between a corresponding track and another track. Specifically, the track reference number indicates an ID (hereinafter, referred to as a track ID) that is unique to a track in other referenced tracks. In Figure 33 the example shown, the track IDs of the base track, the channel audio track, the HOA audio track, the object metadata track, and the object audio track are 1, 2, 3, 4, 10,... respectively. The track reference numbers of the base track are 2, 3, 4, 10,..., and the track reference numbers of the channel audio track / HOA audio track / object metadata track / object audio track are 1, which correspond to the track ID of the base track.
[0365] Therefore, the base track and the channel audio track / HOA audio track / object metadata track / object audio track have a reference relationship. Specifically, the base track is referenced during the playback of the channel audio track / HOA audio track / object metadata track / object audio track.
[0366] (Example syntax for sample entries of a basic track)
[0367] Figure 34 To show in Figure 33 The diagram shows an exemplary syntax of sample entries for the basic track.
[0368] As information about the entire 3D audio, such as Figure 34 The configurationVersion, MPEGHAudioProfile, and MPEGHAudioLevel shown represent the configuration information, profile information, and level information (for normal 3D audio streams) of the entire 3D audio stream, respectively. Additionally, as information about the entire 3D audio stream, such as... Figure 34 The width and height shown represent the number of pixels in the horizontal direction and the number of pixels in the vertical direction of the video content, respectively. As information about the entire 3D audio, θ1, θ2, γ1, and γ2 represent the horizontal angle θv1 at the left end of the image frame, the horizontal angle θv2 at the right end of the image frame, the vertical angle γv1 at the top end of the image frame, and the vertical angle γv2 at the bottom end of the image frame, respectively, in the image frame size information of the video content.
[0369] (Example syntax for sample entries of audio track channels)
[0370] Figure 35 To show in Figure 33 The diagram shows an example syntax of sample entries for the channel audio track (channel audio track).
[0371] Figure 35 The configurationVersion, MPEGHAudioProfile, and MPEGHAudioLevel are displayed, representing the configuration information, profile information, and level information of the audio channel, respectively.
[0372] (Example syntax for sample entries of an object audio track)
[0373] Figure 36 To show in Figure 33 The diagram illustrates an exemplary syntax for sample entries of an object audio track (object audio track).
[0374] In one or more object audio tracks contained within an object audio track, such as Figure 36 The ConfigurationVersion, MPEGHAudioProfile, and MPEGHAudioLevel shown represent configuration information, profile information, and level information, respectively. object_is_fixed indicates whether one or more object audio objects included in the object audio track are fixed. When object_is_fixed is 1, it indicates that the object is fixed; when object_is_fixed is 0, it indicates that the object is displaced. mpegh3daConfig represents the configuration of the identification information for one or more object audio objects included in the object audio track.
[0375] In addition, objectTheta1 / objectTheta2 / objectGamma1 / objectGamma2 / objectRength represent object information for one or more object audio tracks contained within the object audio track. This object information is valid while keeping Object_is_fixed = 1.
[0376] maxobjectTheta1, maxobjectTheta2, maxobjectGamma1, maxobjectGamma2, and maxobjectRength represent the maximum value of object information when one or more object audio objects contained in an object audio track are shifted.
[0377] (Example syntax for sample entries of HOA audio tracks)
[0378] Figure 37 To show in Figure 33 The diagram shows an example syntax of sample entries for a HOA audio track.
[0379] like Figure 37 The ConfigurationVersion, MPEGHAudioProfile, and MPEGHAudioLevel shown represent the configuration information, profile information, and level information of HOA audio, respectively.
[0380] (Example syntax for sample entries in the object metadata track)
[0381] Figure 38 To show in Figure 33 The diagram shows an example syntax of sample entries for the object metadata track (object metadata track).
[0382] likeFigure 38 The ConfigurationVersion shown represents the configuration information for the metadata.
[0383] (First example of the segment structure of an audio file for 3D audio)
[0384] Figure 39 This is a schematic diagram illustrating a first example of the segment structure of an audio file for 3D audio in a first embodiment of the present disclosure.
[0385] In such Figure 39 In the segment structure shown, the initial segment consists of an ftyp box and a moov box. The trak box, used for each track included in the audio file, is arranged within the moov box. The mvex box is arranged within the moov box, and the mvex box contains information indicating the correspondence between the track ID of each track and the level used in the ssix box within the media segment.
[0386] Furthermore, a media segment consists of a sidx box, an ssix box, and one or more sub-segments. Position information indicating the location within the audio file of each sub-segment is placed in the sidx box. The ssix box contains the position information of the audio stream at each level, placed in the mdat box. It should be noted that each level corresponds to each track. Additionally, the position information for the first track is the position information of the data consisting of the audio stream from the moof box and the first track itself.
[0387] Sub-sections are set up for any time length. A pair of moof boxes and mdat boxes, shared by all tracks, are set up within sub-sections. Within the mdat box, the audio streams of all tracks are centrally arranged for any time length. The moof box contains management information for arranging the audio streams. The audio stream of each track arranged within the mdat box is continuous for each track.
[0388] exist Figure 39 In the example, track 1 with track ID 1 is the base track, and tracks 2 to N with track IDs 2 to N are the channel audio track, object audio track, HOA audio track, and object metadata track, respectively. The following descriptions... Figure 40 The same applies to the case of [the other party / entity].
[0389] (Second example of the segment structure of an audio file for 3D audio)
[0390] Figure 40 This is a schematic diagram illustrating a second example of the segment structure of an audio file for 3D audio in the first embodiment of the present disclosure.
[0391] As Figure 40 shown in FIG. 6, the segment structure is different from the segment structure as Figure 39 shown in FIG. 5 in that the moof box and the mdat box are set for each track.
[0392] Specifically, as Figure 40 shown in FIG. 7, the initial segment (Initial segment) is similar to the initial segment as Figure 39 shown in FIG. 6. Like the media segment as Figure 39 shown in FIG. 6, the media segment as Figure 40 shown in FIG. 7 is composed of the sidx box, the ssix box, and one or more sub segments. Further, like the sidx box as Figure 39 shown in FIG. 6, the position information of each sub segment is arranged in the sidx box. The ssix box contains the position information of each level of data composed of the moof box and the mdat box.
[0393] The sub segment is set for any time length. A pair of the moof box and the mdat box is set for each track in the sub segment. Specifically, the audio stream of each track is centrally arranged in any time length (interleaved and stored) in the mdat box of each track, and the management information of the audio stream is arranged in the moof box.
[0394] As Figure 39 and 40 shown, the audio stream for each track is centrally arranged in any time length so that, compared to the case where the audio stream is centrally arranged in a sample unit, the efficiency of acquiring the audio stream via HTTP or the like can be improved.
[0395] (Exemplary description of mvex box)
[0396] Figure 41 A schematic diagram for an exemplary description of the level assignment box arranged in the mvex box as Figure 39 and 40 shown.
[0397] The level assignment box is a box for associating the track ID of each track with the level used in the ssix box. In Figure 41 the example, the base track of the track ID 1 is associated with the level 0, and the channel audio track of the track ID 2 is associated with the level 1. Further, the HOA audio track of the track ID 3 is associated with the level 2, and the object metadata track of the track ID 4 is associated with the level 3. Further, the object audio track of the track ID 10 is associated with the level 4.
[0398] (Exemplary description of MPD file)
[0399] Figure 42 A diagram for illustrating an exemplary description of an MDF file in the first embodiment of the present disclosure.
[0400] As shown in Figure 42 , a "Representation" of a segment of an audio file (3daudio.mp4) for managing 3D audio, a "SubRepresentation" for managing a track contained in the segment, and the like are described in the MPD file.
[0401] In the "Representation" and the "SubRepresentation", a "codecs" is contained, which indicates a type of a codec of a corresponding segment or track in a code defined in the 3D audio file format. Further, an "id", an "associationId", and an "assciationType" are contained in the "Representation".
[0402] The "id" indicates an ID of the "Representation" containing the "id". The "associationId" indicates information indicating a reference relationship between a corresponding track and another track and indicates an "id" of a reference track. The "assciationType" indicates a code indicating a meaning of the reference relationship (correlation relationship) with respect to the reference track. For example, a value identical to a value of a track reference number of the MP4 is used.
[0403] Further, a "level" is contained in the "SubRepresentation", which is a value set in a level assignment box as a value indicating a corresponding track and a corresponding level. A "dependencyLevel" is contained in the "SubRepresentation", which is a value indicating a level corresponding to another track (hereinafter, referred to as a reference track) having a reference relationship (correlation).
[0404] Further, the "SubRepresentation" contains <EssentialProperty schemeIdUri="urn:mpeg:DASH:3daudio:2014" value="audioType, contentkind, priority"> as information required for selecting 3D audio.
[0405] Further, "SubRepresentation" in the object audio track contains <EssentialProperty schemeIdUri="urn:mpeg:DASH:viewingAngle:2014" value="0, y, r">. When the object is fixed corresponding to the "SubRepresentation", 0, y, and r represent a horizontal angle, a vertical angle, and a distance in the object position information, respectively. On the other hand, when the object is displaced, the values 0, y, and r represent a maximum value of a horizontal angle, a maximum value of a vertical angle, and a maximum value of a distance among maximum values of the object position information, respectively.
[0406] Figure 43 To show a schematic diagram of the definition of the essential property shown in Figure 42 .
[0407] In the upper left side of Figure 43 , audioType of <EssentialProperty schemeIdUri="urn:mpeg:DASH:3daudio:2014" value="audioType, contentkind, priority"> is defined. The audioType indicates a type of 3D audio of the corresponding track.
[0408] In the example of Figure 43 , when the audioType indicates 1, it indicates that the audio data of the corresponding track is a channel audio of 3D audio, and when the audioType indicates 2, it indicates that the audio data of the corresponding track is an HOA audio. Further, when the audioType indicates 3, it indicates that the audio data of the corresponding track is an object audio, and when the audioType is 4, it indicates that the audio data of the corresponding track is metadata.
[0409] Further, in the right side of Figure 43 , contentkind of <EssentialProperty schemeIdUri="urn:mpeg:DASH:3daudio:2014" value="audioType, contentkind, priority"> is defined. The contentkind indicates a content of the corresponding audio. For example, in the example of Figure 43 , when the contentkind indicates 3, the corresponding audio is music.
[0410] As Figure 43The priority is defined by 23008-3 and indicates the processing priority of the corresponding object, as shown on the lower left. The value indicating the processing priority of the object is described as "0" only when the value does not change in the course of the audio stream, and is described as "0" when the value changes in the course of the audio stream.
[0411] (Outline of information processing system)
[0412] Figure 44 A schematic diagram showing an outline of an information processing system according to a first embodiment of the present disclosure is shown.
[0413] In Figure 44 the same components as those shown in Figure 1 the same components as those shown in
[0414] The information processing system 140 shown in Figure 44 has a configuration in which a Web server 142 (connected to a file generation device 141) is connected to a video playback terminal 144 via the Internet 13.
[0415] In the information processing system 140, the Web server 142 transmits a video stream (tile streaming) of the video content to the video playback terminal 144 in units of tiles by an MPEG-DASH-compatible method. Further, in the information processing system 140, the Web server 142 transmits an audio stream of object audio, channel audio, or HOA audio corresponding to tiles to be played to the video playback terminal 144.
[0416] The file generation device 141 of the information processing system 140 is similar to the file generation device 11 shown in Figure 11 , except that, for example, the audio file generation unit 56 generates an audio file in the first embodiment and the MPD generation unit 57 generates an MPD file in the first embodiment.
[0417] Specifically, the file generation device 141 acquires image data of the video content and encodes the image data in units of tiles to generate a video stream. The file generation device 141 processes the video stream of each tile into a file format. The file generation device 141 uploads the image file of each tile obtained as a result of the processing to the Web server 142.
[0418] Further, the file generation device 141 acquires 3D audio of the video content and encodes the 3D audio for each type (channel audio / object audio / HOA audio / metadata) of the 3D audio to generate an audio stream. The file generation device 141 assigns a track to the audio stream for each type of the 3D audio. The file generation device 141 generates an MPD file as shown in Figure 39or the segment structure shown in FIG. 40 (in which the audio stream of each track is arranged in sub-segments) and uploads the audio file to the Web server 142.
[0419] The file generation apparatus 141 generates an MPD file containing the image frame size information, tile position information, and object position information. The file generation apparatus 141 uploads the MPD file to the Web server 142.
[0420] The Web server 142 stores the image file, the audio file, and the MPD file uploaded from the file generation apparatus 141.
[0421] In the example of FIG. 43, the Web server 142 stores a segment group formed of image files of a plurality of segments of tile #1 and a segment group formed of image files of a plurality of segments of tile #2. The Web server 142 also stores a segment group formed of an audio file of 3D audio. Figure 44
[0422] The Web server 142 sends the image file, the audio file, the MPD file, and the like stored in the Web server to the video playback terminal 144 in response to a request from the video playback terminal 144.
[0423] The video playback terminal 144 executes the control software 161, the video playback software 162, the access software 163, and the like.
[0424] The control software 161 is software for controlling data streamed from the Web server 142. Specifically, the control software 161 causes the video playback terminal 144 to acquire the MPD file from the Web server 142.
[0425] Further, the control software 161 specifies a tile in the display region based on the display region instructed from the video playback software 162 and the tile position information contained in the MPD file. Then, the control software 161 instructs the access software 163 to send a request for the image file of the tile.
[0426] When the object audio is to be played, the control software 161 instructs the access software 163 to send a request for the image frame size information in the audio file. Further, the control software 161 instructs the access software 163 to send a request for the audio stream of the metadata. The control software 161 specifies the object corresponding to the image in the display region based on the image frame size information and the object position information contained in the audio stream of the metadata, which is sent from the Web server 142 according to the instruction and the display region. Then, the control software 161 instructs the access software 163 to send a request for the audio stream of the object.
[0427] Further, when the channel audio or the HOA audio is to be played, the control software 161 instructs the access software 163 to send a request for an audio stream of the channel audio or the HOA audio.
[0428] The video playback software 162 is software for playing the image file and the audio file acquired from the Web server 142. Specifically, when a display area is designated by a user, the video playback software 162 instructs the control software 161 to send the display area. Further, the video playback software 162 decodes the image file and the audio file acquired from the Web server 142 according to the instruction. The video playback software 162 synthesizes the image data obtained as a result of the decoding in units of tiles and outputs the image data. Further, when necessary, the video playback software 162 synthesizes the object audio, the channel audio, or the HOA audio obtained as a result of the decoding and outputs the audio.
[0429] The access software 163 is software for controlling communication with the Web server 142 via the Internet 13 using HTTP. Specifically, the access software 163 causes the video playback terminal 144 to send a request for image frame size information or a predetermined audio stream in the image file and the audio file in response to an instruction of the control software 161. Further, the access software 163 causes the video playback terminal 144 to receive the image frame size information or the predetermined audio stream in the image file and the audio file sent from the Web server 12 in response to the sent request.
[0430] (Configuration Example of File Generation Apparatus)
[0431] Figure 45 A block diagram showing a configuration example of the file generation apparatus 141 shown in Figure 44
[0432] The same components as those shown in Figure 45 Figure 11 are denoted by the same reference numerals. Repeated explanation is omitted as appropriate.
[0433] The configuration of the file generation apparatus 141 shown in Figure 45 differs from that of the file generation apparatus 11 shown in Figure 11 in that the audio encoding processing unit 171, the audio file generation unit 172, the MPD generation unit 173, and the server upload processing unit 174 are provided in place of the audio encoding processing unit 55, the audio file generation unit 56, the MPD generation unit 57, and the server upload processing unit 58.
[0434] Specifically, the audio encoding processing unit 171 of the file generation apparatus 141 encodes the 3D audio of the video content inputted from the outside for each type (channel audio / object audio / HOA audio / metadata) to generate an audio stream. The audio encoding processing unit 171 provides the audio stream of the 3D audio of each type to the audio file generation unit 172.
[0435] The audio file generation unit 172 allocates a track to the audio stream provided from the audio encoding processing unit 171 for each type of 3D audio. The audio file generation unit 172 generates an audio file of the segment structure as shown in FIG. 40, in which the audio stream of each track is arranged in units of subsegments. At this time, the audio file generation unit 172 stores the image frame size information inputted from the outside in a sample entry. The audio file generation unit 172 provides the generated audio file to the MPD generation unit 173. Figure 39
[0436] The MPD generation unit 173 determines the URL of the Web server 142 storing the image file of each tile provided from the image file generation unit 53 and the like. Further, the MPD generation unit 173 determines the URL of the Web server 142 storing the audio file provided from the audio file generation unit 172 and the like.
[0437] The MPD generation unit 173 arranges the image information provided from the image information generation unit 54 in an "Adaptation Set" for the image of the MPD file. Further, the MPD generation unit 173 arranges the URL of the image file of each tile and the like in a "Segment" for the "Representation" of the tile.
[0438] The MPD generation unit 173 arranges the URL of the audio file and the like in a "Segment" for the "Representation" of the audio file. Further, the MPD generation unit 173 arranges the object position information of each object inputted from the outside and the like in a "SubRepresentation" for the object metadata track of the object. The MPD generation unit 173 provides the MPD file in which various pieces of information are arranged as described above, and the image file and the audio file to the server upload processing unit 174.
[0439] The server upload processing unit 174 uploads the image file, the audio file, and the MPD file of each tile provided from the MPD generation unit 173 to the Web server 142.
[0440] (Explanation of the process of the file generation apparatus)
[0441] Figure 46 To show the flowchart of the file generation process of the file generation apparatus 141 shown in Figure 45
[0442] The processes of steps S191 to S195 shown in Figure 46 are similar to those of steps S11 to S15 shown in Figure 12 , and thus their descriptions are omitted.
[0443] In step S196, the audio encoding processing unit 171 encodes the 3D audio of the video content inputted from the outside for each type (channel audio / object audio / HOA audio / metadata) to generate an audio stream. The audio encoding processing unit 171 provides the audio stream to the audio file generation unit 172 for each type of 3D audio.
[0444] In step S197, the audio file generation unit 172 allocates a track to the audio stream provided from the audio encoding processing unit 171 for each type of 3D audio.
[0445] In step S198, the audio file generation unit 172 generates an audio file of the segment structure shown in Figure 39 or 40 in which the audio stream of each track is arranged in units of subsegments. At this time, the audio file generation unit 172 stores the image frame size information inputted from the outside in a sample entry. The audio file generation unit 172 provides the generated audio file to the MPD generation unit 173.
[0446] In step S199, the MPD generation unit 173 generates an MPD file containing the image information provided from the image information generation unit 54, the URL of each file, and the object position information. The MPD generation unit 173 provides the image file, the audio file, and the MPD file to the server upload processing unit 174.
[0447] In step S200, the server upload processing unit 174 uploads the image file, the audio file, and the MPD file provided from the MPD generation unit 173 to the Web server 142. Then, the process is terminated.
[0448] (Functional configuration example of the video playback terminal)
[0449] Figure 47 To show a block diagram of the configuration example of the stream playback unit, which is implemented in the manner that the video playback terminal 144 executes the control software 161, the video playback software 162, and the access software 163 shown in Figure 44 .
[0450] In Figure 47 the same components as those shown in Figure 13 the same components as those shown in
[0451] As Figure 47 shown in the configuration of the stream playback unit 190 differs from that of the stream playback unit 90 shown in Figure 13 in that the MPD processing unit 191, the audio selection unit 193, the audio file acquisition unit 192, the audio decoding processing unit 194, and the audio synthesis processing unit 195 are provided in place of the MPD processing unit 92, the audio selection unit 94, the audio file acquisition unit 95, the audio decoding processing unit 96, and the audio synthesis processing unit 97, and the meta file acquisition unit 93 is not provided.
[0452] The stream playback unit 190 is similar to the stream playback unit 90 shown in Figure 13 except for, for example, the method of acquiring the audio data to be played of the selected object.
[0453] Specifically, the MPD processing unit 191 of the stream playback unit 190 extracts information such as the URL of the audio file of the segment to be played described in the "Segment" for the audio meta file from the MPD file provided from the MPD acquisition unit 91, and provides the extracted information to the audio file acquisition unit 192.
[0454] The MPD processing unit 191 extracts tile position information described in the "Adaptation Set" for the image from the MPD file, and provides the extracted information to the image selection unit 98. The MPD processing unit 191 extracts information such as the URL described in the "Segment" for the image file of the tile requested from the image selection unit 98 from the MPD file, and provides the extracted information to the image selection unit 98.
[0455] When the object audio is to be played, the audio file acquisition unit 192 requests the Web server 142 to send the initial segment (Initial Segment) of the base track in the audio file specified by the URL based on information such as the URL provided from the MPD processing unit 191, and acquires the initial segment of the base track.
[0456] Further, based on the information such as the URL of the audio file, the audio file acquisition unit 192 requests the Web server 142 to transmit the audio stream of the object metadata track in the audio file designated by the URL and acquires the audio stream of the object metadata track. The audio file acquisition unit 192 provides the audio selection unit 193 with the object position information contained in the audio stream of the object metadata track, the image frame size information contained in the initial fragment of the base track, and the information such as the URL of the audio file.
[0457] Further, when the channel audio is to be played, the audio file acquisition unit 192 requests the Web server 142 to transmit the audio stream of the channel audio track in the audio file designated by the URL based on the information such as the URL of the audio file and acquires the audio stream of the channel audio track. The audio file acquisition unit 192 provides the audio decoding processing unit 194 with the acquired audio stream of the channel audio track.
[0458] When the HOA audio is to be played, the audio file acquisition unit 192 performs a process similar to that performed when the channel audio is to be played. Thus, the audio stream of the HOA audio track is provided to the audio decoding processing unit 194.
[0459] It should be noted that which one of the object audio, the channel audio, and the HOA audio is played is determined, for example, in accordance with an instruction of the user.
[0460] The audio selection unit 193 calculates the position of each object on the image based on the image frame size information and the object position information provided from the audio file acquisition unit 192. The audio selection unit 193 selects the object in the display region designated by the user based on the position of each object on the image. Based on the information such as the URL of the audio file provided from the audio file acquisition unit 192, the audio selection unit 193 requests the Web server 142 to transmit the audio stream of the object audio track of the selected object in the audio file designated by the URL and acquires the audio stream of the object audio track. The audio selection unit 193 provides the audio decoding processing unit 194 with the acquired audio stream of the object audio track.
[0461] The audio decoding processing unit 194 decodes the audio stream of the channel audio track or the HOA audio track provided from the audio file acquisition unit 192, or decodes the audio stream of the object audio track provided from the audio selection unit 193. The audio decoding processing unit 194 provides the audio synthesis processing unit 195 with one of the channel audio, the HOA audio, and the object audio obtained as a result of the decoding.
[0462] When necessary, the audio synthesis processing unit 195 synthesizes the object audio, the channel audio, or the HOA audio supplied from the audio decoding processing unit 194 and outputs the audio.
[0463] (Explanation of the process of the video playback terminal)
[0464] Figure 48 A flowchart showing a channel audio playback process of the stream playback unit 190 shown in Figure 47 is executed, for example, when a user selects channel audio as an object to be played back.
[0465] In Figure 48 S221, the MPD processing unit 191 analyzes the MPD file supplied from the MPD acquisition unit 91 and specifies "SubRepresentation" ("sub representation") of channel audio of a segment to be played back on the basis of the essential attribute and the codec described in "SubRepresentation" ("sub representation"). Further, the MPD processing unit 191 extracts information such as a URL described in "Segment" ("segment") of an audio file for the segment to be played back from the MPD file and supplies the extracted information to the audio file acquisition unit 192.
[0466] In step S222, the MPD processing unit 191 specifies the level of the essential track as a reference track on the basis of the dependencyLevel of the "SubRepresentation" ("sub representation") specified in step S221 and supplies the specified level of the essential track to the audio file acquisition unit 192.
[0467] In step S223, the audio file acquisition unit 192 requests the Web server 142 to send an initial segment of a segment to be played back and acquires the initial segment on the basis of information such as a URL supplied from the MPD processing unit 191.
[0468] In step S224, the audio file acquisition unit 192 acquires a track ID corresponding to the level of the channel audio track and the essential track as a reference track from a level assignment box in the initial segment.
[0469] In step S225, the audio file acquisition unit 192 acquires a sample entry of the initial segment in a track box corresponding to the track ID of the initial segment on the basis of the track ID of the channel audio track and the essential track as a reference track. The audio file acquisition unit 192 supplies codec information contained in the acquired sample entry to the audio decoding processing unit 194.
[0470] In step S226, based on information such as the URL provided from the MPD processing unit 191, the audio file acquisition unit 192 sends a request to the Web server 142 and acquires the sidx box and the ssix box from the header of the audio file of the segment to be played.
[0471] In step S227, the audio file acquisition unit 192 acquires the position information of the reference track and the channel audio track of the segment to be played from the sidx box and the ssix box acquired in step S223. In this case, since the base track as the reference track does not contain any audio stream, there is no position information of the reference track.
[0472] In step S228, the audio file acquisition unit 192, based on the position information of the channel audio track and information such as the URL of the audio file of the segment to be played, requests the Web server 142 to send the audio stream of the channel audio track arranged in the mdat box, and acquires the audio stream of the channel audio track. The audio file acquisition unit 192 provides the audio stream of the channel audio track acquired to the audio decoding processing unit 194.
[0473] In step S229, the audio decoding processing unit 194 decodes the audio stream of the channel audio track based on the codec information provided from the audio file acquisition unit 192. The audio file acquisition unit 192 provides the channel audio obtained as a result of the decoding to the audio synthesis processing unit 195.
[0474] In step S230, the audio synthesis processing unit 195 outputs the channel audio. Then the process is terminated.
[0475] It should be noted that although not shown, the HOA audio playback process for playing the HOA audio by the stream playback unit 190 is executed in a manner similar to the channel audio playback process as shown in FIG. 17. Figure 48
[0476] A flowchart showing the object designation process of the stream playback unit 190 shown in FIG. 16 is shown. The object designation process is executed, for example, when the user selects the object audio as the object to be played and the playback area is changed. Figure 49 Figure 47 In step S251 of FIG. 16, the audio selection unit 193 acquires the display area designated by the user through the user's operation or the like.
[0477] In step S251 of FIG. 16, the audio selection unit 193 acquires the display area designated by the user through the user's operation or the like. Figure 49
[0478] In step S252, the MPD processing unit 191 analyzes the MPD file provided by the MPD acquisition unit 91 and specifies the "SubRepresentation" of the metadata of the segment to be played based on the basic attributes and the encoding / decoding described in the "SubRepresentation". Furthermore, the MPD processing unit 191 extracts information from the MPD file (such as the URL of the audio file of the segment to be played, described in the "Segment" for audio metadata files) and provides the extracted information to the audio file acquisition unit 192.
[0479] In step S253, the MPD processing unit 191 specifies the level of the base track as the reference track based on the dependencyLevel of the “SubRepresentation” specified in step S252, and provides the specified level of the base track to the audio file acquisition unit 192.
[0480] In step S254, the audio file acquisition unit 192 requests the web server 142 to send the initial segment of the segment to be played and acquires the initial segment based on information such as the URL provided by the MPD processing unit 191.
[0481] In step S255, the audio file acquisition unit 192 obtains the track ID corresponding to the level of the object metadata track and the base track used as a reference track from the level assignment box in the initial segment.
[0482] In step S256, the audio file acquisition unit 192 acquires sample entries of the initial segment in the trackbox corresponding to the track ID of the initial segment, based on the object metadata track and the track ID of the base track used as a reference track. The audio file acquisition unit 192 provides the audio selection unit 193 with image frame size information contained in the sample entries of the base track used as a reference track. Furthermore, the audio file acquisition unit 192 provides the initial segment to the audio selection unit 193.
[0483] In step S257, based on information such as the URL provided by the MPD processing unit 191, the audio file acquisition unit 192 sends a request to the web server 142 and acquires the sidx and ssix boxes from the header of the audio file of the segment to be played.
[0484] In step S258, the audio file acquisition unit 192 acquires the position information of the object metadata track of the reference track and the sub-fragments to be played from the sidx box and the ssix box acquired in step S257. In this case, since the base track as the reference track does not contain any audio stream, there is no position information of the reference track. The audio file acquisition unit 192 provides the sidx box and the ssix box to the audio selection unit 193.
[0485] In step S259, the audio file acquisition unit 192 requests the Web server 142 to send the audio stream of the object metadata track arranged in the mdat box based on the position information of the object metadata track and the information such as the URL of the audio file of the fragment to be played, and acquires the audio stream of the object metadata track.
[0486] In step S260, the audio file acquisition unit 192 decodes the audio stream of the object metadata track acquired in step S259 based on the codec information contained in the sample entry acquired in step S256. The audio file acquisition unit 192 provides the object position information contained in the metadata obtained as a result of the decoding to the audio selection unit 193. Further, the audio file acquisition unit 192 provides the information such as the URL of the audio file provided from the MPD processing unit 191 to the audio selection unit 193.
[0487] In step S261, the audio selection unit 193 selects the object in the display area based on the image frame size information and the object position information provided from the audio file acquisition unit 192 and based on the display area specified by the user. Then the process is terminated.
[0488] Figure 50 A flowchart showing the specified object audio playback process executed by the stream playback unit 190 after the object specifying process shown in Figure 49
[0489] In step S281 of Figure 50 , the MPD processing unit 191 analyzes the MPD file provided from the MPD acquisition unit 91, and specifies the "SubRepresentation" ("sub representation") of the object audio of the selected object based on the base attribute and the codec described in the "SubRepresentation".
[0490] In step S282, the MPD processing unit 191 specifies the level of the base track that is the reference track based on the dependencyLevel of the "SubRepresentation" specified in step S281, and supplies the audio file acquisition unit 192 with the specified level of the base track.
[0491] In step S283, the audio file acquisition unit 192 acquires the track ID corresponding to the level of the object audio track and the base track that is the reference track from the Level assignment box in the initial segment, and supplies the audio selection unit 193 with the track ID.
[0492] In step S284, the audio selection unit 193 acquires the sample entry of the initial segment in the track box corresponding to the track ID of the initial segment based on the track ID of the object audio track and the base track that is the reference track. The initial segment is supplied from the audio file acquisition unit 192 in step S256 as shown in Figure 49 The audio selection unit 193 supplies the audio decoding processing unit 194 with the codec information contained in the acquired sample entry.
[0493] In step S285, the audio selection unit 193 acquires the position information of the object audio track of the selected object of the reference track and the sub-segment to be played from the sidx box and the ssix box supplied from the audio file acquisition unit 192 in step S258. In this case, since the base track that is the reference track does not contain any audio stream, there is no position information of the reference track.
[0494] In step S286, the audio selection unit 193 requests the Web server 142 to send the audio stream of the object audio track of the selected object arranged in the mdat box based on the position information of the object audio track and information such as the URL of the audio file of the segment to be played, and acquires the audio stream of the object audio track. The audio selection unit 193 supplies the audio decoding processing unit 194 with the acquired audio stream of the object audio track.
[0495] In step S287, the audio decoding processing unit 194 decodes the audio stream of the object audio track based on the codec information supplied from the audio selection unit 193. The audio selection unit 193 supplies the audio synthesis processing unit 195 with the object audio obtained as a result of the decoding.
[0496] In step S288, the audio synthesis processing unit 195 synthesizes the object audio supplied from the audio decoding processing unit 194 and outputs the object audio. The process is then terminated.
[0497] As described above, in the information processing system 140, the file generation apparatus 141 generates an audio file in which the 3D audio is divided into a plurality of tracks and the tracks are arranged according to the type of the 3D audio. The video playback terminal 144 acquires an audio stream of the predetermined type of 3D audio in the audio file. Thus, the video playback terminal 144 can efficiently acquire an audio stream of the predetermined type of 3D audio. Therefore, it can be said that the file generation apparatus 141 generates an audio file capable of improving the efficiency of acquiring an audio stream of the predetermined type of 3D audio.
[0498] <Second Embodiment>
[0499] (Outline of Track)
[0500] Figure 51 A schematic diagram for illustrating an outline of a track in the second embodiment of the present disclosure is shown.
[0501] As Figure 51 indicated, the second embodiment differs from the first embodiment in that a basic sample is recorded as a sample of a basic track. The basic sample is formed of information that is referred to by a sample of channel audio / object audio / HOA audio / metadata. The samples of channel audio / object audio / HOA audio / metadata that refer to the reference information contained in the basic sample are arranged in the order of arrangement of the reference information, thereby making it possible to generate an audio stream of 3D audio before the 3D audio is divided into tracks.
[0502] (Example Syntax of Sample Entry of Basic Track)
[0503] Figure 52 A schematic diagram for illustrating the example syntax of the sample entry of the basic track shown in Figure 51
[0504] The syntax as Figure 52 indicated is the same as the syntax as Figure 34 indicated, except that "mha2" that describes that the sample entry is a sample entry of the basic track as Figure 51 indicated, instead of "mha1" that describes that the sample entry is a sample entry of the basic track as Figure 33 indicated.
[0505] (Example Structure of Basic Entry)
[0506] Figure 53 A schematic diagram for illustrating the example structure of the basic sample is shown.
[0507] As Figure 53 As shown, the base sample is configured with an extractor of channel audio / object audio / HOA audio / metadata in units of a sample that is a sub-sample of the base sample. The extractor of channel audio / object audio / HOA audio / metadata consists of a type of the extractor and an offset and size of a sub-sample of a corresponding channel audio track / object audio track / HOA audio track / object metadata track. The offset is a difference between a position of the base sample in a file of a sub-sample of the base sample and a position of the channel audio track / object audio track / HOA audio track / object metadata track in a file of the sample. In other words, the offset is information indicating a position within a file of a sample of another track corresponding to the sub-sample of the base sample that contains the offset.
[0508] Figure 54 A diagram for illustrating an example syntax of a base sample.
[0509] As Figure 54 shown, in the base sample, an SCE element for storing object audio in a sample of an object audio track is replaced with an EXT element for storing an extractor.
[0510] Figure 55 A diagram for illustrating an example of extractor data.
[0511] As Figure 55 shown, a type of the extractor and an offset and size of a sub-sample of a corresponding channel audio track / object audio track / HOA audio track / object metadata track are described in the extractor.
[0512] It should be noted that the extractor can utilize a network abstraction layer (NAL) structure extension defined in advanced video coding (AVC) / high efficiency video coding (HEVC) so that the audio element and the configuration information can be stored.
[0513] The information processing system in the second embodiment and the process performed by the information processing system are similar to those of the first embodiment, and thus the description thereof is omitted.
[0514] <Third Embodiment>
[0515] (Overview of Track)
[0516] Figure 56 A diagram for illustrating an overview of a track in which the third embodiment of the present disclosure is applied.
[0517] As Figure 56 shown, the third embodiment differs from the first embodiment in that a base sample and a sample of metadata are recorded as a sample of a base track and an object metadata track is not provided.
[0518] The information processing system and the process executed by the information processing system in the third embodiment are similar to those in the first embodiment, except that an audio stream of a base track instead of an object metadata track is acquired in order to obtain object location information. Therefore, its description is omitted.
[0519] <Fourth Embodiment>
[0520] (Overview of the orbit)
[0521] Figure 57 A schematic diagram illustrating an overview of the track in a fourth embodiment of the present disclosure.
[0522] like Figure 57 As shown, the fourth embodiment differs from the first embodiment in that the tracks are recorded as different files (3da_base.mp4 / 3da_channel.mp4 / 3da_object_1.mp4 / 3da_hoa.mp4 / 3da_meta.mp4). In this case, only the audio data of the desired track can be obtained via HTTP by retrieving the file of the desired track. Therefore, the audio data of the desired track can be effectively obtained via HTTP.
[0523] (Example description of an MPD file)
[0524] Figure 58 This is a schematic diagram illustrating an exemplary description of an MDF file according to a fourth embodiment of the present disclosure.
[0525] like Figure 58 As shown, the "Representation" ("representation") of each segment of the 3D audio file (3da_base.mp4 / 3da_channel.mp4 / 3da_object_1.mp4 / 3da_hoa.mp4 / 3da_meta.mp4) is described in the MPD file.
[0526] The "Representation" field contains "codecs", "id", "associationId", and "assciationType". Additionally, the "Representation" field for channel audio tracks / object audio tracks / HOA audio tracks / object metadata tracks also contains... <EssentialProperty schemeIdUri="urn:mpeg:DASH:3daudio:2014"value="audioType, contentkind,priority"> Additionally, the "Representation" of the object's audio track contains...<EssentialProperty schemeIdUri="urn:mpeg:DASH:viewingAngle:2014" value="θ,γ,r"> .
[0527] (Overview of Information Processing Systems)
[0528] Figure 59 This is a schematic diagram illustrating an overview of an information processing system in which the present disclosure is applied in a fourth embodiment.
[0529] exist Figure 59 The one shown in the middle is the same as Figure 1 Components that are identical in the diagram are indicated by the same reference numeral. Where appropriate, repeated descriptions are omitted.
[0530] like Figure 59 The information processing system 210 shown has the following configuration: a web server 212 connected to a file generation device 211 and a video playback terminal 214 are connected via the Internet 13.
[0531] In the information processing system 210, the web server 212 transmits video streams of video content to the video playback terminal 214 in units of tiles (tile streaming) using an MPEG-DASH compatible method. Furthermore, in the information processing system 210, the web server 212 transmits audio files corresponding to the object audio, channel audio, or HOA audio of the file to be played to the video playback terminal 214.
[0532] Specifically, the file generation device 211 acquires image data of the video content and encodes the image data in units of tiles to generate a video stream. The file generation device 211 processes the video stream of each tile into a file format for each segment. The file generation device 211 uploads the image file of each file obtained as a result of the above processing to the web server 212.
[0533] In addition, the file generation device 211 acquires 3D audio of the video content and encodes the 3D audio for each type (channel audio / object audio / HOA audio / metadata) to generate an audio stream. The file generation device 211 assigns tracks to the audio stream of each type of 3D audio. The file generation device 211 generates an audio file (with the audio stream arranged in the audio file) for each track and uploads the generated audio file to the web server 212.
[0534] The file generation device 211 generates an MPD file, which includes image frame size information, tile position information, and object position information. The file generation device 211 uploads the MPD file to the web server 212.
[0535] Web server 212 stores image files uploaded from file generation device 211, audio files for each type of 3D audio, and MPD files.
[0536] exist Figure 59 In the example, Web server 212 stores a group of segments formed by image files of multiple fragments of tile #1 and a group of segments formed by image files of multiple fragments of tile #2. Web server 212 also stores a group of segments formed by audio files of channel audio and a group of segments of audio files of object #1.
[0537] In response to a request from the video playback terminal 214, the web server 212 transmits image files, audio files of a predetermined type of 3D audio, MPD files, etc., stored in the web server to the video playback terminal 214.
[0538] The video playback terminal 214 executes control software 221, video playback software 222, access software 223, etc.
[0539] Control software 221 is software used to control the data streaming from web server 212. Specifically, control software 221 causes video playback terminal 214 to retrieve MPD files from web server 212.
[0540] Furthermore, the control software 221 specifies the tiles in the MPD file based on the display area commanded by the video playback software 222 and the tile position information contained in the MPD file. Then, the control software 221 commands the access software 223 to send a request for transmitting the image file of that tile.
[0541] When the object audio is to be played, the control software 221 instructs the access software 223 to send a request for sending the audio file of the basic track. Then, the control software 221 instructs the access software 223 to send a request for sending the audio file of the object metadata track. The control software 221 acquires the image frame size information in the audio file of the basic track, which is sent from the Web server 142 according to the instruction, and the object position information contained in the audio file of the metadata. The control software 221 specifies the object corresponding to the image in the display area based on the image frame size information, the object position information, and the display area. In addition, the control software 221 instructs the access software 223 to send a request for sending the audio file of the object.
[0542] In addition, when the channel audio or the HOA audio is to be played, the control software 221 instructs the access software 223 to send a request for sending the audio file of the channel audio or the HOA audio.
[0543] The video playback software 222 is software for playing the image file and the audio file acquired from the Web server 212. Specifically, when the display area is specified by the user, the video playback software 222 gives an instruction on the display area to the control software 221. In addition, the video playback software 222 decodes the image file and the audio file acquired from the Web server 212 according to the instruction. The video playback software 222 synthesizes the image data obtained as a result of the decoding in units of tiles and outputs the image data. In addition, when necessary, the video playback software 222 synthesizes the object audio, the channel audio, or the HOA audio obtained as a result of the decoding and outputs the audio.
[0544] The access software 223 is software for controlling the communication with the Web server 212 via the Internet 13 using HTTP. Specifically, the access software 223 causes the video playback terminal 214 to send a request for sending the image file and the predetermined audio file in response to the instruction from the control software 221. In addition, the access software 223 causes the video playback terminal 214 to receive the image file and the predetermined audio file sent from the Web server 212 according to the transmission request.
[0545] (Configuration example of file generation apparatus)
[0546] Figure 60 is a block diagram of the file generation apparatus 211 shown in Figure 59 .
[0547] is a block diagram of the file generation apparatus 211 shown in Figure 60 . Figure 45 The same components as those shown in FIG. 1 are denoted by the same reference numerals. The repeated explanation is omitted as appropriate.
[0548] As Figure 60 shown in FIG. 21, the configuration of the file generation apparatus 211 is different from that of the file generation apparatus 141 shown in FIG. 17 in that an audio file generation unit 241, an MPD generation unit 242, and a server upload processing unit 243 are provided to replace the audio file generation unit 172, the MPD generation unit 173, and the server upload processing unit 174, respectively. Figure 45
[0549] Specifically, the audio file generation unit 241 of the file generation apparatus 211 allocates a track to an audio stream for each type of 3D audio, the audio stream being provided from the audio encoding processing unit 171. The audio file generation unit 241 generates an audio file in which the audio stream is arranged, for each track. At this time, the audio file generation unit 241 stores the image frame size information inputted from the outside in a sample entry of a base track. The audio file generation unit 241 provides the MPD generation unit 242 with the audio file for each type of 3D audio.
[0550] The MPD generation unit 242 determines a URL or the like of the Web server 212 that stores the image file of each tile provided from the image file generation unit 53. Further, the MPD generation unit 242 determines a URL or the like of the Web server 212 that stores the audio file provided from the audio file generation unit 241, for each type of 3D audio.
[0551] The MPD generation unit 242 arranges the image information provided from the image information generation unit 54 in an "Adaptation Set" for the image of the MPD file. Further, the MPD generation unit 242 arranges the URL or the like of the image file of each tile in a "Segment" for the "Representation" of the tile.
[0552] The MPD generation unit 242 arranges the URL or the like of the audio file in a "Segment" for the "Representation" of the audio file, for each type of 3D audio. Further, the MPD generation unit 242 arranges the object position information or the like of each object inputted from the outside in the "Representation" for the object metadata track of the object. The MPD generation unit 242 provides the server upload processing unit 243 with the MPD file, the image file, and the audio file for each type of 3D audio in which various pieces of information are arranged as described above.
[0553] The server upload processing unit 243 uploads the image file of each tile, the audio file of each type of 3D audio, and the MPD file provided from the MPD generation unit 242 to the Web server 212.
[0554] (Explanation of the process of the file generation apparatus)
[0555] Figure 61 A flowchart showing the file generation process of the file generation apparatus 211 shown in Figure 60
[0556] The processes of steps S301 to S307 shown in Figure 61 are similar to those of steps S191 to S197 shown in Figure 46 , and thus the description thereof is omitted.
[0557] In step S308, the audio file generation unit 241 generates an audio file in which an audio stream is arranged for each track. At this time, the audio file generation unit 241 stores the image frame size information inputted from the outside in a sample entry in the audio file of the base track. The audio file generation unit 241 provides the MPD generation unit 242 with the generated audio file for each type of 3D audio.
[0558] In step S309, the MPD generation unit 242 generates an MPD file containing the image information provided from the image information generation unit 54, the URL of each file, and the object position information. The MPD generation unit 242 provides the server upload processing unit 243 with the image file, the audio file for each type of 3D audio, and the MPD file.
[0559] In step S310, the server upload processing unit 243 uploads the image file, the audio file of each type of 3D audio, and the MPD file provided from the MPD generation unit 242 to the Web server 212. Then, the process is terminated.
[0560] (Functional configuration example of the video playback terminal)
[0561] Figure 62 A block diagram showing a configuration example of the stream playback unit is shown in Figure 59 , which is implemented in such a manner that the video playback terminal 214 executes the control software 221, the video playback software 222, and the access software 223.
[0562] In Figure 62 , the same components as those shown in Figure 13 and 47 are denoted by the same reference numerals. The repeated description is omitted as appropriate.
[0563] As Figure 62 indicated in the configuration of the stream playback unit 260 differs from that of the stream playback unit 90 as indicated in Figure 13 the MPD processing unit 92, the meta file acquisition unit 93, the audio selection unit 94, the audio file acquisition unit 95, the audio decoding processing unit 96 and the audio synthesis processing unit 97 are replaced by the MPD processing unit 261, the meta file acquisition unit 262, the audio selection unit 263, the audio file acquisition unit 264, the audio decoding processing unit 194 and the audio synthesis processing unit 195, respectively.
[0564] Specifically, when object audio is to be played, the MPD processing unit 261 of the stream playback unit 260 extracts information (such as URLs described in "Segment" of the audio file of the object metadata track of the segment to be played) from the MPD file supplied from the MPD acquisition unit 91 and supplies the extracted information to the meta file acquisition unit 262. Further, the MPD processing unit 261 extracts information (such as URLs described in "Segment" of the audio file of the object audio track of the object requested from the audio selection unit 263) from the MPD file and supplies the extracted information to the audio selection unit 263. Further, the MPD processing unit 261 extracts information (such as URLs described in "Segment" of the audio file of the basic track of the segment to be played) from the MPD file and supplies the extracted information to the meta file acquisition unit 262.
[0565] Further, when channel audio or HOA audio is to be played, the MPD processing unit 261 extracts information (such as URLs described in "Segment" of the audio file of the channel audio track or the HOA audio track of the segment to be played) from the MPD file. The MPD processing unit 261 supplies information such as URLs to the audio file acquisition unit 264 via the audio selection unit 263.
[0566] It should be noted that which one of the object audio, the channel audio and the HOA audio is to be played is determined, for example, in accordance with an instruction of a user.
[0567] The MPD processing unit 261 extracts tile position information described in an "Adaptation Set" for an image from an MPD file and supplies the extracted tile position information to the image selection unit 98. The MPD processing unit 261 extracts information such as a URL described in a "Segment" for an image file of a tile requested from the image selection unit 98 from an MPD file and supplies the extracted information to the image selection unit 98.
[0568] Based on information such as a URL supplied from the MPD processing unit 261, the meta file acquisition unit 262 requests the Web server 212 to transmit an audio file of an object metadata track designated by the URL and acquires the audio file of the object metadata track. The meta file acquisition unit 93 supplies object position information contained in an audio meta file of the object metadata track to the audio selection unit 263.
[0569] Further, based on information such as a URL of an audio file, the meta file acquisition unit 262 requests the Web server 142 to transmit an initial segment of an audio file of a base track designated by the URL and acquires the initial segment. The meta file acquisition unit 262 supplies image frame size information contained in a sample entry of the initial segment to the audio selection unit 263.
[0570] The audio selection unit 263 calculates the position of each object on an image based on the image frame size information and the object position information supplied from the meta file acquisition unit 262. The audio selection unit 263 selects an object in a display region designated by a user based on the position of each object on an image. The audio selection unit 263 requests the MPD processing unit 261 to transmit information such as a URL of an audio file of an object audio track of the selected object. The audio selection unit 263 supplies information such as a URL of an audio file of an object audio track, a channel audio track, or an HOA audio track supplied from the MPD processing unit 261 to the audio file acquisition unit 264 in accordance with the request.
[0571] Based on information such as a URL of an audio file of an object audio track, a channel audio track, or an HOA audio track supplied from the audio selection unit 263, the audio file acquisition unit 264 requests the Web server 12 to transmit an audio stream of the audio file designated by the URL and acquires the audio stream of the audio file. The audio file acquisition unit 95 supplies the acquired object-unit audio file to the audio decoding processing unit 194.
[0572] (Explanation of the process of the video playback terminal)
[0573] Figure 63 To show the process of the video playback terminal Figure 62The flowchart shown is a process for playing audio channels in the streaming unit 260. For example, this process is executed when a user selects a channel audio channel as the object to be played.
[0574] exist Figure 63 In step S331, the MPD processing unit 261 analyzes the MPD file provided by the MPD acquisition unit 91 and specifies the "Representation" of the channel audio of the segment to be played based on the basic attributes and the codec described in the "Representation". Furthermore, the MPD processing unit 261 extracts information (such as the URL of the audio file for the channel audio track of the segment to be played, described in the "Segment" contained in the "Representation") and provides the extracted information to the audio file acquisition unit 264 via the audio selection unit 263.
[0575] In step S332, based on the associationId of the “Representation” specified in step S331, the MPD processing unit 261 specifies the “Representation” of the base track as the reference track. The MPD processing unit 261 extracts information (such as the URL of the audio file of the reference track described in the “Segment” contained in the “Representation”) and provides the extracted file to the audio file acquisition unit 264 via the audio selection unit 263.
[0576] In step S333, the audio file acquisition unit 264 requests the Web server 212 to send the initial segment of the audio file of the channel audio track and the reference track of the segment to be played, based on information such as the URL provided by the audio selection unit 263, and acquires the initial segment.
[0577] In step S334, the audio file acquisition unit 264 acquires sample entries in the trak box of the acquired initial segment. The audio file acquisition unit 264 provides the encoding and decoding information contained in the acquired sample entries to the audio decoding processing unit 194.
[0578] In step S335, the audio file acquisition unit 264 sends a request to the web server 142 based on information such as the URL provided by the audio selection unit 263, and acquires the sidx box and ssix box from the header of the audio file of the channel audio track of the segment to be played.
[0579] In step S336, the audio file acquisition unit 264 acquires position information of a subsegment to be played from the sidx box and the ssix box acquired in step S333.
[0580] In step S337, the audio selection unit 263 requests the Web server 142 to transmit an audio stream of a channel audio track arranged in an mdat box of an audio file based on the position information acquired in step S337 and information such as a URL of the audio file of the channel audio track to be played, and acquires the audio stream of the channel audio track. The audio selection unit 263 provides the audio decoding processing unit 194 with the acquired audio stream of the channel audio track.
[0581] In step S338, the audio decoding processing unit 194 decodes the audio stream of the channel audio track provided from the audio selection unit 263 based on codec information provided from the audio file acquisition unit 264. The audio selection unit 263 provides the audio synthesis processing unit 195 with channel audio acquired as a result of the decoding.
[0582] In step S339, the audio synthesis processing unit 195 outputs the channel audio. Then, the process is terminated.
[0583] Although not shown, a HOA audio playback process for playing a HOA audio by the stream playback unit 260 is executed in a manner similar to the channel audio playback process as shown in FIG. 17. Figure 63
[0584] A flowchart showing the object audio playback process of the stream playback unit 260 shown in FIG. 18 is shown in FIG. 19. The object audio playback process is executed, for example, when a user selects an object audio as an object to be played and a playback area is changed. Figure 64 Figure 62 In step S351 of FIG. 18, the audio selection unit 263 acquires a display area designated by a user through the user's operation or the like.
[0585] In step S351 of FIG. 18, the audio selection unit 263 acquires a display area designated by a user through the user's operation or the like. Figure 64
[0586] In step S352, the MPD processing unit 261 analyzes the MPD file supplied from the MPD acquisition unit 91, and specifies the "Representation" ("representation") of the metadata of the segment to be played based on the basic attribute and the codec described in the "Representation" ("representation"). Further, the MPD processing unit 261 extracts information such as the URL of the audio file of the object metadata track of the segment to be played described in the "Segment" ("segment") included in the "Representation" ("representation"), and supplies the extracted information to the meta file acquisition unit 262.
[0587] In step S353, the MPD processing unit 261 specifies the "Representation" ("representation") of the base track that is the reference track based on the associationId of the "Representation" ("representation") specified in step S352. The MPD processing unit 261 extracts information such as the URL of the audio file of the reference track described in the "Segment" ("segment") included in the "Representation" ("representation"), and supplies the extracted information to the meta file acquisition unit 262.
[0588] In step S354, the meta file acquisition unit 262 requests the Web server 212 to transmit the initial segment of the audio file of the object metadata track and the reference track of the segment to be played based on information such as the URL supplied from the MPD processing unit 261, and acquires the initial segment.
[0589] In step S355, the meta file acquisition unit 262 acquires the sample entry in the trak box of the acquired initial segment. The meta file acquisition unit 262 supplies the image frame size information included in the sample entry of the base track that is the reference track to the audio file acquisition unit 264.
[0590] In step S356, the meta file acquisition unit 262 transmits a request to the Web server 142 based on information such as the URL supplied from the MPD processing unit 261, and acquires the sidx box and the ssix box from the header of the audio file of the object metadata track of the segment to be played.
[0591] In step S357, the meta file acquisition unit 262 acquires the position information of the sub-segment to be played from the sidx box and the ssix box acquired in step S356.
[0592] In step S358, the metafile acquisition unit 262, based on the location information obtained in step S357 and information such as the URL of the audio file of the object metadata track of the segment to be played, requests the Web server 142 to transmit the audio stream of the object metadata track arranged in the mdat box of the audio file, and acquires the audio stream of the object metadata track.
[0593] In step S359, the metadata acquisition unit 262 decodes the audio stream of the object metadata track acquired in step S358 based on the encoding / decoding information contained in the sample entries acquired in step S355. The metadata acquisition unit 262 provides the audio selection unit 263 with the object location information contained in the metadata as a result of the decoding.
[0594] In step S360, the audio selection unit 263 selects an object in the display area based on the image frame size information and the object location information provided by the metafile acquisition unit 262, and based on the display area specified by the user. The audio selection unit 263 requests the MPD processing unit 261 to send information such as the URL of the audio file of the object's audio track.
[0595] In step S361, MPD processing unit 261 analyzes the MPD file provided by MPD acquisition unit 91 and specifies the "Representation" of the selected object's audio based on basic attributes and the codec described in the "Representation". Furthermore, MPD processing unit 261 extracts information (such as the URL of the audio file of the selected object's audio track, described in the "Segment" contained in the "Representation"), and provides the extracted information to audio file acquisition unit 264 via audio selection unit 263.
[0596] In step S362, based on the associationId of the “Representation” specified in step S361, the MPD processing unit 261 specifies the “Representation” of the base track as the reference track. The MPD processing unit 261 extracts information (such as the URL of the audio file of the reference track described in the “Segment” contained in the “Representation”) and provides the extracted information to the audio file acquisition unit 264 via the audio selection unit 263.
[0597] In step S363, the audio file acquisition unit 264 requests the Web server 212 to transmit the initial segment of the audio file of the object audio track and the reference track of the segment to be played based on information such as the URL provided from the audio selection unit 263, and acquires the initial segment.
[0598] In step S364, the audio file acquisition unit 264 acquires the sample entry in the trak box of the acquired initial segment. The audio file acquisition unit 264 provides the codec information contained in the sample entry to the audio decoding processing unit 194.
[0599] In step S365, the audio file acquisition unit 264 transmits a request to the Web server 142 based on information such as the URL provided from the audio selection unit 263, and acquires the sidx box and the ssix box from the head of the audio file of the object audio track of the segment to be played.
[0600] In step S366, the audio file acquisition unit 264 acquires the position information of the sub-segment to be played from the sidx box and the ssix box acquired in step S365.
[0601] In step S367, the audio file acquisition unit 264 requests the Web server 142 to transmit the audio stream of the object audio track arranged in the mdat box within the audio file based on the position information acquired in step S366 and information such as the URL of the audio file of the object audio track of the segment to be played, and acquires the audio stream of the object audio track. The audio file acquisition unit 264 provides the acquired audio stream of the object audio track to the audio decoding processing unit 194.
[0602] The processes of steps S368 and S369 are similar to the processes of steps S287 and S288 as Figure 50 shown in FIG. 17, and thus the description thereof is omitted.
[0603] Note that in the above description, the audio selection unit 263 selects all the objects in the display region. However, the audio selection unit 263 can select only the objects having a high processing priority in the display region, or can select only the audio objects of predetermined content.
[0604] Figure 65 A flowchart of the object audio playback process when the audio selection unit 263 selects only the objects having a high processing priority among the objects in the display region.
[0605] The object audio playback process as Figure 65 shown in FIG. 18 is similar to the object audio playback process as Figure 64 shown in FIG. 17, except that Figure 65The process shown in step S390 is performed to replace, for example... Figure 64 Step S360 is shown. Specifically, as shown... Figure 65 The processes shown in steps S381 to S389 and steps S391 to S399 are similar to those shown in the figure. Figure 64 The processes of steps S351 to S359 and steps S361 to S369 are shown. Therefore, only the process of step S390 will be described below.
[0606] In such Figure 65 In step S390, the audio file acquisition unit 264 selects objects with high processing priority in the display area based on image frame size information, object position information, display area, and the priority of each object. Specifically, the audio file acquisition unit 264 specifies each object in the display area based on image frame size information, object position information, and display area. The audio file acquisition unit 264 selects objects from the specified objects whose priority is equal to or higher than a predetermined value. It should be noted that, for example, the MPD processing unit 261 analyzes the MPD file to obtain the priority from the "Representation" of the object audio of the specified object. The audio selection unit 263 requests the MPD processing unit 261 to send information such as the URL of the audio file of the object audio track of the selected object.
[0607] Figure 66 This is a flowchart illustrating the object audio playback process when the audio selection unit 263 selects only audio objects with predetermined content having high processing priority from the objects in the selection display area.
[0608] like Figure 66 The audio playback process of the object shown is similar to that of... Figure 64 The audio playback process of the object shown, in addition to, Figure 66 The process shown in step S420 is performed to replace, for example... Figure 64 Step S360 is shown. Specifically, as shown... Figure 66 The processes shown in steps S381 to S389 and steps S391 to S399 are similar to those shown in the figure. Figure 64 The processes of steps S411 to S419 and steps S421 to S429 are shown. Therefore, only the process of step S420 will be described below.
[0609] In such Figure 66In step S420 shown, the audio file acquisition unit 264 selects an audio object having predetermined content with a high processing priority in the display area based on the image frame size information, the object position information, the display area, the priority of each object, and the content category of each object. Specifically, the audio file acquisition unit 264 specifies each object in the display area based on the image frame size information, the object position information, and the display area. The audio file acquisition unit 264 selects an object having a priority equal to or higher than a predetermined value and having a content category indicated by the predetermined value from among the specified objects.
[0610] Note that, for example, the MPD processing unit 261 analyzes the MPD file to thereby acquire the priority and the content category from the "Representation" of the object audio of the specified object. The audio selection unit 263 requests the MPD processing unit 261 to transmit information such as the URL of the audio file of the object audio track of the selected object.
[0611] Figure 67 A diagram showing an example of selection of objects based on priority.
[0612] In Figure 67 the example, the objects #1 to #4 are objects in the display area, and an object having a priority equal to or lower than 2 is selected from among the objects in the display area. It is assumed that the smaller the value, the higher the processing priority. Further, in Figure 67 the example, the value in the circle indicates the value of the priority of the corresponding object.
[0613] In the example as shown in Figure 67 , when the priorities of the objects #1 to #4 are 1, 2, 3, and 4, respectively, the objects #1 and #2 are selected. Further, when the priorities of the objects #1 to #4 are changed to 3, 2, 1, and 4, respectively, the objects #2 and #3 are selected. Further, when the priorities of the objects #1 to #4 are changed to 3, 4, 1, and 2, respectively, the objects #3 and #4 are selected.
[0614] As described above, only the audio stream of the object audio of the object having a high processing priority is selectively acquired from among the objects in the display area, and the frequency band between the Web server 142 (212) and the video playback terminal 144 (214) is efficiently utilized. The same applies to the case where the objects are selected based on the content category of the objects.
[0615] <5th Embodiment>
[0616] (Overview of Track)
[0617] Figure 68A schematic diagram showing an overview of the tracks in the fifth embodiment of the present disclosure.
[0618] As Figure 68 shown, the fifth embodiment differs from the second embodiment in that the tracks are recorded as different files (3da_base.mp4 / 3da_channel.mp4 / 3da_object_l.mp4 / 3da_hoa.mp4 / 3da_meta.mp4).
[0619] The information processing system according to the fifth embodiment and the process performed by the information processing system are similar to the fourth embodiment, and thus the description thereof is omitted.
[0620] <Sixth Embodiment>
[0621] (Overview of Tracks)
[0622] Figure 69 A schematic diagram showing an overview of the tracks in the sixth embodiment of the present disclosure.
[0623] As Figure 69 shown, the sixth embodiment differs from the third embodiment in that the tracks are recorded as different files (3da_base_meta.mp4 / 3da_channel.mp4 / 3da_object_l.mp4 / 3da_hoa.mp4).
[0624] The information processing system according to the sixth embodiment and the process performed by the information processing system are similar to the fourth embodiment, except that the audio stream of the base track instead of the object metadata track is acquired in order to acquire the object position information. Thus, the description thereof is omitted.
[0625] Note that in the first to third embodiments, the fifth embodiment, and the sixth embodiment, the object in the display region can also be selected based on the priority or the content category of the object.
[0626] Further, in the first to sixth embodiments, the stream playback unit can acquire the audio stream of the object outside the display region and synthesize the object audio of the object and output the object audio, as with the stream playback unit 120 shown in Figure 23
[0627] Further, in the first to sixth embodiments, the object position information is acquired from the metadata, but instead, the object position information can be acquired from the MPD file.
[0628] <Explanation of the Hierarchical Structure of 3D Audio>
[0629] Figure 70 A schematic diagram showing the hierarchical structure of 3D audio.
[0630] As Figure 70 indicated, an audio element (element) different for each audio data is used as the audio data of the 3D audio. As the type of the audio element, there are a single channel element (SCE) and a channel pair element (CPE). The type of the audio element of the audio data for one channel is SCE, and the type of the audio element of the audio data corresponding to two channels is CPE.
[0631] Audio elements of the same audio type (channel / object / SAOC object / HOA) form a group. Examples of the group type (GroupType) include channel, object, SAOC object, and HOA. When necessary, two or more groups can form a switch group or a group preset.
[0632] The switch group defines a group of audio elements to be played individually. Specifically, as Figure 70 indicated, when there are an object audio group for English (EN) and an object audio group for French (FR), one of the groups is to be played. Thus, the switch group is formed of the object audio group for English with a group ID of 2 and the object audio group for French with a group ID of 3. Thus, the object audio for English and the object audio for French are played individually.
[0633] On the other hand, the group preset defines a combination of groups predetermined by a content producer.
[0634] Ext elements (Ext Elements) different for each metadata are used as the metadata of the 3D audio. Examples of the type of the Ext element include object metadata, SAOC 3D metadata, HOA metadata, DRC metadata, SpatialFrame, and SaocFrame. The Ext element of the object metadata is all metadata of the object audio, and the Ext element of the SAOC 3D metadata is all metadata of the SAOC audio. Further, the Ext element of the HOA metadata is all metadata of the HOA audio, and the Ext element of the dynamic range control (DRC) metadata is all metadata of the object audio, the SAOC audio, and the HOA audio.
[0635] As described above, the audio data of the 3D audio is divided in units of the audio element, the group type, the group, the switch group, and the group preset. Thus, the audio data can be divided into the audio element, the group, the switch group, or the group preset, instead of being divided into the track for each group type as described in the first to sixth embodiments (however, in this case, the object audio is divided for each object).
[0636] Further, the metadata of the 3D audio is divided in units of Ext element types (ExtElementType) or in units of audio elements corresponding to the metadata. Thus, the metadata can be divided for each audio element corresponding to the metadata, instead of being divided for each type of Ext element as described in the first to sixth embodiments.
[0637] It is assumed in the following description that the audio data is divided for each audio element; the metadata is divided for each type of Ext element; and the data of different tracks is arranged. The same applies also when using other units of division.
[0638] <Explanation of the first example of the Web server process>
[0639] Figure 71 A schematic diagram showing the first example of the process of the Web server 142 (212).
[0640] In the example of Figure 71 , the 3D audio corresponding to the audio file uploaded from the file generation apparatus 141 (211) is composed of the channel audio of five channels, the object audio of three objects, and the metadata of the object audio (object metadata).
[0641] The channel audio of the five channels is divided into the channel audio of the front center (FC) channel, the channel audio of the front left / front right (FL, FR) channels, and the channel audio of the rear left / rear right (RL, RR) channels, which are arranged as data of different tracks. Further, the object audio of each object is arranged as data of different tracks. Further, the object metadata is arranged as data of one track.
[0642] Further, as shown in Figure 71 , each audio stream of the 3D audio is composed of configuration information and data in units of frames (samples). In the example of Figure 71 , in the audio stream of the audio file, the configuration information of the channel audio of the five channels, the object audio of the three objects, and the object metadata is arranged collectively, and the data items of each frame are arranged collectively.
[0643] In this case, as shown in Figure 71 , the Web server 142 (212) divides the audio stream of the audio file uploaded from the file generation apparatus 141 (211) for each track and generates the audio stream of seven tracks. Specifically, the Web server 142 (212) extracts the configuration information and the audio data of each track from the audio stream of the audio file according to information such as the ssix box, and generates the audio stream of each track. The audio stream of each track is composed of the configuration information of the track and the audio data of each frame.
[0644] Figure 72 A flowchart showing the track division process of the Web server 142 (212) is shown in FIG. 44. This track division process is started when an audio file is uploaded from the file generation apparatus 141 (211), for example.
[0645] In step S441, the Web server 142 (212) stores the audio file uploaded from the file generation apparatus 141. Figure 72
[0646] In step S442, the Web server 142 (212) divides the audio stream constituting the audio file for each track based on information such as the ssix box of the audio file.
[0647] In step S443, the Web server 142 (212) holds the audio stream of each track. Then the process is terminated. The audio stream is transferred from the Web server 142 (212) to the video playback terminal 144 (214) when the audio stream is requested from the audio file acquisition unit 192 (264) of the video playback terminal 144 (214).
[0648] <Explanation of the first example of the process of the audio decoding processing unit>
[0649] Figure 73 A schematic diagram showing the first example of the process of the audio decoding processing unit 194 when the above-described process with reference to Figure 71 and 72 is executed in the Web server 142 (212) is shown in FIG. 45.
[0650] In the example of Figure 73 , the Web server 142 (212) holds the audio stream of each track as shown in Figure 71 . The tracks to be played are the track of the channel audio of the front left / front right channels, the track of the channel audio of the rear left / rear right channels, the track of the object audio of the first object, and the track of the object metadata. The same is true for the case of FIG. 75 described later.
[0651] In this case, the audio file acquisition unit 192 (264) acquires the track of the channel audio of the front left / front right channels, the track of the channel audio of the rear left / right channels, the track of the object audio of the first object, and the track of the object metadata.
[0652] The audio decoding processing unit 194 first extracts the audio stream of the metadata of the object audio of the first object from the audio stream of the track of the object metadata acquired by the audio file acquisition unit 192 (264).
[0653] Next, as shown in Figure 73 As shown, the audio decoding processing unit 194 synthesizes the audio streams of the audio tracks to be played and the audio streams of the extracted metadata. Specifically, the audio decoding processing unit 194 generates an audio stream in which the configuration information items contained in all the audio streams are arranged collectively, and the data items of each frame are arranged collectively. Further, the audio decoding processing unit 194 decodes the generated audio stream.
[0654] As described above, when the audio streams to be played include the audio streams of two or more tracks in addition to the audio stream of one channel audio track, the audio streams of the two or more tracks are to be played. Therefore, the audio streams are synthesized before decoding.
[0655] On the other hand, when only the audio stream of one channel audio track is to be played, the audio stream does not need to be synthesized. Therefore, the audio decoding processing unit 194 directly decodes the audio stream acquired by the audio file acquisition unit 192 (264).
[0656] Figure 74 To show details of a first example of the decoding process of the audio decoding processing unit 194 when the above-described procedure is executed by the Web server 142 (212), a flowchart of the decoding process is shown in FIG. 27. The decoding process is a process in which at least one of the step S229 shown in FIG. 25 and the step S287 shown in FIG. 26 is executed when the tracks to be played include tracks other than one channel audio tracks. Figure 71 and 72 Figure 48 Figure 50
[0657] In the step S461, the audio decoding processing unit 194 sets all the element numbers indicating the number of elements contained in the generated audio stream to "0". In the step S462, the audio decoding processing unit 194 resets (clears) all the element type information indicating the types of elements contained in the generated audio stream. Figure 74 In the step S463, the audio decoding processing unit 194 sets a track among the tracks to be played which is not determined as a track to be processed to the track to be processed. In the step S464, the audio decoding processing unit 194 acquires the number and the type of elements contained in the track to be processed from the audio stream of the track to be processed, for example.
[0658] In the step S465, the audio decoding processing unit 194 adds the acquired number of elements to the total number of elements. In the step S466, the audio decoding processing unit 194 adds the acquired type of elements to all the element type information.
[0659]
[0660] In step S467, the audio decoding processing unit 194 determines whether all the tracks to be played are set as the tracks to be processed. When it is determined in step S467 that not all the tracks to be played are set as the tracks to be processed, the process returns to step S463, and the process of steps S463 to S467 is repeated until all the tracks to be played are set as the tracks to be processed.
[0661] On the other hand, when it is determined in step S467 that all the tracks to be played are set as the tracks to be processed, the process proceeds to step S468. In step S468, the audio decoding processing unit 194 arranges the total number of elements and all the element type information at a predetermined position on the generated audio stream.
[0662] In step S469, the audio decoding processing unit 194 sets the track, which is not determined as the track to be processed, among the tracks to be played, as the track to be processed. In step S470, the audio decoding processing unit 194 sets the element, which is not determined as the element to be processed, among the elements included in the track to be processed, as the element to be processed, when the element is to be processed.
[0663] In step S471, the audio decoding processing unit 194 acquires the configuration information of the element to be processed from the audio stream of the track to be processed and arranges the configuration information on the generated audio stream. At this time, the configuration information items of all the elements of all the tracks to be played are arranged continuously.
[0664] In step S472, the audio decoding processing unit 194 determines whether all the elements included in the track to be processed are set as the elements to be processed. When it is determined in step S472 that not all the elements are set as the elements to be processed, the process returns to step S470, and the process of steps S470 to S472 is repeated until all the elements are set as the elements to be processed.
[0665] On the other hand, when it is determined in step S472 that all the elements are set as the elements to be processed, the process proceeds to step S473. In step S473, the audio decoding processing unit 194 determines whether all the tracks to be played are set as the tracks to be processed. When it is determined in step S473 that not all the tracks to be played are set as the tracks to be processed, the process returns to step S469, and the process of steps S469 to S473 is repeated until all the tracks to be played are set as the tracks to be processed.
[0666] On the other hand, when it is determined in step S473 that all the to-be-played tracks are set as the to-be-processed tracks, the process proceeds to step S474. In step S474, the audio decoding processing unit 194 determines the to-be-processed frame. In the process of step S474 at the first time, the head frame is determined as the to-be-processed frame. In the process of step S474 at the second and subsequent times, the frame immediately following the current frame to-be-processed is determined as the new frame to-be-processed.
[0667] In step S475, the audio decoding processing unit 194 sets the track that is not determined as the to-be-processed track among the to-be-played tracks as the to-be-processed track. In step S476, the audio decoding processing unit 194 sets the element that is not determined as the to-be-processed element among the elements included in the to-be-processed track as the to-be-processed element.
[0668] In step S477, the audio decoding processing unit 194 determines whether the to-be-processed element is an EXT element. When it is determined in step S477 that the to-be-processed element is not an EXT element, the process proceeds to step S478.
[0669] In step S478, the audio decoding processing unit 194 acquires the audio data of the to-be-processed frame of the to-be-processed element from the audio stream of the to-be-processed track and arranges the audio data on the generated audio stream. At this time, the data in the same frame of all the elements of all the to-be-played tracks are arranged continuously. After the process of step S478, the process proceeds to step S481.
[0670] On the other hand, when it is determined in step S477 that the to-be-processed element is an EXT element, the process proceeds to step S479. In step S479, the audio decoding processing unit 194 acquires the metadata of all the objects in the to-be-processed frame of the to-be-processed element from the audio stream of the to-be-processed track.
[0671] In step S480, the audio decoding processing unit 194 arranges the metadata of the to-be-played object among the acquired metadata of all the objects on the generated audio stream. At this time, the data items in the same frame of all the elements of all the to-be-played tracks are arranged continuously. After the process of step S480, the process proceeds to step S481.
[0672] In step S481, the audio decoding processing unit 194 determines whether all the elements included in the to-be-processed track are set as the to-be-processed elements. When it is determined in step S481 that not all the elements are set as the to-be-processed elements, the process returns to step S476, and the processes of steps S476 to S481 are repeated until all the elements are set as the to-be-processed elements.
[0673] On the other hand, when it is determined in step S481 that all the elements are set as the elements to be processed, the process proceeds to step S482. In step S482, the audio decoding processing unit 194 determines whether all the tracks to be played are set as the tracks to be processed. When it is determined in step S482 that not all the tracks to be played are set as the tracks to be processed, the process returns to step S475, and the process of steps S475 to S482 is repeated until all the tracks to be played are set as the tracks to be processed.
[0674] On the other hand, when it is determined in step S482 that all the tracks to be played are set as the tracks to be processed, the process proceeds to step S483.
[0675] In step S483, the audio decoding processing unit 194 determines whether all the frames are set as the frames to be processed. When it is determined in step S483 that not all the frames are set as the frames to be processed, the process returns to step S474, and the process of steps S474 to S483 is repeated until all the frames are set as the frames to be processed.
[0676] On the other hand, when it is determined in step S483 that all the frames are set as the frames to be processed, the process proceeds to step S484. In step S484, the audio decoding processing unit 194 decodes the generated audio stream. Specifically, the audio decoding processing unit 194 decodes the audio stream in which the total number of elements, all the element type information, the configuration information, the audio data, and the metadata of the object to be played are arranged. The audio decoding processing unit 194 provides the audio data (object audio, channel audio, HOA audio) obtained as a result of the decoding to the audio synthesis processing unit 195. Then the process is terminated.
[0677] <Explanation of the second example of the process of the audio decoding processing unit>
[0678] Figure 75 To show the second example of the process of the audio decoding processing unit 194 when the above-described process with reference to Figure 71 and 72 is executed by the Web server 142 (212), a schematic diagram of the second example of the process of the audio decoding processing unit 194.
[0679] As shown in Figure 75 , the second example of the process of the audio decoding processing unit 194 differs from the first example in that the audio streams of all the tracks are arranged on the generated audio stream and the indication of the zero decoding result stream or the flag (hereinafter, referred to as the zero stream) is arranged as the audio stream of the track not to be played.
[0680] Specifically, the audio file acquisition unit 192 (264) acquires the configuration information included in the audio stream of all the tracks held in the Web server 142 (212) and the data of each frame included in the audio stream of the track to be played.
[0681] As shown in Figure 75 , the audio decoding processing unit 194 arranges the configuration information items of all the tracks on the generated audio stream. Further, the audio decoding processing unit 194 arranges the data of each frame of the track to be played and the zero stream as the data of each frame of the track not to be played on the generated audio stream.
[0682] As described above, since the audio decoding processing unit 194 arranges the zero stream as the audio stream of the track not to be played on the generated audio stream, there is also the audio stream of the object not to be played. Therefore, the metadata of the object not to be played can be included in the generated audio stream. This eliminates the need for the audio decoding processing unit 194 to extract the audio stream of the metadata of the object to be played from the audio stream of the track of the object metadata.
[0683] Note that the zero stream can be arranged as the configuration information of the track not to be played.
[0684] Figure 76 To show the details of the second example of the decoding process of the audio decoding processing unit 194 when the Web server 142 (212) performs the above-described procedure with reference to Figure 71 and 72 , a flowchart of the procedure is shown. The decoding process is at least one of the procedure of step S229 shown in Figure 48 and the procedure of step S287 shown in Figure 50 , when the track to be played includes a track other than the one-channel audio track.
[0685] The procedures of steps S501 and S502 shown in Figure 76 are similar to the procedures of steps S461 to S462 shown in Figure 74 , and thus the description thereof is omitted.
[0686] In step S503, the audio decoding processing unit 194 sets the track corresponding to the track not determined as the track to be processed among the tracks of the audio stream held in the Web server 142 (212) as the track to be processed.
[0687] The procedures of steps S504 to S506 are similar to the procedures of steps S464 to S466, and thus the description thereof will be omitted.
[0688] In step S507, the audio decoding processing unit 194 determines whether all the tracks corresponding to the audio stream held in the Web server 142 (212) are set as the tracks to be processed. When it is determined in step S507 that not all the tracks are set as the tracks to be processed, the process returns to step S503, and the process of steps S503 to S507 is repeated until all the tracks are set as the tracks to be processed.
[0689] On the other hand, when it is determined in step S507 that all the tracks are set as the tracks to be processed, the process proceeds to step S508. In step S508, the audio decoding processing unit 194 arranges the total number of elements and all the element type information at a predetermined position on the generated audio stream.
[0690] In step S509, the audio decoding processing unit 194 sets the tracks corresponding to the tracks of the audio stream held in the Web server 142 (212) which are not determined as the tracks to be processed, as the tracks to be processed. In step S510, the audio decoding processing unit 194 sets the elements included in the tracks to be processed which are not determined as the elements to be processed, as the elements to be processed.
[0691] In step S511, the audio decoding processing unit 194 acquires the configuration information of the elements to be processed from the audio stream of the tracks to be processed and generates the configuration information on the generated audio stream. At this time, the configuration information items of all the elements corresponding to all the tracks of the audio stream held in the Web server 142 (212) are arranged continuously.
[0692] In step S512, the audio decoding processing unit 194 determines whether all the elements included in the tracks to be processed are set as the elements to be processed. When it is determined in step S512 that not all the elements are set as the elements to be processed, the process returns to step S510, and the process of steps S510 to S512 is repeated until all the elements are set as the elements to be processed.
[0693] On the other hand, when it is determined in step S512 that all the elements are set as the elements to be processed, the process proceeds to step S513. In step S513, the audio decoding processing unit 194 determines whether all the tracks corresponding to the audio stream held in the Web server 142 (212) are set as the tracks to be processed. When it is determined in step S513 that not all the tracks are set as the tracks to be processed, the process returns to step S509, and the process of steps S509 to S513 is repeated until all the tracks are set as the tracks to be processed.
[0694] On the other hand, when it is determined in step S513 that all the tracks are set as the tracks to be processed, the process proceeds to step S514. In step S514, the audio decoding processing unit 194 determines the frame to be processed. In the process of step S514 at the first time, the head frame is determined as the frame to be processed. In the process of step S514 at the second and subsequent times, the frame immediately following the current frame to be processed is determined as the new frame to be processed.
[0695] In step S515, the audio decoding processing unit 194 sets the track, which is not determined as the track to be processed among the tracks corresponding to the audio streams held in the Web server 142 (212), as the track to be processed.
[0696] In step S516, the audio decoding processing unit 194 determines whether the track to be processed is the track to be played. When it is determined in step S516 that the track to be processed is the track to be played, the process proceeds to step S517.
[0697] In step S517, the audio decoding processing unit 194 sets the element, which is not determined as the element to be processed among the elements included in the track to be processed, as the element to be processed.
[0698] In step S518, the audio decoding processing unit 194 acquires the audio data of the frame to be processed of the element to be processed from the audio stream of the track to be processed and arranges the audio stream on the generated audio stream. At this time, the data items in the same frame of all the elements of all the tracks corresponding to the audio streams held in the Web server 142 (212) are arranged continuously.
[0699] In step S519, the audio decoding processing unit 194 determines whether all the elements included in the track to be processed are set as the elements to be processed. When it is determined in step S519 that not all the elements are set as the elements to be processed, the process returns to step S517, and the process of steps S517 to S519 is repeated until all the elements are set as the elements to be processed.
[0700] On the other hand, when it is determined in step S519 that all the elements are set as the elements to be processed, the process proceeds to step S523.
[0701] Further, when it is determined in step S516 that the track to be processed is not the track to be played, the process proceeds to step S520. In step S520, the audio decoding processing unit 194 sets the element, which is not determined as the element to be processed among the elements included in the track to be processed, as the element to be processed.
[0702] In step S521, the audio decoding processing unit 194 arranges a zero stream of data of the frame to be processed as an element to be processed on the generated audio stream. At this time, data items in the same frame of all elements corresponding to all tracks of the audio stream held in the Web server 142 (212) are arranged continuously.
[0703] In step S522, the audio decoding processing unit 194 determines whether all elements included in the track to be processed are set as elements to be processed. When it is determined in step S522 that not all elements are set as elements to be processed, the process returns to step S520, and the process of steps S520 to S522 is repeated until all elements are set as elements to be processed.
[0704] On the other hand, when it is determined in step S522 that all elements are set as elements to be processed, the process proceeds to step S523.
[0705] In step S523, the audio decoding processing unit 194 determines whether all tracks corresponding to the audio stream held in the Web server 142 (212) are set as tracks to be processed. When it is determined in step S522 that not all tracks are set as tracks to be processed, the process returns to step S515, and the process of steps S515 to S523 is repeated until all tracks to be played are set as tracks to be processed.
[0706] On the other hand, when it is determined in step S523 that all tracks are set as tracks to be processed, the process proceeds to step S524.
[0707] In step S524, the audio decoding processing unit 194 determines whether all frames are set as frames to be processed. When it is determined in step S524 that not all frames are set as frames to be processed, the process returns to step S514, and the process of steps S514 to S524 is repeated until all frames are set as frames to be processed.
[0708] On the other hand, when it is determined in step S524 that all frames are set as frames to be processed, the process proceeds to step S525. In step S525, the audio decoding processing unit 194 decodes the generated audio stream. Specifically, the audio decoding processing unit 194 decodes the audio stream in which the total number of elements, all element type information and configuration information, and data corresponding to all tracks of the audio stream held in the Web server 142 (212) are arranged. The audio decoding processing unit 194 provides audio data (object audio, channel audio, HOA audio) obtained as a result of the decoding to the audio synthesis processing unit 195. Then the process is terminated.
[0709] <Explanation of the second example of the Web server process>
[0710] Figure 77 A schematic diagram showing a second example of the process of the Web server 142 (212) is shown.
[0711] As Figure 77 shown, the second example of the process of the Web server 142 (212) is the same as the first example shown in Figure 71 , except that the object metadata of each object is arranged as data of a different track in the audio file.
[0712] Therefore, as Figure 77 shown, the Web server 142 (212) divides the audio stream of the audio file uploaded from the file generation apparatus 141 (211) for each track and generates audio streams of nine tracks.
[0713] In this case, the track division process of the Web server 142 (212) is similar to the track division process shown in Figure 72 , and thus the description thereof is omitted.
[0714] <Explanation of the third example of the audio decoding processing unit>
[0715] Figure 78 A schematic diagram showing the process of the audio decoding processing unit 194 when the process described above with reference to Figure 77 is executed by the Web server 142 (212) is shown.
[0716] In Figure 78 the example shown in Figure 77 , the Web server 142 (212) holds the audio stream of each track as shown. The tracks to be played are the channel audio of the front left / front right channels, the channel audio of the rear left / rear right channels, the object audio of the first object, and the object metadata of the first object.
[0717] In this case, the audio file acquisition unit 192 (264) acquires the audio stream of the tracks of the channel audio of the front left / front right channels, the channel audio of the rear left / rear right channels, the object audio of the first object, and the object metadata of the first object. The audio decoding processing unit 194 synthesizes the acquired audio streams of the tracks to be played and decodes the generated audio stream.
[0718] As described above, when the object metadata is arranged as data of a different track for each object, the audio decoding processing unit 194 does not need to extract the audio stream of the object metadata of the object to be played. Therefore, the audio decoding processing unit 194 can easily generate the audio stream to be decoded.
[0719] Figure 79 A flowchart showing details of the decoding process of the audio decoding processing unit 194 when the above-described process is executed by the Web server 142 (212) is shown in FIG. 27. The decoding process is one of the process of Step S229 shown in FIG. 22 and the process of Step S287 shown in FIG. 28, when the to-be-played track contains tracks other than one channel audio track. Figure 77 Figure 48 Figure 50
[0720] The decoding process shown in FIG. 26 is similar to the decoding process shown in FIG. 22, except that the process in Steps S477, S479, and S480 is not executed and not only the audio data but also the metadata is arranged in the process of Step S478. Specifically, the process of Steps S541 to S556 shown in FIG. 26 is similar to the process of Steps S461 to S476 shown in FIG. 22. In the process of Step S557 shown in FIG. 26, the data of the to-be-processed frame of the to-be-processed element is arranged as in the process of Step S478. Further, the process of Steps S558 to S561 shown in FIG. 26 is similar to the process of Steps S481 to S484 shown in FIG. 22. Figure 79 Figure 74 Figure 79 Figure 74 Figure 79 Figure 74
[0721] It should be noted that in the above-described process, the video playback terminal 144 (214) generates the audio stream to be decoded, but instead, the Web server 142 (212) can generate a combination of the audio streams of the combination of the tracks to be played. In this case, the video playback terminal 144 (214) can play the audio of the to-be-played tracks only by acquiring the audio stream having the combination of the to-be-played tracks from the Web server 142 (212) and decoding the audio stream.
[0722] Further, the audio decoding processing unit 194 can decode the audio stream of the to-be-played track acquired from the Web server 142 (212) for each track. In this case, the audio decoding processing unit 194 needs to synthesize the audio data and the metadata obtained as a result of decoding.
[0723] <Second Example of Syntax of Configuration Information Arranged in Elementary Sample>
[0724] (Second Example of Syntax of Configuration Information Arranged in Elementary Sample)
[0725] Figure 80 A diagram showing the second example of the syntax of the configuration information arranged in the elementary sample.
[0726] In Figure 80 In the example, the number of elements (numElements) arranged in the base sample is described as configuration information. Furthermore, the type (usacElementType) of each element arranged in the base sample, representing the "ID_USAC_EXT" of the Ext element, is described, as well as the configuration information (mpegh3daExtElementCongfig) for each Ext element is also described.
[0727] Figure 81 To show the use of in Figure 80 The diagram illustrates an exemplary syntax for the configuration information of an Ext element (mpegh3daExtElementCongfig).
[0728] like Figure 81 As shown, "ID_EXT_ELE_EXTRACTOR", representing the extractor of the type of Ext element, is described as being used for... Figure 80 The configuration information for the Ext element is shown (mpegh3daExtElementConfig). Additionally, the configuration information for the extractor is described (ExtractorConfig).
[0729] Figure 82 To show the use of in Figure 81 The diagram illustrates an example syntax for the extractor's configuration information (ExtractorConfig).
[0730] like Figure 82 As shown, as used for Figure 81 The extractor configuration information (ExtractorConfig) shown describes the type of the element to be referenced by this extractor (usac Element Type Extractor). Furthermore, the Ext element type (usacExtElementTypeExtractor) is described when the element type (usac Element Type Extractor) is "ID_USAC_EXT" representing an Ext element. Additionally, the size (configLength) and position (configOffset) of the configuration information for the element (subsample) to be referenced are described.
[0731] (A second example of the data syntax for frame units arranged in the basic sample)
[0732] Figure 83 This is a schematic diagram illustrating a second example of the data syntax arranged in frame units within a basic sample.
[0733] likeFigure 83 As shown, as data arranged in a frame unit in the basic sample, "ID_EXT_ELE_EXTRACTOR" indicating an extractor of a type of an Ext element, which is a data element, is described, and extractor data is also described.
[0734] Figure 84 A schematic diagram showing an example syntax of the extractor data (Extractor Metadata) shown in Figure 83
[0735] As shown, the size (elementLength) and the position (elementOffset) of data of an element to be referenced by the extractor are described as the extractor data (Extractor Metadata) as shown in Figure 84 Figure 83
[0736] <Third example of syntax of a basic sample>
[0737] (Third example of syntax of configuration information arranged in a basic sample)
[0738] Figure 85 A schematic diagram showing a third example of syntax of configuration information arranged in a basic sample.
[0739] In the example of Figure 85 , the number of elements (numElements) arranged in the basic sample is described as configuration information. Further, "1" indicating an extractor is described as a flag Extractor indicating whether the sample in which the configuration information is arranged is an extractor. Further, "1" is described as elementLengthPresent.
[0740] Further, the type of an element to be referenced by the element is described as the type (usacElementType) of each element arranged in the basic sample. When the type of the element (usacElementType) is "ID_USAC_EXT" indicating an Ext element, the type of the Ext element (usacExtElementType) is described. Further, the size (configLength) and the position (configOffset) of the configuration information of the element to be referenced are described.
[0741] (Third example of data syntax in a frame unit arranged in a basic sample)
[0742] Figure 86 A diagram for showing a third example of data syntax arranged in a frame unit in a basic sample.
[0743] As Figure 86 shown, as data arranged in a frame unit in a basic sample, the size (elementLength) and position (elementOffset) of data of an element to be referenced by data are described.
[0744] <Seventh Embodiment>
[0745] (Configuration Example of Audio Stream)
[0746] Figure 87 A diagram for showing a configuration example of an audio stream stored in an audio file in the seventh embodiment of the information processing system to which the present disclosure is applied.
[0747] As Figure 87 shown, in the seventh embodiment, the audio file stores, in units of samples of 3D audio for each group type, encoded data (however, in this case, object audio is stored for each object) and an audio stream (3D audio stream) arranged as a sub-sample.
[0748] Further, the audio file stores a cue stream (3D audio cue stream) in which an extractor containing the size, position, and group type of encoded data in units of samples of 3D audio for each group type is set as a sub-sample. The configuration of this extractor is similar to the above-described configuration, and the group type is described as the type of the extractor.
[0749] (Outline of Track)
[0750] Figure 88 A diagram for showing an outline of a track in the seventh embodiment.
[0751] As Figure 88 shown, in the seventh embodiment, different tracks are respectively assigned to the audio stream and the cue stream. The track ID "2" of the track corresponding to the cue stream is described as the track reference number of the track of the audio stream. Further, the track ID "1" of the track of the corresponding audio stream is described as the track reference number of the track of the cue stream.
[0752] The syntax of the sample entry of the track of the audio stream is the syntax as Figure 34 shown, and the syntax of the sample entry of the track of the cue stream contains the syntax as Figures 35 to 38 shown.
[0753] (Explanation of Process of File Generation Apparatus)
[0754] Figure 89Fig. 8 is a flowchart showing a file generation process of the file generation apparatus in the seventh embodiment.
[0755] It should be noted that the file generation apparatus according to the seventh embodiment is the same as the file generation apparatus 141 as shown in Fig. 1 except for the processes of the audio encoding processing unit 171 and the audio file generation unit 172. Therefore, the file generation apparatus, the audio encoding processing unit, and the audio file generation unit according to the seventh embodiment are hereinafter referred to as a file generation apparatus 301, an audio encoding processing unit 341, and an audio file generation unit 342, respectively. Figure 45
[0756] The processes of steps S601 to S605 as shown in Fig. 8 are similar to the processes of steps S191 to S195 as shown in Fig. 2, and therefore the description thereof is omitted. Figure 89 Figure 46
[0757] In step S606, the audio encoding processing unit 341 encodes the 3D audio of the video content inputted from the outside for each group type and generates an audio stream as shown in Fig. 87. The audio encoding processing unit 341 provides the generated audio stream to the audio file generation unit 342. Figure 87
[0758] In step S607, the audio file generation unit 342 acquires sub-sample information from the audio stream provided from the audio encoding processing unit 341. The sub-sample information indicates the size, position, and group type of the encoded data in units of samples of the 3D audio of each group type.
[0759] In step S608, the audio file generation unit 342 generates a cue stream as shown in Fig. 87 based on the sub-sample information. In step S609, the audio file generation unit 342 multiplexes the audio stream and the cue stream into different tracks and generates an audio file. At this time, the audio file generation unit 342 stores the image frame size information inputted from the outside in a sample entry. The audio file generation unit 342 provides the generated audio file to the MPD generation unit 173.
[0760] The processes of steps S610 and S611 as shown in Fig. 8 are similar to the processes of steps S199 and S200 as shown in Fig. 2, and therefore the description thereof is omitted. Figure 46
[0761] (Explanation of the process of the video playback terminal)
[0762] Figure 90 Fig. 9 is a flowchart showing an audio playback process of the stream playback unit of the video playback terminal in the seventh embodiment.
[0763] It should be noted that the stream playback unit according to the seventh embodiment is the same as the stream playback unit 151 as shown in Fig. 1 except for the processes of the audio decoding processing unit 351 and the audio file generation unit 352. Therefore, the stream playback unit, the audio decoding processing unit, and the audio file generation unit according to the seventh embodiment are hereinafter referred to as a stream playback unit 401, an audio decoding processing unit 451, and an audio file generation unit 452, respectively.Figure 47 The stream playback unit 190 shown is the same except that the processes of the MPD processing unit 191, the audio file acquisition unit 192, and the audio decoding processing unit 194 are different and the audio selection unit 193 is not provided. Therefore, the stream playback unit, the MPD processing unit, the audio file acquisition unit, and the audio decoding processing unit according to the seventh embodiment are hereinafter referred to as a stream playback unit 360, an MPD processing unit 381, an audio file acquisition unit 382, and an audio decoding processing unit 383, respectively.
[0764] In the case where the 3D audio to be played back is object audio, the object to be played back is selected based on the image frame size information and the object position information. Figure 90 In step S621, the MPD processing unit 381 of the stream playback unit 360 analyzes the MPD file provided from the MPD acquisition unit 91, acquires information such as the URL of the audio file of the segment to be played back, and provides the acquired information to the audio file acquisition unit 382.
[0765] In step S622, the audio file acquisition unit 382 requests the Web server to send the initial segment of the segment to be played back and acquires the initial segment based on information such as the URL provided from the MPD processing unit 381.
[0766] In step S623, the audio file acquisition unit 382 acquires the track ID of the track of the audio stream that is the reference track from the sample entry of the track of the hint stream of the moov box in the initial segment (hereinafter referred to as the hint track).
[0767] In step S624, the audio file acquisition unit 382 requests the Web server to send the sidx box and the ssix box from the header of the media segment of the segment to be played back and acquires the sidx box and the ssix box based on information such as the URL provided from the MPD processing unit 381.
[0768] In step S625, the audio file acquisition unit 382 acquires the position information of the hint track from the sidx box and the ssix box acquired in step S624.
[0769] In step S626, the audio file acquisition unit 382 requests the Web server to send the hint stream and acquires the hint stream based on the position information of the hint track acquired in step S625. Further, the audio file acquisition unit 382 acquires the extractor of the group type of the 3D audio to be played back from the hint stream. Note that in the case where the 3D audio to be played back is object audio, the object to be played back is selected based on the image frame size information and the object position information.
[0770] In step S627, the audio file acquisition unit 382 acquires the position information of the reference track from the sidx box and the ssix box acquired in step S624. In step S628, the audio file acquisition unit 382 determines the position information of the audio stream of the group type of the 3D audio to be played based on the position information of the reference track acquired in step S627 and the subsample information contained in the extractor acquired.
[0771] In step S629, the audio file acquisition unit 382 requests the web server to send the audio stream of the group type of the 3D audio to be played based on the position information determined in step S627 and acquires the audio stream. The audio file acquisition unit 382 provides the audio stream acquired to the audio decoding processing unit 383.
[0772] In step S630, the audio decoding processing unit 383 decodes the audio stream provided from the audio file acquisition unit 382 and provides the audio data obtained as a result of the decoding to the audio synthesis processing unit 195.
[0773] In step S631, the audio synthesis processing unit 195 outputs the audio data. Then the process is terminated.
[0774] Note that in the seventh embodiment, the track of the audio stream and the cue track are stored in the same audio file, but can be stored in different files.
[0775] <Embodiment 8>
[0776] (Outline of Track)
[0777] Figure 91 A schematic diagram showing an outline of a track in the eighth embodiment of the information processing system to which the present disclosure is applied.
[0778] The audio file of the eighth embodiment differs from the audio file of the seventh embodiment in that the cue stream stored is a stream for each group type. Specifically, the cue stream of the eighth embodiment is generated for each group type, and an extractor containing the size, position, and group type of the encoded data in units of samples of the 3D audio of each group type is arranged in each cue stream. Note that when the 3D audio contains object audio of a plurality of objects, the extractor is arranged as a subsample for each object.
[0779] Further, as Figure 91 shown, in the eighth embodiment, a different track is assigned to the audio stream and each cue stream. The track of this audio stream is the same as the track of the audio stream as Figure 88 shown, and thus the description thereof is omitted.
[0780] As the track reference number of the cue track of each of the group types "channel", "object", "HOA", and "metadata", the track ID "1" of the track of the corresponding audio stream is described.
[0781] The syntax of the sample entry of the cue track of each of the group types "channel", "object", "HOA", and "metadata" is the same as the syntax shown in Figures 35 to 38 except for the information indicating the type of the sample entry. The information indicating the type of the sample entry of the cue track of each of the group types "channel", "object", "HOA", and "metadata" is the same as the information shown in Figures 35 to 38 except that the number "1" of the information is replaced with "2". The number "2" indicates the sample entry of the cue track.
[0782] (Configuration example of audio file)
[0783] Figure 92 A schematic diagram showing a configuration example of an audio file.
[0784] As shown in Figure 92 , the audio file stores all the tracks shown in Figure 91 . Specifically, the audio file stores the audio stream and the cue stream of each group type.
[0785] The file generation process of the file generation apparatus according to the eighth embodiment is similar to the file generation process shown in Figure 89 except that, contrary to the cue stream shown in Figure 87 , the cue stream is generated for each group type.
[0786] Further, the audio playback process of the stream playback unit of the video playback terminal according to the eighth embodiment is similar to the audio playback process shown in Figure 90 except that the track ID of the cue track of the group type to be played and the track ID of the reference track acquired in step S623; the position information of the cue track of the group type to be played is acquired in step S625; and the cue stream of the group type to be played is acquired in step S626.
[0787] It should be noted that, in the eighth embodiment, the track of the audio stream and the cue track are stored in the same audio file, but can be stored in different files.
[0788] For example, as shown in Figure 93 , the track of the audio stream can be stored in one audio file (3D audio stream MP4 file), and the cue track can be stored in one audio file (3D audio hint stream MP4 file). Further, as shown in Figure 94 , the cue track can be divided into a plurality of audio files to be stored.Figure 94 In the example of the seventh embodiment, the cue track is stored in a different audio file.
[0789] Further, in the eighth embodiment, a cue stream is generated for each group type even when the group type indicates an object. However, when the group type indicates an object, a cue stream can be generated for each object. In this case, a different track is assigned to the cue stream of each object.
[0790] As described above, in the audio files of the seventh and eighth embodiments, the audio streams of 3D audio are stored in one track. Therefore, the video playback terminal can play all the audio streams of 3D audio by acquiring the track.
[0791] Further, the cue stream is stored in the audio files of the seventh and eighth embodiments. Therefore, the video playback terminal can acquire the audio stream of the desired group type among all the audio streams of 3D audio without reference to the moof box in which the table associating a subsample with the size or position of the subsample is described, thus making it possible to play the audio stream.
[0792] Further, in the audio files of the seventh and eighth embodiments, the video playback terminal can be caused to acquire the audio stream of each group type only by storing all the audio streams of 3D audio and the cue stream. Therefore, it is not necessary to prepare the audio stream of 3D audio separately for each group type from all the generated audio streams of 3D audio for the purpose of broadcasting or local storage in order to be able to acquire the audio stream for each group type.
[0793] Note that in the seventh and eighth embodiments, the extractor is generated for each group type, but can be generated in units of audio elements, groups, switch groups, or group presets.
[0794] When the extractor is generated in units of groups, the sample entry of each cue track of the eighth embodiment contains information on the corresponding group. The information on the group is composed of, for example, information indicating the ID of the group and the contents of the data classified as the elements of the group. When the group forms a switch group, the sample entry of the cue track of the group also contains information on the switch group. The information on the switch group is composed of, for example, the ID of the switch group and the IDs of the groups forming the switch group. The sample entry of the cue track of the seventh embodiment contains the information contained in the sample entry of all the cue tracks of the eighth embodiment.
[0795] Further, the segment structure in the seventh and eighth embodiments is the same as the segment structure shown in Figure 39 and 40
[0796] <Ninth Embodiment>
[0797] (Explanation of a computer that applies the present disclosure)
[0798] The series of processes of the Web server described above can also be executed by hardware or software. When the series of processes is executed by software, a program that constitutes the software is installed in a computer. Examples of the computer include a computer that incorporates a dedicated hardware and a general personal computer that is able to execute various functions by installing various programs.
[0799] Figure 95 A block diagram of an example of a hardware configuration of a computer that executes the series of processes of the Web server by using a program is shown.
[0800] In the computer, a central processing unit (CPU) 601, a read only memory (ROM) 602, and a random access memory (RAM) 603 are interconnected via a bus 604.
[0801] The bus 604 is also connected to an input / output interface 605. The input / output interface 605 is connected to each of an input unit 606, an output unit 607, a storage unit 608, a communication unit 609, and a drive 610.
[0802] The input unit 606 is formed of a keyboard, a mouse, a microphone, and the like. The output unit 607 is formed of a display, a speaker, and the like. The storage unit 608 is formed of a hardware, a non-volatile memory, and the like. The communication unit 609 is formed of a network interface and the like. The drive 610 drives a removable medium 611 such as a magnetic disk, a magnetic optical disk, an optical disk, or a semiconductor memory.
[0803] In the computer configured as described above, the CPU 601 loads a program stored in the storage unit 608 in the RAM 603, for example, via the input / output interface 605 and the bus 604 and executes the program, thereby executing the series of processes described above.
[0804] The program executed by the computer (CPU 601) can be provided recorded in the removable medium 611 that serves as a package medium or the like. Further, the program can be provided via a wired or wireless transmission medium such as a local area network, the Internet, or a digital satellite broadcast.
[0805] The program can be installed in the storage unit 608 via the input / output interface 605 by loading the removable medium 611 in the drive 610. Further, the program can be received via the wired or wireless transmission medium by the communication unit 609 and installed in the storage unit 608. Further, the program can be installed in the ROM 602 or the storage unit 608 in advance.
[0806] It should be noted that the program executed by the computer can be a program that executes the process in a time series in the order described in the present description, or can be a program that executes the process, for example, in parallel or at a necessary time when solicited.
[0807] The video playback terminal described above can have a hardware configuration similar to that of a computer as shown in FIG. 6. In this case, for example, the CPU 601 can execute the control software 161 (221), the video playback software 162 (222), and the access software 163 (223). The process of the video playback terminal 144 (214) can be executed by hardware. Figure 95
[0808] In the present description, a system has a plurality of components (such as devices or modules (parts)) and it is not considered whether all of the components are in the same housing. Therefore, the system can be a plurality of devices that can be stored in separate housings and connected through a network and a plurality of modules within a single housing.
[0809] It should be noted that the embodiments of the present disclosure are not limited to the above-described embodiments and various changes can be made without departing from the gist of the present disclosure.
[0810] For example, the file generation device 141 (211) can generate a video stream by multiplexing the encoded data of all tiles to generate one image file instead of generating image files in tile units.
[0811] The present disclosure can be applied not only to MPEG-H 3D audio but also to a general audio codec capable of forming a stream of each object.
[0812] Furthermore, the present disclosure can also be applied to an information processing system that performs broadcasting and local storage playback as well as stream playback.
[0813] Furthermore, the present disclosure can have the following configurations. (1)
[0815] An information processing device including an acquisition unit that acquires audio data of a predetermined track in a file in which a plurality of types of audio data are divided into a plurality of tracks according to the types and the tracks are arranged. (2)
[0817] The information processing device according to the above item (1), in which the type is configured as an element of the audio data, a type of the element, or a group into which the element is classified. (3)
[0819] The information processing device according to the above item (1) or (2), further including a decoding unit that decodes the audio data of the predetermined track acquired by the acquisition unit. (4)
[0821] The information processing apparatus according to the above item (3), wherein, when there are a plurality of predetermined tracks, the decoding unit synthesizes audio data of the predetermined tracks acquired by the acquisition unit and decodes the synthesized audio data. (5)
[0823] The information processing apparatus according to the above item (4), wherein
[0824] The file is configured in such a manner that audio data in units of a plurality of objects is divided into the tracks that differ for each object and the tracks are arranged, and metadata items of all the audio data in units of objects are arranged collectively in a track different from the tracks,
[0825] The acquisition unit is configured to acquire audio data of the tracks of the object to be played as the audio data of the predetermined tracks and acquire the metadata.
[0826] The decoding unit is configured to extract metadata of the object to be played from the metadata acquired by the acquisition unit and synthesize the metadata and the audio data acquired by the acquisition unit. (6)
[0828] The information processing apparatus according to the above item (4), wherein
[0829] The file is configured in such a manner that audio data in units of a plurality of objects is divided into the tracks that differ for each object and the tracks are arranged, and metadata items of all the audio data in units of objects are arranged collectively in a track different from the tracks,
[0830] The acquisition unit is configured to acquire audio data of the tracks of the object to be played as the audio data of the predetermined tracks and acquire the metadata.
[0831] The decoding unit is configured to synthesize zero data and audio data and the metadata acquired by the acquisition unit, the zero data indicating a decoding result of zero as audio data of a track that is not played. (7)
[0833] The information processing apparatus according to the above item (4), wherein
[0834] The file is configured in such a manner that audio data in units of a plurality of objects is divided into the tracks that differ for each object and the tracks are arranged, metadata items of the audio data in units of objects are arranged in the tracks that differ for each object,
[0835] The acquisition unit is configured to acquire audio data of a track of an object to be played as audio data of a predetermined track and acquire metadata of the object to be played, and
[0836] The decoding unit is configured to synthesize the audio data and the metadata acquired by the acquisition unit. (8)
[0838] The information processing apparatus according to any one of the above (1) to (7), wherein the audio data items of the multiple tracks are configured to be arranged in one file. (9)
[0840] The information processing apparatus according to any one of the above (1) to (7), wherein the audio data items of the multiple tracks are configured to be arranged in files different for each track. (10)
[0842] The information processing apparatus according to any one of the above (1) to (9), wherein the file is configured in such a manner that information on the multiple types of the audio data is arranged as a track different from the multiple tracks. (11)
[0844] The information processing apparatus according to the above (10), wherein the information on the multiple types of the audio data is configured to contain image frame size information indicating a size of an image frame of image data corresponding to the audio data. (12)
[0846] The information processing apparatus according to any one of the above (1) to (9), wherein the file is configured in such a manner that, as the audio data of a track different from the multiple tracks, information indicating a position of audio data of another track corresponding to the audio data is arranged. (13)
[0848] The information processing apparatus according to any one of the above (1) to (9), wherein the file is configured in such a manner that, as the data of a track different from the multiple tracks, information indicating a position of audio data of another track of data and metadata of audio data corresponding to the other track is arranged. (14)
[0850] The information processing apparatus according to the above (13), wherein the metadata of the audio data is configured to contain information indicating a position at which the audio data is acquired. (15)
[0852] The information processing apparatus according to any one of the above (1) to (14), wherein the file is configured to contain information indicating a reference relationship between the track and another track. (16)
[0854] The information processing apparatus according to any one of the above (1) to (15), wherein the file is configured to contain codec information of audio data of each track. (17)
[0856] The information processing apparatus according to any one of the above (1) to (16), wherein the predetermined type of audio data is information indicating a position at which another type of audio data is acquired. (18)
[0858] An information processing method including an acquisition step of acquiring, by an information processing apparatus, audio data of a predetermined track in a file in which a plurality of types of audio data are divided into a plurality of tracks according to the types and the tracks are arranged. (19)
[0860] An information processing apparatus including a generation unit that generates a file in which a plurality of types of audio data are divided into a plurality of tracks according to the types and the tracks are arranged. (20)
[0862] An information processing method including a generation step of generating, by an information processing apparatus, a file in which a plurality of types of audio data are divided into a plurality of tracks according to the types and the tracks are arranged.
[0863] List of Reference Signs
[0864] 141 file generation apparatus
[0865] 144 moving image playback terminal
[0866] 172 audio file generation unit
[0867] 192 audio file acquisition unit
[0868] 193 audio selection unit
[0869] 211 file generation apparatus
[0870] 214 moving image playback terminal
[0871] 241 audio file generation unit
[0872] 264 audio file acquisition unit
Claims
1. An information processing apparatus comprising: circuitry configured to acquire an audio file composed of a plurality of audio streams each assigned to a track by type, wherein a plurality of types of audio data is divided into a plurality of tracks according to the type and the tracks are arranged, and wherein each audio stream contains object position information; decode at least one audio stream assigned to a track of the plurality of tracks from the audio file; synthesize the audio data in units of objects based on the object position information; determine audio data for each speaker to be assigned to each object based on position information; synthesize the audio data for each object for each speaker, and output the synthesized audio data for each object for each speaker; and reproduce the at least one audio stream from the audio file as the synthesized audio data for each object for each speaker.
2. The information processing apparatus according to claim 1, wherein the type is configured as an element of audio data, a type of the element, or a group to which the element is classified.
3. The information processing apparatus according to claim 1, further comprising a decoding unit for decoding audio data of a predetermined track acquired by the acquisition unit.
4. The information processing apparatus according to claim 3, wherein In the presence of a plurality of predetermined tracks, the decoding unit synthesizes the audio data of the predetermined tracks acquired by the acquisition unit and decodes the synthesized audio data.
5. The information processing apparatus according to claim 4, wherein the file is configured in such a manner that audio data in units of a plurality of objects is divided into the tracks different for each object and the tracks are arranged, and a metadata item of all the audio data in units of objects is arranged collectively in a track different from the tracks, the acquisition unit is configured to acquire audio data of the track of the object to be played as the audio data of the predetermined track and acquire the metadata, and the decoding unit is configured to extract metadata of the object to be played from the metadata acquired by the acquisition unit and synthesize the metadata and the audio data acquired by the acquisition unit.
6. The information processing apparatus according to claim 4, wherein the file is configured in such a manner that audio data in units of a plurality of objects is divided into the tracks different for each object and the tracks are arranged, and metadata of all the audio data in units of objects is arranged collectively in a track different from the tracks, the acquisition unit is configured to acquire audio data of the track of the object to be played as the audio data of the predetermined track and acquire the metadata, and the decoding unit is configured to synthesize the audio data and the metadata acquired by the acquisition unit with zero data indicating a decoding result of zero as audio data of a track not to be played.
7. The information processing apparatus according to claim 4, wherein the file is configured in such a manner that audio data in units of a plurality of objects is divided into tracks different for each object and the tracks are arranged, metadata of the audio data in units of objects is arranged in the tracks different for each object, the acquisition unit is configured to acquire audio data of the track of the object to be played as the audio data of the predetermined track and acquire the metadata, and the decoding unit is configured to synthesize the audio data and the metadata acquired by the acquisition unit with zero data indicating a decoding result of zero as audio data of a track not to be played. the acquisition unit is configured to acquire audio data in the track of the object to be played as the audio data of the predetermined track and acquire metadata of the object to be played, and the decoding unit synthesizes the audio data and the metadata acquired by the acquisition unit.
8. An information processing method including an acquisition step of acquiring, by an information processing apparatus, audio data of a predetermined track in a file in which a plurality of types of audio data are divided into a plurality of tracks according to the types and the tracks are arranged.
9. An information processing apparatus including a generation unit that generates a file in which a plurality of types of audio data are divided into a plurality of tracks according to the types and the tracks are arranged.
10. An information processing method including a generation step of generating, by an information processing apparatus, a file in which a plurality of types of audio data are divided into a plurality of tracks according to the types and the tracks are arranged.
Citation Information
Patent Citations
Method for creating and accessing a menu for audio content without using a display
CN1735941A