Reproduction device
By embedding the scaling area information during the encoding process, generating a bitstream containing multiple scaling area information, the problem of reproducing 4K or 8K video content on devices with different screen sizes and shapes is solved, and simplified adaptive reproduction is achieved.
Patent Information
- Application Number
- CN202011216551.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2014-10-10
- Filing Date
- 2015-09-28
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2035-09-28
AI Technical Summary
The prior art cannot simplify the reproduction of 4K or 8K video content on devices of different screen sizes and shapes, resulting in the need to prepare separate content for each screen size and shape.
By embedding the scaling area information during the encoding process, a bit stream containing multiple scaling area information is generated, allowing the reproduction device to select a suitable scaling area according to the screen size and shape for cropping or audio conversion processing.
The simplified reproduction of appropriate video and audio content on devices of different screen sizes and shapes is achieved, improving the adaptability and flexibility of the reproduction device.
Smart Images

Figure CN112511833B_ABST
Abstract
Description
[0001] This application is a divisional application of an application with national application number 201580053817.8, international filing date of September 28, 2015, entry date into the country of April 1, 2017, and invention title of "Encoding Device and Method, Reproduction Device and Method, and Program". Technical Field
[0002] The present technology relates to an encoding device, an encoding method, a reproduction device, a reproduction method, and a program, and more particularly, to an encoding device, an encoding method, a reproduction device, a reproduction method, and a program that enable each reproduction device to reproduce appropriate content in a simplified manner. Background Art
[0003] In recent years, high-resolution video content known as 4K or 8K has been known. Such 4K or 8K video content is often generated in consideration of a large viewing angle, i.e., reproduction on a large screen.
[0004] In addition, since 4K or 8K video content has a high resolution, the resolution is also sufficient even when a part of the screen of such video content is cropped. Therefore, such video content can be cropped and reproduced (for example, see Non-Patent Document 1).
[0005] Citation List
[0006] Non-Patent Document
[0007] Non-Patent Document 1: FDR-AX100, [Online], [Retrieved on September 24, 2014], Internet <URL: http: / / www.sony.net / Products / di / en-us / products / j4it / index.html> Summary of the Invention
[0008] Problems to be Solved by the Invention
[0009] At the same time, video reproduction devices have become diversified, and reproduction is considered for various screen sizes from large screens to smart phones (multifunctional mobile phones). However, in the current situation, the same content is reproduced by being enlarged or reduced to match each screen size.
[0010] At the same time, the above-mentioned 4K or 8K video content is often generated in consideration of reproduction on a large screen. Therefore, it is not appropriate to reproduce such video content using a reproduction device with a relatively small screen such as a tablet personal computer (PC) or a smart phone.
[0011] Therefore, for example, for reproduction devices having different screen sizes and the like from each other, in order to provide content suitable for each screen size, screen shape, etc., it is necessary to separately prepare content suitable for each screen size, screen shape, etc.
[0012] This technology takes these situations into consideration and enables each reproduction device to reproduce appropriate content in a simplified manner.
[0013] Solution to the problem
[0014] The reproduction device according to the first aspect of the present technology includes: a decoding unit that decodes the encoded video data or the encoded audio data; a scaling area selection unit that selects one or more pieces of scaling area information from a plurality of pieces of scaling area information specifying the area to be scaled; and a data processing unit that performs a cropping process on the video data obtained by decoding or an audio conversion process on the audio data obtained by decoding based on the selected scaling area information.
[0015] Among the plurality of pieces of scaling area information, there may be included scaling area information specifying the area for each type of reproduction target device.
[0016] Among the plurality of pieces of scaling area information, there may be included scaling area information specifying the area for the rotation direction of each reproduction target device.
[0017] Among the plurality of pieces of scaling area information, there may be included scaling area information specifying the area for each specific video object.
[0018] The scaling area selection unit may be made to select the scaling area information according to the user's operation input.
[0019] The scaling area selection unit may be made to select the scaling area information based on the information related to the reproduction device.
[0020] The scaling area selection unit may be made to select the scaling area information by using at least any one of the information indicating the type of the reproduction device and the information indicating the rotation direction of the reproduction device as the information related to the reproduction device.
[0021] The reproduction method or program according to the first aspect of the present technology includes the following steps: decoding the encoded video data or the encoded audio data; selecting one or more pieces of scaling area information from a plurality of pieces of scaling area information specifying the area to be scaled; and performing a cropping process on the video data obtained by decoding or an audio conversion process on the audio data obtained by decoding based on the selected scaling area information.
[0022] According to a first aspect of the present technology, decode the encoded video data or the encoded audio data; select one or more pieces of scaling region information from a plurality of pieces of scaling region information specifying a region to be scaled; and perform a cropping process on the video data obtained by decoding or perform an audio conversion process on the audio data obtained by decoding based on the selected scaling region information.
[0023] The encoding apparatus according to a second aspect of the present technology includes: an encoding unit that encodes video data or encodes audio data; and a multiplexer that generates a bitstream by multiplexing the encoded video data or the encoded audio data with a plurality of pieces of scaling region information specifying a region to be scaled.
[0024] The encoding method or program according to a second aspect of the present technology includes the following steps: encoding video data or encoding audio data; and generating a bitstream by multiplexing the encoded video data or the encoded audio data with a plurality of pieces of scaling region information specifying a region to be scaled.
[0025] According to a second aspect of the present technology, encode video data or encode audio data; and generate a bitstream by multiplexing the encoded video data or the encoded audio data with a plurality of pieces of scaling region information specifying a region to be scaled.
[0026] Effects of the present invention
[0027] According to the first and second aspects of the present technology, each reproduction device can reproduce appropriate content in a simplified manner.
[0028] Note that the effects of the present technology are not limited to the effects described herein, but may be any effects described in the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 is a diagram showing an example of the configuration of an encoding apparatus.
[0030] Figure 2 is a diagram showing the configuration of encoded content data.
[0031] Figure 3 is a diagram showing scaling region information.
[0032] Figure 4 is a diagram showing the syntax of a scaling region information presence flag.
[0033] Figure 5 is a diagram showing the syntax of scaling region information.
[0034] Figure 6is a diagram showing the syntax of zoom area information.
[0035] Figure 7 is a diagram showing the syntax of zoom area information.
[0036] Figure 8 is a diagram showing the syntax of zoom area information.
[0037] Figure 9 is a diagram showing the syntax of zoom area information.
[0038] Figure 10 is a diagram showing the syntax of zoom area information.
[0039] Figure 11 is a diagram showing zoom area information.
[0040] Figure 12 is a diagram showing zoom area information.
[0041] Figure 13 is a diagram showing the syntax of zoom area information.
[0042] Figure 14 This is a diagram showing the syntax of the zoom area information existence flag, etc.
[0043] Figure 15 is a diagram showing the syntax of zoom area information.
[0044] Figure 16 It is a diagram showing the syntax of zoom area auxiliary information, etc.
[0045] Figure 17 is a diagram showing scaling specifications.
[0046] Figure 18 is a diagram showing an example of reproduced content.
[0047] Figure 19 is a flowchart showing the encoding process.
[0048] Figure 20 is a diagram showing an example of the configuration of a reproduction device.
[0049] Figure 21 is a flowchart showing the reproduction process.
[0050] Figure 22 is a diagram showing an example of the configuration of a reproduction device.
[0051] Figure 23 is a flowchart showing the reproduction process.
[0052] Figure 24 is a diagram showing an example of the configuration of a reproduction device.
[0053] Figure 25 This is a flowchart showing the reproduction process.
[0054] Figure 26 This is a diagram showing an example of the configuration of the reproduction device.
[0055] Figure 27 This is a flowchart showing the reproduction process.
[0056] Figure 28 This is a diagram showing an example of the configuration of the computer. DETAILED DESCRIPTION OF THE INVENTION
[0057] Hereinafter, embodiments of the present technology will be described with reference to the accompanying drawings.
[0058] <First Embodiment>
[0059] <Example of the Configuration of the Encoding Device>
[0060] The present technology enables reproduction devices having different display screen sizes, such as TV receivers and smart phones, to reproduce appropriate content, such as content suitable for such reproduction devices, in a simplified manner. The content described here can be, for example, content formed by video and audio or content formed by either video or audio. Hereinafter, an example of the case where the content formed by video and the audio accompanying the video is used will be continued for description.
[0061] Figure 1 This is a diagram showing an example of the configuration of the encoding device according to the present technology.
[0062] The encoding device 11 encodes the content generated by the content provider and outputs a bitstream (code string) in which the encoded data obtained as a result is stored.
[0063] The encoding device 11 includes: a video data encoding unit 21; an audio data encoding unit 22; a metadata encoding unit 23; a multiplexer 24; and an output unit 25.
[0064] In this example, the video data of the video constituting the content and the audio data of the audio are respectively provided to the video data encoding unit 21 and the audio data encoding unit 22, and the metadata of the content is provided to the metadata encoding unit 23.
[0065] The video data encoding unit 21 encodes the provided video data of the content and provides the encoded video data obtained as a result to the multiplexer 24. The audio data encoding unit 22 encodes the provided audio data of the content and provides the encoded audio data obtained as a result to the multiplexer 24.
[0066] The metadata encoding unit 23 encodes the metadata of the provided content and provides the encoded metadata obtained as a result to the multiplexer 24.
[0067] The multiplexer 24 generates a bitstream by multiplexing the encoded video data provided from the video data encoding unit 21, the encoded audio data provided from the audio data encoding unit 22, and the encoded metadata provided from the metadata encoding unit 23, and provides the generated bitstream to the output unit 25. The output unit 25 outputs the bitstream provided from the multiplexer 24 to a reproduction device or the like.
[0068] Note that hereinafter, the bitstream output from the output unit 25 will also be referred to as the encoded content data.
[0069] <Encoded content data>
[0070] The content encoded by the encoding device 11 is generated as needed considering cropping and reproduction. In other words, the content producer generates the content considering directly reproducing the content or cropping and reproducing a part of the entire area of the video constituting the content.
[0071] For example, the content producer selects a partial area to be cropped and reproduced, that is, an area to be scaled and reproduced by cropping, as a scaling area from the entire area of the video (image) constituting the content.
[0072] Note that, for example, the scaling area for achieving a perspective suitable for the considered reproduction device or the like can be freely determined by the content producer. In addition, the scaling area can be determined based on a scaling purpose, such as magnifying and tracking a specific object, such as a singer or a performer in the video of the content.
[0073] In this way, when the producer side designates a plurality of scaling areas for the content, in the bitstream output from the encoding device 11, that is, in the encoded content data, the scaling area information designating the scaling area is stored as metadata. At this time, when it is desired to designate the scaling area for each predetermined time unit, the scaling area information can be stored in the encoded content data for each of the above time units.
[0074] More specifically, for example, as Figure 2 shown, when storing the content in the bitstream for each frame, the scaling area information can be stored in the bitstream for each frame.
[0075] In Figure 2In the example shown, a header section HD in which header information and the like are stored is arranged at the start of the bitstream, i.e., the encoded content data, and a data section DA in which the encoded video data and the encoded audio data are stored is arranged after the header section HD.
[0076] In the header section HD, a video information header section PHD in which header information related to the video constituting the content is stored, an audio information header section AHD in which header information related to the audio constituting the content is stored, and a meta information header section MHD in which header information related to the metadata of the content is stored are provided.
[0077] In addition, in the meta information header section MHD, a zoom region information header section ZHD in which information related to the zoom region information is stored is provided. For example, in the zoom region information header section ZHD, a zoom region information presence flag indicating whether the zoom region information is stored in the data section DA and the like are stored.
[0078] In addition, in the data section DA, a data section in which the data of the encoded content is stored for each frame of the content is provided. In this example, a data section DAF-1 in which the data of the first frame is stored is provided at the start of the data section DA, and a data section DAF-2 in which the data of the second frame of the content is stored is provided after the data section DAF-1. Also, here, the data sections of the third frame and subsequent frames are not shown in the drawings. Hereinafter, when the data sections DAF-1 and DAF-2 of each frame do not need to be particularly distinguished from each other, each of the data sections DAF-1 and DAF-2 will be simply referred to as the data section DAF.
[0079] In the data section DAF-1 of the first frame, a video information data section PD-1 in which the encoded video data is stored, an audio information data section AD-1 in which the encoded audio data is stored, and a meta information data section MD-1 in which the encoded metadata is stored are provided.
[0080] For example, in the meta information data section MD-1, the position information of the video object and the sound source object included in the first frame of the content and the like are included. In addition, a zoom region information data section ZD-1 in which the encoded zoom region information in the encoded metadata is stored is provided within the meta information data section MD-1. The position information of the video object and the sound source object, the zoom region information, and the like are set as the metadata of the content.
[0081] Similarly, similar to the data section DAF-1, in the data section DAF-2, a video information data section PD-2 in which encoded video data is stored, an audio information data section AD-2 in which encoded audio data is stored, and a meta-information data section MD-2 in which encoded metadata is stored are provided. Additionally, in the meta-information data section MD-2, a zoom area information data section ZD-2 in which encoded zoom area information is stored is provided.
[0082] Furthermore, hereinafter, when the video information data section PD-1 and the video information data section PD-2 do not need to be particularly distinguished from each other, each of the video information data section PD-1 and the video information data section PD-2 will also be simply referred to as the video information data section PD, and when the audio information data section AD-1 and the audio information data section AD-2 do not need to be particularly distinguished from each other, each of the audio information data section AD-1 and the audio information data section AD-2 will also be simply referred to as the audio information data section AD. Additionally, when the meta-information data section MD-1 and the meta-information data section MD-2 do not need to be particularly distinguished from each other, each of the meta-information data section MD-1 and the meta-information data section MD-2 will be simply referred to as the meta-information data section MD, and when the zoom area information data section ZD-1 and the zoom area information data section ZD-2 do not need to be particularly distinguished from each other, each of the zoom area information data section ZD-1 and the zoom area information data section ZD-2 will also be simply referred to as the zoom area information data section ZD.
[0083] In addition, in the case of Figure 2 , in each data section DAF, examples of setting the video information data section PD, the audio information data section AD, and the meta-information data section MD are described. However, the meta-information data section MD can be set in each or one of the video information data section PD and the audio information data section AD. In this case, the zoom area information is stored in the zoom area information data section ZD of the meta-information data section MD set within the video information data section PD or the audio information data section AD.
[0084] Similarly, although examples of setting the video information header section PHD, the audio information header section AHD, and the meta-information header section MHD in the header section HD are described, the meta-information header section MHD can be set in both or any one of the video information header section PHD and the audio information header section AHD.
[0085] In addition, when the zoom region information is the same for each frame of the content, the zoom region information can be configured to be stored in the header section HD. In this case, there is no need to set the zoom region information data section ZD in each data section DAF.
[0086] <Specific Example 1 of Zoom Region Information>
[0087] Subsequently, more specific examples of the zoom region information will be described.
[0088] The above-mentioned zoom region information is information about the zoom region specifying the region to be zoomed. More specifically, the zoom region information is information representing the position of the zoom region. For example, as Figure 3 shown, the zoom region can be specified using the coordinates of the center position of the zoom region, the coordinates of the starting point, the coordinates of the ending point, the vertical width, the horizontal width, etc.
[0089] In Figure 3 the shown case, the region of the entire video (image) of the content is the original region OR, and a rectangular zoom region ZE is specified within the original region OR. In this example, the width of the zoom region ZE in the horizontal (lateral) direction of the figure is the horizontal width XW, and the width of the zoom region ZE in the vertical (longitudinal) direction of the figure is the vertical width YW.
[0090] Here, in the figure, a point in the XY coordinate system with the horizontal (lateral) direction as the X direction and the vertical (longitudinal) direction as the Y direction will be represented as the coordinate (X, Y).
[0091] Now, when the coordinates of the point P11 at the center position (center position) of the zoom region ZE are (XC, YC), the zoom region ZE can be specified using the center coordinates (XC, YC), the horizontal width XW, and the vertical width YW of the zoom region ZE. Therefore, the center coordinates (XC, YC), the horizontal width XW, and the vertical width YW can be set as the zoom region information.
[0092] In addition, for example, when the zoom region ZE is a rectangular region, the upper left vertex P12 of the zoom region ZE in the figure is set as the starting point, and the lower right vertex P13 of the zoom region ZE in the figure is set as the ending point, and the zoom region ZE can also be specified using the coordinates (X0, Y0) of the starting point (vertex P12) and the coordinates (X1, Y1) of the ending point (vertex P13). Therefore, the coordinates (X0, Y0) of the starting point and the coordinates (X1, Y1) of the ending point can be set as the zoom region information.
[0093] More specifically, the coordinates (X0, Y0) of the starting point and the coordinates (X1, Y1) of the ending point are set as the zoom area information. In this case, for example, it can be configured such that, based on the value of the flag indicating the existence of zoom area information, Figure 4 the zoom area information shown is stored in the above-mentioned zoom area information header section ZHD, and Figure 5 the zoom area information shown is stored in each zoom area information data section ZD.
[0094] Figure 4 FIG. is a diagram showing the syntax of the flag indicating the existence of zoom area information. In this example, "hasZoomAreaInfo" represents the flag indicating the existence of zoom area information, and the value of the flag indicating the existence of zoom area information hasZoomAreaInfo is one of "0" and "1".
[0095] Here, when the value of the flag indicating the existence of zoom area information hasZoomAreaInfo is "0", it means that the encoded content data does not include zoom area information. On the contrary, when the value of the flag indicating the existence of zoom area information hasZoomAreaInfo is "1", it means that the encoded content data includes zoom area information.
[0096] In addition, when the value of the flag indicating the existence of zoom area information hasZoomAreaInfo is "1", the zoom area information is stored in the zoom area information data section ZD of each frame. For example, the zoom area information is stored in the zoom area information data section ZD in the syntax shown Figure 5 as follows.
[0097] In Figure 5 ,"ZoomAreaX0" and "ZoomAreaY0" respectively represent the X coordinate X0 and the Y coordinate Y0 of the starting point of the zoom area ZE. In addition, "ZoomAreaX1" and "ZoomAreaY1" respectively represent the X coordinate X1 and the Y coordinate Y1 of the ending point of the zoom area ZE.
[0098] For example, when the video of the content to be encoded is an 8K video, each of the values of "ZoomAreaX0" and "ZoomAreaX1" is set to one of the values from 0 to 7679, and each of the values of "ZoomAreaY0" and "ZoomAreaY1" is set to one of the values from 0 to 4319.
[0099] <Specific Example 2 of Zoom Area Information>
[0100] In addition, for example, also when the center coordinates (XC, YC), the horizontal width XW, and the vertical width YW are set as the zoom area information, Figure 4The flag hasZoomAreaInfo indicating the existence of the zoom area information shown is stored in the zoom area information header section ZHD. When the value of the flag hasZoomAreaInfo indicating the existence of the zoom area information is "1", the zoom area information is stored in the zoom area information data section ZD of each frame. In this case, for example, the zoom area information is stored in the zoom area information data section ZD according to the syntax shown in Figure 6 shown.
[0101] In Figure 6 this case, "ZoomAreaXC" and "ZoomAreaYC" respectively represent the X coordinate XC and the Y coordinate YC of the center coordinates (XC, YC) of the zoom area ZE.
[0102] In addition, "ZoomAreaXW" and "ZoomAreaYW" respectively represent the horizontal width XW and the vertical width YW of the zoom area ZE.
[0103] Also in this example, for example, when the video of the content to be encoded is an 8K video, each value of the values of "ZoomAreaXC" and "ZoomAreaXW" is set to one of the values from 0 to 7679, and each value of the values of "ZoomAreaYC" and "ZoomAreaYW" is set to one of the values from 0 to 4319.
[0104] <Specific Example 3 of Zoom Area Information>
[0105] In addition, for example, when the zoom area is specified using the center coordinates (XC, YC), the horizontal width XW, and the vertical width YW and the horizontal width XW and the vertical width YW are set to fixed values, only the difference of the center coordinates (XC, YC) can be stored as the zoom area information in the zoom area information data section ZD.
[0106] In this case, for example, in the zoom area information data section ZD-1 set in the data section DAF-1 of the first frame, the zoom area information shown in Figure 6 is stored. In addition, in the zoom area information data section ZD set in the data section DAF of each of the second frame and subsequent frames, the zoom area information is stored according to the syntax shown in Figure 7 shown.
[0107] In Figure 7In the case of, "nbits", "ZoomAreaXCshift", and "ZoomAreaYCshift" are stored as zoom area information. "nbits" is the number-of-bits information indicating the number of bits of the information of each of "ZoomAreaXCshift" and "ZoomAreaYCshift".
[0108] In addition, "ZoomAreaXCshift" represents the difference between XC, which is the X coordinate of the center coordinates (XC, YC), and a predetermined reference value. For example, the reference value of the coordinate XC may be the X coordinate of the center coordinates (XC, YC) in the first frame or the X coordinate of the center coordinates (XC, YC) in the previous frame of the current frame.
[0109] "ZoomAreaYCshift" represents the difference between YC, which is the Y coordinate of the center coordinates (XC, YC), and a predetermined reference value. For example, similar to the reference value of the coordinate XC, the reference value of the coordinate YC may be the Y coordinate of the center coordinates (XC, YC) in the first frame or the Y coordinate of the center coordinates (XC, YC) in the previous frame of the current frame.
[0110] Such "ZoomAreaXCshift" and "ZoomAreaYCshift" represent the amount of movement from the reference value of the center coordinates (XC, YC).
[0111] Note that, for example, in the case where the reference value of the center coordinates (XC, YC) is known on the reproduction side of the content, in the case where the reference value of the center coordinates (XC, YC) is stored in the zoom area information header section ZHD or the like, Figure 7 the shown zoom area information can be stored in the zoom area information data section ZD of each frame.
[0112] <Specific Example 4 of Zoom Area Information>
[0113] In addition, for example, in the case where the zoom area is specified using the center coordinates (XC, YC), the horizontal width XW, and the vertical width YW and the center coordinates (XC, YC) are set to fixed values, only the differences, that is, the change amounts of the horizontal width XW and the vertical width YW, can be stored as the zoom area information in the zoom area information data section ZD.
[0114] In this case, for example, in the zoom area information data section ZD-1 set in the data section DAF-1 of the first frame, the Figure 6 shown zoom area information is stored. In addition, in the zoom area information data section ZD set in the data section DAF set in each of the second frame and subsequent frames, the Figure 8 shown syntax is used to store the zoom area information.
[0115] In Figure 8 “nbits”, “ZoomAreaXWshift”, and “ZoomAreaYWshift” are stored as zoom area information. “nbits” is the number of bits information indicating the number of bits of information for each of “ZoomAreaXWshift” and “ZoomAreaYWshift”.
[0116] In addition, “ZoomAreaXWshift” represents the amount of change relative to a predetermined reference value of the horizontal width XW. For example, the reference value of the horizontal width XW can be the horizontal width XW in the first frame or the horizontal width XW of the previous frame of the current frame.
[0117] “ZoomAreaYWshift” represents the amount of change relative to a reference value of the vertical width YW. For example, similar to the reference value of the horizontal width XW, the reference value of the vertical width YW can be the vertical width YW in the first frame or the vertical width YW of the previous frame of the current frame.
[0118] Note that, for example, when the reference values of the horizontal width XW and the vertical width YW are known on the reproduction side of the content, and when the reference values of the horizontal width XW and the vertical width YW are stored in the zoom area information header section ZHD, etc., Figure 8 the zoom area information shown can be stored in the zoom area information data section ZD of each frame.
[0119] <Specific Example 5 of Zoom Area Information>
[0120] In addition, for example, when specifying the zoom area using the center coordinates (XC, YC), the horizontal width XW, and the vertical width YW, as in Figure 7 and Figure 8 the differences of the center coordinates (XC, YC), the horizontal width XW, and the vertical width YW can be stored as zoom area information in the zoom area information data section ZD.
[0121] In this case, for example, in the zoom area information data section ZD-1 set in the data section DAF-1 set in the first frame, the zoom area information shown is stored. In addition, in the zoom area information data section ZD in each of the data sections DAF set in the second frame and subsequent frames, the zoom area information is stored in the syntax shown in Figure 6 Figure 9
[0122] In Figure 9 In the case of, "nbits", "ZoomAreaXCshift", "ZoomAreaYCshift", "ZoomAreaXWshift", and "ZoomAreaYWshift" are stored as zoom area information.
[0123] "nbits" is the number of bits information representing the number of bits of information for each of "ZoomAreaXCshift", "ZoomAreaYCshift", "ZoomAreaXWshift", and "ZoomAreaYWshift".
[0124] As Figure 7 in the case of, "ZoomAreaXCshift" and "ZoomAreaYCshift" respectively represent the differences from the reference values of the X coordinate and Y coordinate of the center coordinates (XC, YC).
[0125] In addition, as Figure 8 in the case of, "ZoomAreaXWshift" and "ZoomAreaYWshift" respectively represent the amounts of change relative to the reference values of the horizontal width XW and the vertical width YW.
[0126] Here, the reference values of the center coordinates (XC, YC), the horizontal width XW, and the vertical width YW can be set to the center coordinates (XC, YC), the horizontal width XW, and the vertical width YW in the first frame or the previous frame of the current frame. In addition, when the reference values of the center coordinates (XC, YC), the horizontal width XW, and the vertical width YW are known on the reproduction side of the content, or when the reference values are stored in the zoom area information header section ZHD, Figure 9 the zoom area information shown in
[0127] <Specific Example 6 of Zoom Area Information>
[0128] In addition, by combining the above Figures 6 to 9 shown examples, for example, the zoom area information can be stored in each zoom area information data section ZD in the Figure 10 shown syntax.
[0129] In this case, Figure 4 the zoom area information presence flag hasZoomAreaInfo shown in Figure 10The syntax shown stores the zoom area information in the zoom area information data section ZD.
[0130] In Figure 10 the case shown, the coding mode information is arranged at the start of the zoom area information, and the coding mode information represents Figures 6 to 9 the format for describing the zoom area information (more specifically, the information specifying the position of the zoom area) in the format shown. In Figure 10 it, "mode" represents the coding mode information.
[0131] Here, the value of the coding mode information mode is set to one of the values 0 to 3.
[0132] For example, when the value of the coding mode information mode is "0", as shown in "case0" and below in the figure, similar to Figure 6 the example shown, "ZoomAreaXC" representing the coordinate XC, "ZoomAreaYC" representing the coordinate YC, "ZoomAreaXW" representing the horizontal width XW, and "ZoomAreaYW" representing the vertical width YW are stored as the zoom area information.
[0133] On the other hand, when the value of the coding mode information mode is "1", as shown in "case 1" and below in the figure, similar to Figure 7 the example shown, "nbits" representing the number of bits information, "ZoomAreaXCshift" representing the difference in the coordinate XC, and "ZoomAreaYCshift" representing the difference in the coordinate YC are stored as the zoom area information.
[0134] When the value of the coding mode information mode is "2", as shown in "case 2" and below in the figure, similar to Figure 8 the example shown, "nbits" representing the number of bits information, "ZoomAreaXWshift" representing the change amount of the horizontal width XW, and "ZoomAreaYWshift" representing the change amount of the vertical width YW are stored as the zoom area information.
[0135] In addition, when the value of the coding mode information mode is "3", as shown in "case3" and below in the figure, similar to Figure 9The example shown is similar. "nbits" representing the number of bits information, "ZoomAreaXCshift" representing the difference in coordinate XC, "ZoomAreaYCshift" representing the difference in coordinate YC, "ZoomAreaXWshift" representing the change amount of the horizontal width XW, and "ZoomAreaYWshift" representing the change amount of the vertical width YW are stored as zoom area information.
[0136] <Specific Example 7 of Zoom Area Information>
[0137] In addition, although the example of storing coordinate information as zoom area information has been described above, angle information specifying the zoom area can be stored as zoom area information in each zoom area information data section ZD.
[0138] For example, as Figure 11 shown, a point located at the same height as the center position CP of the original area OR and separated from the center position CP by a predetermined distance toward the front side in Figure 11 is set as the reference viewing point WP when viewing the content. In addition, it is assumed that the positional relationship between the center position CP and the viewing point WP is constantly the same positional relationship regardless of the frame of the content. Note that in Figure 11 , the same reference numerals are assigned to the parts corresponding to the situation shown in Figure 3 , and their descriptions will not be presented appropriately.
[0139] In Figure 11 , the straight line connecting the center position CP and the viewing point WP is set as the straight line L11. In addition, the midpoint on the left side of the zoom area ZE in the figure is set as the point P21, and the straight line connecting the point P21 and the viewing point WP is set as the straight line L12. Furthermore, the angle formed by the straight line L11 and the straight line L12 is set as the horizontal angle φ 左 .
[0140] Similarly, the midpoint on the right side of the zoom area ZE in the figure is set as the point P22, and the straight line connecting the point P22 and the viewing point WP is set as the straight line L13. In addition, the angle formed by the straight line L11 and the straight line L13 is set as the horizontal angle φ 右 .
[0141] In addition, a point at the position having the same Y coordinate as the center position CP on the right side of the zoom area ZE in the figure is set as the point P23, and the straight line connecting the point P23 and the viewing point WP is set as the straight line L14. In addition, the upper right vertex of the zoom area ZE in the figure is set as the point P24, the straight line connecting the point P24 and the viewing point WP is set as the straight line L15, and the angle formed by the straight line L14 and the straight line L15 is set as the pitch angle θ 顶。
[0142] Similarly, the lower right vertex of the zoom region ZE in the figure is set as point P25, the straight line connecting point P25 and the viewing point WP is set as straight line L16, and the included angle formed by straight line L14 and straight line L16 is set as the pitch angle θ 底 。
[0143] At this time, the horizontal angle φ 左 、horizontal angle φ 右 、pitch angle θ 顶 and pitch angle θ 底 can be used to specify the zoom region ZE. Accordingly, the horizontal angle φ 左 、horizontal angle φ 右 、pitch angle θ 顶 and pitch angle θ 底 can be stored as zoom region information in each zoom region information data section ZD shown in Figure 2 . In addition, some or all of the change amounts of the horizontal angle φ 左 、horizontal angle φ 右 、pitch angle θ 顶 and pitch angle θ 底 can be set as the zoom region information.
[0144] <Specific Example 8 of Zoom Region Information>
[0145] In addition, for example, as shown in Figure 12 , the angle information determined based on the center position CP, the positional relationship between the point P11 located at the center position of the zoom region ZE and the viewing point WP can be set as the zoom region information. Note that in Figure 12 , the same reference numerals are assigned to the parts corresponding to the cases shown in Figure 3 or Figure 11 , and their descriptions will not be presented appropriately.
[0146] In Figure 12 , the straight line connecting the point P11 located at the center position of the zoom region ZE and the viewing point WP is set as straight line L21. In addition, the point having the same X coordinate as the point P11 located at the center position of the zoom region ZE and having the same Y coordinate as the center position CP of the original region OR is set as point P31, and the straight line connecting point P31 and the viewing point WP is set as straight line L22.
[0147] In addition, the midpoint of the upper side of the zoom region ZE in the figure is set as point P32, the straight line connecting point P32 and the viewing point WP is set as straight line L23, the midpoint of the lower side of the zoom region ZE in the figure is set as point P33, and the straight line connecting point P33 and the viewing point WP is set as straight line L24.
[0148] In addition, the included angle formed by the straight line L12 and the straight line L13 is set as the horizontal viewing angle φ W , and the included angle formed by the straight line L11 and the straight line L22 is set as the horizontal angle φ C . In addition, the included angle formed by the straight line L23 and the straight line L24 is set as the vertical viewing angle θ W , and the included angle formed by the straight line L21 and the straight line L22 is set as the pitch angle θ C .
[0149] Here, the horizontal angle φ C and the pitch angle θ C respectively represent the horizontal angle and the pitch angle from the viewing point WP with respect to the point P11 located at the center of the zoom area ZE.
[0150] At this time, the horizontal viewing angle φ W , the horizontal angle φ C , the vertical viewing angle θ W and the pitch angle θ C can be used to specify the zoom area ZE. Therefore, the horizontal viewing angle φ W , the horizontal angle φ C , the vertical viewing angle θ W and the pitch angle θ C or the variation amounts of these angles can be stored as zoom area information in Figure 2 each zoom area information data section ZD shown.
[0151] In this case, for example, Figure 4 the zoom area information presence flag hasZoomAreaInfo shown is stored in the zoom area information header section ZHD. In addition, when the value of the zoom area information presence flag hasZoomAreaInfo is "1", the zoom area information is stored in the zoom area information data section ZD of each frame. For example, the zoom area information is stored in the zoom area information data section ZD in the syntax shown in Figure 13 .
[0152] In Figure 13 shown, coding mode information is arranged at the start of the zoom area information, and the coding mode information represents one of the multiple formats in which the zoom area information (more specifically, the information on the position of the zoom area) is described.
[0153] In Figure 13 , "mode" represents the coding mode information, and the value of the coding mode information mode is set to one of the values from 0 to 3.
[0154] For example, when the value of the encoding mode information mode is "0", as shown in "case0" and below in the figure, it represents the horizontal angle φ C "ZoomAreaAZC" representing the pitch angle θ C "ZoomAreaELC" representing the horizontal viewing angle φ W "ZoomAreaAZW" representing the vertical viewing angle θ W and "ZoomAreaELW" representing the vertical viewing angle θ
[0155] When the value of the encoding mode information is "1", as shown in "case 1" and below in the figure, "nbits" representing the number of bits information, the offset angle of the horizontal angle φ C "ZoomAreaAZCshift" and the offset angle of the pitch angle θ C "ZoomAreaELCshift" are stored as zoom area information.
[0156] Here, the number of bits information nbits is the information representing the number of bits of the information of each of "ZoomAreaAZCshift" and "ZoomAreaELCshift".
[0157] In addition, "ZoomAreaAZCshift" is set to the horizontal angle φ of the previous frame of the current frame C or the horizontal angle φ of a predetermined reference C and the horizontal angle φ of the current frame C The difference between them, "ZoomAreaELCshift" is set to the pitch angle θ of the previous frame of the current frame C or the pitch angle θ of a predetermined reference C and the pitch angle θ of the current frame C The difference between them, and so on.
[0158] When the value of the encoding mode information mode is "2", as shown in "case 2" and below in the figure, "nbits" representing the number of bits information, the change amount of the horizontal viewing angle φ W "ZoomAreaAZWshift" and the change amount of the vertical viewing angle θ W "ZoomAreaELWshift" are stored as zoom area information.
[0159] Here, the number of bits information nbits is the information representing the number of bits of the information of each of "ZoomAreaAZWshift" and "ZoomAreaELWshift".
[0160] In addition, "ZoomAreaAZWshift" is set to the horizontal viewing angle φ of the previous frame of the current frame W or the horizontal viewing angle φ of a predetermined reference W and the horizontal viewing angle φ of the current frame W The difference between them, "ZoomAreaELWshift" is set to the vertical viewing angle θ of the previous frame of the current frame W or the vertical viewing angle θ of a predetermined reference W and the vertical viewing angle θ of the current frame W The difference between them, and so on.
[0161] In addition, when the value of the coding mode information mode is "3", as shown in "case3" and below in the figure, the number of bits information "nbits", the offset angle "ZoomAreaAZCshift" representing the horizontal angle φ C , the offset angle "ZoomAreaELCshift" representing the pitch angle θ C , the change amount "ZoomAreaAZWshift" representing the horizontal viewing angle φ W and the change amount "ZoomAreaELWshift" representing the vertical viewing angle θ W are stored as zoom area information.
[0162] In this case, the number of bits information nbits is the information on the number of bits of the information representing each of "ZoomAreaAZCshift", "ZoomAreaELCshift", "ZoomAreaAZWshift", and "ZoomAreaELWshift".
[0163] Note that the configuration of the zoom area information is not limited to Figure 13 the example shown, and only "ZoomAreaAZC", "ZoomAreaELC", "ZoomAreaAZW", and "ZoomAreaELW" can be set as the zoom area information. In addition, both sides or only one side of "ZoomAreaAZCshift", "ZoomAreaELCshift", "ZoomAreaAZWshift", and "ZoomAreaELWshift" can be set as the zoom area information.
[0164] <Specific Example 9 of Zoom Area Information>
[0165] In addition, although the above describes the case where there is only one piece of zoom area information, multiple pieces of zoom area information can be stored in the zoom area information data section ZD. In other words, by specifying multiple zoom areas for one content, the zoom area information can be stored in the zoom area information data section ZD of each zoom area.
[0166] In this case, for example, each piece of information is stored in the zoom area information header section ZHD in the syntax shown in Figure 14 and the zoom area information is further stored in the zoom area information data section ZD of each frame in the syntax shown in Figure 15 In the example shown in
[0167] In Figure 14 the example shown, "hasZoomAreaInfo" represents the zoom area information presence flag. When the value of the zoom area information presence flag is "1", "numZoomAreas" is stored after the zoom area information presence flag hasZoomAreaInfo.
[0168] Here, "numZoomAreas" represents the zoom area number information, and the zoom area number information represents the number of pieces of zoom area information described in the zoom area information data section ZD, that is, the number of zoom areas set for the content. In this example, the value of the zoom area number information numZoomAreas is one of the values from 0 to 15.
[0169] In the encoded content data, the zoom area information, more specifically, the information specifying the position of each zoom area corresponding to the value obtained by adding 1 to the value of the zoom area number information numZoomAreas is stored in the zoom area information data section ZD.
[0170] Accordingly, for example, when the value of the zoom area number information numZoomAreas is "0", in the zoom area information data section ZD, for one zoom area, the information specifying the position of this zoom area is stored.
[0171] In addition, when the value of the zoom area information presence flag hasZoomAreaInfo is "1", the zoom area information is stored in the zoom area information data section ZD. For example, the zoom area information is described in the zoom area information data section ZD in the syntax shown in Figure 15 In the example shown in
[0172] In Figure 15 the example shown, the zoom area information corresponding to the number represented by the zoom area number information numZoomAreas is stored.
[0173] InFigure 15 In it, "mode[idx]" represents the coding mode information of the zoom area specified by the index idx, and the value of the coding mode information mode[idx] is set to one of the values 0 to 3. Note that the index idx is each value from 0 to numZoomAreas.
[0174] For example, when the value of the coding mode information mode[idx] is "0", as shown in "case0" and below in the figure, "ZoomAreaXC[idx]" representing the coordinate XC, "ZoomAreaYC[idx]" representing the coordinate YC, "ZoomAreaXW[idx]" representing the horizontal width XW, and "ZoomAreaYW[idx]" representing the vertical width YW are stored as the zoom area information of the zoom area specified by the index idx.
[0175] In addition, when the value of the coding mode information mode[idx] is "1", as shown in "case1" and below in the figure, "nbits" as the bit number information, "ZoomAreaXCshift[idx]" representing the difference in the coordinate XC, and "ZoomAreaYCshift[idx]" representing the difference in the coordinate YC are stored as the zoom area information of the zoom area specified by the index idx. Here, the bit number information nbits represents the number of bits of the information of each of "ZoomAreaXCshift[idx]" and "ZoomAreaYCshift[idx]".
[0176] When the value of the coding mode information mode[idx] is "2", as shown in "case 2" and below in the figure, "nbits" representing the bit number information, "ZoomAreaXWshift[idx]" representing the change amount of the horizontal width XW, and "ZoomAreaYWshift[idx]" representing the change amount of the vertical width YW are stored as the zoom area information of the zoom area specified by the index idx. Here, the bit number information nbits represents the number of bits of the information of each of "ZoomAreaXWshift[idx]" and "ZoomAreaYWshift[idx]".
[0177] In addition, when the value of the coding mode information mode[idx] is "3", as shown in "case3" and below in the figure, the number of bits "nbits" as the bit information, the difference in coordinates XC "ZoomAreaXCshift[idx]", the difference in coordinates YC "ZoomAreaYCshift[idx]", the change amount of the horizontal width XW "ZoomAreaXWshift[idx]", and the change amount of the vertical width YW "ZoomAreaYWshift[idx]" are stored as the zoom area information of the zoom area specified by the index idx. Here, the number of bits information nbits represents the number of bits of the information of each of "ZoomAreaXCshift[idx]", "ZoomAreaYCshift[idx]", "ZoomAreaXWshift[idx]", and "ZoomAreaYWshift[idx]".
[0178] In Figure 15 In the example shown, the coding mode information mode[idx] and the zoom area information corresponding to the number of zoom areas are stored in the zoom area information data section ZD.
[0179] Note that alternatively, the zoom area information can consist only of the coordinates XC and YC, the horizontal angle φ C and the pitch angle θ C , the difference in coordinates XC and the difference in coordinates YC, or the difference in the horizontal angle φ C and the difference in the pitch angle θ C .
[0180] In this case, the horizontal width XW and the vertical width YW, and the horizontal viewing angle φ W and the vertical viewing angle θ W can be set on the reproduction side. At this time, the horizontal width XW and the vertical width YW, and the horizontal viewing angle φ W and the vertical viewing angle θ W can be automatically set in the reproduction side device or can be specified by the user.
[0181] In such an example, for example, when the content is a video and audio of a ball game, the coordinates XC and YC representing the position of the ball are set as the zoom area information, and the fixed or user-specified horizontal width XW and vertical width YW are used on the reproduction side device.
[0182] <Zoom area auxiliary information>
[0183] In addition, in the zoom area information header section ZHD, as the zoom area auxiliary information, supplementary information such as the ID representing the reproduction target device or the zoom purpose and other text information can be included.
[0184] In this case, in the zoom area information header section ZHD, for example, the zoom area information presence flag hasZoomAreaInfo and the zoom area auxiliary information are stored in the Figure 16 syntax shown.
[0185] In Figure 16 the example shown, the zoom area information presence flag hasZoomAreaInfo is arranged at the beginning, and in the case where the value of the zoom area information presence flag hasZoomAreaInfo is "1", each piece of information such as the zoom area auxiliary information is stored thereafter.
[0186] In other words, in this example, after the zoom area information presence flag hasZoomAreaInfo, the zoom area number information "numZoomAreas" indicating the number of zoom area information described in the zoom area information data section ZD is stored. Here, the value of the zoom area number information numZoomAreas is set to one of the values 0 to 15.
[0187] In addition, after the zoom area number information numZoomAreas, the information of each zoom area specified by the index idx corresponding to the number indicated by the zoom area number information numZoomAreas is arranged. Here, the index idx is set to each value from 0 to numZoomAreas.
[0188] In other words, "hasExtZoomAreaInfo[idx]" after the zoom area number information numZoomAreas represents an auxiliary information flag, and the auxiliary information flag indicates whether the zoom area auxiliary information of the zoom area specified by the index idx is stored. Here, the value of the auxiliary information flag hasExtZoomAreaInfo[idx] is set to one of "0" and "1".
[0189] In the case where the value of the auxiliary information flag hasExtZoomAreaInfo[idx] is "0", it indicates that the zoom area auxiliary information of the zoom area specified by the index idx is not stored in the zoom area information header section ZHD. On the contrary, in the case where the value of the auxiliary information flag hasExtZoomAreaInfo[idx] is "1", it indicates that the zoom area auxiliary information of the zoom area specified by the index idx is stored in the zoom area information header section ZHD.
[0190] When the value of the auxiliary information flag hasExtZoomAreaInfo[idx] is "1", after the auxiliary information flag hasExtZoomAreaInfo[idx], a specification ID representing the specification of the zoom area specified by the index idx is arranged.
[0191] In addition, "hasZoomAreaCommentary" represents a supplementary information flag, and the supplementary information flag indicates whether there is new supplementary information other than the specification ID for the zoom area specified by the index idx, such as text information including a description of the zoom area.
[0192] For example, when the value of this supplementary information flag hasZoomAreaCommentary is "0", it indicates that there is no supplementary information. On the contrary, when the value of this supplementary information flag hasZoomAreaCommentary is "1", it indicates that there is supplementary information, and after the supplementary information flag hasZoomAreaCommentary, "nbytes" as the number-of-bytes information and "ZoomAreaCommentary[idx]" as the supplementary information are arranged.
[0193] Here, the number-of-bytes information nbytes represents the number of bytes of the information of the supplementary information ZoomAreaCommentary[idx]. In addition, the supplementary information ZoomAreaCommentary[idx] is set to the text information describing the zoom area specified by the index idx.
[0194] More specifically, for example, assume that the content consists of a live video and its audio, and the zoom area specified by the index idx is a zoom area for continuously zooming in on a singer as a video object. In this case, for example, text information such as "Singer zoom" is set as the supplementary information ZoomAreaCommentary[idx].
[0195] In the zoom area information header section ZHD, as needed, the settings of the following items corresponding to the number indicated by the number of zoom areas used information numZoomAreas are stored: auxiliary information flag hasExtZoomAreaInfo[idx], ZoomAreaSpecifiedID[idx] as the specification ID, supplementary information flag hasZoomAreaCommentary, number of bytes information nbytes, and supplementary information ZoomAreaCommentary[idx]. However, for a zoom area whose value of the auxiliary information flag hasExtZoomAreaInfo[idx] is "0", ZoomAreaSpecifiedID[idx], supplementary information flag hasZoomAreaCommentary, number of bytes information nbytes, and supplementary information ZoomAreaCommentary[idx] are not stored. Similarly, for a zoom area whose value of the supplementary information flag hasZoomAreaCommentary is "0", the number of bytes information nbytes and the supplementary information ZoomAreaCommentary[idx] are not stored.
[0196] In addition, ZoomAreaSpecifiedID[idx] as the specification ID is information indicating a zoom specification such as a reproduction target device for a zoom area and a zoom purpose. And for example, as Figure 17 shown, a zoom specification is set for each value of ZoomAreaSpecifiedID[idx].
[0197] In this example, for instance, when the value of ZoomAreaSpecifiedID[idx] is "1", the zoom area representing the zoom specification indicated by the specification ID is a zoom area assuming the reproduction target device is a projector.
[0198] In addition, when the value of ZoomAreaSpecifiedID[idx] is from 2 to 4, these values respectively indicate that the zoom areas representing the zoom specifications indicated by the specification ID are zoom areas assuming the reproduction target device is a television receiver with a screen of over type 50, type 30 to 50, and less than type 30.
[0199] In this way, in the Figure 17 shown example, the zoom area information whose value of ZoomAreaSpecifiedID[idx] is one of "1" to "4" is information indicating the zoom areas set for each type of reproduction target device.
[0200] In addition, for example, when the value of ZoomAreaSpecifiedID[idx] is "7", it means that the zoom area of the zoom specification represented by the specification ID is a zoom area assuming that the reproduction target device is a smartphone and the rotation direction of the smartphone is the vertical direction.
[0201] Here, the rotation direction of the smartphone being the vertical direction means that when the user views the content using the smartphone, the direction of the smartphone is the vertical direction. That is, from the user's perspective, the longitudinal direction of the display screen of the smartphone is the vertical direction (the up / down direction). Therefore, for example, when the value of ZoomAreaSpecifiedID[idx] is "7", the zoom area is considered to be an area that is longer in the vertical direction.
[0202] In addition, for example, when the value of ZoomAreaSpecifiedID[idx] is "8", it means that the zoom area of the zoom specification represented by the specification ID is a zoom area assuming that the reproduction target device is a smartphone and the rotation direction of the smartphone is the horizontal direction. In this case, for example, the zoom area is considered to be an area that is longer in the horizontal direction.
[0203] In this way, in Figure 17 the example shown, the zoom area information whose value of ZoomAreaSpecifiedID[idx] is one of "5" to "8" is information representing the zoom area set for this type of reproduction target device and the rotation direction of the reproduction target device.
[0204] In addition, for example, when the value of ZoomAreaSpecifiedID[idx] is "9", it means that the zoom area of the zoom specification represented by the specification ID is a zoom area having a predetermined zoom purpose set by the content producer. Here, the predetermined zoom purpose is, for example, to display a specific zoom view, such as zooming in on a predetermined video object.
[0205] Therefore, for example, when the value "9" of ZoomAreaSpecifiedID[idx] represents a zoom specification for continuously zooming in on a singer, the supplementary information ZoomAreaCommentary[idx] of the index idx is set to text information such as "Singer zoom". The user can obtain the content of the zoom specification represented by each specification ID based on the specification ID or information related to the specification ID, supplementary information of the specification ID, etc.
[0206] In this way, in Figure 17In the example shown, each piece of zoom area information for which the value of ZoomAreaSpecifiedID[idx] is one of "9" to "15" represents information about an arbitrary zoom area freely set by the content producer side, such as a zoom area set for each specific video object.
[0207] As described above, by setting one or more zoom areas for one piece of content, for example, as Figure 18 shown, content that matches the user's preferences or content suitable for each playback device can be provided in a simplified manner.
[0208] In Figure 18 , image Q11 shows a video (image) of predetermined content. This content is the content of a live video, and image Q11 is a wide-angle image in which the live performers, namely singer M11, guitarist M12, and bassist M13, are projected, and the entire state, the audience, etc. are projected.
[0209] For image Q11 that constitutes such content, the content producer sets one or more zoom areas according to the zoom specifications or zoom purpose of the playback target device.
[0210] For example, in order to display a zoom view in which singer M11, who is a video object, is enlarged, when the area centered on singer M11 on image Q11 is set as the zoom area, image Q12 can be reproduced on the playback side.
[0211] Similarly, for example, in order to display a zoom view in which guitarist M12, who is a video object, is enlarged, when the area centered on guitarist M12 on image Q11 is set as the zoom area, the reproduced image Q13 can be reproduced as content on the playback side.
[0212] In addition, for example, by selecting multiple zoom areas on the playback side and aligning these zoom areas to form one screen, the reproduced image Q14 can be reproduced as content on the playback side.
[0213] In this example, image Q14 is composed of image Q21 with a zoom area having a slightly smaller viewing angle than image Q11, image Q22 with a zoom area in which singer M11 is enlarged, image Q23 with a zoom area in which guitarist M12 is enlarged, and image Q24 with a zoom area in which bassist M13 is enlarged. That is, image Q14 has a multi-view configuration. When multiple zoom areas are preset on the content provider side, on the content playback side, by selecting several zoom areas, the content can be reproduced in a multi-view configuration such as image Q14.
[0214] In addition, for example, when considering a reproduction device such as a tablet PC with a relatively small display screen and setting a viewing angle that is half of the viewing angle of Q11, that is, when setting a region with approximately half the area of the entire image Q11 including the center of the image Q11 as the zoom region, the image Q15 can be reproduced as content on the reproduction side. In this example, on a reproduction device with a relatively small display screen, each performer can also be displayed in a sufficiently large size.
[0215] In addition, for example, when considering a smart phone whose rotation direction is the horizontal direction, that is, the display screen is in a state where it is longer in the horizontal direction, and setting a relatively narrow region that is longer in the horizontal direction and includes the center of the image Q11 within the image Q11 as the zoom region, the image Q16 can be reproduced as content on the reproduction side.
[0216] For example, when considering a smart phone whose rotation direction is the vertical direction, that is, the display screen is in a state where it is longer in the vertical direction, and setting a region that is longer in the vertical direction near the center of the image Q11 as the zoom region, the image Q17 can be reproduced as content on the reproduction side.
[0217] In the image Q17, the singer M11, who is one of the performers, is enlarged and displayed. In this example, since a small display screen that is longer in the vertical direction is considered, instead of displaying all the performers arranged in the horizontal direction, it is appropriate for the reproduction target device to enlarge and display one performer, so such a zoom region is set.
[0218] In addition, for example, when considering a reproduction device with a relatively large display screen such as a large-size television receiver and setting the viewing angle to be slightly smaller than the viewing angle of the image Q11, that is, when setting a relatively large region within the image Q11 that includes the center of the image Q11 as the zoom region, the image Q18 can be reproduced as content on the reproduction side.
[0219] As described above, by setting a zoom region on the content provider side and generating encoded content data including zoom region information indicating the zoom region on the reproduction side, a user who is a viewer of the content can choose to directly reproduce the content or perform zoom reproduction, that is, cropping reproduction, based on the zoom region information.
[0220] Specifically, in the case where there is a plurality of pieces of zoom region information, the user can select zoom reproduction according to specific zoom region information among these plurality of pieces of zoom region information.
[0221] In addition, in the case where the zoom area auxiliary information is stored in the encoded content data, on the reproduction side, by referring to the reproduction target device, the purpose of zooming, and the zoom specifications such as the zoomed content and the auxiliary information, a zoom area suitable for the preferences of the reproduction device or the user can be selected. The selection of the zoom area can be specified by the user or can be automatically performed by the reproduction device
[0222] <Description of the encoding process>
[0223] Next, the specific operation of the encoding device 11 will be described.
[0224] When the video data, audio data, and metadata of the content are provided from the outside, the encoding device 11 performs an encoding process and outputs the encoded content data. Hereinafter, the encoding process performed by the encoding device 11 will be described with reference to Figure 19 the flowchart shown
[0225] In step S11, the video data encoding unit 21 encodes the video data of the provided content and provides the encoded video data obtained as a result to the multiplexer 24.
[0226] In step S12, the audio data encoding unit 22 encodes the audio data of the provided content and provides the encoded audio data obtained as a result to the multiplexer 24.
[0227] In step S13, the metadata encoding unit 23 encodes the metadata of the provided content and provides the encoded metadata obtained as a result to the multiplexer 24.
[0228] Herein, for example, the above-mentioned zoom area information is included in the metadata to be encoded. The zoom area information can be, for example, any information other than the information described with reference to Figures 5 to 10 、 Figure 13 and Figure 15 and so on.
[0229] In addition, the metadata encoding unit 23 also encodes the header information of the zoom area information such as the zoom area information presence flag hasZoomAreaInfo, the number of zoom areas information numZoomAreas, and the zoom area auxiliary information as needed, and provides the encoded header information to the multiplexer 24.
[0230] In step S14, the multiplexer 24 generates a bitstream by multiplexing the encoded video data provided from the video data encoding unit 21, the encoded audio data provided from the audio data encoding unit 22, and the encoded metadata provided from the metadata encoding unit 23, and supplies the generated bitstream to the output unit 25. At this time, the multiplexer 24 also stores the encoded header information of the scaling region information provided from the metadata encoding unit 23 in the bitstream.
[0231] Therefore, for example, Figure 2 the encoded content data shown can be obtained as a bitstream. Note that the configuration of the scaling region information header section ZHD of the encoded content data can be any configuration, such as Figure 4 , Figure 14 or Figure 16 the configurations shown.
[0232] In step S15, the output unit 25 outputs the bitstream provided from the multiplexer 24, and the encoding process ends.
[0233] As described above, the encoding device 11 encodes the metadata including the scaling region information together with the content to generate a bitstream.
[0234] In this way, by generating a bitstream including the scaling region information for specifying the scaling region without preparing the content for each reproduction device, it is possible to provide content that matches the user's preferences or is suitable for each reproduction device in a simplified manner.
[0235] In other words, the content producer can provide content that is considered optimal for the user's preferences, the screen size of the reproduction device, the rotation direction of the reproduction device, etc. in a simplified manner by only specifying the scaling region without preparing the content for each preference or each reproduction device.
[0236] In addition, on the reproduction side, by selecting the scaling region and cropping the content as needed, it is possible to view content that is optimal for the user's preferences, the screen size of the reproduction device, the rotation direction of the reproduction device, etc.
[0237] <Example of the configuration of the reproduction device>
[0238] Next, a reproduction device that receives the bitstream, i.e., the encoded content data, output from the encoding device 11 and reproduces the content will be described.
[0239] Figure 20 is a diagram showing an example of the configuration of a reproduction device according to an embodiment of the present technology.
[0240] In this example, according to requirements, a display device 52 that displays information when selecting a zoom region, a video output device 53 that outputs video content, and an audio output device 54 that outputs audio content are connected to a reproduction device 51.
[0241] Note that the display device 52, the video output device 53, and the audio output device 54 may be provided in the reproduction device 51. Additionally, the display device 52 and the video output device 53 may be the same device.
[0242] The reproduction device 51 includes: a content data decoding unit 61; a zoom region selection unit 62; a video data decoding unit 63; a video segmentation unit 64; an audio data decoding unit 65; and an audio conversion unit 66.
[0243] The content data decoding unit 61 receives the bitstream, i.e., the encoded content data, transmitted from the encoding device 11, and separates the encoded video data, the encoded audio data, and the encoded metadata from the encoded content data.
[0244] The content data decoding unit 61 provides the encoded video data to the video data decoding unit 63 and provides the encoded audio data to the audio data decoding unit 65.
[0245] The content data decoding unit 61 obtains metadata by decoding the encoded metadata and provides the obtained metadata to each unit of the reproduction device 51 as needed. Additionally, when the zoom region information is included in the metadata, the content data decoding unit 61 provides the zoom region information to the zoom region selection unit 62. Furthermore, when the zoom region auxiliary information is stored in the bitstream, the content data decoding unit 61 reads the zoom region auxiliary information, decodes the zoom region auxiliary information as needed, and provides the resulting zoom region auxiliary information to the zoom region selection unit 62.
[0246] The zoom region selection unit 62 selects one piece of zoom region information from one or more pieces of zoom region information provided by the content data decoding unit 61 and provides the selected zoom region information as the selected zoom region information to the video segmentation unit 64 and the audio conversion unit 66. In other words, in the zoom region selection unit 62, the zoom region is selected based on the zoom region information provided by the content data decoding unit 61.
[0247] For example, in the case where the zoom area assist information is provided from the content data decoding unit 61, the zoom area selection unit 62 provides the zoom area assist information to the display device 52 for display on the display device 52. In this way, for example, the following supplementary information is displayed on the display device 52 as the zoom area assist information, such as the purpose and content of the zoom area, the specification ID indicating the zoom specification such as the reproduction target device, the information based on the specification ID, and the text information.
[0248] Then, the user checks the zoom area assist information displayed on the display device 52 and selects a desired zoom area by operating an input unit (not shown in the figure). The zoom area selection unit 62 selects the zoom area based on the signal according to the user's operation provided from the input unit, and outputs the selected zoom area information indicating the selected zoom area. In other words, the zoom area information of the zoom area specified by the user is selected, and the selected zoom area information is output as the selected zoom area information.
[0249] Note that any method can be used to perform the selection of the zoom area. For example, the zoom area selection unit 62 generates information indicating the position and size of each zoom area based on the zoom area information, and displays this information on the display device 52, and the user selects the zoom area based on this display.
[0250] Note that in the case where the selection of the zoom area is not performed, that is, in the case of selecting to reproduce the original content, the selected zoom area information is set to information indicating no cropping or the like.
[0251] In addition, for example, in the case where the reproduction device 51 has previously recorded reproduction device information indicating the type of its own device such as a smart phone or a television receiver, the zoom area information (zoom area) can be selected by using the reproduction device information.
[0252] In this case, for example, the zoom area selection unit 62 obtains the reproduction device information and selects the zoom area information by using the obtained reproduction device information and the zoom area assist information.
[0253] More specifically, the zoom area selection unit 62 selects, as the zoom area assist information, the specification ID indicating that the reproduction target device is the type of device represented by the reproduction device information from the specification IDs. Then, the zoom area selection unit 62 sets the zoom area information corresponding to the selected specification ID, that is, the zoom area information whose index idx is the same as the selected specification ID, as the selected zoom area information.
[0254] In addition, for example, when the reproduction device 51 is a mobile device such as a smart phone or a tablet PC, the zoom area selection unit 62 can obtain direction information indicating the rotation direction of the reproduction device 51 from a gyro sensor (not shown in the figure), etc., and select zoom area information by using this direction information.
[0255] In this case, for example, the zoom area selection unit 62 selects a specification ID indicating that the reproduction target device is a device of the type represented by the reproduction device information, and the assumed rotation direction is the direction represented by the direction information obtained from the specification ID as the zoom area auxiliary information. Then, the zoom area selection unit 62 sets the zoom area information corresponding to the selected specification ID as the selected zoom area information. In this way, in both the state where the user uses the reproduction device 51 in the vertical direction (a screen that is longer in the vertical direction) and the state where the user uses the reproduction device 51 in the horizontal direction (a screen that is longer in the horizontal direction), the zoom area information that is optimal for the current state is selected.
[0256] Note that, in addition to this, either only the reproduction device information or the direction information can be used to select the zoom area information, or any other information related to the reproduction device 51 can be used to select the zoom area information.
[0257] The video data decoding unit 63 decodes the encoded video data provided from the content data decoding unit 61, and provides the video data obtained as a result to the video segmentation unit 64.
[0258] The video segmentation unit 64 crops (segments) the zoom area represented by the selected zoom area information provided from the zoom area selection unit 62 from the video (image) based on the video data provided from the video data decoding unit 63, and outputs the video data obtained as a result to the video output device 53.
[0259] Note that when the selected zoom area information is information indicating that no cropping is to be performed, the video segmentation unit 64 does not perform a cropping process on the video data, and directly outputs the video data to the video output device 53 as the zoom video data.
[0260] The audio data decoding unit 65 decodes the encoded audio data provided from the content data decoding unit 61, and provides the audio data obtained as a result to the audio conversion unit 66.
[0261] The audio conversion unit 66 performs an audio conversion process on the audio data provided from the audio data decoding unit 65 based on the selected zoom area information provided from the zoom area selection unit 62, and provides the zoom audio data obtained as a result to the audio output device 54.
[0262] Here, the audio conversion process is a conversion for audio reproduction suitable for scaling the video of the content.
[0263] For example, according to the cropping process of the scaling region, that is, the splitting process of the scaling region, the distance from the object inside the video to the reference viewing point changes. Therefore, for example, when the audio data is object-based audio, the audio conversion unit 66 converts the position information of the object provided from the content data decoding unit 61 through the audio data decoding unit 65 into metadata based on the selected scaling region information. In other words, the audio conversion unit 66 moves the position of the object as the sound source based on the selected scaling region information, that is, changes the distance from the object.
[0264] Then, the audio conversion unit 66 performs a rendering process based on the audio data in which the position of the object has been moved, and provides the scaled audio data obtained as a result to the audio output device 54 to reproduce the audio.
[0265] Note that such an audio conversion process is described in detail, for example, in PCT / JP2014 / 067508 and the like.
[0266] In addition, when the selected scaling region information is information indicating no cropping, the audio conversion unit 66 does not perform an audio conversion process on the audio data, and directly outputs the audio data as scaled audio data to the audio output device 54.
[0267] <Description of the reproduction process>
[0268] Subsequently, the operation of the reproduction device 51 will be described.
[0269] When receiving the encoded content data output from the encoding device 11, the reproduction device 51 performs a reproduction process in which the received encoded content data is decoded, and reproduces the content. Hereinafter, reference will be made to Figure 21 the flowchart shown to describe the reproduction process performed by the reproduction device 51.
[0270] In step S41, the content data decoding unit 61 separates the encoded video data, the encoded audio data, and the encoded metadata from the received encoded content data, and decodes the encoded metadata.
[0271] Then, the content data decoding unit 61 provides the encoded video data to the video data decoding unit 63, and provides the encoded audio data to the audio data decoding unit 65. In addition, the content data decoding unit 61 provides the metadata obtained by decoding to each unit of the reproduction device 51 as needed.
[0272] At this time, the content data decoding unit 61 provides the zoom area information obtained as metadata to the zoom area selection unit 62. Additionally, when the zoom area auxiliary information, which is the header information as metadata, is stored in the encoded content data, the content data decoding unit 61 reads the zoom area auxiliary information and provides the read zoom area auxiliary information to the zoom area selection unit 62. For example, as the zoom area auxiliary information, the above-mentioned supplementary information ZoomAreaCommentary[idx], ZoomAreaSpecifiedID[idx] as the specification ID, etc. are read.
[0273] In step S42, the zoom area selection unit 62 selects one piece of zoom area information from the zoom area information provided by the content data decoding unit 61, and provides the selected zoom area information to the video segmentation unit 64 and the audio conversion unit 66 according to the selection result.
[0274] For example, when the zoom area information is selected, the zoom area selection unit 62 provides the zoom area auxiliary information to the display device 52 for display on the display device 52, and selects the zoom area information based on the signal provided by the operation input of the user who has seen the display.
[0275] Additionally, as described above, the zoom area information can be selected by using not only the zoom area auxiliary information and the operation input from the user, but also the reproduction device information or the orientation information.
[0276] In step S43, the video data decoding unit 63 decodes the encoded video data provided by the content data decoding unit 61, and provides the video data obtained as a result to the video segmentation unit 64.
[0277] In step S44, the video segmentation unit 64 segments (cuts) the video based on the video data provided by the video data decoding unit 63 for the zoom area indicated by the selected zoom area information provided by the zoom area selection unit 62. In this way, the zoom video data for reproducing the video of the zoom area indicated by the selected zoom area information is obtained.
[0278] The video segmentation unit 64 provides the zoom video data obtained by segmentation to the video output device 53, thereby reproducing the video of the cropped content. The video output device 53 reproduces (displays) the video based on the zoom video data provided by the video segmentation unit 64.
[0279] In step S45, the audio data decoding unit 65 decodes the encoded audio data provided by the content data decoding unit 61, and provides the audio data obtained as a result to the audio conversion unit 66.
[0280] In step S46, the audio conversion unit 66 performs audio conversion processing on the audio data provided by the audio data decoding unit 65 based on the selected zoom area information provided by the zoom area selection unit 62. In addition, the audio conversion unit 66 provides the zoomed audio data obtained through the audio conversion processing to the audio output device 54, thereby outputting the audio. The audio output device 54 reproduces the audio of the content on which the audio conversion processing has been performed based on the zoomed audio data provided by the audio conversion unit 66, and the reproduction processing ends.
[0281] Note that, more specifically, the processes of steps S43 and S44 and the processes of steps S45 and S46 are executed in parallel with each other.
[0282] As described above, the reproduction device 51 selects appropriate zoom area information, performs cropping of the video data and audio conversion processing of the audio data based on the selected result according to the selected zoom area information, and reproduces the content.
[0283] In this way, by selecting the zoom area information, the content that has been appropriately cropped and has converted audio can be reproduced in a simplified manner, such as content that matches the user's preferences or is suitable for the size of the display screen of the reproduction device 51, the rotation direction of the reproduction device 51, etc. In addition, in the case where the user selects a zoom area based on the zoom area assist information presented by the display device 52, the user can select a desired zoom area in a simplified manner.
[0284] Note that, in the reproduction process described with reference to Figure 21 Although the case where both cropping of the video constituting the content and audio conversion processing of the audio constituting the content are performed based on the selected zoom area information is described, only one of them may be performed.
[0285] In addition, even in the case where the content consists only of video or audio, cropping or audio conversion processing is still performed on such video or audio, and the video or audio can be reproduced.
[0286] For example, even in the case where the content consists only of audio, by selecting the zoom area information indicating the area to be zoomed and changing the distance from the sound source object through audio conversion processing according to the selected zoom area information, reproduction of content suitable for the user's preferences, reproduction equipment, etc. can be achieved.
[0287] <Second Embodiment>
[0288] <Example of Configuration of Reproduction Device>
[0289] Note that, although the above describes an example in which the video segmentation unit 64 crops a scaled region from the video of the content according to one piece of selected scaled region information, it can be configured to select a plurality of scaled regions and output such a plurality of scaled regions in a multi-screen layout.
[0290] In this case, for example, the reproduction device 51 is configured as Figure 22 shown. Note that in Figure 22 , the same reference numerals are assigned to the parts corresponding to the case shown in Figure 20 , and their descriptions will not be presented appropriately.
[0291] Figure 22 The reproduction device 51 shown in
[0292] Figure 22 includes: a content data decoding unit 61; a scaled region selection unit 62; a video data decoding unit 63; a video segmentation unit 64; a video layout unit 91; an audio data decoding unit 65; and an audio conversion unit 66.
[0292] Figure 22 The configuration of the reproduction device 51 shown in Figure 20 differs from the reproduction device 51 shown in Figure 20 in that a video layout unit 91 is newly provided at a stage subsequent to the video segmentation unit 64, and is the same as the configuration of the reproduction device 51 shown in Figure 20 in other aspects.
[0293] In this example, the scaled region selection unit 62 selects one or more pieces of scaled region information and provides such scaled region information to the video segmentation unit 64 as selected scaled region information. In addition, the scaled region selection unit 62 selects one piece of scaled region information and provides the scaled region information to the audio conversion unit 66 as selected scaled region information.
[0294] Note that, as in the case of the reproduction device 51 shown in Figure 20 , the selection of the scaled region information performed by the scaled region selection unit 62 can be performed according to a user's input operation, or can be performed based on scaled region auxiliary information, reproduction device information, orientation information, etc.
[0295] Furthermore, the scaled region information that is the selected scaled region information provided to the audio conversion unit 66 can be selected according to a user's input operation, or can be the scaled region information arranged at a predetermined position, such as the start position of the encoded content data. In addition, the scaled region information can be the scaled region information of a representative scaled region, such as a scaled region having the largest size.
[0296] The video segmentation unit 64 crops a scaled region represented by each of one or more pieces of selected scaled region information provided from the scaled region selection unit 62 from a video (image) of the video based on the video data provided from the video data decoding unit 63, thereby generating scaled video data for each scaled region. In addition, the video segmentation unit 64 provides the scaled video data for each scaled region obtained by cropping to the video layout unit 91.
[0297] Note that the video segmentation unit 64 may directly provide the uncropped video data to the video layout unit 91 as one piece of scaled video data.
[0298] The video layout unit 91 generates multi-screen video data to be reproduced as a video based on such multi-screen video data arranged on a plurality of screens from one or more pieces of scaled video data provided from the video segmentation unit 64, and provides the generated multi-screen video data to the video output device 53. Here, the video to be reproduced based on the multi-screen video data is, for example, similar to Figure 18 The image Q14 shown is such a video: in which the videos (images) of the selected scaled regions are arranged to be aligned.
[0299] In addition, the audio conversion unit 66 performs audio conversion processing on the audio data provided from the audio data decoding unit 65 based on the selected scaled region information provided from the scaled region selection unit 62, and provides the scaled audio data obtained as a result to the audio output device 54 as the audio data of the representative audio for the multi-screen layout. In addition, the audio conversion unit 66 may directly provide the audio data provided from the audio data decoding unit 65 to the audio output device 54 as the audio data of the representative audio (scaled audio data).
[0300] <Description of the reproduction process>
[0301] Next, the reproduction process performed by the Figure 23 reproduction device 51 shown will be described with reference to the Figure 22 flowchart shown. Note that the process of step S71 is similar to the process of step S41 shown in Figure 21 so its description is omitted.
[0302] In step S72, the scaled region selection unit 62 selects one or more pieces of scaled region information from the scaled region information provided by the content data decoding unit 61, and provides the selected scaled region information to the video segmentation unit 64 according to the selection result.
[0303] Note that the process of selecting the scaled region information described here is basically similar to the process of step S42 shown in Figure 21 except for the difference in the number of the selected scaled region information.
[0304] In addition, the zoom area selection unit 62 selects the zoom area information of a representative zoom area from the zoom area information provided by the content data decoding unit 61, and provides the selected zoom area information to the audio conversion unit 66 according to the selection result. Here, the selected zoom area information provided to the audio conversion unit 66 is the same as one of the one or more pieces of selected zoom area information provided to the video segmentation unit 64.
[0305] When the zoom area information is selected, thereafter, the processes of steps S73 and S74 are executed, and the decoding of the encoded video data and the cropping of the video from the zoom area are performed. However, such processing is similar to Figure 21 the processing of steps S43 and S44 shown, so the description thereof is omitted. However, in step S74, for each of the one or more pieces of selected zoom area information, cropping (segmentation) of the zoom area represented by the selected zoom area information from the video based on the video data is performed, and the zoomed video data of each zoom area is provided to the video layout unit 91.
[0306] In step S75, the video layout unit 91 performs video layout processing based on the one or more pieces of zoomed video data provided by the video segmentation unit 64. In other words, the video layout unit 91 generates multi-screen video data based on the one or more pieces of zoomed video data, and provides the generated multi-screen video data to the video output device 53, thereby reproducing the video of each zoom area of the content. The video output device 53 reproduces (displays) the video arranged in multiple screens based on the multi-screen video data provided by the video layout unit 91. For example, in the case where multiple zoom areas are selected, the content is reproduced in a multi-screen configuration similar to Figure 18 the image Q14 shown.
[0307] When the video layout processing is executed, thereafter, the processes of steps S76 and S77 are executed, and the reproduction processing ends. However, such processing is similar to Figure 21 the processing of steps S45 and S46 shown, so the description thereof is omitted.
[0308] As described above, the reproduction device 51 selects one or more pieces of zoom area information, and based on the selection result, performs cropping of the video data and audio conversion processing of the audio data based on the selected zoom area information, and reproduces the content.
[0309] In this way, by selecting one or more pieces of zoom area information, appropriate content can be reproduced in a simplified manner, such as content that matches the user's preferences or content that is suitable for the size of the display screen of the reproduction device 51. In particular, when multiple pieces of zoom area information are selected, the content video can be reproduced in a multi-screen display that matches the user's preferences and the like.
[0310] In addition, when the user selects a zoom area based on the zoom area assist information presented by the display device 52, the user can select a desired zoom area in a simplified manner.
[0311] <Third Embodiment>
[0312] <Example of the Configuration of the Reproduction Device>
[0313] In addition, when the above content is sent via a network, the reproduction-side device can be configured to effectively receive only the data necessary for the reproduction of the selected zoom area. In this case, for example, the reproduction device is configured as Figure 24 shown. Note that in Figure 24 , the same reference numerals are assigned to the parts corresponding to the case shown in Figure 20 , and their descriptions will not be presented appropriately.
[0314] In Figure 24 the shown case, the reproduction device 121 for reproducing content receives the provision of the desired encoded video data and encoded audio data from the content data distribution server 122 that records the content and metadata. In other words, the content data distribution server 122 records the content and the metadata of the content in an encoded state or an unencoded state, and distributes the content in response to a request from the reproduction device 121.
[0315] In this example, the reproduction device 121 includes: a communication unit 131; a metadata decoding unit 132; a video / audio data decoding unit 133; a zoom area selection unit 62; a video data decoding unit 63; a video segmentation unit 64; an audio data decoding unit 65; and an audio conversion unit 66.
[0316] The communication unit 131 sends various types of data to the content data distribution server 122 via the network and receives various types of data from the content data distribution server 122.
[0317] For example, the communication unit 131 receives the encoded metadata from the content data distribution server 122 and provides the received encoded metadata to the metadata decoding unit 132, or receives the encoded video data and the encoded audio data from the content data distribution server 122 and provides the received data to the video / audio data decoding unit 133. Further, the communication unit 131 transmits the selected zoom area information provided by the zoom area selection unit 62 to the content data distribution server 122.
[0318] The metadata decoding unit 132 obtains metadata by decoding the encoded metadata provided by the communication unit 131, and provides the obtained metadata to each unit of the playback device 121 as needed.
[0319] In addition, when the zoom area information is included in the metadata, the metadata decoding unit 132 provides the zoom area information to the zoom area selection unit 62. Further, when the zoom area auxiliary information is received from the content data distribution server 122, the metadata decoding unit 132 provides the zoom area auxiliary information to the zoom area selection unit 62.
[0320] When the encoded video data and the encoded audio data are provided by the communication unit 131, the video / audio data decoding unit 133 provides the encoded video data to the video data decoding unit 63 and provides the encoded audio data to the audio data decoding unit 65.
[0321] <Description of the playback process>
[0322] Subsequently, the operation of the playback device 121 will be described.
[0323] The playback device 121 requests the content data distribution server 122 to send the encoded metadata. Then, when the encoded metadata is sent from the content data distribution server 122, the playback device 121 reproduces the content by performing the playback process. Hereinafter, reference will be made to Figure 25 the flowchart shown to describe the playback process performed by the playback device 121.
[0324] In step S101, the communication unit 131 receives the encoded metadata sent from the content data distribution server 122 and provides the received metadata to the metadata decoding unit 132. Note that more specifically, the communication unit 131 also receives the header information of the metadata such as the zoom area number information and the zoom area auxiliary information from the content data distribution server 122 as needed, and provides the received header information to the metadata decoding unit 132.
[0325] In step S102, the metadata decoding unit 132 decodes the encoded metadata provided by the communication unit 131, and provides the metadata obtained by decoding to each unit of the reproduction device 121 as needed. In addition, the metadata decoding unit 132 provides the scaling area information obtained as metadata to the scaling area selection unit 62, and also provides the scaling area auxiliary information to the scaling area selection unit 62 when the scaling area auxiliary information, which is the header information of the metadata, exists.
[0326] In this way, when the metadata is obtained, subsequently, the scaling area information is selected by performing the process of step S103. However, the process of step S103 is similar to Figure 21 the process of step S42 shown, so the description thereof is omitted. However, in step S103, the selected scaling area information obtained by selecting the scaling area information is provided to the video segmentation unit 64, the audio conversion unit 66, and the communication unit 131.
[0327] In step S104, the communication unit 131 sends the selected scaling area information provided by the scaling area selection unit 62 to the content data distribution server 122 via the network.
[0328] The content data distribution server 122 that has received the selected scaling area information performs cropping (segmentation) of the scaling area indicated by the selected scaling area information on the video data of the recorded content, thereby generating scaled video data. The scaled video data obtained in this way is video data for reproducing only the scaling area indicated by the selected scaling area information in the entire video of the original content.
[0329] The content data distribution server 122 sends the encoded video data obtained by encoding the scaled video data that constitutes the content and the encoded audio data obtained by encoding the audio data to the reproduction device 121.
[0330] Note that in the content data distribution server 122, the scaled video data for each scaling area can be prepared in advance. In addition, in the content data distribution server 122, regarding the audio data that constitutes the content, although usually all the audio data is encoded and the encoded audio data is output regardless of the selected scaling area, it can be configured to output only the encoded audio data of a part of the audio data. For example, when the audio data that constitutes the content is the audio data of each object, only the audio data of the object within the scaling area indicated by the selected scaling area information can be encoded and sent to the reproduction device 121.
[0331] In step S105, the communication unit 131 receives the encoded video data and the encoded audio data transmitted from the content data distribution server 122, and provides the encoded video data and the encoded audio data to the video / audio data decoding unit 133. Further, the video / audio data decoding unit 133 provides the encoded video data provided from the communication unit 131 to the video data decoding unit 63, and provides the encoded audio data provided from the communication unit 131 to the audio data decoding unit 65.
[0332] When the encoded video data and the encoded audio data are obtained, thereafter, the processes of steps S106 to S109 are executed, and the reproduction process ends. However, such a process is similar to Figure 21 the processes of steps S43 to S46 shown, and thus the description thereof is omitted.
[0333] However, since the signal obtained by decoding the encoded video data by the video data decoding unit 63 is the scaled video data that has already been cropped, the video segmentation unit 64 basically does not perform the cropping process. Only when additional cropping is required, the video segmentation unit 64 crops the scaled video data provided from the video data decoding unit 63 based on the selected scaled region information provided from the scaled region selection unit 62.
[0334] In this way, when the content is reproduced by the video output device 53 and the audio output device 54 based on the scaled video data and the scaled audio data, the content according to the selected scaled region is reproduced, such as the content shown in Figure 18 the figure.
[0335] As described above, the reproduction device 121 selects appropriate scaled region information, transmits the selected scaled region information to the content data distribution server 122 according to the selection result, and receives the encoded video data and the encoded audio data.
[0336] In this way, by receiving the encoded video data and the encoded audio data according to the selected scaled region information, the appropriate content can be reproduced in a simplified manner, such as the content that matches the user's preference or the content suitable for the size of the display screen of the reproduction device 121, the rotation direction of the reproduction device 121, etc. In addition, only the data that needs to be reproduced in the content can be effectively obtained.
[0337] <Fourth Embodiment>
[0338] <Example of Configuration of Reproduction Device>
[0339] In addition, examples of including zoom region information in the encoded content data have been described above. However, for example, the content may be cropped and reproduced based on the zoom region information publicly available on a network such as the Internet or the zoom region information recorded on a predetermined recording medium, that is, in a manner where the zoom region information is separate from the content. In such a case, for example, the cropping and reproduction may be performed by obtaining the zoom region information generated not only by the content producer but also by a third party other than the content producer, i.e., other users.
[0340] In this way, in the case of separately obtaining the content and the metadata including the zoom region information, for example, the reproduction device is configured as Figure 26 shown. Note that in Figure 26 , the same reference numerals are assigned to the parts corresponding to the case shown in Figure 20 , and their descriptions will not be presented appropriately.
[0341] Figure 26 The reproduction device 161 shown in
[0342] includes: a metadata decoding unit 171; a content data decoding unit 172; a zoom region selection unit 62; a video data decoding unit 63; a video segmentation unit 64; an audio data decoding unit 65; and an audio conversion unit 66.
[0343] The metadata decoding unit 171 obtains, for example, the encoded metadata of the metadata including the zoom region information from a device on the network, a recording medium connected to the reproduction device 161, etc., and decodes the obtained encoded metadata.
[0344] In addition, the metadata decoding unit 171 provides the metadata obtained by decoding the encoded metadata to each unit of the reproduction device 161 as needed, and provides the zoom region information included in the metadata to the zoom region selection unit 62. Further, the unit 171 obtains, as needed, the header information of the metadata such as zoom region auxiliary information together with the encoded metadata, and provides the obtained header information to the zoom region selection unit 62.
[0345] <Description of the reproduction process>
[0346] Subsequently, the operation of the reproduction device 161 will be described.
[0347] When reproduction of content is instructed, the reproduction device 161 performs a reproduction process of obtaining the encoded metadata and the encoded content, and reproduces the content. Hereinafter, the reproduction process performed by the reproduction device 161 will be described with reference to Figure 27 the flowchart shown.
[0348] In step S131, the metadata decoding unit 171 obtains the encoded metadata including the scaling area information, for example, from a device on the network, a recording medium connected to the reproduction device 161, etc. Note that the encoded metadata can be obtained in advance before the start of the reproduction process.
[0349] In step S132, the metadata decoding unit 171 decodes the obtained encoded metadata, and provides the metadata obtained as a result to each unit of the reproduction device 161 as needed. In addition, the metadata decoding unit 171 provides the scaling area information included in the metadata to the scaling area selection unit 62, and also provides the header information of the metadata such as the scaling area auxiliary information to the scaling area selection unit 62 as needed.
[0350] When the metadata is obtained by decoding, the process of step S133 is executed, and the scaling area information is selected. However, the process of step S133 is similar to Figure 21 the process of step S42 shown, so its description is omitted.
[0351] In step S134, the content data decoding unit 172 obtains the encoded video data and the encoded audio data of the content, for example, from a device on the network, a recording medium connected to the reproduction device 161, etc. In addition, the content data decoding unit 172 provides the obtained encoded video data to the video data decoding unit 63, and provides the obtained encoded audio data to the audio data decoding unit 65.
[0352] In this way, when the encoded video data and the encoded audio data of the content are obtained, thereafter, the processes of steps S135 to S138 are executed, and the reproduction process ends. However, such a process is similar to Figure 21 the processes of steps S43 to S46 shown, so its description is omitted.
[0353] As described above, the reproduction device 161 separately obtains the encoded video data and the encoded audio data of the content and the encoded metadata including the scaling area information. Then, the reproduction device 161 selects appropriate scaling area information, and based on the selection result, performs cropping of the video data and audio conversion processing of the audio data according to the selected scaling area information, and reproduces the content.
[0354] In this way, by separately obtaining the encoded metadata including the zoom area information from the encoded video data and the encoded audio data, it is possible to crop and reproduce the zoom areas set not only by the content producer but also by other users or the like.
[0355] Meanwhile, the above-described series of processes can be executed by hardware or software. In the case where the series of processes are executed by software, a program configuring the software is installed in a computer. Here, the computer includes a computer built in dedicated hardware, such as a general-purpose personal computer capable of executing various functions by installing various programs thereto.
[0356] Figure 28 FIG. is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by using a program.
[0357] In the computer, a central processing unit (CPU) 501, a read-only memory (ROM) 502, and a random access memory (RAM) 503 are interconnected via a bus 504.
[0358] In addition, an input / output interface 505 is connected to the bus 504. An input unit 506, an output unit 507, a recording unit 508, a communication unit 509, and a drive 510 are connected to the input / output interface 505.
[0359] The input unit 506 is configured of a keyboard, a mouse, a microphone, an imaging device, and the like. The output unit 507 is configured of a display, a speaker, and the like. The recording unit 508 is configured of a hard disk, a nonvolatile memory, and the like. The communication unit 509 is configured of a network interface, and the like. The drive 510 drives a removable medium 511 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0360] In the computer configured as described above, the CPU 501 loads, for example, a program recorded in the recording unit 508 into the RAM 503 via the input / output interface 505 and the bus 504, and executes the loaded program, thereby executing a series of processes.
[0361] For example, a program executed by the computer (CPU 501) can be provided in the form of being recorded on a removable medium 511 as a packaged medium or the like. In addition, the program can be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.
[0362] In a computer, by loading a removable medium 511 into a drive 510, a program can be installed into a recording unit 508 via an input / output interface 505. In addition, the program can be received by a communication unit 509 via a wired or wireless transmission medium and installed into the recording unit 508. Further, the program can also be pre-installed into a ROM 502 or the recording unit 508.
[0363] Note that a program executed by a computer can be a program that performs processing in a time series according to the order described in this specification, or a program that performs processing in parallel or at a necessary timing, such as when called.
[0364] In addition, embodiments of the present technology are not limited to the above embodiments, and various changes can be made without departing from the concept of the present technology.
[0365] For example, the present technology can adopt a cloud computing configuration in which a function is shared by multiple devices via a network and processed together by all devices.
[0366] In addition, each step described in the above flowchart can be executed not only by one device but also by multiple devices in a shared manner.
[0367] Furthermore, in the case where a step includes multiple processes, the multiple processes included in one step can be executed not only by one device but also by multiple devices in a shared manner.
[0368] In addition, the present technology can adopt the following configuration. [1]
[0370] A reproduction device, comprising:
[0371] A decoding unit that decodes encoded video data or encoded audio data;
[0372] A scaling region selection unit that selects one or more pieces of scaling region information from multiple pieces of scaling region information specifying a region to be scaled; and
[0373] A data processing unit that performs a cropping process on the video data obtained by decoding or an audio conversion process on the audio data obtained by decoding based on the selected scaling region information. [2]
[0375] The reproduction device according to claim 1, wherein, among the multiple pieces of scaling region information, there is included scaling region information specifying the region for each type of reproduction target device. [3]
[0377] The reproduction apparatus according to claim 1, wherein, among the plurality of scaling area information, the scaling area information includes an area that specifies a rotation direction for each reproduction target device. [4]
[0379] The reproduction apparatus according to claim 1, wherein, among the plurality of scaling area information, the scaling area information includes an area that specifies an area for each specific video object. [5]
[0381] The reproduction apparatus according to claim 1, wherein the scaling area selection unit selects the scaling area information according to a user operation input. [6]
[0383] The reproduction apparatus according to claim 1, wherein the scaling area selection unit selects the scaling area information based on information related to the reproduction apparatus. [7]
[0385] The reproduction apparatus according to claim 6, wherein the scaling area selection unit selects the scaling area information by using at least one of information indicating the type of the reproduction apparatus and information indicating the rotation direction of the reproduction apparatus as information related to the reproduction apparatus. [8]
[0387] A reproduction method, comprising the following steps:
[0388] Decoding the encoded video data or the encoded audio data;
[0389] Selecting one or more pieces of scaling area information from a plurality of pieces of scaling area information that specify areas to be scaled; and
[0390] Based on the selected scaling area information, performing a cropping process on the video data obtained by decoding or performing an audio conversion process on the audio data obtained by decoding. [9]
[0392] A program for causing a computer to execute a process including the following steps:
[0393] Decoding the encoded video data or the encoded audio data;
[0394] Selecting one or more pieces of scaling area information from a plurality of pieces of scaling area information that specify areas to be scaled; and
[0395] Based on the selected scaling region information, perform cropping processing on the video data obtained by decoding or perform audio conversion processing on the audio data obtained by decoding.
[10]
[0397] An encoding device, comprising:
[0398] An encoding unit that encodes video data or encodes audio data; and
[0399] A multiplexer that generates a bitstream by multiplexing the encoded video data or the encoded audio data with multiple pieces of scaling region information specifying the region to be scaled.
[11]
[0401] An encoding method, comprising the following steps:
[0402] Encode video data or encode audio data; and
[0403] Generate a bitstream by multiplexing the encoded video data or the encoded audio data with multiple pieces of scaling region information specifying the region to be scaled.
[12]
[0405] A program that causes a computer to execute a process including the following steps:
[0406] Encode video data or encode audio data; and
[0407] Generate a bitstream by multiplexing the encoded video data or the encoded audio data with multiple pieces of scaling region information specifying the region to be scaled.
[0408] List of reference numerals
[0409] 11 Encoding device
[0410] 21 Video data encoding unit
[0411] 22 Audio data encoding unit
[0412] 23 Metadata encoding unit
[0413] 24 Multiplexer
[0414] 25 Output unit
[0415] 51 Reproduction device
[0416] 61 Content data decoding unit
[0417] 62 Scaling region selection unit
[0418] 63 Video data decoding unit
[0419] 64 Video segmentation unit
[0420] 65 Audio data decoding unit
[0421] 66 Audio conversion unit
Claims
1. A reproduction apparatus, comprising: a decoding unit that decodes encoded content data to obtain: video data; audio data; and metadata including a plurality of pieces of scaling region information specifying regions to be scaled; a scaling region selection unit that selects one or more pieces of scaling region information from the plurality of pieces of scaling region information specifying regions to be scaled; and a data processing unit that performs a cropping process on the video data obtained by decoding based on the selected scaling region information, and performs an audio conversion process on the audio data obtained by decoding based on the position of a sound source object and the selected scaling region information, wherein the audio data is object-based audio; wherein the audio conversion process includes moving the position of the sound source object, wherein, among the plurality of pieces of scaling region information, there is scaling region information specifying regions for each type of reproduction target device, and wherein the scaling region selection unit selects the scaling region information based on information related to the type of the reproduction apparatus.
Citation Information
Patent Citations
Microsoft Corporation
CN102244807A
Video retargeting
US20090251594A1