Imaging device, recording device and display control device
By dividing HDR image data into multiple areas and encoding them using HEIF format with specific alignment constraints, the imaging device ensures efficient recording and playback of HDR images across various devices.
Patent Information
- Application Number
- JP2024148279
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-08-30
- Publication Date
- 2026-01-15
- Estimated Expiration
- 2038-02-16
AI Technical Summary
Existing technologies do not provide an optimal recording method for High Dynamic Range (HDR) images, which can result in compatibility issues and increased system scale when encoding large images.
The imaging device divides HDR image data into multiple areas, encoding each part using HEIF format, and records them with varying sizes based on the set recording mode, adhering to specific alignment constraints to enhance playback compatibility.
This approach allows for efficient recording and playback of HDR images, maintaining compatibility across different devices by optimizing the encoding process and adhering to alignment constraints.
Smart Images

Figure 0007799773000001 
Figure 0007799773000002 
Figure 0007799773000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an imaging device, a recording device, and a display control device. [Background technology]
[0002] An imaging device is known as an image processing device that compresses and encodes image data. Such an image processing device acquires a moving image signal using an imaging unit, compresses and encodes the acquired moving image signal, and records the compressed and encoded image file on a recording medium. Conventionally, image data before compression and encoding was expressed in Standard Dynamic Range (SDR), which has a maximum brightness level of 100 nits. However, in recent years, image data has been expressed in High Dynamic Range (HDR), which extends the maximum brightness level to approximately 10,000 nits, and has a brightness range close to the brightness range that humans can perceive.
[0003] Patent Document 1 describes an image data recording device that, when capturing and recording an HDR image, generates and records image data that allows the HDR image to be viewed in detail even on a device that does not support HDR. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Publication No. 2017-139618 Summary of the Invention [Problem to be solved by the invention]
[0005] Patent Document 1 describes recording HDR images, but does not consider the optimal recording method for recording HDR images.
[0006] Therefore, an object of the present invention is to provide a device that records images in a recording format suitable for recording and playback when recording images with a large amount of data, such as HDR images, and a display control device for playing back images recorded in that recording format. [Means for solving the problem]
[0007] In order to solve the above-mentioned problems, an imaging device of the present invention comprises: An imaging device that records HDR (High Dynamic Range) image data obtained by capturing an image, the imaging device comprising: an imaging sensor; an encoding unit that encodes the HDR image data captured by the imaging sensor; and a recording control unit that controls, when recording the HDR image data captured by the imaging sensor as a single image file in HEIF (High Efficiency Image File Format) format on a recording medium, to divide the HDR image data into a plurality of divided HDR image data and encode each of the divided HDR image data by the encoding unit, record the encoded plurality of divided HDR image data as image items in the image file, and record image structure information for combining the plurality of divided HDR image data into a state before the division as a derived image item in the image file; and a setting unit that sets a recording mode, wherein the vertical and horizontal sizes of the divided HDR image data differ depending on the vertical and horizontal sizes of the image corresponding to the recording mode set by the setting unit. [Effects of the Invention]
[0010] According to the present invention, it is possible to provide an imaging device that records images in a recording format suitable for recording and playback when recording images with a large amount of data and high resolution, and a display control device that plays back images recorded in that recording format. [Brief explanation of the drawings]
[0011] [Figure 1] FIG. 1 is a block diagram showing the configuration of an imaging device 100. [Figure 2] A diagram showing the structure of a HEIF file. [Figure 3]10 is a flowchart showing processing in HDR shooting mode. [Figure 4] 10 is a flowchart showing a process for determining a division method for image data in HDR shooting mode. [Figure 5] FIG. 10 is a diagram showing an encoding area and division method for image data when recording HDR image data. [Figure 6] 1 is a flowchart showing a process for constructing a HEIF file. [Figure 7] 10 is a flowchart showing a display process when HDR image data recorded as a HEIF file is played back. [Figure 8] 10 is a flowchart showing an overlay image creation process. [Figure 9] 10 is a flowchart showing a property acquisition process for an image item. [Figure 10] 10 is a flowchart showing a process of acquiring data of an image item. [Figure 11] 10 is a flowchart showing the process of obtaining the item ID of an image that constitutes a main image. [Figure 12] 10 is a flowchart showing image creation processing for one image item. [Figure 13] FIG. 10 is a diagram showing the positional relationship between an overlay image and divided images. DETAILED DESCRIPTION OF THE INVENTION
[0012] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS An embodiment of the present invention will be described in detail below with reference to the accompanying drawings, taking an image pickup device 100 as an example, but the present invention is not limited to the following embodiment.
[0013] <Configuration of imaging device> FIG. 1 is a block diagram showing an imaging device 100. As shown in FIG. 1, the imaging device 100 includes a CPU 101, a memory 102, a nonvolatile memory 103, an operation unit 104, an imaging unit 112, an image processing unit 113, an encoding processing unit 114, a display control unit 115, and a display unit 116. The imaging device 100 also includes a communication control unit 117, a communication unit 118, a recording medium control unit 119, and an internal bus 130. The imaging device 100 forms an optical image of a subject on a pixel array of the imaging unit 112 using a photographing lens 111. The photographing lens 111 may be detachable or non-detachable from the body (housing, main body) of the imaging device 100. The imaging device 100 writes and reads image data to and from a recording medium 120 via the recording medium control unit 119. The recording medium 120 may be detachable or non-detachable from the imaging device 100.
[0014] The CPU 101 executes a computer program stored in the nonvolatile memory 103 to control the operation of each unit (each functional block) of the imaging device 100 via the internal bus .
[0015] The memory 102 is a rewritable volatile memory. The memory 102 temporarily stores computer programs for controlling the operation of each unit of the imaging device 100, information such as parameters related to the operation of each unit of the imaging device 100, information received by the communication control unit 117, etc. The memory 102 also temporarily stores images acquired by the imaging unit 112, and images and information processed by the image processing unit 113, encoding processing unit 114, etc. The memory 102 has a storage capacity sufficient for temporarily storing these.
[0016] The nonvolatile memory 103 is an electrically erasable and recordable memory, such as an EEPROM, etc. The nonvolatile memory 103 stores computer programs that control the operation of each unit of the image capture device 100 and information such as parameters related to the operation of each unit of the image capture device 100. The various operations performed by the image capture device 100 are realized by the computer programs.
[0017] The operation unit 104 provides a user interface for operating the imaging device 100. The operation unit 104 includes various buttons such as a power button, a menu button, and a shooting button, and the various buttons are configured as switches, a touch panel, or the like. The CPU 101 controls the imaging device 100 in accordance with user instructions input via the operation unit 104. Note that, although the description here has been given taking as an example a case where the CPU 101 controls the imaging device 100 based on operations input via the operation unit 104, the present invention is not limited to this. For example, the CPU 101 may control the imaging device 100 based on a request input via the communication unit 118 from a remote controller (not shown), a mobile terminal (not shown), or the like.
[0018] The photographing lens (lens unit) 111 is composed of a group of lenses (not shown) including a zoom lens, a focus lens, etc., a lens control unit (not shown), an aperture (not shown), etc. The photographing lens 111 can function as a zoom unit that changes the angle of view. The lens control unit adjusts the focus and controls the aperture value (F-number) based on a control signal transmitted from the CPU 101. The imaging unit 112 can function as an acquisition unit that sequentially acquires multiple images that constitute a moving image. The imaging unit 112 may be, for example, an area image sensor such as a CCD (charge-coupled device) or a CMOS (complementary metal-oxide semiconductor) element. The imaging unit 112 has a pixel array (not shown) in which photoelectric conversion units (not shown) that convert an optical image of a subject into an electrical signal are arranged in a matrix, i.e., two-dimensionally. An optical image of the subject is formed on the pixel array by the photographing lens 111. The imaging unit 112 outputs the captured image to the image processing unit 113 or the memory 102. The imaging unit 112 can also acquire still images.
[0019] The image processing unit 113 performs predetermined image processing on image data output from the imaging unit 112 or image data read from the memory 102. Examples of the image processing include interpolation, reduction (resizing), and color conversion. The image processing unit 113 also performs predetermined arithmetic processing for exposure control, distance measurement control, and the like, using the image data acquired by the imaging unit 112. The exposure control, distance measurement control, and the like are performed by the CPU 101 based on the calculation results obtained by the arithmetic processing by the image processing unit 113. Specifically, the CPU 101 performs AE (automatic exposure) processing, AWB (auto white balance) processing, AF (auto focus) processing, and the like.
[0020] The encoding processing unit 114 compresses the size of the image data by performing intra-frame predictive coding (intra-screen predictive coding), inter-frame predictive coding (inter-screen predictive coding), or the like on the image data. The encoding processing unit 114 is, for example, a coding device configured with a semiconductor element or the like. The encoding processing unit 114 may be a coding device provided outside the imaging device 100. The encoding processing unit 114 performs encoding processing using, for example, the H.265 (ITU H.265 or ISO / IEC23008-2) method.
[0021] The display control unit 115 controls the display unit 116. The display unit 116 has a display screen (not shown). The display control unit 115 performs resizing, color conversion, and other processes on image data to generate an image that can be displayed on the display screen of the display unit 116, and outputs the image, i.e., an image signal, to the display unit 116. The display unit 116 displays an image on the display screen based on the image signal sent from the display control unit 115. The display unit 116 has an OSD (On Screen Display) function, which is a function for displaying a setting screen such as a menu on the display screen. The display control unit 115 can superimpose an OSD image on the image signal and output the image signal to the display unit 116. The display unit 116 is configured with a liquid crystal display, an organic EL display, or the like, and displays the image signal sent from the display control unit 115. The display unit 116 may be, for example, a touch panel. If the display unit 116 is a touch panel, the display unit 116 can also function as the operation unit 104.
[0022] The communication control unit 117 is controlled by the CPU 101. The communication control unit 117 generates a modulated signal conforming to a predetermined wireless communication standard such as IEEE 802.11 and outputs the modulated signal to the communication unit 118. The communication control unit 117 also receives the modulated signal conforming to the wireless communication standard via the communication unit 118, decodes the received modulated signal, and outputs a signal corresponding to the decoded signal to the CPU 101. The communication control unit 117 includes a register for storing communication settings. The communication control unit 117 can adjust transmission and reception sensitivity during communication under control of the CPU 101. The communication control unit 117 can transmit and receive using a predetermined modulation method. The communication unit 118 outputs the modulated signal supplied from the communication control unit 117 to an external device 127, such as an information communication device, outside the imaging device 100, and also includes an antenna for receiving the modulated signal from the external device 127. The communication unit 118 also includes a communication circuit and the like. Although the case where wireless communication is performed by communication unit 118 has been described as an example here, communication performed by communication unit 118 is not limited to wireless communication. For example, communication unit 118 and external device 127 may be connected by an electrical connection using wiring or the like.
[0023] The recording medium control unit 119 controls the recording medium 120. Based on a request from the CPU 101, the recording medium control unit 119 outputs a control signal for controlling the recording medium 120 to the recording medium 120. For example, a nonvolatile memory or a magnetic disk is used as the recording medium 120. As described above, the recording medium 120 may be removable or non-removable. The recording medium 120 records encoded image data, etc. The image data, etc. are saved as files in a format compatible with the file system of the recording medium 120. Examples of files include MP4 files (ISO / IEC 14496-14:2003) and MXF (Material eXchange Format) files. The functional blocks 101 to 104, 112 to 115, 117, and 119 are mutually accessible via an internal bus 130.
[0024] Here, the normal operation of the imaging device 100 of this embodiment will be described.
[0025] When a user operates the power button on the operation unit 104, a start-up instruction is issued from the operation unit 104 to the CPU 101 of the imaging device 100. In response to this instruction, the CPU 101 controls a power supply unit (not shown) to supply power to each block of the imaging device 100. When power is supplied, the CPU 101 checks, for example, which mode a mode selector switch on the operation unit 104 is in, such as a still image capture mode or a playback mode, based on an instruction signal from the operation unit 102.
[0026] In the normal still image shooting mode, the imaging device 100 performs shooting processing when the user operates the still image recording button on the operation unit 104 while the device is in a shooting standby state. In the shooting processing, image data of a still image shot by the imaging unit 112 is subjected to image processing by the image processing unit 113 and encoding processing by the encoding processing unit 114, and the encoded image data is recorded as an image file by the recording medium control unit 119 on the recording medium 120. Note that in the shooting standby state, a live view image is displayed by shooting an image at a predetermined frame rate by the imaging unit 112, performing image processing for display by the image processing unit, and displaying the image on the display unit 116 by the display control unit 115.
[0027] In playback mode, an image file recorded on the recording medium 120 is read by the recording medium control unit 119, and the image data of the read image file is decoded by the encoding processing unit 114. In other words, the encoding processing unit 114 also has the function of a decoder. Then, the image processing unit 113 performs processing for display, and the display control unit 115 causes the image to be displayed on the display unit 116.
[0028] Normal still image capture and playback are performed as described above, but the imaging device of this embodiment has an HDR shooting mode for capturing not only normal still images but also HDR still images, and can also play back captured HDR still images.
[0029] The process of capturing and playing back an HDR still image will be described below.
[0030] <File structure> First, we will explain the file structure when recording HDR still images.
[0031] Recently, a still image file format called High Efficiency Image File Format (hereinafter referred to as HEIF) was established (ISO / IEC 23008-12:2017).
[0032] Compared to conventional still image file formats such as JPEG, it has the following features: This file format conforms to the ISO Base Media File Format (ISOBMFF) (ISO / IEC 14496-14:2003). -Can store multiple still images, not just a single one. -Can store still images compressed in compression formats used to compress video, such as HEVC / H.265 and AVC / H.264.
[0033] In this embodiment, HEIF is used as a recording file for HDR still images.
[0034] First, we will explain the data stored in HEIF.
[0035] HEIF manages each piece of stored data in units called items.
[0036] In addition to the data itself, each item has an item ID (item_ID) that is an integer value that is unique within the file, and an item type (item_type) that indicates the type of item.
[0037] Items can be divided into image items, whose data represents an image, and metadata items, whose data is metadata.
[0038] Image items include coded image items, which are image data that has been coded, and derived image items, which represent an image that is the result of manipulating one or more other image items.
[0039] An example of a derived image item is an overlay image, which is an overlay type derived image item. This is an overlay image that is the result of arranging any number of image items at any positions and compositing them using the ImageOverlay structure (overlay information).
[0040] An example of a metadata item is Exif data.
[0041] As mentioned above, HEIF can store multiple image items.
[0042] If there are relationships between multiple images, those relationships can be described.
[0043] Examples of relationships between multiple images include the relationship between a derived image item and its constituent image items, the relationship between a main image and a thumbnail image, and so on.
[0044] The relationship between image items and metadata items can also be described in a similar manner.
[0045] The HEIF format is based on the ISOBMFF format, so we will first provide a brief explanation of ISOBMFF.
[0046] The ISOBMFF format manages data in a structure called a box.
[0047] A box is a data structure that begins with a 4-byte data length field and a 4-byte data type field, followed by data of any length.
[0048] The structure of the data section is determined by the data type. The ISOBMFF and HEIF specifications specify several data types and the structures of their data sections.
[0049] A box may also contain other boxes in its data. In other words, boxes can be nested. A box nested in the data part of a box is called a sub-box.
[0050] A box that is not a sub-box is called a file-level box, which is a box that can be accessed sequentially from the beginning of the file.
[0051] Using Figure 2, we will explain HEIF format files.
[0052] First, file-level boxes will be described.
[0053] The file type box, whose data type is 'ftyp', stores information about file compatibility. ISOBMFF-compliant file specifications declare the file structure and data stored in the file using a 4-byte code called a brand, and store these in the file type box. By placing the file type box at the beginning of the file, a file reader can determine the file structure by checking the contents of the file type box without having to further read and interpret the file contents.
[0054] The HEIF specification uses the brand 'mif1' to represent the file structure, and if the encoded image data stored is HEVC compressed data, it is branded as 'heic' or 'heix' according to the HEVC compression profile.
[0055] The metadata box, with a data type of 'meta', contains various sub-boxes that store data about each item, as detailed below.
[0056] The media data box with a data type of 'mdat' stores data for each item, such as coded image data for coded image items and Exif data for metadata items.
[0057] Next, the sub-boxes of the metadata box will be described.
[0058] A handler box with a data type of 'hdlr' stores information that indicates the type of data managed by the metadata box. In the HEIF specification, the handler_type of a handler box is 'pict'.
[0059] A data information box whose data type is 'dinf' specifies the file in which the data targeted by this file exists. In ISOBMFF, it is possible for the data targeted by a file to be stored in a file other than the file itself. In this case, the data reference in the data information box contains a data entry URL box that describes the URL of the file in which the data exists. If the targeted data exists in the same file, a data entry URL box containing only a flag indicating this is stored.
[0060] The primary item box, whose data type is 'pitm', stores the item ID of the image item that represents the main image.
[0061] The item information box whose data type is 'iinf' is a box for storing the following item information entries.
[0062] An item information entry with a data type of 'infe' stores the item ID, item type, and flags for each item.
[0063] The item type of an image item whose encoded image data is HEVC compressed data is 'hvc1', and it is an image item.
[0064] The item type of the overlay type derived image item, that is, the ImageOverlay structure (overlay information), is 'iovl' and is classified as an image item.
[0065] The item type of an Exif format metadata item is 'Exif' and is a metadata item.
[0066] Furthermore, if the lowest bit of the flag field of an image item is set, it can be specified that the image item is a hidden image. When this flag is set, the image item is not treated as a display target during playback and is hidden.
[0067] An item reference box with a data type of 'iref' stores the reference relationship between each item, including the type of reference relationship, the item ID of the referencing item, and the item IDs of one or more referenced items.
[0068] For a derived image item in overlay format, the reference type is set to 'dimg', the item ID of the derived image item is set to the item ID of the referencing item, and the item IDs of each image item that makes up the overlay are set to the item ID of the referenced item.
[0069] In the case of thumbnail images, the reference type is set to 'thmb', the item ID of the thumbnail image is set to the item ID of the referencing item, and the item ID of the original image is set to the item ID of the referenced item.
[0070] An item property box whose data type is 'iprp' is a box for storing the following item property container box and item property association box.
[0071] An item property container box with a data type of 'ipco' is a box that stores individual property data boxes.
[0072] Each image item can have property data that represents the characteristics and attributes of the image.
[0073] The property data box includes the following:
[0074] Decoder configuration and initialization data (for HEVC, the type is 'hvcC') is data used to initialize the decoder. HEVC parameter set data (VideoParameterSet, SequenceParameterSet, PictureParameterSet) is stored here.
[0075] Image spatial extents (type 'ispe') are the dimensions (width, height) of the image.
[0076] Color information (type 'colr') is the color space information of the image.
[0077] Image rotation information (type 'irot') is the direction of rotation when rotating and displaying an image.
[0078] The pixel information of an image (type 'pixi') is information that indicates the number of bits of data that make up the image.
[0079] In addition, in this embodiment, as HDR metadata, mastering display color volume information (type 'MDCV') and content light level information (type 'CLLI') are stored as property data boxes.
[0080] There are other properties besides those mentioned above, but they are not listed here.
[0081] The Item Property Association Box, of data type 'ipma', stores each item-property association in the form of an item ID and an array of indexes in 'ipco' of the associated properties.
[0082] An item data box whose data type is 'idat' is a box that stores item data with a small data size.
[0083] The data for an overlay-format derived image item can store an ImageOverlay structure (overlay information) in 'idat', which contains information such as the position of the constituent images. The overlay information includes the canvas_fill_value parameter, which is background color information, and the output_width and output_height parameters, which indicate the final size of the composited image when composited using overlay. Furthermore, each composite element image has horizontal_offset and vertical_offset parameters, which indicate the horizontal and vertical position coordinates of its placement within the composite image. By using these parameters included in the overlay information, it is possible to create a composite image in which multiple images are placed at any position within a single image with a specified background color and size. The item location box, whose data type is 'iloc', stores the position information for each item's data in the form of an offset reference (construction_method), an offset value from the offset reference, and a length.
[0084] The offset reference is the beginning of the file, or 'idat'.
[0085] There are other metadata boxes besides those mentioned above, but they will not be specified here.
[0086] Figure 2 shows a structure in which two coded image items compose an overlay format image. The number of coded image items that compose an overlay format image is not limited to two. As the number of coded image items that compose an overlay format image increases, the following boxes and items also increase accordingly. -Added an item information entry for coded image items to the item information box. · The item ID of the encoded image item increases in the reference destination item ID of the reference relationship type 'dimg' of the item reference box. · In the item property container box, such as decoder configuration - initialization data, image space range, etc. · In the item property related box, items of the index of the encoded image item and its related properties are added. · In the item location box, an item of the position information of the encoded image item is added. · Encoded image data is added to the media data box.
[0087] <Shooting of HDR Image> Next, the processing in the imaging device 100 when shooting and recording an HDR image will be described.
[0088] The width and height of the image to be recorded have been increasing recently. However, if an overly large size is encoded, compatibility may be lost during decoding on other models, or the scale of the system may increase. Specifically, in this embodiment, an example using H.265 for encoding and decoding will be described. H.265 has parameters such as Profile and Level as standards, and these parameters change as the image encoding method and the image size increase. These values are parameters for determining whether playback is possible on the device during decoding, and there may be cases where playback is rejected as not possible after determining these parameters. In this embodiment, when recording a large image size, the method of dividing it into multiple image display areas and using the overlay method advocated by the HEIF standard to store and manage their encoded data in a HEIF file will be described. By dividing it into multiple image display areas in this way, the size of each image display area becomes smaller and the playback compatibility with other models is enhanced. Hereinafter, it will be described with the size of one side of the encoding area being set to 4096 or less. Note that the encoding format may be other than H.265, or the size of one side may be defined as other than 4096.
[0089] Next, we will explain alignment constraints regarding the start position of the coding area, the width and height of the coding area, and the start position of the playback area, as well as the width and height of the playback area. When encoding a divided image, imaging devices generally have hardware constraints such as vertical and horizontal coding start alignment of the coding area, and playback width and height alignment during playback. In other words, it is necessary to calculate and encode the coding start position of each of one or more coding areas, the start position, width, and height of the playback area. Since imaging devices switch between multiple image sizes for recording in response to user instructions, it is necessary to switch between them for each image size.
[0090] The encoding processing unit 114 described in this embodiment has alignment constraints on the start position in 64-pixel units in the horizontal direction (hereinafter, x-direction) and 1-pixel units in the vertical direction (hereinafter, y-direction) during encoding. The width and height alignment constraints for encoding are also constrained to 32-pixel units and 16-pixel units, respectively. The start position of the playback area is constrained to 2-pixel units in the x direction and 1-pixel units in the y direction. The width and height alignment constraints of the playback area are also constrained to 2-pixel units and 1-pixel units, respectively. Note that although examples of alignment constraints have been described here, other alignment constraints may generally be used.
[0091] Next, the flow from capturing HDR image data in HDR recording mode (HDR shooting mode) to recording the captured HDR image data as an HEIF file will be described with reference to Fig. 3. In HDR mode, the image capture unit 112 and image processing unit 113 perform processing to express the HDR color space using the BT.2020 color gamut, which indicates HDR, and a PQ gamma curve. Recording for HDR will be omitted in this example.
[0092] In S101, CPU 101 determines whether or not the user has pressed SW2. When the user captures a subject and presses SW2 (YES in S101), CPU 101 detects that SW2 has been pressed and proceeds to S102. CPU 101 continues this detection until SW2 is pressed (NO in S101).
[0093] In S102, CPU 101 determines how many parts to divide the encoding area vertically and horizontally (number of divisions N) and how to divide the area in order to save the image with the angle of view captured by the user in a HEIF file. The result is stored in memory 102. S102 will be described later with reference to Figure 4. Proceed to S103.
[0094] In S103, the CPU 101 initializes a variable M in the memory 102 to M=1. Then, the process proceeds to S104.
[0095] In S104, the CPU 101 monitors whether the loop has been performed the number of times equal to the division number N determined in S102. If the number of loops has reached the division number N (YES in S104), the process proceeds to S109; otherwise (NO in S104), the process proceeds to S105.
[0096] In S105, the CPU 101 reads out information about the coding start area and playback area of the divided image M from the memory 102. The process then proceeds to S106.
[0097] In S106, the CPU 101 notifies the encoding processing unit 114 which area of the entire image stored in the memory 102 is to be encoded, and causes the encoding processing unit 114 to encode the area. The process then proceeds to S107.
[0098] In S107, the CPU 101 temporarily stores in the memory 102 the coded data resulting from the coding performed by the coding processing unit 114 and accompanying information generated during coding. Specifically, in the case of H.265, this refers to H.265 standard information such as VPS, SPS, and PPS that must later be stored in the HEIF file. While not directly relevant to this embodiment and therefore not described here, the H.265 VPS, SPS, and PPS represent various pieces of information necessary for decoding, such as the coded data size, bit depth, display / non-display area designation, and frame information. The process then proceeds to S108.
[0099] In S108, the CPU 101 increments the variable M and returns to S104.
[0100] In S109, the CPU 101 generates metadata to be stored in the HEIF file. The CPU 101 extracts information required for playback from the imaging unit 112, image processing unit 113, encoding processing unit 114, etc., and temporarily stores it in the memory 102 as metadata. Specifically, this is data formatted in Exif format. Details of Exif data are already common knowledge, so a detailed explanation will be omitted in this embodiment. Proceed to S110.
[0101] In S110, the CPU 101 encodes the thumbnail. The image processing unit 113 reduces the entire image in the memory 102 to the thumbnail size and temporarily stores it in the memory 102. The CPU 101 instructs the encoding processing unit 114 to encode this thumbnail image in the memory 102. Here, the thumbnail is encoded at a size that satisfies the alignment constraints mentioned above, but a method in which the encoding area and the playback area are different may also be used. This thumbnail is also encoded using H.265. Proceed to S111.
[0102] In S111, the CPU 101 temporarily stores the coded data resulting from the coding of the thumbnail by the coding processing unit 114 and accompanying information generated during coding in the memory 102. As explained above in S107, this refers to H.265 standard information such as VPS, SPS, and PPS. Next, the process proceeds to S112.
[0103] In S112, the CPU 101 has saved various data in the steps up to this point in memory 102. The CPU 101 combines the data saved in memory 102 in order to complete a HEIF file, and saves it in memory 102. The flow for completing this HEIF file will be explained later using Figure 6. Proceed to S113.
[0104] In S113, the CPU 101 instructs the recording medium control unit 119 to write the HEIF file in the memory 102 to the recording medium 120. Then, the process returns to S101.
[0105] After the image is captured in these steps, the HEIF file is recorded on the recording medium 120.
[0106] Next, the division method determination in S102 explained above will be explained using Figures 4 and 5(a). Below, the upper left corner of the sensor image area is taken as the origin (0, 0), and the coordinates are represented as (x, y), with the width and height represented as w and h. Furthermore, the area will be expressed as [x, y, w, h], combining the starting coordinates (x, y) and the width and height (w, h).
[0107] In S121, the CPU 101 obtains the recording mode for shooting from the memory 102. The recording mode determines the recording method, which determines, for example, the image size, image aspect ratio, image compression rate, etc., and the user selects one of these recording modes to shoot. For example, this could be a setting such as L size or 3:2 aspect ratio. Since the shooting angle of view differs depending on the recording mode, the shooting angle of view is determined. The process proceeds to S122.
[0108] In S122, the start coordinate position (x0, y0) of the playback image is acquired from the sensor image area [0, 0, H, V]. Next, the process proceeds to S123.
[0109] In S123, the area [x0, y0, Hp, Vp] of the reproduced image is obtained from the memory 102. Next, the process proceeds to S124.
[0110] In S124, it is necessary to encode the data so that it includes the area of the reproduced image. As mentioned earlier, the alignment of the code start position is 64 pixels in the x direction, so if you want to encode the data so that it includes the area of the reproduced image, you must encode it from the position (Ex0, Ey0). In this way, the size of Lex is determined. If the offset of this Lex is 0 (NO in S124), proceed to S126; if necessary (YES in S124), proceed to S125.
[0111] In S125, Lex is calculated. As mentioned above, Lex is determined by the alignment of the playback image area and the encoding start position. CPU 101 determines Lex as follows: x0 is divided by the horizontal pixel alignment of the encoding start position, which is 64 pixels, and the quotient is multiplied by the previous 64 pixels. Since the calculation does not include a remainder, the coordinate that is obtained is located to the left of the start position of the playback image area. Therefore, Lex is calculated as the difference between the x coordinate of the playback image area and the x coordinate of the encoding start position calculated in the previous calculation. Proceed to S126.
[0112] In S126, the coordinates of the final position of the right edge of the top line of the reproduced image are calculated. These are set as (xN, y0). xN is calculated by adding Hp to x0. Then the process proceeds to S127.
[0113] In S127, the right edge of the coding area is found so that it includes the right edge of the reproduced image. To find the right edge w of the coding area, an alignment constraint on the coding width is required. As mentioned earlier, the coding width and height are each constrained to be multiples of 32 pixels. Based on this constraint and the relationship with (xN, y0), the CPU 101 calculates whether an offset of Rex is required beyond the right edge of the reproduced image. If Rex is required (YES in S127), proceed to S128; if not (NO in S127), proceed to S129.
[0114] In S128, the CPU 101 adds Rex to the right of the right edge coordinates (xN, y0) of the reproduced image so as to align 32 pixels, and determines the end position (ExN, Ey0) of encoding.Then the process proceeds to S129.
[0115] In S129, the CPU 101 obtains the coordinates of the bottom right corner of the display area. Since the size of the reproduced image is obvious, these coordinates are (xN, yN). The process proceeds to S130.
[0116] In S130, the bottom edge of the encoding area is determined so that it includes the bottom edge of the reproduced image. To determine the bottom edge of the encoding area, an alignment constraint on the encoding height is required. As mentioned earlier, the encoding height constraint is a multiple of 16 pixels. Based on the relationship between this constraint and the bottom right coordinates (xN, yN) of the reproduced image, the CPU 101 calculates whether a Vex offset is required beyond the bottom edge of the reproduced image. If Vex is required (YES in S130), proceed to S131; if not (NO in S130), proceed to S132.
[0117] In S131, the CPU 101 calculates Vex based on the relationship between the above constraints and the bottom right coordinate (xN, yN) of the reproduced image, and determines the bottom right coordinate (ExN, EyN) of the encoding area. Vex is calculated by setting it to a position that is a multiple of 16 pixels, which is the encoding height alignment, so that it includes yN from y0, the vertical encoding start position. Specifically, assuming y0 = 0, Vp is divided by 16 pixels, which is the encoding height alignment, and the quotient and remainder are determined. Since the encoding area must be determined so as to include the reproduced area, if there is a remainder, 1 is added to the quotient and multiplied by the previous 16 pixels. This obtains EyN, the y-coordinate of the bottom edge of the encoding area that includes Vp, which is a multiple of 16 pixels. Vex is calculated by subtracting EyN from yN. Note that EyN is a value obtained by offsetting yN downward by Vex, and ExN is the value already calculated in S128. The process proceeds to S132.
[0118] In S132, CPU 101 calculates the size of the coding area, Hp'xVp', as follows: Hp' is calculated by adding Lex and Rex calculated from the alignment constraint to the horizontal size Hp. Vp' is also calculated by adding Vex calculated from the alignment constraint to the vertical size Vp. The size of the coding area, Hp'xVp', is calculated from Hp' and Vp'. The process proceeds to S133.
[0119] In S133, the CPU 101 determines whether the horizontal size Hp' to be coded exceeds 4096 pixels to be divided, and if it does (YES in S133), the process proceeds to S134, and if it does not (NO in S133), the process proceeds to S135.
[0120] In S134, the horizontal size Hp' to be coded is divided into two or more regions so that the size is 4096 pixels or less. For example, if it is 6000 pixels, it is divided into two regions, and if it is 9000 pixels, it is divided into three regions.
[0121] When dividing into two, the image is divided at an approximate center position that satisfies the 32 pixel alignment for the encoding width mentioned earlier. If the divisions are not equal, the left divided image is made larger. When dividing into three, if the divisions are not equal, the image is divided into divided areas A, B, and C while paying attention to the alignment for the encoding width mentioned earlier, so that divided areas A and B are equal in size, and divided area C is slightly smaller. Even when dividing into three or more areas, the division positions and sizes are determined using the same algorithm. Proceed to S135.
[0122] In S135, the CPU 101 determines whether the vertical size Vp' to be coded exceeds 4096 pixels to be divided, and if it does (YES in S135), proceeds to S136, and if it does not (NO in S135), ends the process.
[0123] In S136, Vp' is divided into a plurality of divided regions in the same manner as in S134, and the process ends.
[0124] An example of the division in S102 will be explained using specific numerical values with reference to Figures 5(b), (c), (d), and (e). Figure 5 shows not only the division method of the image data, but also the relationship between the encoding area and the playback area.
[0125] Let's consider Figure 5(b). Figure 5(b) shows a case where the image is divided into two regions, left and right, and Lex, Rex, and Vex are assigned to the left, right, and bottom edges of the playback region. First, the playback image region, [256+48, 0, 4864, 3242], is the region of the final image to be recorded by the imaging device. The coordinates of the upper left corner of this playback image region are (256+48, 0), which is not a multiple of the 64-pixel unit that is the encoding start alignment mentioned earlier. Therefore, Lex is set 48 pixels to the left of this coordinate, and the left edge of the encoding start position is at x=256. Next, the width of the playback region is w=48+4864, but this is not a multiple of 32 pixels, the encoding width alignment. In consideration of the width constraint alignment during encoding, 16 pixels of Rex are assigned to the right edge of the playback image region. Similarly, the vertical direction is calculated to be Vex=4 pixels. Because the horizontal size of this reproduced image exceeds 4096, it is divided horizontally into two parts approximately at the center, taking into account the alignment of the encoding width. As a result, the encoding area of divided image 1 is [256, 0, 2496, 3242+8], and the encoding image area of divided image 2 is [256+2496, 0, 2432, 3242+8].
[0126] Let us now explain Figure 5(c). Figure 5(c) shows a case where the image is divided into two regions, left and right, and Rex and Vex are assigned to the right and bottom edges of the playback region. Since the x position of the playback image region is x=0, it also coincides with the 64th pixel, which is the starting alignment of the encoding region, so Lex=0. From here on, the calculation method is the same as described above, and as a result, the encoding region of divided image 1 is obtained as [0, 0, 2240, 32456+8], and the encoding image region of divided image 2 is obtained as [2240, 0, 2144, 2456+8].
[0127] Let us now consider Figure 5(d). Figure 5(d) shows a case where the image is divided into two regions, left and right, and Lex and Rex are assigned to the left and right of the playback region. The vertical size of the playback image region is h=3648, which is a multiple of 16 pixels, the image height constraint during encoding, so Vex=0. The calculation method is the same as described above, and as a result, the encoding region of segmented image 1 is determined to be [256+48, 0, 2240, 3648], and the encoding image region of segmented image 2 is determined to be [256+2240, 0, 2144, 3648].
[0128] Now, let us consider Figure 5(d). Figure 5(d) shows a case where an image is divided into four regions, top, bottom, left, and right, and Lex and Rex are assigned to the left and right of the playback region. In this case, the coding region exceeds 4096 in both the vertical and horizontal directions, so the division is performed in a square pattern. The calculation method is the same as described above, and as a result, the coding region of divided image 1 is [0, 0, 4032, 3008], the coding image region of divided image 2 is [4032, 0, 4032, 3008], the coding image region of divided image 3 is [0, 3008, 4032, 2902], and the coding image region of divided image 4 is [4032, 3008, 4032, 2902].
[0129] In this way, the CPU 101 determines the area of the reproduced image according to the recording mode, and must calculate the divided areas for each. Alternatively, these calculations can be performed in advance and the results can be stored in the memory 102. The CPU 101 can also read information about the divided areas and reproduction areas according to the recording mode from the memory 102 and set it in the encoding processing unit 114.
[0130] When a captured image is divided into multiple split images and recorded as in this embodiment, multiple display areas must be combined to restore the captured image from the recorded image. When generating a combined image, it is more efficient if there are no overlapping areas between the split images, since this eliminates the need to encode or record extra data. Therefore, in this embodiment, each playback area and each encoding area is set so that there are no overlapping areas (overlapping areas) at the boundaries where the split images meet. However, overlapping areas may be provided between the split images depending on conditions such as hardware alignment constraints.
[0131] Next, the method of constructing an HEIF file mentioned above in S112 will be described with reference to Fig. 6. Since an HEIF file has the structure shown in Fig. 2, the CPU 101 constructs the file in memory 102 in order from the beginning.
[0132] In S141, an 'ftyp' box is created. This box is placed at the beginning of the file to determine file compatibility. CPU 101 stores 'heic' in the major brand of this box, and 'mif1' and 'heic' in the compatible brands. In this embodiment, these values are used, but other values may also be used. Proceed to S142.
[0133] In S142, a 'meta' box is generated. This box stores the multiple boxes described below. After generating this box, the CPU 101 proceeds to S143.
[0134] In S143, an 'hdlr' box is generated to be stored in the 'iinf' box. This box indicates the attributes of the 'meta' box mentioned earlier. CPU 101 stores 'pict' in this type and proceeds to S144. Note that 'pict' is information indicating the type of data managed by the metadata box, and the HEIF specification stipulates that the handler_type of the handler box is 'pict'.
[0135] In S144, a 'dinf' box is generated to be stored in the 'iinf' box. This box indicates the location of the data that this file targets. After generating this box, CPU 101 proceeds to S145.
[0136] In S145, a 'pitm' box is generated to be stored in the 'iinf' box. This box stores the item ID of the image item representing the main image. Since the main image is an overlay image synthesized by overlay, CPU 101 stores the overlay information item ID as the item ID of the main image. Then, proceed to S146.
[0137] In S146, an 'iinf' box is generated to be stored in the 'meta' box. This box is used to manage a list of items. Here, an initial value is entered in the data length field, 'iinf' is saved in the data type field, and the process proceeds to S147.
[0138] In S147, an 'infe' box is generated to be stored in the 'iinf' box. This 'infe' box is a box for registering item information for each item stored in the file, and an 'infe' box is generated for each item. For split images, each split image is registered as a single item in this box. In addition, overlay information for composing the main image from multiple split images, Exif information, and thumbnail images are each registered as items. In this case, as described above, the overlay information, split images, and thumbnail images are registered as image items. A flag indicating a hidden image can be added to an image item, and adding this flag prevents the image from being displayed during playback. In other words, adding this hidden image flag to an image item allows you to specify an image to be hidden during playback. Therefore, a flag indicating a hidden image is set for split images, and the flag is not set for overlay information items. As a result, the individual split images themselves are not displayed, and only the image resulting from the overlay composition of multiple split images using the overlay information is displayed. An 'infe' box is created for each item individually and stored in the previous 'iinf'. Proceed to S148.
[0139] In S148, an 'iref' box is generated to be stored in the 'iinf' box. This box stores information indicating the relationship between the image (main image) configured as an overlay and the divided images that make up that image. Because the division method was determined in S102 described above, CPU 101 generates this box based on the determined division method. Proceed to S149.
[0140] In S149, an 'iprp' box is generated to be stored in the 'iinf' box. This box stores the properties of the item, and stores the 'ipco' box generated in S150 and the 'ipma' box generated in S151. Proceed to S150.
[0141] In S150, an 'ipco' box is generated to be stored in the 'iprp' box. This box is a property container box for the item and stores various properties. There are multiple properties, and CPU 101 generates the following property container boxes and stores them in the 'ipco' box. Some properties are generated for each image item, while others are generated commonly to multiple image items. A 'colr' box is generated as information common to the overlay image composed of overlay information for the main image and the divided images. The 'colr' box stores color information such as the HDR gamma curve as color space information for the main image (overlay image) and the divided images. For the divided images, an 'hvcC' box and an 'ispe' box are generated, respectively. In S107 and S111 described above, CPU 101 reads the accompanying information from the encoding stored in memory 102 and generates an 'hvcC' property container box. The associated information stored in this 'hvcC' includes not only information about when the coding area was coded and the size (width, height) of the coding area, but also the size (width, height) of the playback area and the position of the playback area within the coding area. This associated information is Golomb-compressed and recorded in the property container box of 'hvcC'. Golomb compression is a common compression method, so a detailed explanation is omitted here. The 'ispe' box also stores the size (width, height) of the playback area of the divided image (e.g., 2240x2450). An 'irot' box stores information indicating the rotation of the overlay image (the main image), and a 'pixi' box indicates the number of bits of the image data. Separate 'pixi' boxes may be generated for the main image (overlay image) and divided images; however, in this embodiment, only one box is generated because the overlay image and divided images have the same number of bits, i.e., 10 bits. 'CLLI' and 'MDCV' boxes, which store HDR supplementary information, are also generated.Furthermore, as properties of the thumbnail image, aside from the main image (overlay image), a 'colr' box storing color space information, an 'hvcC' box storing encoding information, an 'ispe' box storing image size, a 'pixi' box storing information on the number of bits in the image data, and a 'CLLI' box storing supplementary information on HDR are generated. These property container boxes are generated and stored in the 'ipco' box. Proceed to S151.
[0142] In S151, an 'ipma' box is generated. This box indicates the relationship between items and properties, and indicates which of the properties mentioned above each item is related to. The CPU 101 determines the relationship between items and properties from various data stored in the memory 102, and generates this box. The process proceeds to S152.
[0143] In S152, an 'idat' box is generated. This box stores overlay information that indicates how the playback area of each divided image is arranged to generate an overlay image. The overlay information includes a canvas_fill_value parameter, which is background color information, and output_width and output_height parameters, which indicate the overall size of the overlay image. Additionally, for each divided image that will become a composite element, horizontal_offset and vertical_offset parameters indicate the horizontal and vertical position coordinates for combining the divided images. The CPU 101 enters this information into these parameters based on the division method determined in S102. Specifically, the CPU 101 enters the overall size of the playback area into the output_width and output_height parameters, which indicate the size of the overlay image. The horizontal_offset and vertical_offset parameters, which indicate the position information of each divided image, respectively, enter the offset values in the width and height directions from the starting coordinate position (x0, y0) of the upper left corner of the playback area to the upper left corner of the divided image. By generating overlay information in this manner, when an image is played back, the divided images are positioned based on the horizontal_offset and vertical_offset parameters and then combined, allowing the image to be played back in its original state before being split. In this embodiment, since no overlapping areas are provided in the divided images, position information is written so that the divided images are positioned so that they do not overlap. Then, by specifying the area to be displayed using the output_width and output_height parameters from the image obtained by combining the divided images, it is possible to play back the image with only the playback area as the display target. An 'idat' box is generated that stores the overlay information generated in this manner. Proceed to S153.
[0144] In S153, an 'iloc' box is generated. This box indicates where in the file various data will be placed. Since various information is stored in memory 102, this box is generated from the size of this information. Specifically, information indicating the overlay is stored in the aforementioned 'idat', and this is position and size information within this 'idat'. Furthermore, thumbnail data and code data 12 are stored in the 'mdat' box, and this position and size information is also stored. Proceed to S154.
[0145] In S154, an 'mdat' box is generated. This box contains multiple boxes described below. After generating the 'mdat' box, the CPU 101 proceeds to S155.
[0146] In S155, the Exif data is stored in the 'mdat' box. Since the Exif metadata was saved in the memory 102 in the previous S109, the CPU 101 reads it out of the memory 102 and adds it to the 'mdat' box. The process proceeds to S156.
[0147] In S156, the thumbnail data is stored in the 'mdat' box. Since the thumbnail data was saved in the memory 102 in S110, the CPU 101 reads it from the memory 102 and adds it to the 'mdat' box. The process proceeds to S157.
[0148] In S157, the data for encoded image 1 is stored in the 'mdat' box. Since the data for divided image 1 was saved in memory 102 during the first loop in S106, CPU 101 reads this data from memory 102 and adds it to the 'mdat' box. This process is repeated up to encoded image N, and all encoded images 1-N are added to the 'mdat' box. Through these steps, CPU 101 constructs a HEIF file.
[0149] Through the above processing, an HDR image can be recorded on the recording medium 120 in the HEIF format as shown in FIG.
[0150] <Playback of HDR Images> Next, the processing in the imaging device 100 when playing back an HDR image file recorded in the HEIF format on the recording medium 120 will be described. In this embodiment, the case of playing back on the imaging device 100 will be described, but in an image processing device without an imaging unit, when playing back an HDR image recorded in the HEIF format on the recording medium 120, similar processing may be realized.
[0151] Using FIG. 7, the playback (display) processing of HEIF when the main image is an overlay image will be described.
[0152] In step S701, the CPU 101 uses the recording medium control unit 119 to read the head portion of the specified file existing on the recording medium 120 into the memory 102. Then, it checks whether a file type box with a correct structure exists in the head portion of the read file, and further checks whether'mif1' representing HEIF exists in the brand therein.
[0153] Also, when a brand corresponding to a unique file structure is recorded, the existence check of that brand is performed. As long as that brand guarantees a specific structure, by checking this brand, some subsequent structure checks, for example, step S703 and step S704, can be omitted. <00In step S704, CPU 101 searches for a data information box from the metadata box read in step S702 and checks the structure. In this embodiment, it is assumed that data exists in the same file, so it checks whether a flag indicating this is set in the data entry URL box.
[0157] In step S705, the CPU 101 searches for a primary item box from the metadata box read in step S702, and acquires the item ID of the main image.
[0158] In step S706, CPU 101 searches for an item information box from the metadata box read in step S702 and obtains an item information entry corresponding to the item ID of the main image obtained in step S705. In the case of an HEIF-format image file recorded by the HDR image shooting process described above, overlay information is specified as the main image. In other words, the item information entry corresponding to the item ID of the main image has an item type of 'iovl', which indicates overlay.
[0159] In step S707, CPU 101 executes an overlay image creation process, which will be described later with reference to FIG.
[0160] In step S708, the CPU 101 displays the overlay image created in step S707 on the display unit 116 via the display control unit 115.
[0161] When displaying the created overlay image, it may be necessary to specify the image's color space information. The image's color space information is specified by the color_primaries of the color space property 'colr', where the color gamut is specified, and the transfer characteristics (equivalent to gamma) are specified by transfer_characteristics, so these values are used. For example, an HDR image might have a color gamut of Rec.ITU-R BT.2020 and transfer characteristics of Rec.ITU-R BT.2100(PQ).
[0162] Furthermore, if HDR metadata such as 'MDCV' or 'CLLI' exists in the item property container box, it is also possible to use it.
[0163] Next, the overlay image creation process will be described with reference to FIG.
[0164] In step S801, CPU 101 executes a process for acquiring properties of an overlay image, which will be described later with reference to FIG.
[0165] In step S802, CPU 101 executes a process for acquiring overlay information for the overlay image. Since the overlay information is recorded in the 'idat' box, the data stored in the 'idat' box is acquired.
[0166] The overlay information includes the image size of the overlay image (output_width, output_height parameters), position information (horizontal_offset, vertical_offset) of the individual image items (divided image items) that make up the overlay image, and the like.
[0167] Overlay information is handled as image item data for an image item whose item type is overlay 'iovl'.
[0168] This acquisition process will be described later with reference to FIG.
[0169] In step S803, the CPU 101 executes a process for obtaining item IDs for all divided image items, which will be described later with reference to FIG.
[0170] In step S804, CPU 101 uses the overlay image size in the overlay information acquired in step S802 to allocate memory for storing image data of that size. This memory area will be called a canvas.
[0171] In step S805, the CPU 101 initializes a divided image counter n to 0.
[0172] In step S806, CPU 101 checks whether divided image counter n is equal to the number of divided image items. If they are equal, the process proceeds to step S2210. If they are not equal, the process proceeds to step S807.
[0173] In step S807, the CPU 101 executes image creation processing for one image item for the n-th divided image item, as will be described later with reference to FIG.
[0174] In step S808, CPU 101 places the image (divided image) of the n-th divided image item created in step S807 on the canvas secured in step S804 in accordance with the overlay information acquired in step S802.
[0175] In an overlay, the divided images that make up the overlay can be placed anywhere on the canvas, and the position is described in the overlay information. However, the area outside the overlay image is not displayed. Therefore, when placing a divided image on the canvas, only the area of the coded area of the divided image that overlaps with the area of the overlay image, i.e., the playback area of the divided image, is placed.
[0176] The following description will be given with reference to FIG.
[0177] The coordinate system has the upper left corner of the overlay image as the origin (0,0), with the X coordinate increasing to the right and the Y coordinate increasing downward.
[0178] Let the size of the overlay image be width Wo and height Ho (0 < Wo, 0 < Ho).
[0179] Therefore, the overlay image has top - left coordinates (0, 0) and bottom - right coordinates (Wo - 1, Ho - 1).
[0180] Also, let the size of the divided image be width Wn and height Hn (0 < Wn, 0 < Hn).
[0181] Let the top - left position of the divided image be (Xn, Yn).
[0182] Therefore, the divided image has top - left coordinates (Xn, Yn) and bottom - right coordinates (Xn + Wn - 1, Yn + Hn - 1).
[0183] The overlapping area between the divided image and the overlay image (canvas) can be obtained by the following method.
[0184] In the following cases, there is no overlapping area. Wo - 1 < Xn (the left end of the divided image is to the right of the right end of the overlay image) Xn + Wn - 1 < 0 (the right end of the divided image is to the left of the left end of the overlay image) Ho - 1 < Yn (the upper end of the divided image is below the lower end of the overlay image) Yn + Hn - 1 < 0 (the lower end of the divided image is above the upper end of the overlay image) In this case, this image is not a processing target.
[0185] Also, in the following case, the entire divided image becomes the overlapping area. 0 <= Xn and (Xn + Wn - 1) <= (Wo - 1) and 0 <= Yn and (Yn + Hn - 1) <= (Ho - 1) In this case, place the entire divided image at the specified position (Xn, Yn) on the canvas.
[0186] In cases other than the above, a part of the divided image is the placement target. The upper left coordinate of the overlapping area is (Xl, Yt), and the lower right coordinate is (Xr, Yb). The leftmost Xl is as follows. if(0<=Xn) Xl=Xn; else Xl=0; The rightmost Xr is as follows: if(Xn+Wn-1<=Wo-1) Xr=Xn+Wn-1; else Xr=Wo-1; The upper limit Yt is as follows: if(0<=Yn) Yt=Yn; else Yt=0; The lower limit Yb is as follows: if(Yn+Hn-1<=Ho-1) Yb = Yn + Hn - 1; else Yb=Ho-1; The size of this overlapping area is Xr-Xl+1 in width and Yb-Yt+1 in height.
[0187] As mentioned above, the upper left coordinates (Xl, Yt) and lower right coordinates (Xr, Yb) of this overlapping area are in a coordinate system with the upper left corner of the canvas as the origin (0,0).
[0188] The top left coordinates (Xl, Yt) of the overlapping area can be expressed as follows in a coordinate system with the top left of the divided image as the origin: Xl' = Xl - Xn; Yt' = Yt - Yn;
[0189] To summarize, the following is the case.
[0190] A rectangle with width (Xr-Xl+1) and height (Yb-Yt+1) is cut out from the top left of the divided image at a distance (Xl', Yt') from the top left and placed at position (Xl, Yt) on the canvas.
[0191] This allows the playback area of the divided image to be positioned at an appropriate position on the canvas.
[0192] In step S809, the CPU 101 adds 1 to the divided image counter n, and the process returns to step S806.
[0193] In step S810, CPU 101 checks whether a rotation property 'irot' exists among the properties of the overlay image acquired in step S801, and if so, checks the rotation angle. If the rotation property does not exist, or if the rotation property exists but the rotation angle is 0, processing ends. If the rotation property exists and the rotation angle is other than 0, the process proceeds to step S811.
[0194] In step S811, CPU 101 rotates the canvas image created in steps S806 to S809 by the angle acquired in step S810, and sets the rotated image as a created image.
[0195] This allows for the creation of overlay images.
[0196] In the above description, the processes of steps S806 to S809 are repeated in a loop for the number of divided images, but in an environment where parallel processing is possible, the processes of steps S806 to S809 may be processed in parallel for the number of divided images.
[0197] Next, the property acquisition process for an image item will be described with reference to FIG.
[0198] In step S901, CPU 101 searches for the entry of the specified image item ID from the item property related box in the item property box in the metadata box, and obtains the array of property indexes stored therein.
[0199] In step S902, the CPU 101 initializes an array counter n to zero.
[0200] In step S903, CPU 101 checks whether array counter n is equal to the number of array elements. If they are equal, the process ends. If they are not equal, the process proceeds to step S904.
[0201] In step S904, CPU 101 obtains the property at the index of the nth element in the array from the item property container box in the item property box in the metadata box.
[0202] In step S905, the CPU 101 adds 1 to the array counter n, and the process returns to step S903.
[0203] Next, the image item data acquisition process will be described with reference to FIG.
[0204] In step S1001, the CPU 101 searches for an entry of the specified image item ID from the item location box in the metadata box, and obtains the offset reference (construction_method), offset, and length.
[0205] In step S1002, CPU 101 checks the offset reference acquired in step S1001. A value of 0 for the offset reference indicates an offset from the beginning of the file, and a value of 1 indicates an offset within the item data box. If the offset reference is 0, the process proceeds to step S1003. If the offset reference is 1, the process proceeds to step S1004.
[0206] In step S1003, CPU 101 reads length bytes from the offset byte position from the beginning of the file into memory 102.
[0207] In step S1004, CPU 101 reads length bytes from the offset byte position from the beginning of the data portion of the item data box in the metadata box into memory 102.
[0208] Next, the process of obtaining the item ID of the image that constitutes the main image will be described with reference to FIG.
[0209] In step S1101, CPU 101 searches the item reference box in the metadata box for an entry whose reference type is 'ding' and whose referencing item ID is the item ID of the main image.
[0210] In step S1102, CPU 101 acquires the array of reference destination item IDs of the entry acquired in step S1101.
[0211] Next, the image creation process for one encoded image item will be described with reference to FIG.
[0212] In step S1201, CPU 101 acquires the properties of the image item, as explained above with reference to FIG.
[0213] In step S1202, CPU 101 acquires the image item data, as explained above with reference to Fig. 10. The image item data is encoded image data.
[0214] In step S1203, the CPU 101 initializes the decoder using the decoder configuration and initialization data from the properties acquired in step S1201.
[0215] In step S1204, the CPU 101 decodes the coded image data acquired in step S1202 using a decoder, and acquires the decoding result.
[0216] In step S1205, CPU 101 checks whether pixel format conversion is required.
[0217] If it is not necessary, the process ends. If it is necessary, the process proceeds to step S1206.
[0218] If the pixel format of the decoder output data is different from the pixel format of the image supported by the display device, pixel format conversion is required.
[0219] For example, if the pixel format of the decoder output data is YCbCr (luminance-chrominance) format and the pixel format of the image on the display device is RGB format, pixel format conversion from YCbCr format to RGB format is required. Also, pixel format conversion is required if the same YCbCr format has a different bit depth (8 bits, 10 bits, etc.) or chrominance samples (4:2:0, 4:2:2, etc.).
[0220] Note that the matrix_coefficients of the color space information property 'colr' specifies the coefficients used when converting from RGB format to YCbCr format, so to convert from YCbCr format to RGB format, the reciprocal of those coefficients can be used.
[0221] In step S1206, the CPU 101 converts the decoder output data acquired in step S1204 into a desired pixel format.
[0222] This allows for the creation of an image of the specified coded image item.
[0223] Through the above processing, it is possible to play back HDR images recorded on the recording medium 120 in the HEIF format as shown in FIG.
[0224] <Other embodiments> The present invention has been described above based on the embodiments, but the present invention is not limited to these embodiments, and various forms within the scope of the gist of the present invention are also included in the present invention.
[0225] The functions of the above-described embodiments may be implemented as a control method, which may be executed by an image processing device. Alternatively, a program having the functions of the above-described embodiments may be implemented as a control program, which may be executed by a computer included in the image processing device. The control program may be recorded, for example, on a computer-readable storage medium.
[0226] The present invention can be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and by having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.
Claims
1. An imaging device that records HDR (high dynamic range) image data obtained by shooting, an imaging sensor; an encoding unit that encodes HDR image data obtained by capturing an image with the imaging sensor; a recording control means for controlling, when recording HDR image data obtained by photographing with the imaging sensor as a single image file in HEIF (High Efficiency Image File Format) format, the HDR image data to be divided into a plurality of divided HDR image data and encoded by the encoding unit, the plurality of encoded divided HDR image data to be recorded as image items in the image file, and image structure information for combining the plurality of divided HDR image data to a state before division to be recorded as a derived image item in the image file; a setting means for setting a recording mode, the vertical and horizontal sizes of the divided HDR image data differ depending on the vertical and horizontal sizes of the image corresponding to the recording mode set by the setting means; An imaging device characterized by:
2. the setting means allows a user to select and set an image size and an image aspect ratio as a recording mode; The vertical and horizontal sizes of the image corresponding to the recording mode set by the setting means are sizes corresponding to the size and aspect ratio of the image set by user selection.
2. The imaging device according to claim 1.
3. 3. The imaging device according to claim 2, wherein the image sizes that can be set by the setting unit include an L size.
4. 4. The imaging device according to claim 2, wherein the aspect ratio of the image that can be set by the setting unit includes 3:
2.
5. 5. The imaging device according to claim 1, wherein the recording control means determines the vertical and horizontal sizes of the divided HDR image data and the number of vertical and horizontal divisions according to the vertical and horizontal sizes of the image corresponding to the recording mode set by the setting means.
6. 6. The imaging device according to claim 1, wherein the recording control means controls the encoding unit to divide the HDR image data obtained by capturing an HDR image using the imaging sensor into a plurality of divided HDR image data, encode each of the divided HDR image data using the encoding unit, record the encoded divided HDR image data as image items in the image file, and record image structure information for combining the divided HDR image data into a state before the division as a derived image item in the image file, in response to a user operation for capturing an HDR image.
7. The imaging device described in any one of claims 1 to 6, characterized in that the recording control means controls the HDR image data that can be captured by the imaging sensor and that is larger than the vertical and horizontal sizes of the image corresponding to the recording mode set by the setting means to be divided into the plurality of divided HDR image data and encoded respectively by the encoding unit, and the encoded plurality of divided HDR image data to be recorded in the image file as image items.
8. 8. The imaging device according to claim 1, wherein the recording control means controls the plurality of divided HDR image data and the image structure information to be recorded in an overlay format of HEIF.
9. a horizontal size of the plurality of divided HDR image data is a multiple of a first size; a vertical size of the plurality of divided HDR image data is a multiple of a second size different from the first size; 9. The imaging device according to claim 1, wherein the imaging device is a lens.
10. 10. The imaging device according to claim 9, wherein the first size and the second size are different sizes.
11. 11. The imaging device according to claim 1, wherein the recording control means controls the encoding unit to encode the HDR image data obtained by shooting without dividing it, and record the encoded HDR image data in the image file in HEIF format, when the horizontal and vertical sizes of the image corresponding to the recording mode set by the setting means are equal to or smaller than a predetermined size.
12. 12. The imaging device according to claim 11, wherein the predetermined size is 4096 pixels.
13. 13. The imaging device according to claim 1, wherein the HDR image data is an image having a bit depth greater than 8 bits.
14. 14. The imaging device according to claim 13, wherein the HDR image data is image data with a bit depth of 10 bits.
15. 15. The imaging device according to claim 1, wherein the HDR image data is image data with a color gamut of BT.2020.
16. 16. The imaging device according to claim 1, wherein the HDR image data is image data generated by processing HDR image data obtained by shooting using a PQ gamma curve.
17. 17. The imaging device according to claim 1, wherein the encoding unit encodes each of the plurality of divided HDR image data in accordance with the High Efficiency Video Coding (HEVC) standard.
18. an operation unit for setting a shooting mode; When the HDR shooting mode for shooting HDR still images is set, the recording control means controls the recording of the HDR image data obtained by the shooting in HEIF format, not in JPEG format.
18. The imaging device according to claim 1, wherein the imaging device is a lens.
19. A program for causing a computer to function as each of the means of the imaging device according to any one of claims 1 to 18.
Citation Information
Patent Citations
Image pickup device
JP2013141074A
Reception device and reception method
JP2017073761A
Image data generating apparatus, image data generating method, and program
JP2017139618A
Information processing device and method
WO2015008775A1
Information processing device, information recording medium, information processing method, and program
WO2016103968A1