Control method for imaging device

By segmenting and encoding HDR image data in HEIF format, the method addresses compatibility and size challenges, ensuring high-quality playback across diverse devices.

JP2026069792APending Publication Date: 2026-04-24CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
CANON KK
Filing Date
2025-12-24
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing imaging devices struggle to efficiently record and playback High Dynamic Range (HDR) images due to compatibility issues and increased data size, leading to potential loss of image quality and compatibility with various devices.

Method used

The method involves dividing HDR image data into segmented data, encoding each segment in High Efficiency Image File Format (HEIF), and recording it as a single file with metadata for reconstruction, ensuring alignment with device capabilities and compatibility.

Benefits of technology

This approach allows for efficient recording and playback of HDR images with high resolution, maintaining image quality and compatibility across different devices by adhering to standard encoding formats.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026069792000001_ABST
    Figure 2026069792000001_ABST
Patent Text Reader

Abstract

The purpose is to record images in a recording format suitable for recording and playback, especially when recording images with large data volumes, such as HDR images. [Solution] A method for controlling an imaging device, comprising: an encoding step of dividing HDR image data obtained by capturing from an imaging sensor into a plurality of segmented HDR image data and encoding each of the plurality of segmented HDR image data; a recording control step of controlling the device to record the encoded plurality of segmented HDR image data as image items in HEIF format into a single image file, and to record image structure information for combining the plurality of segmented HDR image data into their pre-division state as a derived image item into the image file; and a setting step of setting a recording mode, wherein the vertical and horizontal sizes of the segmented HDR image data differ according to the vertical and horizontal sizes of the image corresponding to the recording mode set in the setting step.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] , , , , , , , ,

[0006] , , , ,

[0005] , , , , , , ,

[0001] The present invention relates to a method for controlling an imaging device.

Background Art

[0002] An imaging device is known as an image processing device that compresses and encodes image data. Such an image processing device acquires a moving image signal by an imaging unit, compresses and encodes the acquired moving image signal, and records an image file subjected to compression encoding on a recording medium. Conventionally, the image data before compression encoding was expressed in SDR (Standard Dynamic Range) with a luminance level of 100 nits as the upper limit. However, in recent years, image data having a luminance range close to the luminance range that can be perceived by humans has been provided, which is expressed in HDR (High Dynamic Range) in which the upper limit of the luminance level is extended to about 10,000 nits.

[0003] In Patent Document 1, there is a description of an image data recording device that generates and records image data that allows detailed confirmation of an HDR image even in a device that does not support HDR when shooting and recording an HDR image. ​​​​​​​​​​​​​​​​​​​​​​​​​Therefore, the present invention aims to provide a device for recording images in a recording format suitable for recording and playback when recording images with a large amount of data, such as HDR images, and a display control device for playing back images recorded in that recording format. [Means for solving the problem]

[0007] To solve the above-mentioned problems, the control method of the imaging device of the present invention is as follows: A control method for an imaging device that records HDR (High Dynamic Range) image data obtained by capturing images with an imaging sensor, comprising: an encoding step of dividing the HDR image data obtained by capturing images with an imaging sensor into a plurality of segmented HDR image data and encoding each of the plurality of segmented HDR image data; a recording control step of controlling the device to record the encoded plurality of segmented HDR image data as image items in HEIF (High Efficiency Image File Format) format into a single image file, and to record image structure information for combining the plurality of segmented HDR image data into their pre-division state as a derived image item into the image file; and a setting step of setting a recording mode. A key feature is that the vertical and horizontal dimensions of the segmented HDR image data differ depending on the vertical and horizontal dimensions of the image corresponding to the recording mode set during the setup process. [Effects of the Invention]

[0008] According to the present invention, it is possible to provide an imaging device that records images in a recording format suitable for recording and playback when recording images with a large amount of data and high resolution, and a display control device for playback of images recorded in that recording format. [Brief explanation of the drawing]

[0009] [Figure 1] A block diagram showing the configuration of the imaging device 100. [Figure 2] A diagram showing the structure of a HEIF file. [Figure 3] A flowchart illustrating the processing in HDR shooting mode. [Figure 4] A flowchart illustrating the process for determining how to segment image data in HDR shooting mode. [Figure 5] A diagram showing the encoding region and division method of image data when recording HDR image data. [Figure 6] A flowchart illustrating the process of constructing a HEIF file. [Figure 7] A flowchart illustrating the display process when playing back HDR image data recorded as a HEIF file. [Figure 8] A flowchart illustrating the process of creating an overlay image. [Figure 9] A flowchart illustrating the process of retrieving the properties of an image item. [Figure 10] A flowchart illustrating the data acquisition process for image items. [Figure 11] A flowchart illustrating the process of obtaining item IDs for the images that make up the main image. [Figure 12] A flowchart illustrating the image creation process for a single image item. [Figure 13] A diagram showing the positional relationship between the overlay image and the segmented image. [Modes for carrying out the invention]

[0010] Embodiments of the present invention will be described in detail below with reference to the attached drawings, using the imaging device 100 as an example, but the present invention is not limited to the following embodiments.

[0011] <Configuration of the imaging device> FIG. 1 is a block diagram showing an imaging device 100. As shown in FIG. 1, the imaging device 100 includes a CPU 101, a memory 102, a non-volatile memory 103, an operation unit 104, an imaging unit 112, an image processing unit 113, an encoding processing unit 114, a display control unit 115, and a display unit 116. Further, the imaging device 100 includes a communication control unit 117, a communication unit 118, a recording medium control unit 119, and an internal bus 130. The imaging device 100 forms an optical image of a subject on the pixel array of the imaging unit 112 using a photographing lens 111. The photographing lens 111 may be non-detachable or detachable from the body (housing, main body) of the imaging device 100. Also, the imaging device 100 performs writing and reading of image data to and from a recording medium 120 via the recording medium control unit 119. The recording medium 120 may be detachable or non-detachable from the imaging device 100.

[0012] The CPU 101 controls the operations of each part (each functional block) of the imaging device 100 via the internal bus 130 by executing a computer program stored in the non-volatile memory 103.

[0013] The memory 102 is a rewritable volatile memory. The memory 102 temporarily records a computer program for controlling the operations of each part of the imaging device 100, information such as parameters related to the operations of each part of the imaging device 100, information received by the communication control unit 117, etc. Also, the memory 102 temporarily records an image acquired by the imaging unit 112, an image and information processed by the image processing unit 113, the encoding processing unit 114, etc. The memory 102 has a storage capacity sufficient to temporarily record these.

[0014] The non-volatile memory 103 is a memory that can be electrically erased and recorded, and for example, an EEPROM or the like is used. The non-volatile memory 103 stores a computer program for controlling the operations of each part of the imaging device 100 and information such as parameters related to the operations of each part of the imaging device 100. Various operations performed by the imaging device 100 are realized by this computer program.

[0015] The operation unit 104 provides a user interface for operating the imaging device 100. The operation unit 104 includes various buttons such as a power button, a menu button, and a shooting button, and the various buttons are constituted by a switch, a touch panel, etc. The CPU 101 controls the imaging device 100 according to the user's instructions input via the operation unit 104. Here, the case where the CPU 101 controls the imaging device 100 based on the operations input via the operation unit 104 has been described as an example, but it is not limited thereto. For example, the CPU 101 may control the imaging device 100 based on a request input via the communication unit 118 from a remote controller (not shown), a mobile terminal (not shown), etc.

[0016] The photographing lens (lens unit) 111 is constituted by a lens group (not shown) including a zoom lens, a focus lens, etc., a lens control unit (not shown), an aperture (not shown), etc. The photographing lens 111 can function as zoom means for changing the angle of view. The lens control unit controls the adjustment of focus and the aperture value (F value) according to a control signal transmitted from the CPU 101. The imaging unit 112 can function as acquisition means for sequentially acquiring a plurality of images constituting a moving image. As the imaging unit 112, for example, an area image sensor using a CCD (charge-coupled device), a CMOS (complementary metal oxide semiconductor) element, etc. is used. The imaging unit 112 has a photoelectric conversion unit (not shown) that converts the optical image of the subject into an electrical signal and a pixel array (not shown) arranged in a matrix, that is, two-dimensionally. The optical image of the subject is formed on the pixel array by the photographing lens 111 at the imaging unit 112. The imaging unit 112 outputs the captured image to the image processing unit 113 or the memory 102. Note that the imaging unit 112 can also acquire a still image.

[0017] The image processing unit 113 performs predetermined image processing on image data output from the imaging unit 112 or image data read from the memory 102. Examples of such image processing include interpolation, reduction (resizing), and color conversion. The image processing unit 113 also performs predetermined calculations for exposure control, distance measurement control, etc., using the image data acquired by the imaging unit 112. Based on the calculation results obtained by the image processing unit 113, exposure control, distance measurement control, etc., are performed by the CPU 101. Specifically, AE (automatic exposure) processing, AWB (auto white balance) processing, AF (autofocus) processing, etc., are performed by the CPU 101.

[0018] The encoding processing unit 114 compresses the size of the image data by performing intra-frame predictive coding (in-frame predictive coding), inter-frame predictive coding (inter-frame predictive coding), etc. on the image data. The encoding processing unit 114 is, for example, an encoding device composed of semiconductor elements. The encoding processing unit 114 may be an encoding device provided outside the imaging device 100. The encoding processing unit 114 performs encoding processing, for example, using the H.265 (ITU H.265 or ISO / IEC 23008-2) method.

[0019] The display control unit 115 controls the display unit 116. The display unit 116 has a display screen (not shown). The display control unit 115 generates an image that can be displayed on the display screen of the display unit 116 by performing resizing, color conversion, etc., on the image data, and outputs the image, i.e., the image signal, to the display unit 116. The display unit 116 displays the image on the display screen based on the image signal sent from the display control unit 115. The display unit 116 has an OSD (On Screen Display) function, which is a function that displays setting screens such as menus on the display screen. The display control unit 115 can superimpose an OSD image on the image signal and output the image signal to the display unit 116. The display unit 116 is composed of a liquid crystal display, an organic EL display, etc., and displays the image signal sent from the display control unit 115. The display unit 116 may be, for example, a touch panel. If the display unit 116 is a touch panel, the display unit 116 can also function as an operation unit 104.

[0020] The communication control unit 117 is controlled by the CPU 101. The communication control unit 117 generates a modulated signal conforming to a predetermined wireless communication standard such as IEEE 802.11 and outputs the modulated signal to the communication unit 118. The communication control unit 117 also receives the modulated signal conforming to the wireless communication standard via the communication unit 118, decodes the received modulated signal, and outputs a signal corresponding to the decoded signal to the CPU 101. The communication control unit 117 is equipped with a register for storing communication settings. The communication control unit 117 can adjust the transmission and reception sensitivity during communication under control from the CPU 101. The communication control unit 117 can transmit and receive using a predetermined modulation scheme. The communication unit 118 is equipped with an antenna that outputs the modulated signal supplied from the communication control unit 117 to an external device 127 such as an information and communication device located outside the imaging device 100, and also receives modulated signals from the external device 127. The communication unit 118 is also equipped with communication circuits and the like. Although this explanation uses the example of wireless communication performed by the communication unit 118, the communication performed by the communication unit 118 is not limited to wireless communication. For example, the communication unit 118 and the external device 127 may be connected by electrical connections using wiring or the like.

[0021] The recording medium control unit 119 controls the recording medium 120. Based on a request from the CPU 101, the recording medium control unit 119 outputs control signals to the recording medium 120 for controlling the recording medium 120. For example, non-volatile memory or magnetic disks can be used as the recording medium 120. As described above, the recording medium 120 may be removable or non-removable. The recording medium 120 records encoded image data, etc. The image data, etc. is saved as a file in a format compatible with the file system of the recording medium 120. Examples of files include MP4 files (ISO / IEC 14496-14:2003) and MXF (Material eXchange Format) files. Each of the functional blocks 101-104, 112-115, 117, and 119 is accessible from each other via the internal bus 130.

[0022] Here, the normal operation of the imaging device 100 in this embodiment will be described.

[0023] When a user operates the power button on the control unit 104 of the imaging device 100, the control unit 104 sends a startup instruction to the CPU 101. Upon receiving this instruction, the CPU 101 controls a power supply unit (not shown) to supply power to each block of the imaging device 100. Once power is supplied, the CPU 101 checks, for example, which mode the mode selector switch on the control unit 104 is currently in, based on an instruction signal from the control unit 102, such as still image shooting mode or playback mode.

[0024] In normal still image shooting mode, the imaging device 100 performs the shooting process when the user operates the still image recording button on the operation unit 104 while in shooting standby mode. During the shooting process, the image processing unit 113 processes the image data of the still image captured by the imaging unit 112, the encoding processing unit 114 encodes it, and the recording medium control unit 119 records the encoded image data as an image file on the recording medium 120. In shooting standby mode, the imaging unit 112 captures an image at a predetermined frame rate, the image processing unit performs image processing for display, and the display control unit 115 displays it on the display unit 116 to display a live view image.

[0025] In playback mode, the recording medium control unit 119 reads the image files recorded on the recording medium 120, and the encoding processing unit 114 decodes the image data of the read image files. In other words, the encoding processing unit 114 also functions as a decoder. Then, the image processing unit 113 performs processing for display, and the display control unit 115 displays it on the display unit 116.

[0026] While the capture and playback of normal still images are performed as described above, the imaging device in this embodiment has an HDR shooting mode for capturing not only normal still images but also HDR still images, and can also play back the captured HDR still images.

[0027] The following sections will explain the process of capturing and playing back HDR still images.

[0028] <File structure> First, let's explain the file structure when recording HDR still images.

[0029] Recently, a still image file format called High Efficiency Image File Format (hereinafter referred to as HEIF) was established. (ISO / IEC 23008-12:2017) This has the following characteristics compared to conventional still image file formats such as JPEG. • This file format conforms to the ISO-based media file format (hereinafter referred to as ISOBMFF). (ISO / IEC 14496-14:2003) • It can store not only a single still image, but also multiple still images. It can store still images compressed using video compression formats such as HEVC / H.265 and AVC / H.264.

[0030] In this implementation example, HEIF is used as the recording file for HDR still images.

[0031] First, let's explain the data that HEIF stores.

[0032] HEIF manages the individual data it stores in units called items.

[0033] Each item, in addition to the data itself, has a unique integer item ID (item_ID) within the file and an item type (item_type) that indicates the type of item.

[0034] Items can be divided into image items, where the data represents an image, and metadata items, where the data is metadata.

[0035] Image items include coded image items, which are image data with encoded data, and derived image items, which represent images resulting from manipulating one or more other image items.

[0036] An example of a derived image item is an overlay image, which is a derived image item in overlay format. This is an overlay image that is created by compositing an arbitrary number of image items by placing them at arbitrary positions using an ImageOverlay structure (overlay information).

[0037] Exif data can be stored as an example of a metadata item.

[0038] As mentioned earlier, HEIF can store multiple image items.

[0039] If there are relationships between multiple images, those relationships can be described.

[0040] Examples of relationships between multiple images include the relationship between a derived image item and the image items that comprise it, and the relationship between the main image and a thumbnail image.

[0041] Similarly, the relationship between image items and metadata items can also be described in the same way.

[0042] The HEIF format conforms to the ISOBMFF format. Therefore, let's first briefly explain ISOBMFF.

[0043] The ISOBMFF format manages data using a structure called boxes.

[0044] A box is a data structure that begins with a 4-byte data length field and a 4-byte data type field, followed by data of arbitrary length.

[0045] The structure of the data section is determined by the data type. The ISOBMFF and HEIF specifications define several data types and the structure of their data sections.

[0046] Furthermore, a box can contain data from another box; in other words, boxes can be nested. Here, a box nested within the data section of another box is called a subbox.

[0047] Furthermore, boxes that are not subboxes are called file-level boxes. These are boxes that can be accessed sequentially from the beginning of the file.

[0048] Figure 2 will be used to explain HEIF format files.

[0049] First, let's discuss file-level boxes.

[0050] A file type box with data type 'ftyp' stores information about file compatibility. ISOBMFF-compliant file specifications declare the file structure and the data the file contains using a 4-byte code called a brand, and store these in file type boxes. By placing a file type box at the beginning of a file, a file reader can understand the file structure by checking the contents of the file type box without having to read and interpret the file contents further.

[0051] The HEIF specification uses the brand 'mif1' to represent the file structure. Furthermore, if the encoded image data stored is HEVC compressed, it is represented by the brand 'heic' or 'heix' according to the HEVC compression profile.

[0052] A metadata box with data type 'meta' contains various subboxes that store data about each item. The breakdown will be explained later.

[0053] Media data boxes with data type 'mdat' store data for each item. For example, encoded image data for encoded image items or Exif data for metadata items.

[0054] Next, we will explain the subboxes of the metadata box.

[0055] A handler box with data type 'hdlr' stores information representing the type of data managed by the metadata box. In the HEIF specification, the handler_type of a handler box is 'pict'.

[0056] A data information box with data type 'dinf' specifies the file where the data targeted by this file resides. ISOBMFF allows a file to store data targeted by a particular file in a file other than its own. In this case, the data reference within the data information box contains a data entry URL box that describes the URL of the file where the data resides. If the target data exists in the same file, a data entry URL box containing only a flag indicating this is stored.

[0057] A primary item box with data type 'pitm' stores the item ID of the image item that represents the main image.

[0058] An item information box with data type 'iinf' is a box for storing the following item information entries.

[0059] Item information entries with data type 'infe' store the item ID, item type, and flags for each item.

[0060] The item type of an image item whose encoded image data is HEVC compressed is 'hvc1', and it is an image item.

[0061] The item type of the derived image item in the overlay format, i.e., the ImageOverlay structure (overlay information), is 'iovl', and it is classified as an image item.

[0062] The item type of an Exif-formatted metadata item is 'Exif', and it is a metadata item.

[0063] Additionally, if the least significant bit in the flag field of an image item is set, it can be specified that the image item is a hidden image. When this flag is set, the image item will not be treated as a display target during playback and will be hidden.

[0064] An item reference box with data type 'iref' stores the reference relationships between each item, including the type of reference relationship, the item ID of the referencing item, and the item IDs of one or more referenced items.

[0065] For overlay-format derived image items, the type of reference relationship is 'dimg', the item ID of the derived image item is stored in the item ID of the source item, and the item ID of each image item that makes up the overlay is stored in the item ID of the referenced item.

[0066] For thumbnail images, the type of reference relationship is 'thmb', the item ID of the thumbnail image is stored in the item ID of the source item, and the item ID of the main image is stored in the item ID of the referenced item.

[0067] An item property box with data type 'iprp' is a box for storing the following item property container boxes and item property related boxes.

[0068] An item property container box with data type 'ipco' is a box that stores individual property data boxes.

[0069] Each image item can have property data that represents the characteristics and attributes of that image.

[0070] The property data box includes the following:

[0071] Decoder configuration and initialization data (type 'hvcC' in the case of HEVC) is data used to initialize the decoder. HEVC parameter set data (VideoParameterSet, SequenceParameterSet, PictureParameterSet) is stored here.

[0072] Image spatial extents (type 'ispe') are the dimensions (width, height) of the image.

[0073] Color space information (type 'colr') is the color space information of an image.

[0074] Image rotation information (type 'irot') indicates the direction of rotation when displaying an image after rotation.

[0075] The pixel information of an image (type 'pixi') indicates the number of bits in the data that makes up the image.

[0076] In this embodiment, the Mastering display color volume (type 'MDCV') and Contents light level information (type 'CLLI') are stored as property data boxes as HDR metadata.

[0077] Other properties exist besides those listed above, but they will not be specified here.

[0078] The item property association box with data type 'ipma' stores the association between each item and property in the form of the item ID and an array of indices in 'ipco' for the associated property.

[0079] An item data box with data type 'idat' is a box that stores item data with a small data size.

[0080] Data for derived image items in overlay format can be stored in an ImageOverlay structure (overlay information) in 'idat', which contains position information of the constituent images. The overlay information includes parameters such as canvas_fill_value, which is background color information, and output_width and output_height, which are the final composite image size when composited with an overlay. Furthermore, for each image that becomes a composite element, there are horizontal_offset and vertical_offset parameters that indicate the horizontal and vertical position coordinates of the composite image. By using these parameters included in the overlay information, it is possible to reproduce a composite image in which multiple images are placed at arbitrary positions within a single image of a specified background color and size. Item location boxes with data type 'iloc' store the position information of each item's data in the format of offset base (construction_method), offset value from the offset base, and length.

[0081] The offset reference point is either the beginning of the file or 'idat'.

[0082] Note that there are other metadata boxes besides those mentioned above, but they will not be listed here.

[0083] Figure 2 illustrates a structure in which two encoded image items constitute an overlay image. The number of encoded image items that make up an overlay image is not limited to two. As the number of encoded image items that make up an overlay image increases, the following boxes and items increase accordingly. • An item information entry for encoded image items has been added to the item information box. · The item ID of the encoded image item increases in the reference item ID of the reference relationship type 'dimg' of the item reference box. · In the item property container box, decoder configuration, initialization data, image space range, etc. · In the item property related box, items of the index of the encoded image item and its related properties are added. · In the item location box, an item of the position information of the encoded image item is added. · Encoded image data is added to the media data box.

[0084] <Shooting of HDR Image> Next, the processing in the imaging device 100 when shooting and recording an HDR image will be described.

[0085] The width and height of the image to be recorded have been increasing recently. However, if an overly large size is encoded, compatibility may be lost during decoding on other devices, or the scale of the system may increase. Specifically, in this embodiment, an example using H.265 for encoding and decoding will be described. H.265 has parameters such as Profile and Level as standards, and these parameters change as the image encoding method and the image size increase. These values are parameters for determining the feasibility of playback on the device during decoding, and there may be cases where the playback is rejected as unplayable by determining this parameter. In this embodiment, when recording a large image size, it is divided into multiple image display areas, and the overlay method advocated by the HEIF standard is used, and the encoded data is stored and managed in a HEIF file. By dividing it into multiple image display areas in this way, the size of each image display area becomes smaller, and the playback compatibility with other devices is enhanced. Hereinafter, the description will be made with the size of one side of the encoding area set to 4096 or less. Note that the encoding format may be other than H.265, or the size of one side may be defined as other than 4096.

[0086] Next, we will explain the alignment constraints regarding the start position, width, and height of the encoded region, as well as the start position, width, and height of the playback region. For imaging devices, hardware constraints generally exist when encoding the divided images, such as the vertical and horizontal encoding start alignment of the encoded region, and the playback width and height alignment during playback. In other words, it is necessary to calculate and encode the encoding start position for one or more encoded regions, as well as the start position, width, and height of the playback region. Since imaging devices switch between multiple image sizes for recording based on user instructions, these constraints must be switched for each image size.

[0087] The encoding processing unit 114 described in this embodiment has alignment constraints on the starting position during encoding, with a starting position of 64 pixels in the horizontal direction (hereafter x-direction) and 1 pixel in the vertical direction (hereafter y-direction). Furthermore, there are alignment constraints of 32 pixels and 16 pixels, respectively, for the width and height to be encoded. The starting position of the playback area has alignment constraints of 2 pixels in the x-direction and 1 pixel in the y-direction. Additionally, the width and height of the playback area have alignment constraints of 2 pixels and 1 pixel, respectively. While examples of alignment constraints are given here, other alignment constraints may generally be used.

[0088] Next, Figure 3 will explain the flow from capturing HDR image data in HDR recording mode (HDR shooting mode) to recording the captured HDR image data as a HEIF file. In HDR mode, the imaging unit 112 and the image processing unit 113 perform processing to represent the HDR color space using the HDR color gamut BT.2020 and the PQ gamma curve. The recording process for HDR will be omitted in this embodiment.

[0089] In S101, the CPU 101 determines whether the user has pressed SW2. If the user has captured a subject and pressed SW2 (YES in S101), the CPU 101 detects that SW2 has been pressed and proceeds to S102. The CPU 101 continues this detection until SW2 is pressed (NO in S101).

[0090] In S102, the CPU 101 determines how many divisions (number of divisions N) to divide the encoding area into horizontally and vertically, and how to divide it, in order to save the image of the field of view captured by the user into a HEIF file. This result is saved in memory 102. S102 will be described later with reference to Figure 4. Proceed to S103.

[0091] In S103, CPU101 initializes the variable M in memory 102 to M=1. Then the program proceeds to S104.

[0092] In S104, CPU101 monitors whether it has completed the N-fold loop determined in S102. If the number of loops reaches the N-fold (YES in S104), proceed to S109; otherwise (NO in S104), proceed to S105.

[0093] In S105, the CPU 101 reads the information of the coding start region and playback region of the divided image M from the memory 102. The process then proceeds to S106.

[0094] In S106, the CPU 101 informs the encoding processing unit 114 which region of the entire image stored in memory 102 should be encoded, and the CPU 101 then proceeds to S107.

[0095] In S107, the encoding processing unit 114 encodes the resulting encoded data, and the CPU 101 temporarily stores the associated information generated during encoding in memory 102. Specifically, in the case of H.265, this refers to H.265 standard information such as VPS, SPS, and PPS, which must later be stored in the HEIF file. Although not directly related to this embodiment, the H.265 VPS, SPS, and PPS are various pieces of information necessary for decoding, such as the size of the encoded data, bit depth, display / hide area specification, and frame information. The process then proceeds to S108.

[0096] In S108, CPU101 increments the variable M and returns to S104.

[0097] In S109, the CPU 101 generates metadata to be stored in the HEIF file. The CPU 101 extracts the information necessary for playback from the imaging unit 112, image processing unit 113, encoding processing unit 114, etc., and temporarily stores it in memory 102 as metadata. Specifically, this is data formatted in Exif format. Details of Exif data are already common knowledge, so an explanation will be omitted in this embodiment. Proceed to S110.

[0098] In S110, the CPU 101 encodes the thumbnail. The image processing unit 113 reduces the entire image in memory 102 to thumbnail size and temporarily places it in memory 102. The CPU 101 instructs the encoding processing unit 114 to encode this thumbnail image in memory 102. Here, the thumbnail is encoded to a size that satisfies the alignment constraints mentioned earlier, but the encoding area and the playback area may be different. This thumbnail is also encoded in H.265. The process proceeds to S111.

[0099] In S111, the encoding processing unit 114 encodes the thumbnail, and the CPU 101 temporarily stores the encoded data and any associated information generated during encoding in memory 102. This refers to H.265 standard information such as VPS, SPS, and PPS, as explained in S107 above. Next, we proceed to S112.

[0100] In S112, the CPU 101 has saved various data to memory 102 in the steps up to this point. The CPU 101 then sequentially combines the data stored in memory 102 to complete the HEIF file and saves it back to memory 102. The flow for completing this HEIF file will be explained later using Figure 6. Proceed to S113.

[0101] In S113, the CPU 101 instructs the recording medium control unit 119 to write the HEIF file located in memory 102 to the recording medium 120. The process then returns to S101.

[0102] After the shooting is completed in these steps, the HEIF file is recorded on the recording medium 120.

[0103] Next, we will explain the method for determining the division in S102, as described earlier, using Figures 4 and 5(a). Hereafter, the upper left corner of the sensor image region will be considered the origin (0,0), and the coordinates will be (x,y), while the width and height will be represented by w and h, respectively. The region itself will be represented as [x,y,w,h], combining the starting coordinates (x,y) and the width and height w and h.

[0104] In S121, the CPU 101 retrieves the recording mode for shooting from memory 102. The recording mode determines the recording method, such as image size, aspect ratio, and compression ratio. The user selects one of these recording modes to take a picture. For example, this could be settings such as L size and 3:2 aspect ratio. Since the shooting angle of view differs depending on the recording mode, the shooting angle of view is determined. The process then proceeds to S122.

[0105] In S122, the starting coordinate position (x0, y0) of the playback image is obtained from the sensor image region [0, 0, H, V]. Then the process proceeds to S123.

[0106] In S123, the region [x0, y0, Hp, Vp] of the playback image is obtained from memory 102. Then the process proceeds to S124.

[0107] In S124, it is necessary to encode so that the region of the reconstructed image is included. As mentioned earlier, the alignment of the encoding start position is 64 pixels in the x direction, so in order to encode so that the region of the reconstructed image is included, encoding must start from the position (Ex0, Ey0). In this way, the size of Lex is determined. If the offset of this Lex is 0, proceed to S126 (NO in S124), or to S125 if necessary (YES in S124).

[0108] In S125, Lex is calculated. As mentioned earlier, Lex is determined by the alignment of the reconstructed image region and the coding start position. The CPU 101 determines Lex as follows: x0 is divided by the 64 pixels of the horizontal coding start position's pixel alignment, and the quotient is multiplied by the aforementioned 64 pixels. Since the remainder is not included in the calculation, the coordinates obtained are those located to the left of the starting position of the reconstructed image region. Therefore, Lex is obtained as the difference between the x-coordinate of the reconstructed image region and the x-coordinate of the coding start position obtained by the previous calculation. Proceed to S126.

[0109] In S126, the coordinates of the final position of the rightmost edge of the topmost line of the regenerated image are calculated. Let this be (xN, y0). xN is obtained by adding Hp to x0. Proceed to S127.

[0110] In S127, the right edge of the encoding region is determined so that it includes the right edge of the reconstructed image. To determine the right edge w of the encoding region, an alignment constraint on the encoding width is necessary. As mentioned earlier, the constraints on the encoding width and height are each multiples of 32 pixels. Based on the relationship between this constraint and (xN, y0), CPU 101 calculates whether an offset of Rex is needed before the right edge of the reconstructed image. If Rex is needed (YES in S127), proceed to S128; otherwise (NO in S127), proceed to S129.

[0111] In S128, CPU101 adds Rex to the right of the rightmost coordinate (xN, y0) of the regenerated image so that it is aligned by 32 pixels, and then determines the end position of the encoding (ExN, Ey0). Proceed to S129.

[0112] In S129, CPU101 calculates the coordinates of the bottom right corner of the display area. Since the size of the playback image is obvious, these coordinates are (xN, yN). Proceed to S130.

[0113] In S130, the lower edge of the encoding region is determined so that it includes the lower edge of the reconstructed image. To determine the lower edge of the encoding region, an alignment constraint on the encoding height is required. As mentioned earlier, the constraint on the encoding height is that each part must be a multiple of 16 pixels. Based on the relationship between this constraint and the lower-right coordinates (xN, yN) of the reconstructed image, CPU 101 calculates whether a Vex offset is needed before the lower edge of the reconstructed image. If Vex is needed (YES in S130), the process proceeds to S131; otherwise (NO in S130), the process proceeds to S132.

[0114] In S131, CPU101 calculates Vex from the relationship between the previous constraints and the lower-right coordinates (xN, yN) of the reconstructed image, and finds the lower-right coordinates (ExN, EyN) of the encoded region. The method for calculating Vex is to set it at a position that is a multiple of 16 pixels, which is the height alignment of the encoding, so that it includes yN from y0, which is the vertical encoding start position. Specifically, assuming y0=0, Vp is divided by 16 pixels, which is the height alignment of the encoding, and the quotient and remainder are found. Since the encoded region must be taken so as to include the reconstructed region, if there is a remainder, 1 is added to the quotient and multiplied by the aforementioned 16 pixels. This gives EyN, the y-coordinate of the lower edge of the encoded region that is a multiple of 16 pixels and includes Vp. Vex is found as the difference between EyN and yN. Note that EyN is the value obtained by offsetting yN downward by Vex, and ExN is the value already found in S128. Proceed to S132.

[0115] In S132, CPU101 calculates the size of the encoding region Hp'xVp' as follows: Hp' is calculated by adding Lex and Rex, obtained from the alignment constraint, to the horizontal size Hp. Vp' is calculated by adding Vex, obtained from the alignment constraint, to the vertical size Vp. The size of the encoding region Hp'×Vp' is then obtained from Hp' and Vp'. The process proceeds to S133.

[0116] In S133, the CPU 101 determines whether the horizontal size Hp' to be encoded exceeds the 4096 pixels to be divided. If it exceeds the limit (YES in S133), the process proceeds to S134; otherwise, it proceeds to S135.

[0117] In S134, the horizontal size Hp' to be encoded is divided into two or more regions so that it is 4096 pixels or less. For example, if it is 6000 pixels, it will be divided into two regions, and if it is 9000 pixels, it will be divided into three regions.

[0118] In the case of a two-part division, the division is made at an approximate center position that satisfies the 32-pixel alignment regarding the width of the encoding mentioned earlier. If the divisions are not equal, the left divided image is made larger. Similarly, when dividing into three parts and the divisions are not equal, the division areas are divided into A, B, and C, paying attention to the alignment regarding the width of the encoding mentioned earlier. The division areas A and B are made equal in size, and the division area C is made slightly smaller. The same algorithm is used to determine the division positions and sizes even for three or more divisions. Proceed to S135.

[0119] In S135, the CPU 101 determines whether the vertical size Vp' to be encoded exceeds the 4096 pixels to be divided. If it exceeds the limit (YES in S135), the process proceeds to S136; otherwise, it terminates (NO in S135).

[0120] In S136, similar to S134, Vp' is divided into multiple partitioned regions and the process is terminated.

[0121] An example of segmentation in S102 will be explained with specific numerical values ​​using Figures 5(b), (c), (d), and (e). Figure 5 shows not only the method of segmenting the image data, but also the relationship between the encoded region and the re-encoded region.

[0122] Let's explain Figure 5(b). Figure 5(b) shows a case where the image is divided into two regions, left and right, and Lex, Rex, and Vex are assigned to the left, right, and bottom edges of the playback region. First, the playback image region is [256+48, 0, 4864, 3242], which is the region of the final image that the imaging device wants to record. The coordinates of the top left of this playback image region are (256+48, 0), which is not a multiple of the 64-pixel unit that is the starting alignment for encoding mentioned earlier. Therefore, Lex is placed 48 pixels to the left of this point, and the left edge of the encoding start position is at x=256. Next, the width of the playback region is w=48+4864, but this is not a multiple of the 32-pixel width alignment for encoding, so considering the width constraint alignment during encoding, 16 pixels of Rex are assigned to the right edge of the playback image region. Similarly, calculating the vertical direction gives Vex=4 pixels. Since the horizontal size of this reconstructed image exceeds 4096, we divide it horizontally into two parts at approximately the center, taking into account the alignment of the encoding width. As a result, the encoded region of divided image 1 is [256, 0, 2496, 3242+8], and the encoded image region of divided image 2 is [256+2496, 0, 2432, 3242+8].

[0123] Let's explain Figure 5(c). Figure 5(c) shows a case where the image is divided into two regions, left and right, and Rex and Vex are assigned to the right and bottom edges of the playback region. Since the x-position of the playback image region is x=0, it coincides with the 64 pixels that are the starting alignment of the encoded region, so Lex=0. From here on, the calculation method is the same as described above, and as a result, the encoded region of divided image 1 is obtained as [0, 0, 2240, 32456+8], and the encoded image region of divided image 2 is obtained as [2240, 0, 2144, 2456+8].

[0124] Let's explain Figure 5(d). Figure 5(d) shows a case where the image is divided into two regions, left and right, and Lex and Rex are assigned to the left and right of the playback region. The vertical size of the playback image region is h=3648, which is a multiple of 16 pixels, a constraint on the image height during encoding, so Vex=0. From here on, the calculation method is the same as described above, and as a result, the encoded region of divided image 1 is [256+48, 0, 2240, 3648], and the encoded image region of divided image 2 is [256+2240, 0, 2144, 3648].

[0125] Let's explain Figure 5(d). Figure 5(d) shows a case where the image is divided into four regions vertically and horizontally, and Lex and Rex are assigned to the left and right of the playback region. In this case, the encoded region exceeds 4096 in both the vertical and horizontal directions, so a grid-like division is performed. The calculation method is the same as described above, and as a result, the encoded region of divided image 1 is [0, 0, 4032, 3008], the encoded image region of divided image 2 is [4032, 0, 4032, 3008], the encoded image region of divided image 3 is [0, 3008, 4032, 2902], and the encoded image region of divided image 4 is [4032, 3008, 4032, 2902].

[0126] In this way, the CPU 101 determines the region of the playback image according to the recording mode and needs to calculate the division region for each. Alternatively, these calculations can be performed in advance and the results stored in memory 102. The CPU 101 may also read the division region and playback region information according to the recording mode from memory 102 and set it in the encoding processing unit 114.

[0127] When recording a captured image by dividing it into multiple segmented images, as in this embodiment, it is necessary to combine multiple display areas to reconstruct the captured image from the recorded images. When generating the combined image, it is more efficient if there are no overlapping areas in each segmented image, as this avoids the need to encode or record extra data. Therefore, in this embodiment, each playback area and each encoding area was set so that there are no overlapping parts (overlapping areas) at the boundaries where the segmented images meet. However, overlapping areas may be provided between segmented images depending on conditions such as hardware alignment constraints.

[0128] Next, the method for constructing the HEIF file mentioned in S112 above will be explained using Figure 6. Since the HEIF file has the structure shown in Figure 2, the CPU 101 constructs it in memory 102 sequentially from the beginning of the file.

[0129] In S141, the 'ftyp' box is created. This box is used to determine file compatibility and is placed at the beginning of the file. CPU101 stores 'heic' as the major brand and 'mif1' and 'heic' as the compatible brands in this box. These values ​​are used in this embodiment, but other values ​​may also be used. Proceed to S142.

[0130] In S142, a 'meta' box is created. This box is used to store the multiple boxes described below. CPU101 creates this box and proceeds to S143.

[0131] In S143, an 'hdlr' box is created to be stored within the 'iinf' box. This box indicates the attributes of the 'meta' box mentioned earlier. CPU101 stores 'pict' in this type and proceeds to S144. Note that 'pict' is information representing the type of data managed by the metadata box, and the HEIF specification defines the handler_type of the handler box as 'pict'.

[0132] In S144, a 'dinf' box is created to be stored within the 'iinf' box. This box indicates the location of the data targeted by this file. CPU101 creates this box and proceeds to S145.

[0133] In S145, a 'pitm' box is created to be stored within the 'iinf' box. This box stores the item ID of the image item representing the main image. Since the main image is an overlay image created by overlay, CPU 101 stores the overlay information item ID as the item ID of the main image. Proceed to S146.

[0134] In S146, an 'iinf' box is created to be stored within the 'meta' box. This box manages the list of items. Here, an initial value is entered in the data length field, 'iinf' is saved in the data type field, and the process proceeds to S147.

[0135] In S147, an 'infe' box is created to be stored within the 'iinf' box. This 'infe' box is used to register item information for each item stored in the file, and an 'infe' box is created for each item. For split images, each split image is registered as a single item in this box. In addition, overlay information for constructing the main image from multiple split images, Exif information, and thumbnail images are each registered as items. As mentioned above, overlay information, split images, and thumbnail images are registered as image items. For image items, it is possible to add a flag to indicate a hidden image, and when this flag is added, the image will not be displayed during playback. In other words, by adding this hidden image flag to an image item, it is possible to specify which images should be hidden during playback. Therefore, the flag indicating a hidden image is set for split images, but not for the overlay information item. As a result, the individual split images themselves are not displayed, and only the image resulting from the overlay synthesis of multiple split images by the overlay information is displayed. Create an individual 'infe' box for each item and store them in the aforementioned 'iinf'. Proceed to S148.

[0136] In S148, an 'iref' box is created to be stored within the 'iinf' box. This box stores information indicating the relationship between the image composed of the overlay (main image) and the segmented images that make up that image. Since the segmentation method was determined in S102 as explained earlier, the CPU 101 creates this box based on the determined segmentation method. Proceed to S149.

[0137] In S149, an 'iprp' box is created to be stored inside the 'iinf' box. This box stores the item's properties and will contain the 'ipco' box created in S150 and the 'ipma' box created in S151. Proceed to S150.

[0138] In S150, an 'ipco' box is created to be stored within the 'iprp' box. This box is the item's property container box and stores various properties. Multiple properties exist, and the CPU 101 creates the following property container boxes and stores them within the 'ipco' box. Some properties are created for each image item, while others are created for multiple image items. A 'colr' box is created as information common to both the overlay image, which is composed of the overlay information that forms the main image, and the split image. The 'colr' box stores color information such as the HDR gamma curve as color space information for the main image (overlay image) and the split image. For the split image, an 'hvcC' box and an 'ispe' box are created, respectively. In S107 and S111, as explained earlier, the CPU 101 reads the associated information from the encoding process stored in memory 102 and creates the 'hvcC' property container box. The supplementary information stored in 'hvcC' includes not only information about when the encoding region was encoded and the size (width, height) of the encoding region, but also the size (width, height) of the playback region and the position information of the playback region within the encoding region. This supplementary information is Golomb compressed and recorded in the property container box of 'hvcC'. Golomb compression is a common compression method, so an explanation is omitted here. The 'ispe' box stores the size (width, height) information of the playback region of the divided image (e.g., 2240x2450). An 'irot' box is generated to store information representing the rotation of the overlay image, which is the main image, and a 'pixi' box is generated to indicate the number of bits in the image data. The 'pixi' box may be generated separately for the main image (overlay image) and the divided images, but in this embodiment, the overlay image and the divided images have the same number of bits, i.e., 10 bits, so only one is generated. 'CLLI' and 'MDCV' boxes, which store supplementary HDR information, are also generated.Furthermore, as properties of the thumbnail image, separate from the main image (overlay image), a 'colr' box is generated to store color space information, an 'hvcC' box to store encoding information, an 'ispe' box to store image size, a 'pixi' box to store the bit depth of the image data, and a 'CLLI' box to store supplementary HDR information. These property container boxes are generated and stored in the 'ipco' box. Proceed to S151.

[0139] In S151, the 'ipma' box is generated. This box shows the relationship between items and properties, indicating which of the previously mentioned properties each item is associated with. The CPU 101 determines the relationship between items and properties from the various data stored in memory 102 and generates this box. Proceed to S152.

[0140] In S152, the 'idat' box is generated. This box stores overlay information that indicates how to arrange the playback area of ​​each divided image to generate the overlay image. The overlay information includes the canvas_fill_value parameter, which is the background color information, and the output_width and output_height parameters, which are the overall size of the overlay image. In addition, for each divided image that will be used as a composite element, there are horizontal_offset and vertical_offset parameters that indicate the horizontal and vertical position coordinates where the divided images will be composited. The CPU 101 writes this information to these parameters based on the division method determined in S102. Specifically, the output_width and output_height parameters, which indicate the size of the overlay image, are set to the size of the entire playback area. The horizontal_offset and vertical_offset parameters, which indicate the position information of each divided image, are set to the width and height offset values, respectively, from the starting coordinate position (x0, y0) of the top left of the playback area to the top left position of the divided image. By generating overlay information in this way, when playing back an image, the divided images are positioned based on the horizontal_offset and vertical_offset parameters and combined, allowing the image to be played back in its pre-division state. In this embodiment, since there are no overlapping areas in the divided images, positional information is provided so that each divided image is positioned without overlapping. Then, by specifying the area to be displayed using the output_width and output_height parameters from the image obtained by combining the divided images, only the playback area can be displayed during playback. An 'idat' box containing the overlay information generated in this way is created. Proceed to S153.

[0141] Generate an 'iloc' box at S153. This box indicates the positions in the file where various data are placed. Since various information is stored in the memory 102, generate this box based on these sizes. Specifically, the information indicating the overlay is stored in the previous 'idat', which is the position and size information inside this 'idat'. Also, the thumbnail data and signature data 12 are stored in the'mdat' box, and this position and size information are stored. Proceed to S154.

[0142] Generate an'mdat' box at S154. This box contains a plurality of boxes described below. The CPU 101 generates the'mdat' box and proceeds to S155.

[0143] Store the Exif data in the'mdat' box at S155. In the previous S109, the Exif of the metadata was stored in the memory 102, so the CPU 101 reads this from the memory 102 and appends it to the'mdat' box. Proceed to S156.

[0144] Store the thumbnail data in the'mdat' box at S156. In the previous S110, the thumbnail data was stored in the memory 102, so the CPU 101 reads this from the memory 102 and appends it to the'mdat' box. Proceed to S157.

[0145] Store the data of the encoded image 1 in the'mdat' box at S157. In the first loop of the previous S106, the data of the divided image 1 was stored in the memory 102, so the CPU 101 reads this from the memory 102 and appends it to the'mdat' box. Repeat this until the encoded image N to append all the encoded images 1 - N to the'mdat' box. In such steps, the CPU 101 constructs a HEIF file.

[0146] Through the above processing, an HDR image can be recorded on the recording medium 120 in the HEIF format as shown in Figure 2. <Playback of HDR Image> Next, the processing performed by the imaging device 100 when playing back an HDR image file recorded in HEIF format on the recording medium 120 will be described. In this embodiment, the process described is when the imaging device 100 plays back the image, but similar processing may be implemented when an image processing device without an imaging unit plays back an HDR image recorded in HEIF format on the recording medium 120.

[0147] Figure 7 illustrates the HEIF playback (display) process when the main image is an overlay image.

[0148] In step S701, the CPU 101 uses the recording medium control unit 119 to read the beginning portion of the specified file present on the recording medium 120 into memory 102. The CPU then checks whether a file type box with the correct structure exists at the beginning of the read file, and whether the brand within it contains 'mif1', which represents HEIF.

[0149] Furthermore, if a brand corresponding to a unique file structure is recorded, a check for the existence of that brand is performed. As long as that brand guarantees a specific structure, this brand check may allow for the omission of several subsequent structure checks, such as steps S703 and S704.

[0150] In step S702, the CPU 101 reads the metadata box of the specified file from the recording medium 120 into memory 102.

[0151] In step S703, CPU101 searches for the handler box in the metadata box read in step S702 and checks its structure. If it is HEIF, the handler type must be 'pict'.

[0152] In step S704, the CPU 101 searches for the data information box in the metadata box read in step S702 and checks its structure. In this embodiment, it is assumed that the data exists within the same file, so it checks that the flag indicating this is set in the data entry URL box.

[0153] In step S705, CPU 101 searches for the primary item box in the metadata box read in step S702 and obtains the item ID of the main image.

[0154] In step S706, the CPU 101 searches for an item information box in the metadata box read in step S702 and obtains an item information entry corresponding to the item ID of the main image obtained in step S705. In the case of an HEIF format image file recorded by the HDR image capture process described above, the overlay information is specified as the main image. In other words, the item information entry corresponding to the item ID of the main image has an item type of 'iovl', which represents an overlay.

[0155] In step S707, the CPU 101 executes the overlay image creation process. This will be explained later using Figure 8.

[0156] In step S708, the CPU 101 displays the overlay image created in step S707 on the display unit 116 via the display control unit 115.

[0157] When displaying the created overlay image, it may be necessary to specify the image's color space information. The color space information of the image is specified by the color space property 'colr', where the color gamut is specified by `colour_primaries` and the conversion characteristics (equivalent to gamma) are specified by `transfer_characteristics`, so these values ​​should be used. For example, HDR images may use a color gamut of Rec.ITU-R BT.2020 and conversion characteristics of Rec.ITU-R BT.2100(PQ).

[0158] Additionally, if HDR metadata such as 'MDCV' or 'CLLI' exists in the item properties container box, it can be used.

[0159] Next, we will explain the overlay image creation process using Figure 8.

[0160] In step S801, the CPU 101 performs the process of acquiring the properties of the overlay image. This will be explained later using Figure 9.

[0161] In step S802, the CPU 101 executes the process of acquiring overlay information for the overlay image. Since the overlay information is recorded in the 'idat' box, the CPU 101 will acquire the data stored in the 'idat' box.

[0162] Overlay information includes the image size of the overlay image (output_width, output_height parameters), and the position information of the individual image items (divided image items) that make up the overlay image (horizontal_offset, vertical_offset).

[0163] Overlay information is treated as image item data for image items whose item type is overlay 'iovl'.

[0164] This acquisition process will be explained later using Figure 10.

[0165] In step S803, the CPU 101 performs the process of obtaining all item IDs of the divided image items. This will be explained later using Figure 11.

[0166] In step S804, the CPU 101 uses the overlay image size from the overlay information obtained in step S802 to allocate memory to store image data of that size. This memory area will be called the canvas.

[0167] In step S805, the CPU 101 initializes the segmented image counter n to 0.

[0168] In step S806, the CPU 101 checks if the segmented image counter n is equal to the number of segmented image items. If they are equal, the process proceeds to step S2210. If they are not equal, the process proceeds to step S807.

[0169] In step S807, the CPU 101 performs the image creation process for the nth segmented image item. This will be explained later using Figure 12.

[0170] In step S808, the CPU 101 places the image of the nth divided image item (divided image), created in step S807, onto the canvas reserved in step S804, according to the overlay information obtained in step S802.

[0171] In an overlay, the segmented images that make up the overlay can be placed at any position on the canvas, and their positions are described in the overlay information. However, the area outside the overlay image is not displayed. Therefore, when placing the segmented images on the canvas, only the region of the segmented image's encoding area that overlaps with the overlay image's area, i.e., the playback area of ​​the segmented image, is placed.

[0172] The following explanation will use Figure 13.

[0173] The coordinate system is defined as having the top-left corner of the overlay image as the origin (0,0), with the X-coordinate increasing to the right and the Y-coordinate increasing downwards.

[0174] Let the size of the overlay image be width Wo and height Ho (0 < Wo, 0 < Ho).

[0175] Therefore, the overlay image has its upper left coordinate at (0, 0) and its lower right coordinate at (Wo - 1, Ho - 1).

[0176] Also, let the size of the divided image be width Wn and height Hn (0 < Wn, 0 < Hn).

[0177] Let the upper left position of the divided image be (Xn, Yn).

[0178] Therefore, the divided image has its upper left coordinate at (Xn, Yn) and its lower right coordinate at (Xn + Wn - 1, Yn + Hn - 1).

[0179] The overlapping area between the divided image and the overlay image (canvas) can be obtained by the following method.

[0180] In the following cases, there is no overlapping area. Wo - 1 < Xn (the left end of the divided image is to the right of the right end of the overlay image) Xn + Wn - 1 < 0 (the right end of the divided image is to the left of the left end of the overlay image) Ho - 1 < Yn (the upper end of the divided image is below the lower end of the overlay image) Yn + Hn - 1 < 0 (the lower end of the divided image is above the upper end of the overlay image) In this case, this image is not a processing target.

[0181] Also, in the following cases, the entire divided image becomes the overlapping area. 0 <= Xn and (Xn + Wn - 1) <= (Wo - 1) and 0 <= Yn and (Yn + Hn - 1) <= (Ho - 1) In this case, place the entire divided image at the specified position (Xn, Yn) on the canvas.

[0182] In cases other than the above, a part of the divided image is the placement target.

[0183] Let the top-left coordinate of the overlapping region be (Xl, Yt) and the bottom-right coordinate be (Xr, Yb).

[0184] The leftmost Xl is as follows: if(0<=Xn) Xl = Xn; else Xl=0; The rightmost Xr is as follows: if(Xn+Wn-1<=Wo-1) Xr = Xn + Wn - 1; else Xr=Wo-1; The upper limit Yt is as follows: if(0<=Yn) Yt = Yn; else Yt=0; The lower limit Yb is as follows: if(Yn+Hn-1<=Ho-1) Yb = Yn + Hn - 1; else Yb=Ho-1; The size of this overlapping region is width Xr-Xl+1 and height Yb-Yt+1.

[0185] The top-left coordinates (Xl, Yt) and bottom-right coordinates (Xr, Yb) of this overlapping region are, as mentioned above, part of a coordinate system with the top-left corner of the canvas as the origin (0,0).

[0186] The top-left coordinates (Xl, Yt) of the overlapping region can be expressed in a coordinate system with the top-left corner of the divided image as the origin, as follows:

[0187] Xl' = Xl - Xn; Yt' = Yt - Yn; In summary, it is as follows:

[0188] A rectangle with width (Xr-Xl+1) and height (Yb-Yt+1) is cut out from the top left of the divided image at a distance of (Xl', Yt'), and placed at the (Xl, Yt) position on the canvas.

[0189] This makes it possible to position the playback area of ​​the divided image in the appropriate location on the canvas.

[0190] In step S809, the CPU 101 increments the segmented image counter n by 1 and returns to step S806.

[0191] In step S810, the CPU 101 checks if a rotation property 'irot' exists in the properties of the overlay image obtained in step S801, and if it exists, it checks the rotation angle. If the rotation property does not exist, or if the rotation property exists but the rotation angle is 0, the process terminates. If the rotation property exists and the rotation angle is not 0, the process proceeds to step S811.

[0192] In step S811, the CPU 101 rotates the canvas image created in steps S806 to S809 by the angle obtained in step S810, and uses that as the created image.

[0193] This makes it possible to create overlay images.

[0194] In the above explanation, steps S806 to S809 are repeated in a loop for the number of divided images. However, in environments where parallel processing is possible, steps S806 to S809 may be processed in parallel for the number of divided images.

[0195] Next, we will explain the process of obtaining the properties of an image item using Figure 9.

[0196] In step S901, CPU101 searches for an entry for the specified image item ID in the Item Property Related Box within the Item Property Box within the Metadata Box, and retrieves the array of property indices stored therein.

[0197] In step S902, CPU101 initializes array counter n to 0.

[0198] In step S903, CPU101 checks if the array counter n is equal to the number of array elements. If they are equal, it terminates. If they are not equal, it proceeds to step S904.

[0199] In step S904, CPU101 retrieves the property of the index of the nth element of the array from the item property container box within the item property box within the metadata box.

[0200] In step S905, CPU101 increments array counter n by 1 and returns to step S903.

[0201] Next, we will explain the data acquisition process for image items using Figure 10.

[0202] In step S1001, CPU101 searches for the entry for the specified image item ID in the item location box within the metadata box and retrieves the offset basis (construction_method), offset, and length.

[0203] In step S1002, CPU 101 checks the offset criterion obtained in step S1001. The offset criterion represents the offset from the beginning of the file (value 0) and the offset within the item data box (value 1). If the offset criterion is 0, proceed to step S1003. If the offset criterion is 1, proceed to step S1004.

[0204] In step S1003, the CPU 101 reads a number of bytes from the beginning of the file, starting from the offset byte position, into memory 102.

[0205] In step S1004, the CPU 101 reads the data portion of the item data box within the metadata box, starting from the offset byte position and ending in bytes, into memory 102.

[0206] Next, we will explain the process of obtaining the item IDs of the images that make up the main image, using Figure 11.

[0207] In step S1101, CPU101 searches the item reference box within the metadata box for an entry where the reference type is 'ding' and the source item ID is the item ID of the main image.

[0208] In step S1102, CPU 101 obtains an array of referenced item IDs for the entries obtained in step S1101.

[0209] Next, we will explain the image creation process for a single encoded image item using Figure 12.

[0210] In step S1201, CPU 101 retrieves the properties of the image item. This has been explained using Figure 9.

[0211] In step S1202, the CPU 101 acquires the data for the image item. This has already been explained using Figure 10. The data for the image item is the encoded image data.

[0212] In step S1203, the CPU 101 initializes the decoder using the decoder configuration and initialization data among the properties obtained in step S1201.

[0213] In step S1204, the CPU 101 decodes the encoded image data acquired in step S1202 using a decoder and obtains the decoding result.

[0214] In step S1205, CPU101 checks whether a pixel format conversion is required.

[0215] If unnecessary, terminate. If necessary, proceed to step S1206.

[0216] If the pixel format of the decoder's output data differs from the pixel format of the image supported by the display device, pixel format conversion is required.

[0217] For example, if the pixel format of the decoder's output data is YCbCr (luminance-to-chrominance) format, and the pixel format of the display device's image is RGB format, then pixel format conversion from YCbCr to RGB format is necessary. Furthermore, even within the same YCbCr format, pixel format conversion is required if the bit depth (8-bit, 10-bit, etc.) or color difference sampling (4:2:0, 4:2:2, etc.) differs.

[0218] Furthermore, the matrix_coefficients property of the color space information property 'colr' specifies the coefficients used when converting from RGB format to YCbCr format. Therefore, to convert from YCbCr format to RGB format, you can use the reciprocal of those coefficients.

[0219] In step S1206, the CPU 101 converts the decoder output data acquired in step S1204 into a desired pixel format.

[0220] This makes it possible to create an image of the specified encoded image item.

[0221] Through the above process, it is possible to play back the HDR image recorded on the recording medium 120 in HEIF format as shown in Figure 2.

[0222] <Other Embodiments> Although the present invention has been described above based on embodiments, the present invention is not limited to these embodiments, and various forms that do not depart from the spirit of the invention are also included in the present invention.

[0223] Alternatively, the functions of the above embodiment may be controlled by an image processing device, and this control method may be executed by the image processing device. Alternatively, a program having the functions of the above embodiment may be used as a control program, and this control program may be executed by a computer provided in the image processing device. The control program may be recorded, for example, on a storage medium readable by the computer.

[0224] The present invention provides a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium. It can also be implemented by a process in which one or more processors in the computer of the system or device read and execute the program. Furthermore, it can also be implemented by a circuit (e.g., an ASIC) that implements one or more functions.

Claims

[Claim 1] A control method for an imaging device that records HDR (high dynamic range) image data obtained by capturing images with an imaging sensor, The HDR image data captured by the aforementioned imaging sensor is divided into a plurality of segmented HDR image data, and each of the plurality of segmented HDR image data is encoded in an encoding step, A recording control step that controls the recording of the multiple encoded segmented HDR image data as image items in HEIF (High Efficiency Image File Format) format into a single image file, and the recording of image structure information for combining the multiple segmented HDR image data into their pre-segmentation state as a derived image item in the image file. The setup process for setting the recording mode, It has, The vertical and horizontal dimensions of the segmented HDR image data differ depending on the vertical and horizontal dimensions of the image corresponding to the recording mode set in the setting step. A control method for an imaging device, characterized by the following:

Citation Information

Patent Citations

  • Image data generating apparatus, image data generating method, and program

    JP2017139618A