Moving picture compression apparatus

The video compression device optimizes the compression process for stacked imaging elements by setting search areas based on varied imaging conditions, improving efficiency and accuracy in handling frames with diverse settings.

JP2026031648APending Publication Date: 2026-02-24NIKON CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025216757
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2017-09-29
Filing Date
2025-12-02
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing video compression technologies do not effectively handle frames captured under multiple imaging conditions in stacked imaging elements, where each area of the image sensor can have different settings.

Method used

A video compression device and program that utilize a setting unit to determine a search area based on multiple imaging conditions and generate motion vectors by detecting specific areas within reference frames, optimizing the compression process for stacked imaging elements with varied imaging conditions.

Benefits of technology

Enhances the efficiency and accuracy of video compression by adapting to different imaging conditions across the image sensor, reducing processing load and maintaining high compression quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026031648000001_ABST
    Figure 2026031648000001_ABST
Patent Text Reader

Abstract

To optimize moving image compression when a plurality of imaging conditions are set.SOLUTION: A moving image compression device compresses moving image data which are a series of frames output from an imaging element having a plurality of imaging regions for imaging a subject and capable of setting imaging conditions for each imaging region. The moving image compression device includes a setting unit configured to set, based on the plurality of imaging conditions, a search area in a reference frame used for processing of detecting a specific area from the reference frame based on a compression target area, and a generation unit configured to generate a motion vector by detecting the specific area based on the processing using the search area set by the setting unit.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

Incorporation by Reference

[0001] This application claims priority from Japanese Patent Application No. 2017-192105, filed on September 29, 2017, the contents of which are incorporated herein by reference. [Technical Field]

[0002] The present invention relates to a video compression device, an electronic device, and a video compression program. [Background technology]

[0003] An electronic device has been proposed that includes an imaging element in which a back-illuminated imaging chip and a signal processing chip are stacked (hereinafter referred to as a stacked imaging element) (see Patent Document 1). In the stacked imaging element, the back-illuminated imaging chip and the signal processing chip are stacked so that they are connected via microbumps in each predetermined area. However, when a stacked imaging element allows multiple imaging conditions to be set within the imaging area, frames captured under the multiple imaging conditions are output, and video compression of such frames has not been considered in the past. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2006-49361 Summary of the Invention

[0005] A video compression device that is one aspect of the technology disclosed in the present application is a video compression device that compresses video data that is a series of frames output from an image sensor that has a plurality of imaging areas for imaging a subject and in which imaging conditions can be set for each of the imaging areas, and has: a setting unit that sets a search area in the reference frame that is used in a process of detecting a specific area from a reference frame based on the area to be compressed, based on the plurality of imaging conditions; and a generation unit that generates a motion vector by detecting the specific area based on the process using the search area set by the setting unit.

[0006] An electronic device that is one aspect of the technology disclosed in the present application includes an image sensor that has a plurality of imaging areas for imaging a subject, imaging conditions can be set for each of the imaging areas, and outputs video data that is a series of frames; a setting unit that sets a search area in the reference frame based on the plurality of imaging conditions, which is used in a process of detecting a specific area from a reference frame based on a region to be compressed; and a generation unit that generates a motion vector by detecting the specific area based on the process using the search area set by the setting unit.

[0007] A video compression program that is one aspect of the technology disclosed in the present application is a video compression program that causes a processor to compress video data, which is a series of frames output from an image sensor that has multiple imaging areas for imaging a subject and for which imaging conditions can be set for each imaging area, and causes the processor to set, based on the multiple imaging conditions, a search area within the reference frame that is used in a process of detecting a specific area from a reference frame based on an area to be compressed, and generates a motion vector by detecting the specific area based on the process using the search area. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a cross-sectional view of a stacked imaging device. [Figure 2] FIG. 2 is a diagram illustrating the pixel arrangement of the imaging chip. [Figure 3]FIG. 3 is a circuit diagram of the imaging chip. [Figure 4] FIG. 4 is a block diagram showing an example of the functional configuration of the imaging element. [Figure 5] FIG. 5 is an explanatory diagram illustrating an example of a block configuration of an electronic device. [Figure 6] FIG. 6 is an explanatory diagram showing an example of the structure of a moving image file. [Figure 7] FIG. 7 is an explanatory diagram showing the relationship between the imaging surface and the subject image. [Figure 8] FIG. 8 is an explanatory diagram showing a specific example of the structure of a moving image file. [Figure 9] FIG. 9 is an explanatory diagram showing an example of block matching. [Figure 10] FIG. 10 is a block diagram illustrating an example of the configuration of the control unit illustrated in FIG. [Figure 11] FIG. 11 is a block diagram showing an example of the configuration of the compression unit. [Figure 12] FIG. 12 is an explanatory diagram showing an example of a search range, a search area, and a search window. [Figure 13] FIG. 13 is an explanatory diagram showing a first scanning example at the boundary between different imaging conditions. [Figure 14] FIG. 14 is an explanatory diagram showing a second scanning example at the boundary between different imaging conditions. [Figure 15] FIG. 15 is an explanatory diagram showing a third scanning example at the boundary between different imaging conditions. [Figure 16] FIG. 16 is an explanatory diagram showing an example of region enlargement / reduction of the imaging conditions. [Figure 17] FIG. 17 is an explanatory diagram showing a fourth scanning example at the boundary between different imaging conditions. [Figure 18] FIG. 18 is an explanatory diagram showing a fifth scanning example at the boundary between different imaging conditions. [Figure 19] FIG. 19 is an explanatory diagram showing a sixth scanning example at the boundary between different imaging conditions. [Figure 20] FIG. 20 is a flowchart illustrating an example of a preprocessing procedure performed by the preprocessing unit. [Figure 21]FIG. 21 is a flowchart illustrating a first example of a procedure for motion detection processing by the motion detection unit. [Figure 22] FIG. 22 is a flowchart illustrating a second example of a procedure for motion detection processing by the motion detection unit. [Figure 23] FIG. 23 is an explanatory diagram showing an example of block matching at different pixel precisions. [Figure 24] FIG. 24 is a flowchart illustrating an example of a procedure for detecting a motion vector at different pixel precisions by the motion detection unit. DETAILED DESCRIPTION OF THE INVENTION

[0009] <Example of imaging element configuration> First, we will explain a stacked imaging element to be mounted on an electronic device. This stacked imaging element is described in Japanese Patent Application No. 2012-139026, previously filed by the applicant of the present application. The electronic device is, for example, an imaging device such as a digital camera or a digital video camera.

[0010] 1 is a cross-sectional view of a stacked imaging element 100. The stacked imaging element (hereinafter simply referred to as "imaging element") 100 includes a back-illuminated imaging chip (hereinafter simply referred to as "imaging chip") 113 that outputs pixel signals corresponding to incident light, a signal processing chip 111 that processes the pixel signals, and a memory chip 112 that stores the pixel signals. The imaging chip 113, signal processing chip 111, and memory chip 112 are stacked and electrically connected to each other by conductive bumps 109 such as Cu.

[0011] As shown in FIG. 1, incident light is mainly incident in the positive direction of the Z axis, as indicated by the white arrow. In this embodiment, the surface of the imaging chip 113 on which the incident light is incident is referred to as the back surface. As shown by the coordinate axes, the left direction on the paper, perpendicular to the Z axis, is the positive X axis, and the front direction on the paper, perpendicular to the Z axis and the X axis, is the positive Y axis. In the following figures, the coordinate axes are displayed so that the orientation of each figure can be understood, based on the coordinate axes in FIG. 1.

[0012] An example of the imaging chip 113 is a back-illuminated MOS (Metal Oxide Semiconductor) image sensor. A PD (Photodiode) layer 106 is arranged on the back side of a wiring layer 108. The PD layer 106 has a plurality of PDs 104 arranged two-dimensionally and accumulating electric charges according to incident light, and transistors 105 provided corresponding to the PDs 104.

[0013] A color filter 102 is provided on the incident side of the PD layer 106, on which incident light is incident, via a passivation film 103. There are multiple types of color filters 102 that transmit different wavelength ranges, and each has a specific arrangement corresponding to each PD 104. The arrangement of the color filters 102 will be described later. A set of a color filter 102, a PD 104, and a transistor 105 forms one pixel.

[0014] A microlens 101 is provided corresponding to each pixel on the incident light side of the color filter 102. The microlens 101 condenses the incident light toward the corresponding PD 104.

[0015] The wiring layer 108 has wiring 107 that transmits pixel signals from the PD layer 106 to the signal processing chip 111. The wiring 107 may be multi-layered, and may be provided with passive elements and active elements.

[0016] A plurality of bumps 109 are arranged on the surface of the wiring layer 108. The plurality of bumps 109 are aligned with a plurality of bumps 109 provided on the opposing surface of the signal processing chip 111, and the imaging chip 113 and the signal processing chip 111 are pressed together, whereby the aligned bumps 109 are bonded together and electrically connected.

[0017] Similarly, a plurality of bumps 109 are arranged on the opposing surfaces of the signal processing chip 111 and the memory chip 112. These bumps 109 are aligned with each other, and the signal processing chip 111 and the memory chip 112 are pressed together, whereby the aligned bumps 109 are bonded together and electrically connected.

[0018] The bonding between the bumps 109 is not limited to Cu bump bonding by solid-phase diffusion, but may also employ micro-bump bonding by solder melting. For example, it is sufficient to provide approximately one bump 109 per block, as will be described later. Therefore, the size of the bumps 109 may be larger than the pitch of the PDs 104. Furthermore, in a peripheral region other than the pixel region where the pixels are arranged, bumps larger than the bumps 109 corresponding to the pixel region may also be provided.

[0019] The signal processing chip 111 has through-silicon vias (TSVs) 110 that connect circuits provided on the front and back surfaces of the chip to each other. The TSVs 110 are preferably provided in the peripheral region. The TSVs 110 may also be provided in the peripheral region of the imaging chip 113 and the memory chip 112.

[0020] 2 is a diagram illustrating the pixel arrangement of the imaging chip 113. In particular, it shows the imaging chip 113 observed from the back side. (a) is a plan view schematically showing an imaging surface 200, which is the back side of the imaging chip 113, and (b) is an enlarged plan view of a partial region 200a of the imaging surface 200. As shown in (b), a large number of pixels 201 are arranged two-dimensionally on the imaging surface 200.

[0021] Each pixel 201 has a color filter (not shown). The color filters are of three types: red (R), green (G), and blue (B), and the notations "R," "G," and "B" in (b) indicate the type of color filter that the pixel 201 has. As shown in (b), the pixels 201 equipped with such color filters are arranged in a so-called Bayer array on the imaging surface 200 of the image sensor 100.

[0022] The pixel 201 having a red filter photoelectrically converts light in the red wavelength band of the incident light and outputs a light reception signal (photoelectric conversion signal). Similarly, the pixel 201 having a green filter photoelectrically converts light in the green wavelength band of the incident light and outputs a light reception signal. Furthermore, the pixel 201 having a blue filter photoelectrically converts light in the blue wavelength band of the incident light and outputs a light reception signal.

[0023] The image sensor 100 is configured so that each unit group 202, each consisting of four pixels 201 (2 pixels by 2 pixels) adjacent to each other, can be individually controlled. For example, when charge accumulation starts simultaneously in two different unit groups 202, charge readout, i.e., light reception signal readout, occurs 1 / 30 seconds after the start of charge accumulation in one unit group 202, and charge readout occurs 1 / 15 seconds after the start of charge accumulation in the other unit group 202. In other words, the image sensor 100 can set a different exposure time (charge accumulation time, so-called shutter speed) for each unit group 202 in one image capture.

[0024] In addition to the exposure time described above, the image sensor 100 can also vary the amplification factor of the image signal (so-called ISO sensitivity) for each unit group 202. The image sensor 100 can change the timing for starting charge accumulation and the timing for reading out the light reception signal for each unit group 202. In other words, the image sensor 100 can change the frame rate for capturing a moving image for each unit group 202.

[0025] To summarize the above, the image sensor 100 is configured to be able to vary imaging conditions such as exposure time, amplification factor, frame rate, and resolution for each unit group 202. For example, if a readout line (not shown) for reading out imaging signals from a photoelectric conversion unit (not shown) included in the pixel 201 is provided for each unit group 202 and imaging signals can be read out independently for each unit group 202, it is possible to vary the exposure time (shutter speed) for each unit group 202.

[0026] Furthermore, if an amplifier circuit (not shown) that amplifies the image signal generated by the photoelectrically converted charge is provided independently for each unit group 202 and the amplification factor of the amplifier circuit is configured to be controllable independently for each amplifier circuit, the signal amplification factor (ISO sensitivity) can be made different for each unit group 202.

[0027] Furthermore, imaging conditions that can be varied for each unit group 202 include, in addition to the imaging conditions described above, the frame rate, gain, resolution (thinning rate), the number of rows or columns for adding pixel signals, the charge accumulation time or number of accumulations, the number of digitization bits, etc. Furthermore, the control parameters may be parameters for image processing after image signals are acquired from the pixels.

[0028] In addition, the imaging conditions can be controlled by, for example, providing the image sensor 100 with a liquid crystal panel having sections (each section corresponding to one unit group 202) that can be controlled independently for each unit group 202, and using this as a neutral density filter that can be turned on and off, thereby making it possible to control the brightness (aperture value) for each unit group 202.

[0029] The number of pixels 201 constituting the unit group 202 does not have to be the above-mentioned 2×2=4 pixels. The unit group 202 only needs to have at least one pixel 201, and conversely, the unit group 202 may have more than four pixels 201.

[0030] 3 is a circuit diagram of the imaging chip 113. In FIG. 3, a rectangle surrounded by a dotted line representatively represents a circuit corresponding to one pixel 201. Furthermore, a rectangle surrounded by a dashed line corresponds to one unit group 202 (202-1 to 202-4). Note that at least a part of the transistors described below corresponds to the transistor 105 in FIG. 1.

[0031] As described above, the reset transistors 303 of the pixels 201 are turned on / off for each unit group 202. The transfer transistors 302 of the pixels 201 are also turned on / off for each unit group 202. In the example shown in Fig. 3, a reset wiring 300-1 is provided for turning on / off the four reset transistors 303 corresponding to the upper left unit group 202-1, and a TX wiring 307-1 is also provided for supplying transfer pulses to the four transfer transistors 302 corresponding to the same unit group 202-1.

[0032] Similarly, a reset wiring 300-3 for turning on / off the four reset transistors 303 corresponding to the lower left unit group 202-3 is provided separately from the reset wiring 300-1. Also, a TX wiring 307-3 for supplying transfer pulses to the four transfer transistors 302 corresponding to the same unit group 202-3 is provided separately from the TX wiring 307-1.

[0033] Similarly, for the upper right unit group 202-2 and the lower right unit group 202-4, a reset line 300-2 and a TX line 307-2, and a reset line 300-4 and a TX line 307-4 are provided in each unit group 202, respectively.

[0034] The 16 PDs 104 corresponding to each pixel 201 are connected to the corresponding transfer transistors 302. A transfer pulse is supplied to the gate of each transfer transistor 302 via the TX wiring for each unit group 202. The drain of each transfer transistor 302 is connected to the source of the corresponding reset transistor 303, and a so-called floating diffusion FD between the drain of the transfer transistor 302 and the source of the reset transistor 303 is connected to the gate of the corresponding amplification transistor 304.

[0035] The drains of the reset transistors 303 are commonly connected to a Vdd wiring 310 to which a power supply voltage is supplied. A reset pulse is supplied to the gate of each reset transistor 303 via the reset wiring for each unit group 202.

[0036] The drains of the amplifier transistors 304 are commonly connected to a Vdd line 310 to which a power supply voltage is supplied. The source of each amplifier transistor 304 is connected to the drain of the corresponding selection transistor 305. The gate of each selection transistor 305 is connected to a decoder line 308 to which a selection pulse is supplied. The decoder line 308 is provided independently for each of the 16 selection transistors 305.

[0037] The sources of the selection transistors 305 are connected to a common output wiring 309. A load current source 311 supplies a current to the output wiring 309. That is, the output wiring 309 for the selection transistors 305 is formed by a source follower. The load current source 311 may be provided on the imaging chip 113 side or on the signal processing chip 111 side.

[0038] Here, the flow from the start of charge accumulation to pixel output after accumulation is completed will be described. When a reset pulse is applied to the reset transistor 303 through the reset wiring for each unit group 202, and at the same time a transfer pulse is applied to the transfer transistor 302 through the TX wiring for each unit group 202 (202-1 to 202-4), the potentials of the PD 104 and the floating diffusion FD are reset for each unit group 202.

[0039] When the transfer pulse is released, each PD 104 converts the incident light it receives into electric charges and stores them. After that, when the transfer pulse is applied again without the reset pulse being applied, the stored electric charges are transferred to the floating diffusion FD, and the potential of the floating diffusion FD changes from the reset potential to the signal potential after the electric charges are stored.

[0040] When a selection pulse is applied to the selection transistor 305 through the decoder wiring 308, a fluctuation in the signal potential of the floating diffusion FD is transmitted to the output wiring 309 via the amplification transistor 304 and the selection transistor 305. As a result, a pixel signal corresponding to the reset potential and the signal potential is output from the unit pixel to the output wiring 309.

[0041] As described above, the reset wiring and TX wiring are common to the four pixels forming the unit group 202. That is, the reset pulse and transfer pulse are each applied simultaneously to the four pixels in the unit group 202. Therefore, all the pixels 201 forming a certain unit group 202 start and end charge accumulation at the same timing. However, pixel signals corresponding to the accumulated charges are selectively output from the output wiring 309 by sequentially applying selection pulses to the respective selection transistors 305.

[0042] In this way, the charge accumulation start timing can be controlled for each unit group 202. In other words, different unit groups 202 can capture images at different timings.

[0043] 4 is a block diagram showing an example of the functional configuration of the image sensor 100. An analog multiplexer 411 sequentially selects the 16 PDs 104 that form a unit group 202 and outputs the pixel signals from each of the 16 PDs 104 to the output wiring 309 provided corresponding to the unit group 202. The multiplexer 411 is formed on the image sensor chip 113 together with the PDs 104.

[0044] The pixel signals output via the multiplexer 411 undergo correlated double sampling (CDS) and analog-to-digital (A / D) conversion by a signal processing circuit 412 that performs CDS and A / D conversion and is formed in the signal processing chip 111. The A / D converted pixel signals are passed to a demultiplexer 413 and stored in pixel memories 414 corresponding to the respective pixels. The demultiplexer 413 and pixel memories 414 are formed in the memory chip 112.

[0045] The arithmetic circuit 415 processes the pixel signals stored in the pixel memory 414 and passes them to a downstream image processing unit. The arithmetic circuit 415 may be provided in the signal processing chip 111 or in the memory chip 112. Note that although Fig. 4 shows connections for four unit groups 202, in reality, these exist for every four unit groups 202 and operate in parallel.

[0046] However, it is not necessary for there to be an arithmetic circuit 415 for each of the four unit groups 202; for example, one arithmetic circuit 415 may process sequentially by referring to the values ​​of the pixel memories 414 corresponding to each of the four unit groups 202 in order.

[0047] As described above, output wiring 309 is provided corresponding to each unit group 202. Since the imaging element 100 has the imaging chip 113, the signal processing chip 111, and the memory chip 112 stacked on top of each other, by using the bumps 109 for electrical connection between the chips for these output wiring 309, it is possible to route the wiring without increasing the size of each chip in the planar direction.

[0048] <Example of electronic device block configuration> 5 is an explanatory diagram showing an example block configuration of an electronic device. The electronic device 500 is, for example, a lens-integrated camera. The electronic device 500 includes an imaging optical system 501, an imaging element 100, a control unit 502, an LCD monitor 503, a memory card 504, an operation unit 505, a DRAM 506, a flash memory 507, and an audio recording unit 508. The control unit 502 includes a compression unit that compresses video data, as described below. Therefore, the electronic device 500 including at least the control unit 502 serves as a video compression device.

[0049] The imaging optical system 501 is made up of a plurality of lenses, and forms a subject image on the imaging surface 200 of the image sensor 100. For convenience, the imaging optical system 501 is illustrated as a single lens in FIG.

[0050] The imaging element 100 is, for example, an imaging element such as a CMOS (Complementary Metal Oxide Semiconductor) or a CCD (Charge Coupled Device), and captures an image of a subject formed by an imaging optical system 501 and outputs an imaging signal. The control unit 502 is an electronic circuit that controls each unit of the electronic device 500, and is composed of a processor and its peripheral circuits.

[0051] A predetermined control program is written in advance in flash memory 507, which is a non-volatile storage medium. Control unit 502 controls each unit by reading and executing the control program from flash memory 507. This control program uses DRAM 506, which is a volatile storage medium, as a working area.

[0052] The liquid crystal monitor 503 is a display device that uses a liquid crystal panel. The control unit 502 causes the image sensor 100 to repeatedly capture an image of the subject at predetermined intervals (for example, 1 / 60th of a second). Then, various types of image processing are performed on the image signal output from the image sensor 100 to create a so-called through image, which is displayed on the liquid crystal monitor 503. In addition to the through image, the liquid crystal monitor 503 also displays, for example, a setting screen for setting image capturing conditions.

[0053] The control unit 502 creates an image file (described later) based on the imaging signal output from the imaging element 100, and records the image file on a portable recording medium, such as a memory card 504. The operation unit 505 has various operation members such as push buttons, and outputs operation signals to the control unit 502 in response to the operation of these operation members.

[0054] Recording unit 508 is configured by, for example, a microphone, and converts environmental sounds into audio signals and inputs them to control unit 502. Note that control unit 502 may record the video file on a recording medium (not shown) built into electronic device 500, such as a hard disk, instead of recording the video file on memory card 504, which is a portable recording medium.

[0055] <Video file configuration example> 6 is an explanatory diagram showing an example of the structure of a moving image file. A moving image file 600 is generated during compression processing by a compression unit 902 (described later) in the control unit 502, and is stored in the memory card 504, DRAM 506, or flash memory 507. The moving image file 600 is composed of two blocks: a header section 601 and a data section 602. The header section 601 is the block located at the beginning of the moving image file 600. The header section 601 stores a file basic information area 611, a mask area 612, and an imaging information area 613 in the order described above.

[0056] The file basic information area 611 records, for example, the size and offset of each section (header section 601, data section 602, mask area 612, imaging information area 613, etc.) in the video file 600. The mask area 612 records imaging condition information and mask information, which will be described later. The imaging information area 613 records information related to imaging, such as the model name of the electronic device 500 and information about the imaging optical system 21 (for example, information about optical characteristics such as aberration). The data section 602 is a block located after the header section 601, and records image information, audio information, etc.

[0057] <Relationship between the imaging screen and the subject image> 7 is an explanatory diagram showing the relationship between the imaging surface and a subject image. (a) schematically shows the imaging surface 200 (imaging range) of the image sensor 100 and a subject image 701. In (a), the control unit 502 captures the subject image 701. The imaging in (a) may also serve as imaging performed to create, for example, a live view image (a so-called through image).

[0058] The control unit 502 executes a predetermined image analysis process on the subject image 701 obtained by capturing (a). The image analysis process is a process of detecting a main subject region and a background region, for example, by using a well-known subject detection technology (a technology that calculates feature amounts to detect an area where a predetermined subject exists). The image analysis process divides the imaging plane 200 into a main subject region 702 where the main subject exists and a background region 703 where the background exists.

[0059] In (a), the main subject region 702 is shown as a region that roughly includes the subject image 701, but the main subject region 702 may have a shape that follows the outline of the subject image 701. In other words, the main subject region 702 may be set so as to include as little as possible of anything other than the subject image 701.

[0060] The control unit 502 sets different imaging conditions for each unit group 202 in the main subject region 702 and each unit group 202 in the background region 703. For example, a faster shutter speed is set for each of the former unit groups 202 than for each of the latter unit groups 202. In this way, image blurring is less likely to occur in the main subject region 702 when imaging (c) is performed after imaging (a).

[0061] Furthermore, when the main subject region 702 is backlit due to the influence of a light source such as the sun present in the background region 703, the control unit 502 sets a relatively high ISO sensitivity and a slow shutter speed for each of the former unit groups 202. Furthermore, the control unit 502 sets a relatively low ISO sensitivity and a fast shutter speed for each of the latter unit groups 202. In this way, in the image capture of (c), it is possible to prevent crushed shadows in the main subject region 702 that is backlit and blown out highlights in the background region 703 that is bright.

[0062] The image analysis process may be a process different from the process of detecting the main subject region 702 and the background region 703 described above. For example, it may be a process of detecting parts of the entire imaging surface 200 that are brighter than a certain level (parts that are too bright) and parts that are less than a certain level (parts that are too dark). When the image analysis process is such a process, the control unit 502 sets the shutter speed and ISO sensitivity for the unit groups 202 included in the former region so that the exposure value (Ev value) is lower than that of the unit groups 202 included in the other regions.

[0063] Furthermore, the control unit 502 sets the shutter speed and ISO sensitivity for the unit groups 202 included in the latter region so that the exposure value (Ev value) is higher than that of the unit groups 202 included in the other regions. In this way, the dynamic range of the image obtained by capturing (c) can be wider than the original dynamic range of the image sensor 100.

[0064] 7(b) shows an example of mask information 704 corresponding to the imaging plane 200 shown in (a). A "1" is stored at the position of the unit group 202 belonging to the main subject region 702, and a "2" is stored at the position of the unit group 202 belonging to the background region 703.

[0065] The control unit 502 performs image analysis processing on the image data of the first frame to detect a main subject region 702 and a background region 703. As a result, the frame captured in (a) is divided into a main subject region 702 and a background region 703 as shown in (c). The control unit 502 sets different imaging conditions for each unit group 202 in the main subject region 702 and each unit group 202 in the background region 703, performs the imaging in (c), and creates image data. An example of mask information 704 at this time is shown in (d).

[0066] The mask information 704 (b) corresponding to the imaging result of (a) and the mask information 704 (d) corresponding to the imaging result of (b) are captured at different times (there is a time difference), and therefore, for example, if the subject is moving or if the user moves the electronic device 500, these two pieces of mask information 704 will have different contents. In other words, the mask information 704 is dynamic information that changes over time. Therefore, different imaging conditions are set for each frame in a certain unit group 202.

[0067] <Example of 600 video files> 8 is an explanatory diagram showing a specific example of the configuration of the video file 600. In the mask area 612, identification information 801, image capture condition information 802, and mask information 704 are recorded in the order described above.

[0068] The identification information 801 indicates that this moving image file 600 was created using a multi-imaging condition moving image capturing function. The multi-imaging condition moving image capturing function is a function for capturing moving images using the image sensor 100 with multiple imaging conditions set.

[0069] The imaging condition information 802 is information that indicates what kind of use (purpose, role) exists in the unit group 202. For example, as described above, when the imaging plane 200 (FIG. 7(a)) is divided into a main subject region 702 and a background region 703, each unit group 202 belongs to either the main subject region 702 or the background region 703.

[0070] In other words, the imaging condition information 802 is information that indicates that when this moving image file 600 was created, the unit group 202 had two uses, for example, "shooting a moving image of the main subject area at 60 fps" and "shooting a moving image of the background area at 30 fps," and that a unique number was assigned to each of these uses. For example, the number 1 is assigned to the use of "shooting a moving image of the main subject area at 60 fps," and the number 2 is assigned to the use of "shooting a moving image of the background area at 30 fps."

[0071] The mask information 704 is information that indicates the use (purpose, role) of each unit group 202. The mask information 704 is defined as "information that expresses the numbers assigned to the imaging condition information 802 in the form of a two-dimensional map in accordance with the positions of the unit groups 202." In other words, when the unit groups 202 arranged two-dimensionally are specified by two-dimensional coordinates (x, y) using two integers x and y, the use of the unit group 202 present at the position (x, y) is expressed by the number present at the position (x, y) of the mask information 704.

[0072] For example, if the number "1" is entered at the coordinate (3, 5) position in the mask information 704, it can be seen that the unit group 202 located at the coordinate (3, 5) has been assigned the purpose of "capturing the main subject area at 60 [fps]." In other words, it can be seen that the unit group 202 located at the coordinate (3, 5) belongs to the main subject area 702.

[0073] It should be noted that the mask information 704 is dynamic information that changes for each frame, and is therefore recorded for each frame, that is, for each data block Bi (described later), during the compression process (not shown).

[0074] Data blocks B1 to Bn are stored as moving image data in the order of capture for each frame F (F1 to Fn) in the data section 602. Each data block Bi (i is an integer satisfying the condition 1≦i≦n) includes mask information 704, image information 811, a Tv value map 812, an Sv value map 813, a Bv value map 814, Av value information 815, audio information 816, and additional information 817.

[0075] Image information 811 is information in which the imaging signal output from the imaging element 100 by the imaging of FIG. 7C is recorded in a form before various image processing is performed, and is so-called RAW image data.

[0076] The Tv value map 812 is information in which the Tv values ​​indicating the shutter speeds set for each unit group 202 are expressed in the form of a two-dimensional map according to the positions of the unit groups 202. For example, the shutter speed set for the unit group 202 located at the coordinates (x, y) can be determined by checking the Tv value stored at the coordinates (x, y) of the Tv value map 812.

[0077] The Sv value map 813 is information in which the Sv value indicating the ISO sensitivity set for each unit group 202 is expressed in the form of a two-dimensional map, similar to the Tv value map 812 .

[0078] The Bv value map 814 is information that represents the subject brightness measured for each unit group 202 during the imaging of Figure 7 (c), i.e., the Bv value representing the brightness of the subject light incident on each unit group 202, in the form of a two-dimensional map, similar to the Tv value map 812.

[0079] The Av value information 815 is information that represents the aperture value at the time of capturing the image in (c) of Fig. 7. Unlike the Tv value, Sv value, and Bv value, the Av value is not a value that exists for each unit group 202. Therefore, unlike the Tv value, Sv value, and Bv value, the Av value stores only a single value, and is not information in which multiple values ​​are mapped two-dimensionally.

[0080] To facilitate video playback, the audio information 816 is divided into information for each frame, multiplexed with the data block Bi, and stored in the data section 602. Note that the audio information 816 may be multiplexed not for each frame, but for each predetermined number of frames. Note that the audio information 816 does not necessarily have to be included.

[0081] The additional information 817 is information that expresses, in the form of a two-dimensional map, the frame rate set for each unit group 202 when capturing the image of (c) in Fig. 7. The setting of the additional information 817 will be described later with reference to Figs. 14 and 15. The additional information 817 may be stored in the frame F, or may be stored in a cache memory of the processor 1001, which will be described later. In particular, when performing compression processing in real time, it is preferable to use a cache memory from the viewpoint of high-speed processing.

[0082] As described above, the control unit 502 performs imaging using such a video imaging function, and records on the memory card 504 a video file 600 in which image information 811 generated by the imaging element 100, in which imaging conditions can be set for each unit group 202, and data related to the imaging conditions for each unit group 202 (imaging condition information 802, mask information 704, Tv value map 812, Sv value map 813, Bv value map 814, etc.) are associated.

[0083] <Block matching example> Next, block matching when multiple imaging conditions are set for one frame will be described. In this embodiment, when there are differences in imaging conditions within the search range used in block matching, the differences are utilized to optimize block matching, thereby reducing the processing load of block matching and suppressing a decrease in block matching accuracy.

[0084] 9 is an explanatory diagram showing an example of block matching. Electronic device 500 has the above-mentioned image sensor 100 and control unit 502. Control unit 502 includes a pre-processing unit 900, an image processing unit 901, and a compression unit 902. As described above, image sensor 100 has multiple imaging regions for capturing images of a subject.

[0085] The imaging area is a set of at least one pixel, for example, one or more unit groups 202 described above. For the imaging area, imaging conditions can be set for each unit group 202. As described above, the imaging conditions specifically include, for example, exposure time (shutter speed), amplification factor (ISO sensitivity), and resolution.

[0086] The image sensor 100 captures an image of a subject and outputs video data 910 including multiple frames to a pre-processing unit 900 in the control unit 502. Within a frame F, an area of ​​image data captured in a certain imaging area of ​​the image sensor 100 is referred to as an image area.

[0087] For example, when an imaging area is made up of one unit group 202 (2×2 pixels), the size of the corresponding image area is also the size of the unit group 202. Similarly, when an imaging area is made up of 2×2 unit groups 202 (4×4 pixels), the size of the corresponding image area is also the size of the 2×2 unit group 202.

[0088] In FIG. 9, it is assumed that the main subject (for example, the focused subject) among the subjects captured in frame F was captured under imaging condition A, and the background area was captured under imaging condition B. Here, imaging conditions A and B are of the same type but have different values. For example, if imaging conditions A and B are exposure times, imaging condition A is 1 / 500 [second] and imaging condition B is 1 / 60 [second]. When the imaging condition is exposure time, the imaging condition is stored in the Tv value map 812 of the video file 600 shown in FIG. 8.

[0089] Similarly, if the imaging condition is ISO sensitivity, the imaging condition is stored in the Sv value map 813 of the moving image file 600 shown in Fig. 8. If the imaging condition is frame rate, the imaging condition is stored in the additional information 817 of the moving image file 600 shown in Fig. 8.

[0090] The pre-processing unit 900 performs pre-processing of the image processing by the image processing unit 901 on the video data 910. Specifically, for example, when the pre-processing unit 900 receives the video data 910 (here, a collection of RAW image data) from the image sensor 100, it detects a specific subject such as a main subject using a well-known subject detection technique. The pre-processing unit 900 also predicts the imaging area of ​​the specific subject and sets the imaging conditions of the imaging area in the image sensor 100 to specific imaging conditions.

[0091] For example, when imaging condition B is set over the entire imaging surface 200, when a specific subject such as a main subject is detected and imaged, the pre-processing unit 900 outputs to the imaging element 100 to set the imaging area of ​​the imaging element 100 that captured the specific subject to imaging condition A. As a result, the imaging area of ​​the specific subject is set to imaging condition A, and the other imaging areas are set to imaging condition B.

[0092] Furthermore, the pre-processing unit 900 may specifically detect a motion vector of the specific subject from the difference between the imaging area in which the specific subject is detected in the input frame and the imaging area in which the specific subject is detected in the previously input frame, and identify the imaging area of ​​the specific subject in the next input frame. In this case, the pre-processing unit 900 outputs an instruction to the image sensor 100 to change the imaging condition A for the identified imaging area. As a result, the imaging area of ​​the specific subject is set to imaging condition A, and the other imaging areas are set to imaging condition B.

[0093] The image processing unit 901 performs image processing such as demosaic processing, white balance adjustment, noise reduction, and debayering on video data 910 input from the image sensor 100. The compression unit 902 compresses the video data 910 input from the image processing unit 901. The compression unit 902 performs compression using hybrid coding that combines motion compensation interframe prediction (MC), discrete cosine transform (DCT), and entropy coding, for example.

[0094] The compression unit 902 performs block matching when detecting motion. Block matching is an example of a process for detecting a specific region. Block matching is a technique in which a certain block in a frame F1 to be compressed is set as a target block b1, which is the region to be compressed, and a frame F2 input temporally earlier (or later) than the frame F1 is used as a reference frame to detect a block b2 within a search range SR in the frame F2 that has the highest correlation with the target block b1 as the specific region. The difference between the coordinate position of the block b2 detected by block matching and the coordinate position of the target block b1 is then detected as a motion vector mv. Squared error or absolute error is generally used as an evaluation value for the degree of correlation.

[0095] When the image data of a target block b1 in frame F1 is captured under imaging condition A, the compression unit 902 sets the same position as the target block b1 in frame F2 as a search window w. The shape of the search window w is not limited to a rectangle, as long as it is a polygon. The search range SR is a predetermined range centered on the search window w. The range of the search range SR may be determined as a compression standard. The compression unit 902 sets the overlapping area of ​​the search range SR and the range of imaging condition A in frame F2 as a search area SA, scans the search window w (represented by an arrow) within the search area SA, detects the block b2 that has the highest correlation with the target block b1, and generates a motion vector mv. Note that scanning can be performed not only in pixel units, but also in any unit, such as half-pixel units or quarter-pixel units.

[0096] The main subject is captured under imaging condition A. Target block b1 is part of the main subject. Block b2, which has the highest correlation with target block b1, exists only under imaging condition A, so it is sufficient to search only search area SA, which is the overlapping area between search range SR and the range of imaging condition A. In this way, since the search range SR can be narrowed down to search area SA in advance, the search process in block matching can be sped up. Furthermore, since the search range SR can be narrowed down to search area SA, which has the same imaging condition as that of target block b1, a decrease in the matching accuracy of block matching can be suppressed.

[0097] The control unit 502 may perform the compression process of the video data 910 from the image sensor 100 in real time or in batch. For example, the control unit 502 may temporarily store the video data 910 from the image sensor 100, the preprocessing unit 900, or the image processing unit 901 in the memory card 504, the DRAM 506, or the flash memory 507, and then read out the video data 910 and cause the compression unit 902 to perform the compression process, either automatically or when triggered by a user operation.

[0098] <Configuration example of control unit 502> Fig. 10 is a block diagram showing an example of the configuration of the control unit 502 shown in Fig. 5. The control unit 502 has a preprocessing unit 900, an image processing unit 901, an acquisition unit 1020, and a compression unit 902, and is composed of a processor 1001, a memory 1002, an integrated circuit 1003, and a bus 1004 connecting these.

[0099] The preprocessing unit 900, image processing unit 901, acquisition unit 1020, and compression unit 902 may be realized by having the processor 1001 execute a program stored in the memory 1002, or may be realized by an integrated circuit 1003 such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field-Programmable Gate Array). The processor 1001 may use the memory 1002 as a work area. The integrated circuit 1003 may use the memory 1002 as a buffer for temporarily storing various data including image data.

[0100] The pre-processing unit 900 performs pre-processing of the image processing by the image processing unit 901 on video data 910 from the image sensor 100. Specifically, for example, the pre-processing unit 900 has a detection unit 1011 and a setting unit 1012. The detection unit 1011 detects a specific subject using the well-known subject detection technique described above.

[0101] The setting unit 1012 assigns additional information 817 to each frame constituting video data 910 from the image sensor 100. The setting unit 1012 also changes the imaging area in the imaging plane 200 of the image sensor 100 where a specific subject is detected. Specifically, for example, the setting unit 1012 detects a motion vector of the specific subject from the difference between the imaging area in which the specific subject is detected in the input frame and the imaging area in which the specific subject is detected in the previously input frame, and predicts the imaging area of ​​the specific subject in the next input frame. The setting unit 1012 outputs an instruction to the image sensor 100 to change the predicted imaging area to a specific imaging condition (for example, imaging condition A).

[0102] The image processing unit 901 performs image processing on each frame of the video data 910 output from the pre-processing unit 900. Specifically, for example, as described above, the image processing unit 901 performs known image processing such as demosaic processing and white balance adjustment.

[0103] The acquisition unit 1020 stores the video data 910 output from the image processing unit 901 in the memory 1002, and outputs the multiple frames included in the video data 910 one by one in chronological order to the compression unit 902 at a predetermined timing.

[0104] The compression unit 902 compresses the input video data 910 as shown in Fig. 9. Specifically, for example, as described above, the compression unit 902 sets the overlapping area of ​​the search range SR and the range of the imaging condition A in frame F2 as the search area SA, and scans the search window w (represented by the arrow) within the search area SA to detect block b2.

[0105] <Configuration example of compression unit 902> 11 is a block diagram showing an example configuration of the compression unit 902. As described above, the compression unit 902 compresses each frame of video data 910 by hybrid coding, which combines motion-compensated inter-frame prediction (MC), discrete cosine transform (DCT), and entropy coding.

[0106] The compression unit 902 includes a subtraction unit 1101, a DCT unit 1102, a quantization unit 1103, an entropy coding unit 1104, a code amount control unit 1105, an inverse quantization unit 1106, an inverse DCT unit 1107, a generation unit 1108, a frame memory 1109, a motion detection unit 1110, and a motion compensation unit 1111. The subtraction unit 1101 to the frame memory 1109 and the motion compensation unit 1111 have the same configuration as existing compressors.

[0107] Specifically, for example, the subtraction unit 1101 subtracts a predicted frame from an input frame, the predicted frame being output from the motion compensation unit 1111, and outputs the difference data. The DCT unit 1102 performs a discrete cosine transform on the difference data from the subtraction unit 1101.

[0108] The quantization unit 1103 quantizes the differential data that has been subjected to the discrete cosine transform. The entropy coding unit 1104 entropy codes the quantized differential data and also entropy codes the motion vectors from the motion estimation unit 1110.

[0109] The code amount control unit 1105 controls the quantization by the quantization unit 1103. The inverse quantization unit 1106 inverse quantizes the difference data quantized by the quantization unit 1103 to generate differential data that has been subjected to discrete cosine transform. The inverse DCT unit 1107 performs inverse discrete cosine transform on the inverse quantized differential data.

[0110] The generation unit 1108 adds the inverse discrete cosine transformed difference data to the predicted frame from the motion compensation unit 1111 to generate a reference frame to be referenced by a frame input temporally after the input frame. The frame memory 1109 holds the reference frame obtained from the generation unit 1108.

[0111] The motion detection unit 1110 detects a motion vector using an input frame and a reference frame. The motion detection unit 1110 has a region setting unit 1121 and a motion generation unit 1122. The region setting unit 1121 sets a search region for block matching when detecting motion.

[0112] As described above, block matching is a technique in which a certain block in a frame F1 to be compressed is designated as a target block b1, a block b2 having the highest correlation with the target block b1 is detected from a search range SR in a frame F2 input temporally earlier than (or later than) frame F1, and the difference between the coordinate position of the block b2 and the coordinate position of the target block b1 is detected as a motion vector mv (see FIG. 9). Generally, squared error or absolute error is used as an evaluation value of the correlation.

[0113] When a target block b1 in frame F1 is image data captured under imaging condition A, the region setting unit 1121 sets the same position as the target block b1 in frame F2 as a search window w. The search range SR is a range centered around the search window w. The region setting unit 1121 sets the overlapping region of the search range SR and the range of imaging condition A in frame F2 as a search region SA.

[0114] The motion generation unit 1122 generates a motion vector mv from the current block b1 and block b2. Specifically, for example, the motion generation unit 1122 scans a search window w (represented by an arrow) within the search area SA set by the area setting unit 1121 to detect a block b2 that has the highest correlation with the current block b1. The motion generation unit 1122 generates the difference between the coordinate position of the block b2 and the coordinate position of the current block b1 as the motion vector mv.

[0115] In this way, the search range SR can be narrowed down to the search area SA in advance, which speeds up the search process in block matching. Also, the search range SR can be narrowed down to the search area SA that has the same imaging conditions as the target block b1, which makes it possible to suppress a decrease in the matching accuracy of block matching.

[0116] The motion compensation unit 1111 generates a predicted frame using a reference frame and the motion vector mv. Specifically, for example, the motion compensation unit 1111 performs motion compensation using a specific reference frame from among multiple reference frames stored in the frame memory 1109 and the motion vector mv.

[0117] By using a specific reference frame as the reference frame, it is possible to suppress high-load motion compensation that uses reference frames other than the specific reference frame. Also, by using a single reference frame obtained from the frame immediately preceding the input frame as the specific reference frame, it is possible to avoid high-load motion compensation and reduce the processing load of motion compensation.

[0118] <Example of scanning a search window> Next, an example of scanning the search window w will be described with reference to Fig. 12 to Fig. 19. Here, the frame F2 shown in Fig. 9 will be used for the description.

[0119] Fig. 12 is an explanatory diagram showing an example of a search range, search area, and search window. Fig. 12 is a diagram focusing on a part of frame F2, specifically, for example, the boundary between imaging conditions A and B. An image area 1200 is an area within frame F2 that corresponds to the imaging area of ​​image sensor 100, and corresponds to, for example, 4x4 pixels, i.e., a 2x2 unit group 202. In Fig. 12, frame F2 is composed of a 4x5 image area 1200.

[0120] 12, as an example, the imaging conditions are set in imaging area units corresponding to the image area 1200, that is, in 2×2 unit groups 202, but the imaging conditions may be set in imaging units consisting of one unit group 202, or may be set in units larger than the group 202, other than 2×2. The search window w has a size of 3×3 pixels.

[0121] When scanning the search window w within the search area SA, the movement generation unit 1122 may scan the search window w at the boundary between imaging conditions A and B so that it does not include even one pixel of the area of ​​imaging condition B, or may scan the areas of both imaging conditions A and B, or may scan the area of ​​only imaging condition B that is a predetermined number of pixels away from imaging condition A. These will be explained in order below. Note that the legend of FIG. 12 applies to FIGS. 13 to 19.

[0122] 13 to 19, the description will be focused on a 4×4 image area 1200 (the upper left image area 1200 is image area 1201, the upper right image area 1200 is image area 1202, the lower left image area 1200 is image area 1203, and the lower right image area 1200 is image area 1204) that includes the boundary between imaging conditions A and B. The search window w is scanned from the upper left to the right of frame F2 (thick white arrow), and when it reaches the right end, it is shifted downward by one pixel and scanned from the left end to the right end (raster scan), but a diamond scan that scans radially from the center may also be used.

[0123] The scanning width from the left end to the right end of the search window w need only be within the number of pixels (3 pixels in this example) in the scanning direction width of the search window w. Similarly, the downward shift width of the search window w is not limited to 1 pixel, but need only be within the number of pixels (3 pixels in this example) in the shift direction width of the search window w.

[0124] Fig. 13 is an explanatory diagram showing scanning example 1 at the boundary between different imaging conditions. Fig. 13 shows (a) to (e) in chronological order. In scanning example 1, the movement generation unit 1122 scans the search window w so as to include only the area of ​​imaging condition A. Therefore, in (a) to (e), the search window w does not include the area of ​​imaging condition B.

[0125] FIG. 14 is an explanatory diagram showing scanning example 2 at the boundary between different imaging conditions. FIG. 14 shows (a) to (d) in chronological order. In scanning example 2, the search window w is at the boundary between imaging conditions A and B, and the motion generation unit 1122 scans the search window w so that the area of ​​imaging condition A is always larger than the area of ​​imaging condition B. Of the nine pixels in the search window w, the area of ​​imaging condition A is six pixels in (a), six pixels in (b), five pixels in (c), and six pixels in (d). From (d) onwards, the scanning is the same as (d) and (e) in FIG.

[0126] FIG. 15 is an explanatory diagram showing scanning example 3 at the boundary between different imaging conditions. FIG. 15 shows (a) to (d) in chronological order. In scanning example 2, the search window w is at the boundary between imaging conditions A and B, and the motion generation unit 1122 scans the search window w so that the area of ​​imaging condition B is as large as possible than the area of ​​imaging condition A. Of the nine pixels in the search window w, the area of ​​imaging condition B is six pixels in (a), six pixels in (b), and six pixels in (c). At the scanning position (d), which is one row down from (c), the area of ​​imaging condition B is three pixels at the bottom of the image area 1201. From (d) onwards, the scanning is the same as (d) and (e) in FIG. 13.

[0127] 16 is an explanatory diagram showing an example of area enlargement / reduction under different imaging conditions. (a) shows an example of enlargement, and (b) shows an example of reduction. When image areas under different imaging conditions are adjacent to each other, the video compression device enlarges or reduces the search area SA from the viewpoint of reducing the processing load of block matching or preventing a decrease in accuracy of block matching.

[0128] For example, let the imaging conditions of adjacent image areas be imaging conditions A and B. The image area under imaging condition A is the image area of ​​a specific subject, and the image area under imaging condition B is the image area of ​​the background. The imaging conditions are also exposure time (shutter speed).

[0129] For example, if the exposure time of imaging condition A is 1 / 30 [second] and the exposure time of imaging condition B is 1 / 60 [second], the difference in exposure time is small, so there is a possibility that the main subject is present in the image area of ​​imaging condition B. Therefore, if the difference between imaging conditions A and B is equal to or less than threshold T1, the area setting unit 1121 expands the search area SA of imaging condition A. This makes it possible to suppress a decrease in the accuracy of block matching.

[0130] Furthermore, for example, if the exposure time under imaging condition A is 1 / 30 [second] and the exposure time under imaging condition B is 1 / 1000 [second], the difference in exposure time is large, and therefore it is highly likely that the image of the main subject will not be present in the image area under imaging condition B. Therefore, if the difference between imaging conditions A and B exceeds threshold T2 (T2≧T1), area setting unit 1121 reduces the search area SA under imaging condition A. This reduces the processing load of block matching.

[0131] Furthermore, for example, if the subject includes a dark space and a bright space and the main subject is in the dark space, the exposure time needs to be longer for the dark space than for the bright space. Therefore, the area setting unit 1121 sets the exposure time of imaging condition A for the dark space and the exposure time of imaging condition B for the bright space (imaging condition A is longer than imaging condition B). Then, the area setting unit 1121 reduces the search area SA for imaging condition A. This reduces the processing load of block matching.

[0132] Although the case where the imaging condition is exposure time has been described here, the same applies when the imaging condition is ISO sensitivity or resolution. An example of area enlargement / reduction of the imaging condition will now be described in detail.

[0133] In (a), the area setting unit 1121 expands the outer edge 1600 of the area under imaging condition A outward, that is, toward the area under imaging condition B, to set it as outer edge 1601. However, the imaging condition between outer edge 1600 and outer edge 1601 remains the same as that of imaging condition B. The expanded search area SA1 is an area within the search range SR that includes the area under imaging condition A and the area between outer edge 1600 and outer edge 1601. As a result, since the search area SA has been expanded to search area SA1, the area outside the search area SA can also be subject to block matching, and the accuracy of block matching can be improved compared to when searching the search area SA.

[0134] In (b), the area setting unit 1121 reduces the outer edge 1600 of the area under imaging condition A inward, i.e., toward the area under imaging condition A, to set it as outer edge 1602. However, the imaging condition between outer edge 1600 and outer edge 1602 remains the same as that of imaging condition A. The search area SA2 after reduction is the area obtained by excluding the area between outer edge 1600 and outer edge 1602 from search area SA. As a result, because search area SA has been reduced to search area SA2, the area under imaging condition A outside search area SA2 can be excluded from block matching, and the block matching process can be performed faster than when searching search area SA.

[0135] In Figure 16, the area setting unit 1121 enlarges the search area SA after it is set, but when setting the search area SA, the search area SA1 may also be set to include the image area of ​​imaging condition B that surrounds the image area of ​​imaging condition A.

[0136] FIG. 17 is an explanatory diagram showing scanning example 4 at the boundary between different imaging conditions. FIG. 17 shows (a) to (d) in chronological order. In scanning example 4, the movement generation unit 1122 scans the search window w in the search area SA1 shown in (a) of FIG. 16. In scanning example 4, the movement generation unit 1122 scans the search window w so that it touches the boundary between imaging conditions A and B as much as possible and includes the area of ​​imaging condition B. At the scanning position (c), which is one row down from (b), the search window w includes the area of ​​imaging condition A. Similarly, at the scanning position (d), which is one row down from (c), the search window w includes the area of ​​imaging condition A.

[0137] Fig. 18 is an explanatory diagram showing scanning example 5 at the boundary between different imaging conditions. Fig. 18 shows (a) to (d) in chronological order. Scanning example 5 is an example of scanning by the movement generation unit 1122 of the search window w in the search area SA1 shown in Fig. 16(a). In scanning example 5, the movement generation unit 1122 scans the search window w so as to avoid contact with the boundary between imaging conditions A and B as much as possible and to include the area of ​​imaging condition B.

[0138] In (a), the search window w is located one pixel to the left and above imaging condition A. In (b), the search window w is located one pixel to the left from imaging condition A. At the scanning position (c), which is one row down from (b), the search window w includes the area of ​​imaging condition A. Similarly, at the scanning position (d), which is one row down from (c), the search window w includes the area of ​​imaging condition A.

[0139] 19 is an explanatory diagram showing scanning example 6 at the boundary between different imaging conditions. FIG. 19 shows (a) to (d) in chronological order. In scanning example 6, the motion generator 1122 scans the search window w in the search area SA2 shown in FIG. 16(b). In scanning example 6, the motion generator 1122 always scans the search window w so as not to include the area of ​​imaging condition B.

[0140] In this way, it is possible to adjust the reduction in processing load for block matching and the prevention of accuracy degradation at the boundary of image areas under different imaging conditions A and B. In particular, the motion generation unit 1122 can specialize in reducing the processing load for block matching at the boundary of image areas under different imaging conditions by performing block matching so that the search window w includes only pixels under imaging condition A.

[0141] In addition, the motion generation unit 1122 performs block matching so that the number of pixels under imaging condition A within the search window w is greater than the number of pixels under imaging condition B, thereby prioritizing reduction in processing load for block matching at the boundary of image areas with different imaging conditions while suppressing accuracy degradation.

[0142] In addition, the motion generation unit 1122 performs block matching so that at least one pixel of imaging condition A exists within the search window w, thereby suppressing accuracy degradation at the boundary of image areas with different imaging conditions while maintaining a reduced processing load for block matching.

[0143] <Example of pre-processing procedure> Fig. 20 is a flowchart showing an example of a preprocessing procedure by the preprocessing unit 900. Fig. 20 illustrates an example in which imaging condition B is set in advance in the image sensor 100, and the image area under imaging condition A is tracked using the subject detection technology of the detection unit 1011 and fed back to the image sensor 100. Note that the image areas under imaging conditions A and B may be fixed at all times.

[0144] The preprocessing unit 900 waits for input of frames constituting the video data 910 (step S2001: No), and if a frame is input (step S2001: Yes), it determines whether a specific subject such as a main subject has been detected by the detection unit (step S2002). If a specific subject has not been detected (step S2002: No), the process proceeds to step S2001.

[0145] On the other hand, if a specific subject is detected (step S2002: Yes), the preprocessing unit 900 causes the detection unit 1011 to compare the input frame with the previous frame in time (for example, a reference frame) to detect a motion vector, predict the image area of ​​imaging condition A in the next input frame, and output this to the image sensor 100 (step S2003), before proceeding to step S2001. As a result, the image sensor 100 sets the imaging condition of the unit groups 202 that constitute the imaging area corresponding to the predicted image area to imaging condition A, sets the imaging condition of the remaining unit groups 202 to imaging condition B, and images the subject.

[0146] Then, the process returns to step S2001. If no frames are input (step S2001: No) and input of all frames constituting the video data 910 has been completed, the process ends.

[0147] <Motion vector detection processing procedure> Next, a description will be given of an example of a detection process procedure for a motion vector mv by the motion detection unit 1110. The following flowchart shows an example of a detection process procedure for a motion vector mv under imaging condition A where a specific subject image may exist.

[0148] FIG. 21 is a flowchart showing a first example of a motion detection procedure performed by the motion detection unit 1110. First, the motion detection unit 1110 acquires an input frame to be compressed and a reference frame in a frame memory (step S2101). The motion detection unit 1110 sets a search range SR for imaging condition A in the reference frame (step S2102). Specifically, for example, the motion detection unit 1110 sets a target block b1 in the input frame from the image area of ​​imaging condition A, and sets a search window w at the same position as the target block b1 in the reference frame. Then, the motion detection unit 1110 sets a search range SR in the reference frame centered on the search window w (see FIG. 12).

[0149] Next, the motion detection unit 1110 identifies a search area SA that is within the search range SR and is an image area under the imaging condition A (step S2103). Then, the motion detection unit 1110 causes the motion generation unit 1122 to perform block matching of the target block b1 by scanning a search window w in the identified search area SA (step S2104), and generates a motion vector mv from block b2 to the target block b1 (step S2105).

[0150] The motion detection unit 1110 uses block matching to detect, for example, block b2 that has the highest correlation with the current block b1 from the reference frame, and generates a motion vector mv that is the difference between the coordinate position of block b2 and the coordinate position of the current block b1. For example, a squared error or absolute error is used as an evaluation value of the correlation. This completes the series of processes.

[0151] In this way, the search range SR can be narrowed down to the search area SA in advance, which speeds up the search process in block matching. Also, the search range SR can be narrowed down to the search area SA that has the same imaging conditions as the target block b1, which makes it possible to suppress a decrease in the matching accuracy of block matching.

[0152] Fig. 22 is a flowchart showing a second example of a motion detection process procedure by the motion detection unit 1110. In the second example of the motion detection process procedure in Fig. 22, a process example of expanding and contracting a search area SA of an imaging condition will be described. Note that the imaging conditions A and B may be set by a user operating the operation unit 505, or may be automatically set by the electronic device 500 according to the amount of light received by each unit group 202 in the imaging element 100. Furthermore, the same process contents as those in Fig. 21 are assigned the same step numbers, and their description will be omitted.

[0153] After step S2103, the motion detection unit 1110 uses the region setting unit 1121 to identify an image region of imaging condition B that is within the search range SR and adjacent to imaging condition A (step S2204). Then, the motion detection unit 1110 uses the region setting unit 1121 to enlarge or reduce the search region SA identified in step S2103 based on the image region of imaging condition A and the image region of imaging condition B that is adjacent to the image region (step S2205). Specifically, for example, the region setting unit 1121 enlarges or reduces the search region SA as shown in FIG. 16.

[0154] As in steps S2104 and S2105, the motion detection unit 1110 performs block matching of the target block b1 by scanning the search window w in the enlarged or reduced search area SA using the motion generation unit 1122 (step S2206), and generates a motion vector mv from block b2 to the target block b1 (step S2207).

[0155] This makes it possible to selectively suppress a decrease in accuracy of block matching or reduce the processing load of block matching depending on the expansion or contraction of the search area SA.

[0156] <Examples of block matching with different pixel accuracy> Next, an example of block matching at different pixel precisions will be described. In the above example, regardless of the type of imaging condition, the motion detection unit 1110 performs block matching on the search area SA (including after scaling) of imaging condition A within the search range SR, but does not perform block matching on the remaining image area of ​​the search range SR. Here, an example of performing block matching at different pixel precisions within the search area SA will be described.

[0157] Figure 23 is an explanatory diagram showing an example of block matching at different pixel accuracy. The same components as in Figure 9 are assigned the same reference numerals and their description will be omitted. In Figure 23, pixel accuracy is used, and the motion detection unit 1110 performs block matching at a certain pixel accuracy PA1 for the search area SA (which may include scaling; hereinafter referred to as the first search area SA10) of imaging condition A within the search range SR, and performs block matching at a pixel accuracy PA2 lower than the pixel accuracy PA1 for the remaining image area of ​​the search range SR (hereinafter referred to as the second search area SA20).

[0158] For example, the pixel accuracy PA1 in the first search area SA10 is 1 / 2 pixel accuracy, and the pixel accuracy PA2 in the second search area SA20 is integer pixel accuracy. The combination of pixel accuracy PA1 and PA2 is not limited to the above, as long as the pixel accuracy PA1 is higher in accuracy than the pixel accuracy PA2. For example, the pixel accuracy PA1 in the first search area SA10 may be 1 / 4 pixel accuracy, and the pixel accuracy PA2 in the second search area SA20 may be 1 / 2 pixel accuracy.

[0159] Furthermore, the motion detection unit 1110 may determine the pixel accuracy of each of the first search area SA10 and the second search area SA20 based on the respective imaging conditions A and B (or the difference between A and B). For example, 1-pixel accuracy, 1 / 2-pixel accuracy, and 1 / 4-pixel accuracy are applicable as pixel accuracy. If the exposure time of imaging condition A is 1 / 30 [second] and the exposure time of imaging condition B is 1 / 60 [second], the difference in exposure time is small, so there is a possibility that the main subject is present in the image area of ​​imaging condition B.

[0160] Therefore, if the difference between the imaging conditions A and B is equal to or smaller than the threshold value T1, the area setting unit 1121 sets the pixel accuracy PA1 of the first search area SA10 to 1 / 2 pixel accuracy and the pixel accuracy PA2 of the second search area SA20 to 1 pixel accuracy, thereby preventing a decrease in the accuracy of block matching.

[0161] For example, if the exposure time under imaging condition A is 1 / 30 [second] and the exposure time under imaging condition B is 1 / 1000 [second], the difference in exposure time is large, and it is highly likely that the main subject image will not exist in the image area under imaging condition B. Therefore, if the difference between imaging conditions A and B exceeds threshold value T2 (T2≧T1), the area setting unit 1121 sets the pixel accuracy PA1 of the first search area SA to 1 / 4 pixel accuracy and the pixel accuracy PA2 of the second search area SA to 1 pixel accuracy. The area setting unit 1121 sets the pixel accuracy PA1 of the first search area SA10 to 1 / 2 pixel accuracy and the pixel accuracy PA2 of the second search area SA20 to 1 pixel accuracy. This reduces the processing load of block matching.

[0162] In this way, by lowering the pixel accuracy of the second search area SA20 below that of the first search area SA10, the processing load for block matching in the second search area SA20 can be reduced compared to the first search area SA10, while suppressing a decrease in block matching accuracy compared to when block matching in the second search area SA20 is not performed.

[0163] <Motion vector detection processing procedure with different pixel accuracy> 24 is a flowchart showing an example of a motion vector detection processing procedure at different pixel accuracies by the motion detection unit 1110. The pixel accuracies in the first search area SA10 and the second search area SA20 may be set by a user operating the operation unit 505, or may be set automatically by the electronic device 500 according to the amount of light received by each unit group 202 in the image sensor 100. The same processing contents as those in FIGS. 21 and 22 are denoted by the same step numbers, and their description will be omitted.

[0164] After step S2204, the motion detection unit 1110 determines pixel accuracies PA1 and PA2 for block matching of the first search area SA10 and the second search area SA20 using the area setting unit 1121 based on the imaging conditions A and B (step S2405).

[0165] As in steps S2104 and S2105, the motion detection unit 1110 performs block matching of the target block b1 by scanning the search window w in the first search area SA10 and the second search area SA20 after determining the pixel accuracy using the motion generation unit 1122 (step S2406), and generates a motion vector mv from block b2 to the target block b1 (step S2407).

[0166] This makes it possible to suppress deterioration in block matching accuracy and optimize reduction in block matching processing load according to the pixel accuracy of the search area SA. Even when motion vectors are detected with different pixel accuracy, the motion detection unit 1110 may use the area setting unit 1121 to enlarge (in this case, reduce the second search area SA20) or reduce (in this case, enlarge the second search area SA20) the first search area SA10, as shown in motion detection processing procedure example 2 in Figure 22.

[0167] This makes it possible to more effectively suppress a decrease in accuracy of block matching and optimize the reduction of the processing load of block matching in accordance with the pixel accuracy and enlargement / reduction of the search area SA.

[0168] (1) As described above, the above-described video compression device compresses video data, which is a series of frames output from an image sensor 100 that has multiple imaging regions for capturing images of a subject and allows for setting imaging conditions for each imaging region. This video compression device includes a region setting unit 1121 and a motion generation unit 1122. The region setting unit 1121 sets a search region SA in a reference frame (e.g., frame F2) based on multiple imaging conditions and is used in a process (e.g., block matching) for detecting a specific region (e.g., block b2) from a reference frame (e.g., frame F2) based on a compression target region (e.g., target block b1). The motion generation unit 1122 generates a motion vector mv by detecting the specific region (e.g., block b2) based on a process (e.g., block matching) using the search region SA set by the region setting unit 1121.

[0169] This allows the range of the search area SA to be set in consideration of a plurality of imaging conditions.

[0170] (2) In addition, in the above (1), the area setting unit 1121 may set the search area SA to a specific image area captured under a specific imaging condition (for example, imaging condition A) in which a compression target area exists, among image areas captured under each of multiple imaging conditions.

[0171] This limits the search area SA according to specific imaging conditions, thereby reducing the processing load of the specific area detection process (for example, block matching).

[0172] (3) In the above (2), the area setting unit 1121 may set the search area SA to a specific image area and an image area under another imaging condition (for example, imaging condition B) surrounding the specific image area.

[0173] In this way, by setting the search area SA so as to include the periphery of a specific image area, it is possible to reduce the processing load of motion vector detection and prevent a decrease in the accuracy of motion vector detection.

[0174] (4) In the above (2), the region setting unit 1121 may enlarge or reduce the search region SA based on the relationship between a plurality of imaging conditions.

[0175] This makes it possible to selectively suppress a decrease in the accuracy of the specific area detection process or reduce the processing load depending on the expansion or contraction of the search area SA.

[0176] (5) In the above (4), the area setting unit 1121 may enlarge or reduce the search area SA based on the difference between values ​​indicated by a plurality of imaging conditions (for example, the difference between ISO sensitivities).

[0177] (6) Also, in the above 4, the area setting unit 1121 reduces the search area when a specific imaging condition (for example, imaging condition A) is a specific exposure time, and another imaging condition (for example, imaging condition B) among the multiple imaging conditions other than the specific imaging condition is an exposure time shorter than the specific exposure time.

[0178] This makes it possible to reduce the processing load of the specific area detection process even when the exposure time for the main subject is so-called long seconds.

[0179] (7) Also, in the above (1), the area setting unit 1121 may set a specific image area, among the image areas for which multiple imaging conditions are set, for which specific imaging conditions are set and in which a compression target area exists, as the first search area SA10, and may set other image areas among the image areas other than the specific image area as the second search area SA20, and the motion generation unit 1122 may generate a motion vector mv by performing specific area detection processing with different pixel accuracy for the first search area SA10 set by the area setting unit 1121 and the second search area SA20 set by the area setting unit 1121.

[0180] This makes it possible to execute the specific area detection process with pixel accuracy according to the imaging conditions.

[0181] (8) In the above (7), the motion generation unit 1122 may perform the specific area detection process in the first search area SA10 with higher pixel accuracy than the specific area detection process in the second search area SA20.

[0182] In this way, by lowering the pixel accuracy of the second search area SA20 below that of the first search area SA10, the processing load for block matching in the second search area SA20 can be reduced compared to that in the first search area SA10, while suppressing a decrease in the accuracy of the specific area detection process compared to when the specific area detection process is not performed in the second search area SA20.

[0183] (9) Furthermore, in the above (2), the motion generation unit 1122 may execute the specific region detection process outside the search area SA based on the result of the specific region detection process in the search area SA. Specifically, for example, if the motion generation unit 1122 does not detect a block b2 that matches the target block b1 in the search area SA, the motion generation unit 1122 executes the specific region detection process in the remaining image area within the search range SR excluding the search area SA, that is, the image area under the imaging condition B.

[0184] This allows the processing load of the specific area detection process to be reduced by prioritizing searching the search area SA, and also prevents a decrease in the accuracy of the specific area detection process even if block b2 is not detected from the search area SA.

[0185] (10) In addition, in the above (2), the motion generation unit 1122 may perform a specific area detection process for a search target range (e.g., a search window w) in the search area SA that is to be matched with the compression target area, based on the ratio of specific pixels included in the search target range and a specific image area (e.g., an image area under imaging condition A) to other pixels included in other image areas other than the search target range and the specific image area (e.g., an image area under imaging condition B).

[0186] This makes it possible to adjust the reduction in processing load and the prevention of accuracy degradation in the specific area detection process at the boundary of image areas with different imaging conditions.

[0187] (11) In the above (10), the motion generation unit 1122 may execute a specific region detection process so that the search target range is limited to specific pixels.

[0188] This makes it possible to specialize in reducing the processing load for the specific area detection process at the boundary of image areas with different imaging conditions.

[0189] (12) In the above (10), the motion generation unit 1122 may execute the specific region detection process so that the number of specific pixels in the search target range is greater than the number of other pixels.

[0190] This makes it possible to suppress a decrease in accuracy while prioritizing reduction in the processing load for specific area detection processing at the boundary between image areas with different imaging conditions.

[0191] (13) In the above (10), the motion generation unit 1122 may execute the specific region detection process so that at least one specific pixel exists within the search target range.

[0192] This makes it possible to suppress a decrease in accuracy while maintaining a reduction in the processing load for specific area detection processing at the boundary of image areas with different imaging conditions.

[0193] (14) The electronic device described above also includes an image sensor 100, an area setting unit 1121, and a motion generation unit 1122. The image sensor 100 has multiple imaging areas for capturing images of a subject, and imaging conditions can be set for each imaging area. The image sensor 100 outputs video data, which is a series of frames. The area setting unit 1121 sets a search area SA in a reference frame (e.g., frame F2) based on multiple imaging conditions. The search area SA is used in a process (e.g., block matching) for detecting a specific area (e.g., block b2) from a reference frame (e.g., frame F2) based on a compression target area (e.g., target block b1). The motion generation unit 1122 generates a motion vector mv by detecting the specific area (e.g., block b2) based on a process (e.g., block matching) using the search area SA set by the area setting unit 1121.

[0194] This makes it possible to realize electronic device 500 that can set the range of search area SA to a range that takes into consideration multiple imaging conditions. Note that examples of the above-mentioned electronic device 500 include digital cameras, digital video cameras, smartphones, tablets, surveillance cameras, drive recorders, and drones.

[0195] (15) The above-described video compression program also causes the processor 1001 to compress video data, which is a series of frames output from the image sensor 100, which has multiple imaging regions for capturing images of a subject and allows imaging conditions to be set for each imaging region. The video compression program also causes the processor 1001 to set, based on multiple imaging conditions, a search area SA in a reference frame, which is used in a process (e.g., block matching) for detecting a specific region (e.g., block b2) from a reference frame (e.g., frame F2) based on a compression target region (e.g., target block b1). The video compression program also causes the processor to generate a motion vector mv by detecting the specific region (e.g., block b2) based on a process (e.g., block matching) using the set search area SA.

[0196] This allows the range of the search area SA to be set by software, taking into account multiple imaging conditions. This video compression program may be recorded on a portable recording medium such as a CD-ROM, a DVD-ROM, a flash memory, or a memory card 504. This video compression program may also be recorded on a server that can be downloaded to the video compression device or electronic device 500. [Explanation of symbols]

[0197] PA1, PA2 pixel accuracy, SA, SA1, SA2, SA10, SA20 search area, b1 target block, b2 block, mv motion vector, 100 image sensor, 202 unit group, 600 video file, 900 preprocessing unit, 901 image processing unit, 902 compression unit, 910 video data, 1001 processor, 1011 detection unit, 1012 setting unit, 1110 motion detection unit, 1111 motion compensation unit, 1121 area setting unit, 1122 motion generation unit

Claims

[Claim 1] 1. A moving image compression device that compresses moving image data that is a series of frames output from an image sensor that has a plurality of imaging areas for capturing images of a subject and in which imaging conditions can be set for each of the imaging areas, a setting unit that sets a search area in the reference frame, which is used in a process of detecting a specific area from the reference frame based on a compression target area, based on a plurality of the imaging conditions; a generation unit that generates a motion vector by detecting the specific area based on the process using the search area set by the setting unit; A video compression device having the above configuration.

Citation Information

Patent Citations

  • Semiconductor module and MOS solid-state imaging device

    JP2006049361A