Image processing device, image processing method, and imaging system

The image processing apparatus simplifies high-frame-rate video decoding by synthesizing codewords from keyframes and subframes, addressing computational and power challenges in existing systems.

WO2026070197A1PCT designated stage Publication Date: 2026-04-02SONY SEMICON SOLUTIONS CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing image processing systems face difficulties in decoding high-frame-rate videos due to the need for frame interpolation using motion vectors, which is computationally intensive and power-consuming.

Method used

An image processing apparatus that generates codewords by encoding keyframes and motion vectors from higher-frame-rate subframes, synthesizing them to create a synthesized codeword, enabling easier decoding with standard decoding devices.

Benefits of technology

This approach reduces power consumption and bandwidth requirements while facilitating the playback of high-frame-rate videos using general decoding devices, improving usability and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025030423_02042026_PF_FP_ABST
    Figure JP2025030423_02042026_PF_FP_ABST
Patent Text Reader

Abstract

An information processing device according to the present disclosure comprises a control unit. The control unit encodes a first frame image to generate a first code word. The control unit generates a motion vector from a second frame image having a higher frame rate than the first frame image. The control unit generates a second code word of an interpolation frame image for interpolating the first frame image according to the motion vector. The control unit synthesizes the first code word and the second code word to generate a synthesized code word.
Need to check novelty before this filing date? Find Prior Art

Description

Image Processing Apparatus, Image Processing Method, and Imaging System

[0001] The present disclosure relates to an image processing apparatus, an image processing method, and an imaging system.

[0002] In recent years, with the increase in the resolution and frame rate of imaging devices, there has been a problem that the data volume of frame data has increased. In response to this, a technique for reducing the data volume of frame data by using frame interpolation is known.

[0003] In the prior art, an encoding device transmits a coded word obtained by encoding an image frame and information for generating an interpolation frame of the image frame to a decoding device. The decoding device can obtain a high-frame-rate video by generating an interpolation frame using the image frame obtained by decoding the coded word and side information.

[0004] Japanese Unexamined Patent Application Publication No. 2010-124369, Japanese Unexamined Patent Application Publication No. 2015-159442

[0005] S. Tulyakov et al., “Time Lens++: Event-based Frame Interpolation with Parametric Non-linear Flow and Multi-scale Fusion”, Computer Vision and Pattern Recognition (CVPR), New Orleans, 2022.

[0006] However, in the prior art, the decoding device needs to generate an interpolation frame using the image frame and side information. For example, when the side information is a coded word obtained by encoding a motion vector, the decoding device needs to decode the coded word of the motion vector.

[0007] As described above, in the prior art, there has been a problem that the decoding device needs to perform frame interpolation using a motion vector, and it is difficult for the decoding device to easily perform decoding processing.

[0008] Therefore, the present disclosure proposes an image processing apparatus, an image processing method, and an imaging system that can generate a codeword that enables a decoding apparatus to more easily decode a high-frame-rate moving image.

[0009] Note that the above problem or objective is only one of a plurality of problems or objectives that can be solved or achieved by the plurality of embodiments disclosed in this specification.

[0010] The image processing apparatus of the present disclosure includes a control unit. The control unit encodes a first frame image and generates a first codeword. The control unit generates a motion vector from a second frame image having a higher frame rate than the first frame image. The control unit generates a second codeword of an interpolated frame image that interpolates the first frame image according to the motion vector. The control unit synthesizes the first codeword and the second codeword to generate a synthesized codeword.

[0011] This figure shows an example of the image processing flow according to the first embodiment of this disclosure. This figure shows an example of the codeword generation process according to the first embodiment of this disclosure. This block diagram shows a schematic configuration example of the imaging device according to the first embodiment of this disclosure. This figure shows an example of the stacked structure of the solid-state imaging device according to the first embodiment of this disclosure. This block diagram shows a schematic configuration example of the solid-state imaging device according to the first embodiment of this disclosure. This circuit diagram shows an example of the circuit configuration of a pixel according to the first embodiment of this disclosure. This figure is for explaining an example of a keyframe image according to the first embodiment of this disclosure. This figure is for explaining an example of a subframe image according to the first embodiment of this disclosure. This block diagram shows another example of the schematic configuration of the solid-state imaging device according to the first embodiment of this disclosure. This block diagram shows a schematic configuration example of a pixel block according to the first embodiment of this disclosure. This circuit diagram shows an example of the circuit configuration of an event pixel according to the first embodiment of this disclosure. This block diagram shows a schematic configuration example of an address event detection circuit according to the first embodiment of this disclosure. This block diagram shows an example of the configuration of an image processing device according to the first embodiment of this disclosure. This figure shows an overview of the image processing according to the first embodiment of this disclosure. This block diagram shows an example of the configuration of a prediction unit according to the first embodiment of this disclosure. This figure shows an example of a keyframe according to the first embodiment of this disclosure. This figure shows an example of a region according to the first embodiment of this disclosure. This figure shows an example of a moving object region according to the first embodiment of this disclosure. This figure shows an example of a region according to the first embodiment of the present disclosure. This figure shows an example of a motion vector according to the first embodiment of the present disclosure. This figure shows another example of a motion vector according to the first embodiment of the present disclosure. This figure shows another example of a motion vector according to the first embodiment of the present disclosure. This figure shows another example of a motion vector according to the first embodiment of the present disclosure. This figure shows an example of a reference direction according to the first embodiment of the present disclosure. This figure shows that the second determinant divides the composite target using blocks of different sizes, but the composite target may be uniformly divided using blocks of a single size. This figure shows an example of the configuration of the encoding unit according to the first embodiment of the present disclosure. This is a block diagram showing an example of the configuration of the rate control according to the first embodiment of the present disclosure. This is a block diagram showing an example of the configuration of the generation unit according to the first embodiment of the present disclosure.This figure shows an example of an image frame after high frame rate enhancement according to the first embodiment of this disclosure. This figure shows an example of a composite codeword according to the first embodiment of this disclosure. This flowchart shows an example of the image processing flow according to the first embodiment of this disclosure. This figure shows another example of the configuration of the control unit according to the first embodiment of this disclosure. This figure illustrates an example of processing in the control unit according to the first embodiment of this disclosure. This block diagram shows an example of the configuration of an image processing apparatus according to the second embodiment of this disclosure. This figure shows an overview of the image processing according to the second embodiment of this disclosure. This figure shows an example of local image generation by the image generation unit according to the second embodiment of this disclosure. This figure shows an example of prediction results by the prediction unit according to the second embodiment of this disclosure. This figure shows an example of the configuration of the encoding unit according to the second embodiment of this disclosure. This figure shows an example of the configuration of the generation unit according to the second embodiment of this disclosure. This figure shows an example of a locally decoded image according to the second embodiment of this disclosure. This flowchart shows an example of the image processing flow according to the second embodiment of this disclosure. This figure shows another example of the configuration of the control unit according to the second embodiment of this disclosure. This figure shows an example of an image processing system according to an application example of this disclosure. This is a hardware configuration diagram showing an example of a computer that realizes the functions of an image processing apparatus according to the proposed technology of this disclosure.

[0012] Embodiments of this disclosure will be described in detail below with reference to the attached drawings. In this specification and the drawings, components having substantially the same functional configuration are denoted by the same reference numerals, and redundant descriptions will be omitted.

[0013] Furthermore, in this specification and drawings, similar components of embodiments may be distinguished by adding at least one different alphabet and number after the same reference numeral. However, if there is no need to particularly distinguish each of the similar components, only the same reference numeral will be used.

[0014] The one or more embodiments (including examples, modifications, and applications) described below can each be implemented independently. On the other hand, at least some of the embodiments described below may be implemented in appropriate combination with at least some of the other embodiments. These embodiments may contain novel features that differ from each other. Therefore, these embodiments may contribute to solving different objectives or problems and may produce different effects.

[0015] <<1. First Embodiment>> <1-1. Overview> Figure 1 is a diagram showing an example of the image processing flow according to the first embodiment of this disclosure. The image processing according to this embodiment is performed by the image processing system 10.

[0016] As shown in Figure 1, the image processing system 10 comprises an imaging system 11 and a display system 12.

[0017] The imaging system 11 captures imaging light and generates a frame image. The imaging system 11 encodes the frame image to generate a codeword and transmits the generated codeword to the display system 12. The imaging system 11 comprises an imaging device 100 and an image processing device 200.

[0018] The display system 12 decodes the received codeword to generate a display image and, for example, displays the display image to a user (not shown). The display system 12 comprises a playback device 300 and a display device 400.

[0019] As shown in Figure 1, when the image processing device 200 acquires a keyframe from the imaging device 100, it encodes the keyframe image (hereinafter also referred to as the key image) and generates a key codeword (step S1). The key codeword is data encoded from a series of key images.

[0020] When the image processing device 200 acquires a subframe from the imaging device 100, it predicts a motion vector (MV) from the image of the subframe (hereinafter also referred to as a subimage) (step S2). The MV is data that predicts the motion between key images.

[0021] The image processing device 200 generates codewords for interpolated frames (hereinafter also referred to as subcodewords) from the MV (step S3). The key codeword and subcodeword are generated as codewords conforming to the same video codec, for example.

[0022] Here, the keyframe image is, for example, an image having a brightness value corresponding to the amount of light received by the imaging device 100. For example, the keyframe image is an RGB image. The subframe image has, for example, a higher frame rate than the keyframe image.

[0023] For example, the subframe image may be an image with luminance values ​​(e.g., an RGB image), similar to the keyframe image. In this case, the subframe image may have a lower resolution than the keyframe image. Alternatively, the subframe image may be an image containing event data that indicates the content of an event, which is a change in light luminance.

[0024] As shown in Figure 1, the image processing device 200 synthesizes a key codeword and a subcodeword to generate code data (also referred to as a synthesized codeword) (step S4). Since the key codeword and subcodeword are generated as codewords conforming to the same video codec, for example, the image processing device 200 can synthesize the key codeword and subcodeword to generate codeword data (an example of a synthesized codeword). The codeword data is transmitted from the image processing device 200 to the playback device 300.

[0025] The playback device 300 decodes the acquired codeword data (step S5). The image data obtained by decoding the codeword data is output to the display device 400. The display device 400 displays the image data on a display, for example, not shown.

[0026] As described above, the codeword data includes a key codeword and subcodewords. The key codeword and subcodewords are generated as codewords conforming to the same video codec. Therefore, the playback device 300 can decode the codeword data in accordance with this video codec.

[0027] For example, if the video codec is one of the standard specifications for video encoding schemes such as H265 / HEVC, the image processing system 10 can play back the video (corresponding to the video data mentioned above) using a general playback device 300 that conforms to this standard specification.

[0028] Furthermore, the image data includes data decoded from key codewords that encode keyframes, and data decoded from sub-codewords that encode interpolated frames. In other words, the image data includes keyframes and interpolated frames. Therefore, the image data becomes a moving image with a higher frame rate compared to keyframes.

[0029] For example, conventional encoding devices transmitted both the encoded video data corresponding to keyframes and the encoded motion vector data to the decoding device. Therefore, the decoding device needed to interpolate the keyframes from the decoded motion vectors, which was difficult for general decoding devices conforming to standard specifications.

[0030] In contrast, the image processing system 10 according to this embodiment can play back moving images using a general decoding device (corresponding to the playback device 300) that conforms to standard specifications, making it easier to decode high-frame-rate moving images. As a result, the image processing system 10 can further improve usability.

[0031] Furthermore, the image processing apparatus 200 according to this embodiment generates subcodewords directly from motion vectors. As a result, compared to, for example, a case where interpolated frames are generated from motion vectors and those interpolated frames are encoded, the image processing apparatus 200 can generate codeword data with lower power consumption.

[0032] Figure 2 shows an example of a codeword generation process according to the first embodiment of this disclosure. Figure 2 shows a case where the imaging device 100 is equipped with a Hybrid Sensor that captures RGB moving images and event data. Note that the numerical values ​​shown in Figure 2 are examples and are not limited thereto.

[0033] In the example shown in Figure 2, the imaging device 100 includes a Hybrid Sensor and an ISP (Image Signal Processor). The Encoder in Figure 2 corresponds to the image processing device 200 in Figure 1.

[0034] The Hybrid Sensor outputs RGB 4K pixel data in Xfps format (X=15 in Figure 2). The ISP generates an RGB 4K image (corresponding to a keyframe image) from this 4K pixel data and outputs it to the Encoder.

[0035] Furthermore, the Hybrid Sensor outputs event data (corresponding to subframe images) in Yfps (Y=1k in Figure 2) from its onboard EVS (Event-based Vision Sensor) to the Encoder.

[0036] The encoder, for example, uses its built-in ME NN (Motion Estimation Neural Network) to predict motion vectors using event data. The encoder generates Zfps (Z=60 in Figure 2) codeword data from the motion vectors and Xfps RGB 4K images. Figure 2 shows an example where the encoder generates codeword data using the H265 video encoding scheme, but the video encoding scheme used by the encoder is not limited to H265.

[0037] In this way, the encoder can generate codeword data using RGB 4K images and event data, thereby generating codeword data with a higher frame rate (Zfps, Z>X) than the frame rate (Xfps) of the RGB 4K images.

[0038] One possible method is to generate a Zfps RGB 4K image from Xfps 4K pixel data by performing frame interpolation using event data. For example, if the ISP is equipped with a frame interpolation unit (e.g., video frame interpolation (VFI) NN), and this frame interpolation unit generates interpolated frames from event data, the ISP can output a Zfps RGB 4K image to the encoder.

[0039] The Encoder can generate Zfps codeword data by encoding 4K RGB images in Zfps.

[0040] Thus, even if the image processing system 10 performs frame interpolation processing before encoding processing, it can generate codeword data with a high frame rate. In this case, compared to the case where codewords for interpolated frames are generated directly from motion vectors, power consumption increases due to the implementation of frame interpolation processing.

[0041] Furthermore, when frame interpolation is performed by the ISP, the ISP needs to output a 4K RGB image in Zfps format to the encoder, requiring a wide bandwidth between the ISP and the encoder.

[0042] On the other hand, as described above, the image processing system 10 according to this embodiment predicts motion vectors from event data and generates subcodewords that encode interpolated frames using these motion vectors.

[0043] As a result, the imaging system 11 does not need to perform frame interpolation, and the increase in power consumption due to frame interpolation can be suppressed. In addition, since the imaging system 11 does not generate 4K RGB images in Zfps, the increase in bandwidth in the image signal processing pipeline of the imaging system 11 can be suppressed.

[0044] <1-2. Configuration Example> <1-2-1. Configuration Example of Imaging Device> Figure 3 is a block diagram showing a schematic configuration example of an imaging device 100 according to the first embodiment of this disclosure. As shown in Figure 3, the imaging device 100 comprises an optical system 110, a solid-state imaging device 1200, a recording unit 120, a control unit 130, and an external interface (I / F) 140. The imaging device 100 is envisioned to be a camera mounted on an information device such as a smartphone, a camera mounted on an industrial robot, or an in-vehicle camera.

[0045] The optical system 110 includes, for example, lenses, and forms an image of the incident light on the light-receiving surface of the solid-state imaging device 1200.

[0046] The solid-state imaging device 1200 converts incident light into photoelectric data to capture a frame image. The solid-state imaging device 1200 generates image data (hereinafter also referred to as grayscale image data) with brightness values ​​corresponding to the amount of incident light. The solid-state imaging device 1200 generates, for example, two frame images with different frame rates.

[0047] For example, the solid-state imaging device 1200 generates high-frame-rate frame images as keyframes (also referred to as keyframe images). The solid-state imaging device 1200 generates low-frame-rate frame images as subframes (also referred to as subframe images). The subframe images may have a lower resolution than the keyframe images.

[0048] For example, the solid-state imaging device 1200 may input at least one of the keyframe image and the subframe image to the recording unit 120, or it may output it to the image processing device 200 or the like via the external I / F 140.

[0049] The external I / F 140 may be a communication adapter for establishing communication with an external device such as an image processing device 200 via a communication network compliant with any standard, such as wireless LAN (Local Area Network), wired LAN, CAN (Controller Area Network), LIN (Local Interconnect Network), FlexRay®, MIPI (Mobile Industry Processor Interface), or LVDS (Low Voltage Differential Signaling).

[0050] The recording unit 120 is composed of, for example, a non-volatile memory such as flash memory, and records keyframe images, subframe images, and various other data input from the solid-state imaging device 1200.

[0051] The control unit 130 is composed of an information processing device such as a CPU (Central Processing Unit), and controls the solid-state imaging device 1200 to acquire keyframe images and subframe images.

[0052] (Example of a stacked configuration of the solid-state imaging device 1200) Figure 4 is a diagram showing an example of a stacked structure of the solid-state imaging device 1200 according to the first embodiment of the present disclosure. As shown in Figure 4, the solid-state imaging device 1200 has a stacked chip structure in which a light-receiving chip 1201 and a detection chip 1202 are stacked vertically. For joining the light-receiving chip 1201 and the detection chip 1202, for example, a so-called direct bonding method can be used, in which the respective joining surfaces are flattened and the two are bonded together by electron-electron force. However, it is not limited to this, and for example, a so-called Cu-Cu bonding method can be used, in which copper (Cu) electrode pads formed on each other's joining surfaces are bonded together, or other methods such as bump bonding can also be used.

[0053] Furthermore, the light-receiving chip 1201 and the detection chip 1202 are electrically connected via a connection such as a TSV (Through-Silicon Via) that penetrates the semiconductor substrate. For connections using TSVs, for example, a so-called twin TSV method can be employed, in which two TSVs, one provided on the light-receiving chip 1201 and another provided from the light-receiving chip 1201 to the detection chip 1202, are connected on the surface of the chip, or a so-called shared TSV method can be employed, in which the two are connected by a TSV that penetrates from the light-receiving chip 1201 to the detection chip 1202.

[0054] However, if a Cu-Cu junction or bump junction is used to connect the light-receiving chip 1201 and the detection chip 1202, the two are electrically connected via the Cu-Cu junction or bump junction.

[0055] (Schematic Configuration Example of Solid-State Imaging Device 1200) Figure 5 is a block diagram showing a schematic configuration example of a solid-state imaging device 1200 according to the first embodiment of the present disclosure. As shown in Figure 5, the solid-state imaging device 1200 comprises a drive circuit 1211, a signal processing unit 1212, a column ADC (conversion unit) 1220, and a pixel array unit 1300.

[0056] The pixel array unit 1300 has a configuration in which a plurality of pixel blocks 1310 are arranged in a two-dimensional grid (also called a matrix). Hereinafter, a set of pixel blocks arranged in the horizontal direction will be referred to as a "row," and a set of pixel blocks arranged in a direction perpendicular to the row will be referred to as a "column." The position of each pixel block 1310 in the pixel array unit 1300 in the row direction is identified by the X address, and the position in the column direction is identified by the Y address.

[0057] Each pixel block 1310 generates an analog pixel signal with a voltage value corresponding to the amount of incident light by converting the incident light into photoelectric energy.

[0058] Here, an example of the configuration of a pixel 1320 included in the pixel block 1310 will be described using Figure 6. Figure 6 is a circuit diagram showing an example of the circuit configuration of a pixel 1320 according to the first embodiment of this disclosure.

[0059] As shown in Figure 6, the pixel 1320 comprises a photoelectric conversion element 1321, a transfer transistor 1322, a floating diffusion layer 1323, a reset transistor 1324, an amplification transistor 1325, and a selection transistor 1326, and generates an analog signal of voltage corresponding to the photocurrent as the pixel signal Vsig. The components of the pixel 1320 other than the photoelectric conversion element 1321 are also referred to as the pixel circuit. The transfer transistor 1322, reset transistor 1324, amplification transistor 1325, and selection transistor 1326 may be, for example, N-type MOS (Metal-Oxide-Semiconductor) transistors.

[0060] The photoelectric conversion element 1321 is composed of, for example, a photodiode, and generates electric charge by photoelectric conversion of incident light. The transfer transistor 1322 transfers the charge from the photoelectric conversion element 1321 to the floating diffusion layer 1323 according to the transfer signal TRG from the drive circuit 1211.

[0061] The floating diffusion layer 1323 is a charge storage unit that generates a voltage corresponding to the amount of charge it has stored. The reset transistor 1324 releases (initializes) the charge in the floating diffusion layer 1323 according to the reset signal RST from the drive circuit 1211. The amplification transistor 1325 amplifies the voltage in the floating diffusion layer 1323. The selection transistor 1326 causes the amplified voltage signal to appear on the vertical signal line 1308 as a pixel signal Vsig according to the selection signal SEL from the drive circuit 1211. The pixel signal Vsig that appears on the vertical signal line 1308 is read out by, for example, the column ADC 1220 and converted into a digital pixel signal.

[0062] Returning to Figure 5, the drive circuit 1211 drives each of the pixel blocks 1310 that output the detection signal, causing a pixel signal with a voltage value corresponding to the amount of light incident on the photoelectric conversion element 1321 to appear on the vertical signal line 1308 to which each of the pixel blocks 1310 is connected.

[0063] The column ADC 1220 reads out the pixel signals in parallel across the column by converting the analog pixel signals appearing on the vertical signal lines 1308 of each column into digital pixel signals for each row. The column ADC 1220 then supplies the read-out digital pixel signals to the signal processing unit 1212.

[0064] The signal processing unit 1212 performs predetermined signal processing, such as CDS (Correlated Double Sampling) processing, on the pixel signals from the column ADC 1220, and outputs grayscale image data (keyframe image and subframe image) consisting of the processed pixel signals to the outside.

[0065] Figure 7 is a diagram illustrating an example of a keyframe image according to the first embodiment of the present disclosure. Figure 7 shows a portion of the pixel block 1310 used to generate the keyframe image.

[0066] As shown in Figure 7, for example, the pixel block 1310 includes an R pixel block that receives light in the R (red) wavelength band, a G pixel block that receives light in the G (green) wavelength band, and a B pixel block that receives light in the B (blue) wavelength band. Note that the arrangement of the pixel blocks shown in Figure 7 is just one example and is not limited thereto.

[0067] The solid-state imaging device 1200 generates a keyframe image, for example, by reading pixel signals from all pixel blocks 1310.

[0068] Figure 8 is a diagram illustrating an example of a subframe image according to the first embodiment of the present disclosure. Figure 8 shows a portion of the pixel block 1310 used to generate the subframe image.

[0069] For example, the signal processing unit 1212 of the solid-state imaging device 1200 takes the pixel signals read from the pixel block 1310 and converts the resulting image, which has 1 / 16th the number of pixels, into a subframe image, taking pixel addition into consideration.

[0070] For example, the signal processing unit 1212 sets the average brightness value of the R pixel blocks in a 16-pixel block (a 4x4 pixel block in the example of Figure 8) as the brightness value of one R pixel block in the subframe image. Also, for example, it sets the average brightness value of the G pixel blocks in a 16-pixel block as the brightness value of one G pixel block in the subframe image. The average brightness value of the B pixel blocks in a 16-pixel block as the brightness value of one B pixel block in the subframe image.

[0071] As a result, the signal processing unit 1212 generates a subframe image with a lower resolution than the keyframe image.

[0072] Furthermore, the signal processing unit 1212 generates keyframe images at a first frame rate (e.g., Xfps) and generates subframe images at a second frame rate higher than the first frame rate (e.g., Yfps (Y>X)).

[0073] In this way, the signal processing unit 1212 generates keyframes and subframes that have a higher fps and lower resolution than the keyframes.

[0074] As described above, the subframe image may also be event data. In other words, the solid-state imaging device 1200 may be a hybrid sensor that outputs RGB images and event data. Below, an example of a solid-state imaging device 1200 that is a hybrid sensor will be described using Figures 9 to 12. Note that the description may be omitted for configurations that are the same as the solid-state imaging device 1200 shown in Figures 4 to 8.

[0075] Figure 9 is a block diagram showing another example of the schematic configuration of a solid-state imaging device 1200 according to the first embodiment of the present disclosure. As shown in Figure 9, the solid-state imaging device 1200 includes a drive circuit 1211, a signal processing unit 1212, a Y arbiter (arbitration unit) 1213, a column ADC (conversion unit) 1220, an event encoder 1250, and a pixel array unit 1300.

[0076] Each pixel block 1310 generates an analog pixel signal with a voltage value corresponding to the amount of incident light by converting the incident light into photoelectric energy. The pixel block 1310 also detects whether or not an address event has been fired based on whether or not the amount of change in the amount of incident light exceeds a predetermined threshold.

[0077] When pixel block 1310 detects the firing of an address event, it outputs a request to Y arbiter 1213. Upon receiving a response to the request from Y arbiter 1213, pixel block 1310 sends a detection signal indicating the address event detection result to drive circuit 1211 and column ADC 1220.

[0078] The Y arbiter 1213 mediates requests from the pixel blocks 1310, determining the read order for the rows to which each pixel block 1310 that sent the request belongs, and then returns a response to all pixel blocks 1310 included in the rows to which each pixel block 1310 that sent the request belongs, based on the determined read order. In the following explanation, mediating requests and determining the read order is referred to as "mediating the read order."

[0079] The event encoder 1250 generates data for each row in the pixel array section 1300 indicating which pixel block 1310 experienced an on-event and which pixel block 1310 experienced an off-event. For example, when the event encoder 1250 receives a request from a certain pixel block 1310, it generates event detection data (corresponding to event data) that includes information indicating whether an on-event or off-event occurred in that pixel block 1310, and the X and Y addresses that identify the location of that pixel block 1310 in the pixel array section 1300.

[0080] In this process, the event encoder 1250 also includes information (timestamp) regarding the time when an on-event or off-event was detected in the event detection data. The event encoder 1250 then outputs the generated event detection data to the outside.

[0081] (Example of Pixel Block 1310 Configuration) Figure 10 is a block diagram showing a schematic configuration example of a pixel block 1310 according to the first embodiment of the present disclosure. As shown in Figure 10, the pixel block 1310 includes a pixel 1320 for generating a pixel signal which is grayscale information (for example, brightness value), a pixel 1330 for detecting whether or not an address event has been fired, and an address event detection circuit (detection unit) 1400 for detecting whether or not an address event has been fired based on the photocurrent from the pixel 1330.

[0082] Hereafter, a pixel 1320 that generates a pixel signal containing grayscale information (for example, brightness value) will be referred to as a grayscale pixel 1320. Similarly, a pixel 1330 that detects whether or not an address event has been triggered will be referred to as an event pixel 1330.

[0083] In Figure 10, one pixel block 1310 includes one tone pixel 1320 and one event pixel 1330, but the number of at least one of the tone pixels 1320 and event pixels 1330 included in one pixel block 1310 is not limited to one. For example, one pixel block 1310 may include three tone pixels 1320 and one event pixel 1330. In this case, the three tone pixels 1320 and the one event pixel 1330 may be arranged in a single column or in a 2x2 matrix.

[0084] Furthermore, although Figure 10 shows the case where the grayscale pixels 1320, event pixels 1330, and address event detection circuit 1400 are arranged on a plane (for example, on the same chip), they may also be arranged in a stacked configuration.

[0085] In the pixel block 1310, for example, the grayscale pixels 1320 and the event pixels 1330 may be arranged on the light-receiving chip 1201, and the address event detection circuit 1400 may be arranged on the detection chip 1202. In this way, the pixel block 1310 may be formed by stacking.

[0086] However, the stacking method of the pixel blocks 1310 is not limited to this, and various modifications are possible, such as arranging a part of the circuit configuration of the grayscale pixels 1320 on the detection chip 1202.

[0087] (Example of circuit configuration of event pixel 1330) Figure 11 is a circuit diagram showing an example of the circuit configuration of an event pixel 1330 according to the first embodiment of the present disclosure. Note that the configuration of the grayscale pixel 1320 is the same as in Figure 6, so its description is omitted. As shown in Figure 11, the event pixel 1330 includes a photoelectric conversion element 1331.

[0088] The photoelectric conversion element 1331, like the photoelectric conversion element 1321, is composed of, for example, a photodiode, and generates an electric charge by photoelectrically converting incident light. The electric charge generated by the photoelectric conversion of the photoelectric conversion element 1331 is supplied to the address event detection circuit 1400 as a photocurrent.

[0089] (Example of the function of the address event detection circuit 1400) The address event detection circuit 1400 shown in Figure 11 detects whether or not an address event has been triggered based on whether or not the change in the amount of photocurrent flowing out from the photoelectric conversion element 1331 exceeds a predetermined threshold. This address event consists of, for example, an on-event indicating that the change in the amount of photocurrent corresponding to the amount of incident light has exceeded an upper threshold, and an off-event indicating that the change has fallen below a lower threshold. In other words, an address event is detected when the change in the amount of incident light is outside a predetermined range from the lower limit to the upper limit. The detection signal for the address event consists of, for example, one bit indicating the detection result of the on-event and one bit indicating the detection result of the off-event. The address event detection circuit 1400 can also detect only on-events.

[0090] When an address event occurs, the address event detection circuit 1400 sends a request to the Y arbiter 1213 to send a detection signal. Upon receiving a response to the request from the Y arbiter 1213, the address event detection circuit 1400 sends detection signals DET+ and DET- to the drive circuit 1211 and the column ADC 1220. Here, the detection signal DET+ is a signal indicating the detection result of whether or not an on-event has occurred, and is sent to the column ADC 1220, for example, via the detection signal line 1306. The detection signal DET- is a signal indicating the detection result of whether or not an off-event has occurred, and is sent to the column ADC 1220, for example, via the detection signal line 1307.

[0091] Furthermore, the address event detection circuit 1400 sets the column enable signal ColEN to enable in synchronization with the selection signal SEL, and transmits this signal to the column ADC 1220 via the enable signal line 1309. Here, the column enable signal ColEN is a signal for enabling or disabling the AD (Analog to Digital) conversion for the pixel signal of the corresponding column.

[0092] When an address event is detected in a row, the drive circuit 1211 drives that row using a selection signal SEL or the like. Each of the pixel blocks 1310 in the driven row causes a pixel signal Vsig to appear on the vertical signal line 1308. The pixel signals Vsig that appear on the vertical signal line 1308 are read out by the column ADC 1220 and converted into digital pixel signals.

[0093] Furthermore, pixel blocks 1310 that have detected an address event among the driven rows send a column enable signal ColEN, which is set to enable, to the column ADC 1220. On the other hand, the column enable signal ColEN of pixel blocks 1310 that have not detected an address event is set to disable.

[0094] (Example of the configuration of the address event detection circuit 1400) Figure 12 is a block diagram showing a schematic configuration example of the address event detection circuit 1400 according to the first embodiment of the present disclosure. As shown in Figure 12, the address event detection circuit 1400 comprises a current-voltage conversion unit 1410, a buffer 1420, a subtractor 1430, a quantizer 1440, and a transfer unit 1450.

[0095] The current-voltage conversion unit 1410 converts the photocurrent from the event pixel 1330 into a logarithmic voltage signal. The current-voltage conversion unit 1410 then supplies the voltage signal to the buffer 1420.

[0096] The buffer 1420 outputs the voltage signal from the current-voltage conversion unit 1410 to the subtractor 1430. This buffer 1420 improves the driving force that drives the subsequent stage. Furthermore, the buffer 1420 ensures noise isolation associated with the switching operation of the subsequent stage.

[0097] The subtractor 1430 reduces the level of the voltage signal from the buffer 1420 according to the row drive signal from the drive circuit 1211. The subtractor 1430 then supplies the reduced voltage signal to the quantizer 1440.

[0098] The quantizer 1440 quantizes the voltage signal from the subtractor 1430 into a digital signal and outputs it to the transfer unit 1450 as a detection signal.

[0099] The transfer unit 1450 transfers the detection signal from the quantizer 1440 to the signal processing unit 1212, etc. When an address event is detected, the transfer unit 1450 sends a request to the Y arbiter 1213 and the event encoder 1250 to send a detection signal. When the transfer unit 1450 receives a response to the request from the Y arbiter 1213, it supplies the detection signals DET+ and DET- to the drive circuit 1211 and the column ADC 1220. Also, when the selection signal SEL is transmitted, the transfer unit 1450 sends the column enable signal ColEN, which is set to enable, to the column ADC 1220.

[0100] In this description, the configuration of the imaging device 100 is described as a case where a single solid-state imaging device 1200 outputs both keyframe images and subframe images, but the configuration of the imaging device 100 is not limited to this.

[0101] For example, the imaging device 100 may be configured to include a first solid-state imaging device that outputs keyframe images and a second solid-state imaging device that outputs subframe images. Alternatively, the imaging system 11 may include a first imaging device equipped with a first solid-state imaging device that outputs keyframe images and a second imaging device equipped with a second solid-state imaging device that outputs subframe images.

[0102] Thus, the device that outputs the keyframe image and the device that outputs the subframe image may be different. In this case, it is desirable that the imaging ranges of the keyframe image and the subframe image are the same or similar (close). This allows the image processing system 10 to perform frame interpolation with higher accuracy.

[0103] <1-2-2. Example of Image Processing Device Configuration> Figure 13 is a block diagram showing an example of the configuration of an image processing device 200 according to the first embodiment of the present disclosure. As shown in Figure 13, the image processing device 200 includes an external I / F 210, a storage unit 220, a communication unit 230, and a control unit 240.

[0104] Note that the configuration shown in Figure 13 is a functional configuration and may differ from the hardware configuration. Also, the functions of the image processing device 200 may be distributed and implemented across multiple physically separated devices.

[0105] Alternatively, the functions of the image processing device 200 may be implemented in the same physical device as other functions (for example, at least one function of the imaging device 100, the playback device 300, and the display device 400).

[0106] (External I / F 210) The external I / F 210 may be a communication adapter for establishing communication with an external device such as an imaging device 100 via a communication network compliant with any standard such as wireless LAN, wired LAN, CAN, LIN, FlexRay, MIPI, LVDS, etc.

[0107] (Storage Unit 220) The storage unit 220 is implemented by, for example, semiconductor memory elements such as RAM (Random Access Memory) and flash memory, or storage devices such as hard disks and optical discs. The storage unit 220 stores various types of data. For example, the storage unit 220 stores various programs, including image processing programs.

[0108] (Communication Unit 230) The communication unit 230 is implemented by, for example, a network interface controller or a NIC (Network Interface Card). The communication unit 230 may also be a USB interface consisting of a USB (Universal Serial Bus) host controller, a USB port, etc. The communication unit 230 may also be a wired interface or a wireless interface. For example, the communication unit 230 may be a wireless communication interface using a wireless LAN method or a cellular communication method. The communication unit 230 functions as a communication means or transmission means of the image processing device 200. For example, the communication unit 230 is connected to the network by wire or wireless and transmits and receives information with other devices such as the playback device 300 via the network.

[0109] (Control Unit 240) The control unit 240 is implemented, for example, by a CPU (Central Processing Unit) or MPU (Micro Processing Unit) executing a program (for example, an image processing program) stored inside the memory unit 220 using RAM (Random Access Memory) or the like as the working area. The control unit 240 is also a controller and may be implemented by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array).

[0110] Before describing the details of the control unit 240, an overview of the image processing performed by the control unit 240 will be explained using Figure 14. Figure 14 is a diagram showing an overview of the image processing according to the first embodiment of this disclosure.

[0111] The control unit 240 acquires keyframes (a group of at least one keyframe image captured at a first frame rate) and subframes (a group of at least one subframe image captured at a second frame rate) from the imaging device 100 via the external I / F 210.

[0112] In the example shown in Figure 14, the control unit 240 acquires keyframe images F100 and F102. These keyframe images F100 and F102 are images output from the sensor (imaging device 100).

[0113] The control unit 240 generates information for the playback device 300 to generate interpolated frames, such as codewords (subcodewords) for the interpolated frames, based on the keyframes and subframes.

[0114] In the example shown in Figure 14, the control unit 240 generates a subcode word for the playback device 300 to generate a composite target T101 located at a time between the capture times of keyframe images F100 and F102 (for example, exactly midway between the capture times of keyframe images F100 and F102). The composite target T101 is an interpolated frame image, which is not output from the sensor (imaging device 100) but is generated by the playback device 300 based on keyframes, etc.

[0115] The control unit 240 generates subcodewords from motion vectors (MV) generated based on keyframes and subframes, for example. The control unit 240 predicts motion vectors MV104 and MV105 from keyframe images F100 and F102 and subframes, for example.

[0116] Here, motion vector MV104 is the motion vector from keyframe image F100 to composite target T101. Motion vector MV105 is the motion vector from keyframe image F102 to composite target T101.

[0117] The composite target T101 is reconstructed by motion compensation using motion vectors MV104 and MV105. At this time, the image of the composite target T101 can be reconstructed by one of the following three methods: - Motion compensation using motion vector MV104 - Motion compensation using motion vector MV105 - Motion compensation and synthesis using both motion vectors MV104 and MV105

[0118] For example, when the composite target T101 is reconstructed by motion compensation using the motion vector MV104, the composite target T101 is reconstructed by warping the keyframe image F100 using the motion vector MV104.

[0119] For example, when the composite target T101 is reconstructed by motion compensation using the motion vector MV105, the composite target T101 is reconstructed by warping the keyframe image F102 using the motion vector MV105.

[0120] For example, suppose the composite target T101 is reconstructed by motion compensation and synthesis using both motion vectors MV104 and MV105. In this case, for example, the composite target T101 is reconstructed by synthesizing an image obtained by warping keyframe image F100 using motion vector MV104 and an image obtained by warping keyframe image F102 using motion vector MV105.

[0121] The control unit 240 does not add a residual signal to the predicted frame P106 obtained by motion compensation. That is, the subcodeword contains information indicating the motion vector and the reference method (one of the three methods described above), but does not contain a residual signal. The residual signal is the difference signal between the composite target T101 and the predicted frame P106.

[0122] As described above, the composite target T101 is not output from the sensor (imaging device 100), nor is it generated by the control unit 240. Therefore, the control unit 240 does not generate a residual signal.

[0123] The control unit 240 encodes the keyframe images F100 and F102 and generates a key codeword. The control unit 240 then transmits a composite codeword, which is a combination of the key codeword and a subcodeword, to the playback device 300 via the communication unit 230.

[0124] Returning to Figure 13, the control unit 240 includes a prediction unit 241, an encoding unit 242, a generation unit 243, and a synthesis unit 244, and realizes or executes the information processing functions and operations described below. Note that the internal configuration of the control unit 240 is not limited to the configuration shown in Figure 13, and other configurations are also acceptable as long as they perform the information processing described later.

[0125] (Prediction Unit 241) The prediction unit 241 calculates motion vectors (MV) from keyframes and subframes. The prediction unit 241 also generates the reference direction of the motion vectors and the method for dividing the image region for calculating the motion vectors from the keyframes and subframes.

[0126] Motion vectors can be generated on a pixel-by-pixel basis or on a local image region basis. The reference direction of the motion vector indicates which frame image (e.g., keyframe images F100, F102) of the current frame image (e.g., synthesis target T101) the motion vector refers to. The method for dividing the image region for calculating the motion vector is determined by the prediction unit 241, for example, according to the encoding scheme of the key codeword and subcodeword.

[0127] Here, the prediction unit 241 calculates motion vectors, etc., from keyframes and subframes, but the information used by the prediction unit 241 to calculate motion vectors, etc., is not limited to this. For example, the prediction unit 241 may generate motion vectors, the reference direction of the motion vectors, and a method for dividing the image region for calculating motion vectors from subframes. In this case, for example, the prediction unit 241 may not receive keyframes as input, but only subframes.

[0128] Figure 15 is a block diagram showing an example configuration of a prediction unit 241 according to the first embodiment of the present disclosure. The prediction unit 241 shown in Figure 15 comprises a first decision-maker 2411 and a second decision-maker 2412. Note that the internal configuration of the prediction unit 241 is not limited to the configuration shown in Figure 15, and other configurations are also possible as long as they perform the information processing described later.

[0129] (First Decision Maker 2411) The first decision maker 2411 determines the motion vector. The first decision maker 2411 generates the motion vector (MV) from the keyframes and subframes. The first decision maker 2411 may also generate the motion vector from the subframes. In this case, the input of keyframes to the first decision maker 2411 may be omitted.

[0130] Below, an example of how the first decision-maker 2411 calculates the motion vector will be explained using Figures 16 to 21. Note that the method for calculating the motion vector by the first decision-maker 2411 is not limited to the method described below and is arbitrary.

[0131] Figure 16 shows an example of a keyframe according to the first embodiment of the present disclosure. As described above, the imaging device 100 outputs keyframe images F100 and F102. The playback device 300 generates a composite target T101 as an interpolated frame image between the keyframe images F100 and F102, and generates image data for display.

[0132] First, we will explain how to calculate the motion vector from the composite target T101 to the keyframe image F100.

[0133] The first decision-maker 2411 divides the composite target T101 into multiple sub-regions, for example. For example, the first decision-maker 2411 may divide it into three regions: a background region, a moving object region, and an occluded region.

[0134] Figure 17 shows an example of a region according to the first embodiment of this disclosure. Region R200 is a background region (shown by hatching in Figure 17) in which no motion exists within the composite target T101 and keyframe image F100.

[0135] Region R200 is a region where no motion exists, so the motion vector of this region R200 is the zero vector. Region R200 can also be described as the region excluding the moving body region and the occluded region, which will be discussed later.

[0136] Figure 18 shows an example of a moving body region according to the first embodiment of the present disclosure. Region R201 is a position within the composite target T101 that is presumed to be a moving body region where a moving body (in this case, a vehicle) exists.

[0137] This position can be calculated, for example, from a subframe image taken at the timing of the composite target T101, in other words, from a subframe image corresponding to the composite target T101, and from subframe images before and after that subframe image.

[0138] It is assumed that the vehicle in keyframe image F100 moves to the predicted position in composite target T101. Therefore, the motion vector of region R201 is a vector indicating the position where the moving object exists in keyframe image F100, derived from the predicted position of the moving object within composite target T101.

[0139] Figure 19 shows an example of a region according to the first embodiment of this disclosure. Region R202 is the location of the occluded region in the keyframe image F100 where a moving object (in this case, a vehicle) exists.

[0140] This position can be calculated, for example, from a subframe image captured at the timing of keyframe image F100, in other words, from a subframe image corresponding to keyframe image F100, and from subframe images before and after that subframe image. Alternatively, this position may be calculated, for example, from keyframe images F100, F102, etc.

[0141] Region R202 corresponds to the background area within the composite target T101. This region is occluded by the moving object in keyframe image F100. Therefore, there is no appropriate reference point in keyframe image F100 that corresponds to region R202. For this reason, "no reference point" may be selected as the motion vector for the occluded region.

[0142] If "No reference" is selected as the motion vector, and the motion vector for region R202 is not output, there is a risk that both the forward (keyframe image F100) and backward (keyframe image F102) motion vectors will be absent.

[0143] If motion vectors in both reference directions are absent, there is a risk that interpolation processing may not be possible in the frame interpolation process performed by the playback device 300. In other words, there is a risk that there will be regions in the synthesis target T101 where interpolation processing cannot be performed.

[0144] To avoid this, it is desirable that the first decision-maker 2411 outputs some motion vector when it selects a motion vector with "no reference" in region R202. For example, the first decision-maker 2411 outputs a hypothetical vector shown by the dotted line in Figure 19 in region R202.

[0145] In Figure 18, the first decision-maker 2411 is shown to use a hypothetical vector of a predetermined size as the motion vector of region R202, but the hypothetical vector may also be the zero vector.

[0146] In this way, the first decision-maker 2411 generates motion vectors for the region R200 to R202.

[0147] Figure 20 shows an example of a motion vector according to the first embodiment of the present disclosure. In Figure 20, the motion vector from the composite target T101 to the keyframe image F100 is shown.

[0148] As shown in Figure 20, the first decision-maker 2411 calculates a zero vector as the motion vector in region R200. The first decision-maker 2411 calculates a vector that references the moving object as the motion vector in region R201. The first decision-maker 2411 selects "no reference" as the motion vector in region R202. Alternatively, the first decision-maker 2411 may calculate a hypothetical vector as the motion vector in region R202.

[0149] The first decision-maker 2411 similarly calculates a motion vector from the composite target T101 to the keyframe image F102.

[0150] Figure 21 shows another example of a motion vector according to the first embodiment of the present disclosure. In Figure 21, the motion vector from the composite target T101 to the keyframe image F102 is shown.

[0151] As shown in Figure 21, the first decision-maker 2411 calculates a zero vector as the motion vector in region R210. The first decision-maker 2411 calculates a vector that references the moving object as the motion vector in region R201. The first decision-maker 2411 selects "no reference" as the motion vector in region R203. Alternatively, the first decision-maker 2411 may calculate a hypothetical vector as the motion vector in region R203.

[0152] Returning to Figure 15, the first decision-maker 2411 outputs the calculated motion vector to the second decision-maker 2412.

[0153] (Second Decision Maker 2412) The second decision maker 2412 determines the reference method and division method from the motion vector determined by the first decision maker 2411. The reference direction is information indicating the target of the motion vector's reference, that is, whether the motion vector refers to the keyframe image before or after the current frame (for example, the composite target T101).

[0154] The division method is a method for dividing the image region for calculating motion vectors. This division method is specified according to the encoding scheme of the key codeword and subcodeword.

[0155] (Method for determining the reference direction) An example of a method for determining the reference direction will be explained using Figures 22 to 24. Here, an example of a method for determining the reference direction will be explained when the first determinant 2411 calculates the motion vector at the composite target T101 between the keyframe images F100 and F102 described above.

[0156] In the example described above, the first decision-maker 2411 calculates motion vectors in regions R200-R203 and R210 (see Figures 20 and 21). Here, in addition to these regions, the first decision-maker 2411 also calculates motion vectors (in this case, "no reference") in region R205.

[0157] Figure 22 shows another example of a motion vector according to the first embodiment of the present disclosure. In Figure 22, the motion vector from the composite target T101 to the keyframe image F100 is shown.

[0158] As shown in Figure 22, the first decision-maker 2411 selects "no reference" as the motion vector in region R205, in addition to the motion vector shown in Figure 20. Alternatively, the first decision-maker 2411 may calculate a hypothetical vector as the motion vector in region R205.

[0159] Figure 23 shows another example of a motion vector according to the first embodiment of the present disclosure. In Figure 23, the motion vector from the composite target T101 to the keyframe image F102 is shown.

[0160] As shown in Figure 23, the first decision-maker 2411 selects "no reference" as the motion vector in region R205, in addition to the motion vector shown in Figure 21. Alternatively, the first decision-maker 2411 may calculate a hypothetical vector as the motion vector in region R205.

[0161] The second decision-maker 2412 determines the reference direction for each region based on the motion vectors shown in Figures 22 and 23. For example, the second decision-maker 2412 determines one of forward prediction, backward prediction, and bidirectional prediction as the reference direction, depending on the reference keyframe image.

[0162] Figure 24 shows an example of a reference direction according to the first embodiment of this disclosure. The second decision-maker 2412 determines the reference direction in the regions R201 to R203, R205, and R220 from which the first decision-maker 2411 calculated the motion vector. Here, region R220 is a region that includes both regions R200 and R210. In other words, region R220 is a region that does not include regions R201 to R203 and R205.

[0163] For example, in region R220, the first decisioner 2411 calculates both the motion vector from the composite target T101 to the keyframe image F100 and the motion vector from the composite target T101 to the keyframe image F102. Therefore, the second decisioner 2412 sets the reference direction of region R220 to "bidirectional prediction".

[0164] Furthermore, region R220 is the background region, and the first decisionator 2411 calculates the zero vector as the motion vector of region R220. Therefore, the second decisionator 2412 sets the reference destination of keyframe image F100 ("F100 reference destination") as the "zero vector" and the reference destination of keyframe image F102 ("F102 reference destination") as the "zero vector".

[0165] For example, in region R201, the first decisioner 2411 calculates both the motion vector from the composite target T101 to the keyframe image F100 and the motion vector from the composite target T101 to the keyframe image F102. Therefore, the second decisioner 2412 sets the reference direction of region R201 to "bidirectional prediction".

[0166] Furthermore, region R201 is a dynamic region, and the first decision-maker 2411 calculates the motion vector from region R201 to the dynamic positions of keyframe images F100 and F102. Therefore, the second decision-maker 2412 sets the reference destination of keyframe image F100 ("F100 reference destination") to "the dynamic position of F100" and the reference destination of keyframe image F102 ("F102 reference destination") to "the dynamic position of F102".

[0167] Thus, the second decision-maker 2412 selects bidirectional prediction as the reference direction in the region where motion vectors are calculated for both keyframe images F100 and F102.

[0168] For example, region R202 is the background region in keyframe image F100. Therefore, in region R202, the first decisionator 2411 calculates a motion vector from the composite target T101 to the keyframe image F100. On the other hand, region R202 is the region in keyframe image F102 where a moving object exists. Therefore, in region R202, the first decisionator 2411 does not calculate a motion vector from the composite target T101 to the keyframe image F102. For this reason, the second decisionator 2412 sets the reference direction of region R202 to "forward prediction".

[0169] Furthermore, the second decisionator 2412 sets the reference destination of keyframe image F100 ("F100 reference destination") to the "zero vector" and the reference destination of keyframe image F102 ("F102 reference destination") to "none" (or may set it to a temporary vector).

[0170] For example, region R203 is the background region in keyframe image F102. Therefore, in region R203, the first decisionator 2411 calculates a motion vector from the composite target T101 to the keyframe image F102. On the other hand, region R203 is the region where the moving object exists in keyframe image F100. Therefore, in region R203, the first decisionator 2411 does not calculate a motion vector from the composite target T101 to the keyframe image F100. For this reason, the second decisionator 2412 sets the reference direction of region R203 to "backward prediction".

[0171] Furthermore, the second decision-maker 2412 sets the reference destination of keyframe image F102 ("F102 reference destination") to the "zero vector" and the reference destination of keyframe image F100 ("F100 reference destination") to "none" (or may set it to a temporary vector).

[0172] Thus, the second decision-maker 2412 selects either forward prediction or backward prediction as the reference direction in the region where the motion vector to one of the keyframe images F100 or F102 is calculated.

[0173] For example, in region R205, the first decisioner 2411 calculates both the motion vector from the composite target T101 to the keyframe image F100 and the motion vector from the composite target T101 to the keyframe image F102. Therefore, the second decisioner 2412 sets the reference direction of region R205 to "bidirectional prediction".

[0174] Furthermore, the first decision-maker 2411 calculates a "no reference direction," i.e., a provisional vector, as the motion vector from the composite target T101 to the keyframe image F100. Similarly, the first decision-maker 2411 calculates a "no reference direction," i.e., a provisional vector, as the motion vector from the composite target T101 to the keyframe image F102. Therefore, the second decision-maker 2412 sets the reference destination of the keyframe image F100 ("F100 reference destination") as a "provisional vector" and the reference destination of the keyframe image F102 ("F102 reference destination") as a "provisional vector."

[0175] In this way, the second determinator 2412 determines, for example, the motion vector and reference direction for each pixel according to the region within the composite target T101.

[0176] (Method for determining the division method) An example of a method for determining the division method will be explained using Figure 25. Here, an example of a method for determining the block division method when the second decision-maker 2412 calculates the reference direction in the synthesis target T101 will be explained.

[0177] The second decision-maker 2412, once it has determined the reference direction, determines the block partitioning method. Block partitioning is preferably performed in a way that divides regions with the same reference target and motion vector into blocks of the largest possible size.

[0178] In the example shown in Figure 25, the composite target T101 is divided into multiple blocks of different sizes. For example, the second decisionator 2412 divides the region containing areas R201-R203 and R205 into four blocks of the same size. The second decisionator 2412 also divides the region R220 into blocks of different sizes.

[0179] The second decision-maker 2412 performs block partitioning according to, for example, the prediction constraints of the coding scheme implemented by the control unit 240.

[0180] Note that the block division method shown in Figure 25 is just one example, and the second decision-maker 2412 can divide the synthesis target T101 into multiple blocks in various ways. The second decision-maker 2412 only needs to perform block division according to the encoding scheme implemented by the control unit 240, and the block division method is arbitrary.

[0181] For example, in Figure 25, the second decision-maker 2412 divides the composite target T101 using blocks of different sizes, but the composite target T101 may also be uniformly divided using blocks of a single size. By using a single block size, subsequent calculations are simplified.

[0182] (Encoding Unit 242) Returning to Figure 13, the encoding unit 242 encodes the keyframe and generates a key codeword. The encoding unit 242 acquires the keyframe from the imaging device 100 via the external I / F 210. The encoding unit 242 also acquires the amount of generated bits for the subcodeword generated by the generation unit 243 from the generation unit 243.

[0183] The encoding unit 242 refers to the acquired amount of generated bits, encodes the keyframe, and generates a key codeword. The encoding unit 242 outputs the generated key codeword to the synthesis unit 244.

[0184] Figure 26 shows an example configuration of the encoding unit 242 according to the first embodiment of the present disclosure. The encoding unit 242 in Figure 26 includes a subtractor 2421, a DCT (Discrete Cosine Transform) 2422, a quantizer 2423, an encoder 2424, a rate control 2425, and a bit accumulator 2426. The encoding unit 242 also includes an inverse quantizer 2427, an inverse DCT 2428, an adder 2429, a loop filter 2430, a reference buffer 2431, and an image generator 2432.

[0185] The subtractor 2421 calculates the difference between the input data (in this case, the keyframe image) input to the encoding unit 242 and the predicted image generated by the image generator 2432, and generates a residual (for example, the difference signal between the keyframe image and the predicted image). The subtractor 2421 outputs the generated residual to the DCT 2422.

[0186] The DCT 2422 applies DCT to the residual generated by the subtractor 2421 to generate DCT coefficients. The DCT 2422 outputs the generated DCT coefficients to the quantizer 2423. The quantizer 2423 quantizes the DCT coefficients using the quantized values ​​obtained from the rate control 2425, and outputs the quantized DCT coefficients to the encoder 2424 and the inverse quantizer 2427.

[0187] The encoder 2424 encodes the header information and DCT coefficients to generate a codeword. This codeword is output to the subsequent synthesis unit 244 as a key codeword. The encoder 2424 outputs the bit amount generated for the key codeword to the rate control 2425 and the bit amount accumulator 2426.

[0188] The rate control 2425 determines the quantization value according to the amount of bits generated for the key codeword, the amount of bits generated for the sub-codeword, and the cumulative amount of bits generated for the key codeword.

[0189] The rate control 2425 obtains the amount of bits generated for subcodewords from the prediction unit 241. The amount of bits generated for subcodewords is, for example, the cumulative amount of bits generated for subcodewords. The rate control 2425 obtains the amount of bits generated for key codewords from the encoder 2424. The rate control 2425 obtains the cumulative amount of bits generated for key codewords from the bit amount accumulator 2426.

[0190] The rate control 2425 outputs the determined quantization value to the quantizer 2423.

[0191] Figure 27 is a block diagram showing an example configuration of a rate control 2425 according to the first embodiment of the present disclosure. The rate control 2425 in Figure 27 comprises a third determinator 2425A and a fourth determinator 2425B.

[0192] The third decision-maker 2425A determines the assignment rate according to a preset target rate (set value), the amount of bits generated for sub-codewords (more specifically, the cumulative amount of bits generated), and the cumulative amount of bits generated for key codewords. Here, the assignment rate is the bit rate assigned to the current Slice, which is a set of coding blocks CU (coding units). The third decision-maker 2425A outputs the determined assignment rate to the fourth decision-maker 2425B.

[0193] The fourth decision-maker 2425B determines the quantization value of CU from the amount of bits generated and the assignment rate of the key codeword. The amount of bits generated of the key codeword is the amount of bits generated of the current CU. The fourth decision-maker 2425B outputs the determined quantization value to the quantizer 2423.

[0194] In this way, the rate control 2425 calculates the assignment rate, which is the upper limit output data rate assigned to the current Slice, from the cumulative amount of generated bits (cumulative number of generated bits) of the key codeword and sub-codeword. The rate control 2425 also generates the quantized value of the CU from the assignment rate and the generated bits of the current CU.

[0195] This allows the rate control 2425 to control the encoding parameters in order to limit the output bit rate of the key codeword to the target rate.

[0196] In this embodiment, encoding is performed by both the encoder 2424 and the generation unit 243. Therefore, the rate control 2425 obtains the cumulative amount of generated bits from both the encoder 2424 and the generation unit 243.

[0197] Furthermore, the method for determining the assignment rate and the quantization value by rate control 2425 is not particularly limited and is arbitrary. For example, the algorithm of MPEG-2 Test Model 5 (TM5) may be used as a method for determining these values.

[0198] Returning to Figure 26, the bit amount accumulator 2426 calculates the cumulative value of the number of bits generated in one frame. The bit amount accumulator 2426 outputs the calculated cumulative value to the rate control 2425 as the cumulative number of bits generated for the key codeword.

[0199] The inverse quantizer 2427 inversely quantizes the DCT coefficients quantized by the quantizer 2423 and outputs them to the inverse DCT 2428. The inverse DCT 2428 inversely quantizes the DCT coefficients inversely quantized by the inverse quantizer 2427 and generates residual components. The inverse DCT 2428 outputs the residual components to the adder 2429.

[0200] The adder 2429 adds the residual component and the predicted image generated by the image generator 2432 to generate an added image. The adder 2429 outputs the added image to the loop filter 2430. The loop filter 2430 applies loop filtering to the added image. The reference buffer 2431 stores the filtered added image to which the loop filtering has been applied.

[0201] The image generator 2432 generates a predicted image according to the keyframes and the summed images stored in the reference buffer 2431. The image generator 2432 is, for example, ME (Motion Estimation) / MC (Motion Compensation) / intraPrediction.

[0202] The image generator 2432 outputs the predicted image to the subtractor 2421 and the adder 2429.

[0203] (Generation Unit 243) Returning to Figure 13, the generation unit 243 generates subcodewords using motion vectors (MV), reference direction, and division method. The subcodewords consist of codewords obtained by interframe predictive coding. The subcodewords include codewords for motion vectors. The subcodewords do not include codewords for residual signals (difference images between the composite target T101 and the predicted frames obtained by motion compensation).

[0204] The generation unit 243 outputs the subcodeword to the synthesis unit 244. The generation unit 243 also generates a cumulative value of the number of bits generated for the subcodeword and outputs it to the encoding unit 242.

[0205] Figure 28 is a block diagram showing an example configuration of a generation unit 243 according to the first embodiment of the present disclosure. The generation unit 243 in Figure 28 comprises a generation controller 2435 and a bit amount accumulator 2436.

[0206] The generation controller 2435 generates subcodewords and generated bit amounts according to the motion vector, reference direction, and division method. For example, the generation controller 2435 generates subcodewords using interframe predictive coding. The generation controller 2435 outputs the generated subcodewords to the synthesis unit 244 and outputs the generated bit amounts (more specifically, the cumulative generated bit amounts) to the encoding unit 242.

[0207] Furthermore, the generation controller 2435 includes a first generator 2435A and a second generator 2435B. The generation controller 2435 performs control such as driving the first generator 2435A and the second generator 2435B at appropriate timings.

[0208] The first generator 2435A generates a subcodeword using interframe predictive coding according to the motion vector, reference direction, division method, and header information generated by the second generator 2435B. This subcodeword is transmitted to the playback device 300, thereby transmitting the motion vector, reference direction, and division method to the playback device 300. The residual signal is not transmitted to the playback device 300. The second generator 2435B generates the header information for the subcodeword.

[0209] The bit amount accumulator 2436 accumulates the amount of bits generated by subcodewords to generate the amount of bits generated in one frame.

[0210] (Combination Unit 244) Returning to Figure 13, the combination unit 244 combines the key codeword and subcodewords to generate a single combined codeword. The combination unit 244 transmits the generated combined codeword to the regeneration device 300 via the communication unit 230.

[0211] In this way, the synthesis unit 244 reconstructs the key codeword and sub-codewords as a single codeword (composite codeword).

[0212] An example of a reconstruction method will be explained using Figures 29 and 30. For simplicity, this example uses HEVC, a type of standard video codec, as the encoding method; however, the encoding method is not limited to HEVC, and other encoding methods may be used.

[0213] Figure 29 shows an example of an image frame after high frame rate processing according to the first embodiment of the present disclosure. For example, the image frame is obtained by the playback device 300 decoding a composite codeword. Figure 29 shows the frame number and encoding type of the image frame.

[0214] For example, suppose an image frame is a frame that has been processed to a high frame rate (e.g., 60 fps) using a 15 fps keyframe and subframes acquired at a faster rate.

[0215] In this case, frame images #0, #4, and #8 correspond to keyframe images. Frame images #0, #4, and #8 are obtained by decoding the key codeword that was encoded and transmitted using I or P-slice.

[0216] Frame images #1-#3 and #5-#7 are interpolated frames obtained by interpolating keyframes. Frame images #1-#3 and #5-#7 are obtained by decoding subcodewords that were encoded and transmitted using B-slice.

[0217] Subcodewords encoded and transmitted using B-slice do not contain residual signals. Therefore, decoding of subcodewords is performed by referring to the frame images of I-slice or P-slice (frame images #0, #4, and #8 in Figure 29).

[0218] For example, to obtain frame images #1 through #3, frame images #0 and #4 are referenced.

[0219] Therefore, the synthesis unit 244 reconstructs the synthesized codeword so that it can decode the frame images encoded with B-slice. As described above, in order to obtain frame images #1 to #3, frame images #0 and #4 are referenced. In other words, in order to decode frame images #1 to #3, the synthesized codeword is reconstructed so that frame images #0 and #4 have already been decoded.

[0220] Figure 30 shows an example of a composite codeword according to the first embodiment of this disclosure. As described above, in order to decode frame images #1 to #3, the composite codeword is reconstructed so that frame images #0 and #4 have already been decoded.

[0221] In other words, the compositing unit 244 transmits frames #0 and #4 before transmitting frames #1 to #3. Similarly, the compositing unit 244 transmits frames #4 and #8 before transmitting frames #5 to #7.

[0222] Therefore, the order in which the frames are combined by the combining unit 244 is, for example, frame #0 of the I-slice, frame #4 of the P-slice, frames #1 to #3 of the B-slice, frame #8 of the P-slice, and frames #5 to #7 of the B-slice.

[0223] <1-3. Processing Example> Figure 31 is a flowchart showing an example of the image processing flow according to the first embodiment of this disclosure. The image processing in Figure 31 is repeatedly performed by the image processing device 200, for example, while the imaging device 100 is taking images. That is, the image processing is repeatedly performed, for example, while the image processing device 200 is acquiring keyframes and subframes.

[0224] As shown in Figure 31, the image processing device 200 generates a key codeword from the keyframe (step S101). The image processing device 200 generates the key codeword according to a predetermined encoding scheme, such as HEVC.

[0225] The image processing device 200 determines the motion vector of the composite target to interpolate between keyframes from the subframes (step S102). The image processing device 200, for example, divides the composite target into at least one region and calculates the motion vector for each pixel in each region.

[0226] The image processing device 200 determines the reference direction of the motion vector for each pixel based on the motion vector (step S103). The image processing device 200 selects, for example, at least one of forward prediction, backward prediction, and bidirectional prediction as the reference direction depending on the motion vector.

[0227] The image processing device 200 determines a method for dividing the composite target according to the motion vector and the reference direction (step S104). For example, the image processing device 200 determines a division method (e.g., division blocks) according to the region from which the motion vector was calculated.

[0228] The image processing device 200 encodes at least one of the motion vector, the reference direction, and the division method to generate a subcodeword (step S105). The subcodeword is generated, for example, according to the same encoding format as the key codeword.

[0229] Thus, the image processing apparatus 200 according to this embodiment does not generate a composite target which is an interpolated frame, but rather generates information for generating a composite target in the playback device 300 (for example, a subcode word which encodes at least one of the motion vector, the reference direction, and the division method).

[0230] As a result, the image processing device 200 can reduce the processing load compared to the case where interpolated frames are generated, high-frame-rate image frames are generated, and then these image frames are encoded. Therefore, the image processing device 200 can further reduce the power consumption and bandwidth when generating codewords.

[0231] As shown in Figure 31, the image processing device 200 generates a composite codeword by combining the key codeword and subcodewords (step S106). For example, the image processing device 200 generates a composite codeword by reconstructing the key codeword and subcodewords in the frame order corresponding to the decoding order at the playback device 300.

[0232] The composite codeword is transmitted, for example, to the playback device 300 via the communication unit 230. Alternatively, or in addition to this, the composite codeword may be stored in the storage unit 220.

[0233] As described above, the image processing device 200 according to this embodiment encodes a keyframe image (an example of a first frame image) and generates a key codeword (an example of a first codeword). The image processing device 200 generates a motion vector from a subframe image (an example of a second frame image) with a higher frame rate than the keyframe image. The image processing device 200 generates a codeword (a subcodeword; an example of a second codeword) for a composite target (an example of an interpolated frame) that interpolates the keyframe image according to the motion vector. For example, the image processing device 200 encodes the motion vector and generates a subcodeword. The image processing device 200 then synthesizes the key codeword and the subcodeword to generate a composite codeword.

[0234] As a result, the playback device 300 (an example of a decoding device) can more easily reproduce (decode) image frames with a higher frame rate than the keyframe by decoding the composite codeword, without having to perform processing such as frame interpolation. For example, if the composite codeword is encoded with a standard video codec, the image processing system 10 can reproduce the composite codeword with a standard playback device 300.

[0235] <1-4. Other Configuration Examples> Figure 32 shows another configuration example of the control unit 240 according to the first embodiment of the present disclosure. The control unit 240 in Figure 32 includes, for example, a Motion Estimation 241X, a Keyframe Encoder 242X, an Intermediate Frame Encoder 243X, and a Stream Synthesis 244X.

[0236] Motion Estimation 241X is, for example, an example of the prediction unit 241 in Figure 13. Keyframe Encoder 242X is, for example, an example of the encoding unit 242 in Figure 13. Intermediate Frame Encoder 243X is, for example, an example of the generation unit 243 in Figure 13. Stream Synthesis 244X is, for example, an example of the synthesis unit 244 in Figure 13.

[0237] The Keyframe Encoder 242X includes a subtractor 2421X, a DCT 2422X, a Quantization 2423X, an Encoding 2424X, a Rate Control 2425X, and a Bit Accmulator 2426X. The Keyframe Encoder 242X also includes an Inverse Quantization 2427X, an Inverse DCT 2428X, an adder 2429X, a Loop Filter 2430X, a Buffer 2431X, and an Intra / Inter Prediction 2432X.

[0238] Subtractor 2421X is, for example, an example of subtractor 2421 in Figure 26. DCT 2422X is, for example, an example of DCT 2422 in Figure 26. Quantization 2423X is, for example, an example of quantizer 2423 in Figure 26. Encoding 2424X is, for example, an example of encoder 2424 in Figure 26. Rate Control 2425X is, for example, an example of rate control 2425 in Figure 26. Bit Accumulator 2426X is, for example, an example of bit accumulator 2426 in Figure 26.

[0239] Inverse Quantization 2427X is, for example, an example of the inverse quantizer 2427 in Figure 26. Inverse DCT 2428X is, for example, an example of the inverse DCT 2428 in Figure 26. Adder 2429X is, for example, an example of the adder 2429 in Figure 26. Loop Filter 2430X is, for example, an example of the loop filter 2430 in Figure 26. Buffer 2431X is, for example, an example of the reference buffer 2431 in Figure 26. Intra / Inter Prediction 2432X is, for example, an example of Figure 26.

[0240] The Intermediate Frame Encoder 243X comprises a Header generator 2435X1, a Stream generator 2435X2, and a Bit Accmulator 2436X.

[0241] The Header generator 2435X1 is, for example, an example of the first generator 2435A in FIG. 28. The Stream generator 2435X2 is, for example, an example of the second generator 2435B in FIG. 28. The Bit Accmulator 2436X is an example of the bit amount accumulator 2436 in FIG. 28.

[0242] FIG. 33 is a diagram for explaining an example of the processing in the control unit 240 according to the first embodiment of the present disclosure. Here, three intermediate frames are generated from two key frames and an event stream. Each intermediate frame is encoded only by motion information.

[0243] In the encoding process (Encording work) executed in the control unit 240, for example, the first key frame I t and the second key frame I t+1 and the event (events) E t→t+1 are given. The encoding process here is an example of at least part of the above-described image processing.

[0244] The first key frame I t and the second key frame I t+1 correspond to the key frames in FIG. 26. For example, the first key frame I t is the key frame at time t, and the second key frame I t+1 corresponds to the key frame at time t + 1. The event E t→t+1 corresponds to a sub-frame. For example, the event E t→t+1 corresponds to the sub-frame acquired between time t and time t + 1.

[0245] In the encoding process, bitstreams S t 、S t+1 、S τ and the intermediate frame I τ are generated. The intermediate frame I τ may be generated at any time (timestamps) τ between the key frames.

[0246] The bitstream S τF is a motion vector. τ→t F τ→t+1 It is generated using the motion vector F. τ→t F τ→t+1 Event E t→τ , E τ→t+1 It is calculated by each of the following.

[0247] Bitstream S t S τ These are combined to form the final bitstream S all This is generated.

[0248] (Keyframe Encoder 242X) Keyframe Encoder 242X is given by the bitstream S given by the following equation (1) t Generates.

[0249]

[0250] P t Q indicates the predicted signal. t This indicates the quantized value. Predicted signal P t This is given by the following equation (2). Note that the predicted signal is an example of a predicted image generated by the image generator 2432 in Figure 26.

[0251]

[0252] ^I t This is a local decoded signal. t This corresponds, for example, to the summation image stored in the reference buffer 2431. Quantization value Q in Figure 26 t This is given by equation (3) below.

[0253]

[0254] BC t This indicates the bits used in the current coding unit (CU). BC Key Acc This indicates the bits used across keyframes in the current GOP. BC Inter Acc This indicates the bits used across intermediate frames within the current GOP.

[0255] BC t This is an example of the amount of bits generated by the encoder 2424 in Figure 26. BC Key Acc This is an example of the cumulative bit amount output by the bit amount accumulator 2426. BC Inter Acc This is an example of the amount of bits generated output by the bit amount accumulator 2436 in Figure 28.

[0256] (Motion Estimation 241X) Motion Estimation 241X is the motion vector F given by the following equation (4). τ→t and F τ→t+1 We estimate this.

[0257]

[0258] (Intermediate Frame Encoder 243X) The Intermediate Frame Encoder 243X is given by the bitstream S given by the following equation (5). τ Generates.

[0259]

[0260] The main difference from conventional encoders (which may include Keyframe Encoder 242X) is that it does not refer to frames captured by the sensor, but instead generates a bitstream S for intermediate frames. τ This is what is generated.

[0261] This method degrades the quality of the reconstructed video and encodes the signal without performing calculations for residuals (an example of the residual signal mentioned above) that are represented as zero.

[0262] On the other hand, this method only performs encoding of motion vectors, and does not involve other processes such as DCT or quantization. This reduces the bandwidth and power consumption within the encoder.

[0263] When optimizing power and bandwidth, the accuracy of motion estimation is crucial for improving video quality. Events captured by a hybrid event camera (an example of an imaging device 100 that outputs subframe images) provide efficient information between keyframes and have high temporal resolution, contributing to improved motion estimation accuracy.

[0264] Bitstream S τ This is generated for the motion vector. Therefore, bitstream S for intermediate frames τ This follows interpicture prediction, meaning the image type of the intermediate frame is either P or B.

[0265] The encoder uses the intermediate frame in the current GOP for bit BC Inter Acc Record bit BC Inter Acc This is supplied to the Keyframe Encoder 242X and used to determine the quantization value, thereby allowing control of the bitrate.

[0266] (Stream Synthesis 244X) Stream Synthesis 244X provides bitstream S for keyframes and intermediate frames respectively. t S τ The bitstream S is then synthesized to form the final bitstream S. all Generates bitstream S. all This is a bitstream S that is rearranged according to the image configuration. t S τ It is a sequence of columns.

[0267] For example, as shown in Figure 29 above, the sequence is IBBBPBBBP, and the referencing directions of the P and B images are indicated by arrows. Since the decoding of the P image is performed before the decoding of the B image, the decoding order is IPBBBPBB.

[0268] <<2. Second Embodiment>> In the first embodiment described above, subcodes were generated using inter-frame predictive coding. However, the coding format used to generate subcodes is not limited to inter-frame predictive coding. For example, the image processing device 200 may generate subcodes using intra-frame predictive coding. In the second embodiment, an image processing device 200A that generates subcodes using at least one of inter-frame predictive coding and intra-frame predictive coding will be described.

[0269] In addition, in the image processing system 10 according to the second embodiment, components and processes that are the same as those in the first embodiment may be denoted by the same reference numerals and their descriptions may be omitted.

[0270] <2-1. Example of Image Processing Device Configuration> Figure 34 is a block diagram showing an example of the configuration of an image processing device 200A according to the second embodiment of the present disclosure. The image processing device 200A in Figure 34 differs from the image processing device 200 in Figure 13 in that it includes a control unit 240A.

[0271] Before describing the details of the control unit 240A, an overview of the image processing performed by the control unit 240A will be explained using Figure 35. Figure 35 is a diagram showing an overview of the image processing according to the second embodiment of this disclosure.

[0272] The control unit 240A acquires keyframes (a group of at least one keyframe image captured at a first frame rate) and subframes (a group of at least one subframe image captured at a second frame rate) from the imaging device 100 via the external I / F 210.

[0273] In the example shown in Figure 35, the control unit 240A acquires keyframe images F300 and F302. These keyframe images F300 and F302 are images output from the sensor (imaging device 100).

[0274] The control unit 240A generates information for the playback device 300 to generate interpolated frames, such as codewords (subcodewords) for the interpolated frames, based on the keyframes and subframes.

[0275] In the example shown in Figure 35, the control unit 240A generates a subcodeword for the playback device 300 to generate a composite target T301 located at a time between the capture times of keyframe images F300 and F302 (for example, exactly midway between the capture times of keyframe images F300 and F302). The composite target T301 is an interpolated frame image, which is not output from the sensor (imaging device 100) but is generated by the playback device 300 based on keyframes, etc.

[0276] The control unit 240A generates subcodewords from motion vectors (MV) generated based on keyframes and subframes, for example. The control unit 240A predicts motion vectors MV304 and MV305 from keyframe images F300 and F302 and subframes, for example.

[0277] Here, motion vector MV304 is the motion vector from the composite target T301 to the keyframe image F300. Motion vector MV305 is the motion vector from the composite target T301 to the keyframe image F302.

[0278] The synthetic target T301 is reconstructed by motion compensation using motion vectors MV304 and MV305.

[0279] Here, for example, there are video codecs that support translation but do not support motion compensation for motion involving rotation. For instance, encoding schemes known as standard video codecs, such as VVC / HEVC / AVC, do not support motion compensation for motion involving rotation.

[0280] For example, as shown in keyframe images F300 and F302, if the keyframes include motion involving rotation, it may be difficult for the playback device 300 to accurately reconstruct the composite target T301 in inter-frame prediction corresponding to translation.

[0281] Therefore, in this embodiment, the control unit 240A performs motion compensation using intra-frame coding instead of, or in addition to, inter-frame coding. As a result, the composite target T301 is reconstructed with greater accuracy in the playback device 300.

[0282] In Figure 35, for example, the control unit 240A reconstructs an image of region R404 of the synthesis target T301 (hereinafter also referred to as the local image). The control unit 240A calculates a motion vector from this local image to region R402 and generates a subcodeword by applying intra-frame coding.

[0283] Furthermore, the background region including regions R400, R401, and R403 is encoded using inter-frame coding, similar to the first embodiment, to generate subcodewords. Thus, the encoding scheme for the synthesis target T301 is a mixture of inter-frame coding and intra-frame coding. In this way, the control unit 240A generates subcodewords using multiple different encoding schemes.

[0284] Returning to Figure 34, the control unit 240A differs from the image processing apparatus 200 in Figure 13 in that it includes an encoding unit 242A, a generation unit 243A, and an image generation unit 245.

[0285] (Image generation unit 245) The image generation unit 245 generates a local image of the composite target based on the keyframes, subframes, the reference direction of the motion vector, and the method for dividing the composite target.

[0286] An example of a local image generated by the image generation unit 245 will be explained using Figures 36 and 37. Figure 36 is a diagram showing an example of local image generation by the image generation unit 245 according to the second embodiment of this disclosure. Figure 37 is a diagram showing an example of prediction results by the prediction unit 241 according to the second embodiment of this disclosure.

[0287] As shown in Figure 36, the prediction unit 241 uses subframes and keyframes to predict the motion vectors of regions R400 to R403 of the composite target T301. For example, the prediction unit 241 selects regions R401 to R403 that include moving objects and region R400 that does not include moving objects from subframes, etc.

[0288] The prediction unit 241 calculates, for example, the translational motion vectors for each region R400 to R403.

[0289] Region R400 is the background region of keyframe images F300 and F301 and the composite target T301. The prediction unit 241 calculates a zero vector for each pixel in region R400 as a motion vector from the composite target T301 to keyframe image F300. The prediction unit 241 also calculates a zero vector for each pixel in region R400 as a motion vector from the composite target T301 to keyframe image F302.

[0290] Since the prediction unit 241 calculates motion vectors for both keyframe images F300 and F301 in region R400, the reference direction for region R400 is both forward and backward. Therefore, as shown in Figure 37, in region R400, the reference direction is bidirectional interpretation, and both motion vectors of keyframe images F300 and F301 (referenced by F300 and F302 in Figure 37) become zero vectors.

[0291] Region R401 is the background region in keyframe image F300 and the occluded region where a moving object exists in keyframe image F302. The prediction unit 241 calculates a zero vector for each pixel in region R401 as the motion vector from the composite target T301 to the keyframe image F300. The prediction unit 241 also calculates a hypothetical vector (without a reference) for each pixel in region R401 as the motion vector from the composite target T301 to the keyframe image F302.

[0292] Since the prediction unit 241 calculates the motion vector using the keyframe image F300 in region R401, the reference direction of region R401 is forward. Therefore, as shown in Figure 37, the reference direction in region R401 is forward interpretation. Also, the motion vector of keyframe image F300 (referenced by F300 in Figure 37) becomes the zero vector, and the motion vector of keyframe image F302 (referenced by F302 in Figure 37) becomes the dummy vector.

[0293] Region R402 is the background region of keyframe images F300 and F301, and is the moving area in the composite target T101. The prediction unit 241 calculates a provisional vector for each pixel in region R402 as the motion vector from the composite target T301 to keyframe image F300. The prediction unit 241 also calculates a provisional vector for each pixel in region R400 as the motion vector from the composite target T301 to keyframe image F302.

[0294] Since the prediction unit 241 does not calculate motion vectors for both keyframe images F300 and F301 in region R402, region R402 has no reference points in both directions. Therefore, as shown in Figure 37, in region R402, the reference direction (specifically, the encoding method) is intra-prediction, and both the motion vectors of keyframe images F300 and F301 (the reference points for F300 and F302 in Figure 37) become provisional vectors.

[0295] Region R403 is an occluded region where a moving object exists in keyframe image F300, and a background region in keyframe image F302. The prediction unit 241 calculates a hypothetical vector (without a reference) for each pixel in region R403 as the motion vector from the composite target T301 to keyframe image F300. The prediction unit 241 also calculates a zero vector for each pixel in region R403 as the motion vector from the composite target T301 to keyframe image F302.

[0296] Since the prediction unit 241 calculates the motion vector in region R403 using the keyframe image F302, the reference direction for region R403 is backward. Therefore, as shown in Figure 37, in region R403, the reference direction is backward interpretation. Also, the motion vector of keyframe image F300 (referenced by F300 in Figure 37) becomes a temporary vector, and the motion vector of keyframe image F302 (referenced by F302 in Figure 37) becomes the zero vector.

[0297] As shown in Figure 36, the image generation unit 245 reconstructs the composite target T302 of region R404, which includes region R402 that the prediction unit 241 has predicted intra (no references in either forward or backward directions), as a local image (corresponding to the reconstructed image in Figure 36).

[0298] Intra coding is a method that performs prediction and coding by referring to the vicinity of the current frame (here, the synthesis target T301) without referring to the preceding or succeeding frames (here, keyframe images F300 and F302).

[0299] As described above, the composite target T301 is not output from the sensor (corresponding to the imaging device 100). Therefore, when performing intra-coding of the composite target T301, the control unit 240A needs to reconstruct the composite target T301.

[0300] In this case, if the image generation unit 245 of the control unit 240A reconstructs the entire synthesis target T301, the processing load for reconstruction will increase. Therefore, in this embodiment, based on the prediction results of the prediction unit 241, the image of a local region of the synthesis target T301 (local image) is reconstructed. This allows the image generation unit 245 to further reduce the increase in the processing load for reconstructing the synthesis target T301.

[0301] The image generation unit 245 generates a local image in any way, for example, using keyframe images F300, F302 and subframes. The image generation unit 245 outputs the generated local image to the generation unit 243A.

[0302] Returning to Figure 34, the encoding unit 242A generates a key codeword from the keyframe, similar to the encoding unit 242 in Figure 13, and outputs the quantized value to the generation unit 243A.

[0303] Figure 38 shows an example of the configuration of the encoding unit 242A according to the second embodiment of the present disclosure. The encoding unit 242A shown in Figure 38 differs from the encoding unit 242 in Figure 26 in that the rate control 2425 adds the quantized value to the quantizer 2463 and outputs it to an external generation unit 243A.

[0304] As described above, the encoding unit 242A in this embodiment outputs a quantized value to the generation unit 243A. This quantized value is used by the generation unit 243A to generate subcodewords using intra encoding (also referred to as intra predictive encoding).

[0305] Furthermore, if the generation unit 243A does not use the quantized values ​​generated by the encoding unit 242A, for example, by performing intra encoding using predetermined quantized values, the encoding unit 242A may choose not to output the quantized values ​​to the generation unit 243A. In this case, the configuration of the encoding unit 242A is, for example, the same as the encoding unit 242 in Figure 26.

[0306] Returning to Figure 34, the generation unit 243A obtains prediction results from the prediction unit 241, local images from the image generation unit 245, and quantization values ​​from the encoding unit 242A. The prediction results include motion vectors, reference directions, and division methods.

[0307] The generation unit 243A generates subcodewords using the acquired prediction results, local images, and quantization values. The subcodewords consist of codewords encoded either within a frame or between frames. For example, within-frame encoding is applied to the region corresponding to the local image, and between-frame encoding is applied to the other regions.

[0308] The generation unit 243A performs intra encoding based, for example, on the quantization values ​​obtained from the encoding unit 242A. The codewords generated by interframe encoding include the motion vector codewords but do not include the residual signal codewords, similar to the first embodiment. The subcodewords are output to the synthesis unit 244.

[0309] Furthermore, the generation unit 243A generates the amount of bits to be generated for the subcodeword and outputs it to the encoding unit 242A.

[0310] Figure 39 shows an example configuration of a generation unit 243A according to a second embodiment of the present disclosure. The generation unit 243A in Figure 39 comprises a subtractor 2461, a DCT 2462, a quantizer 2463, an encoder 2464, and a bit accumulator 2465. The generation unit 243A also comprises an inverse quantizer 2466, an inverse DCT 2467, an adder 2468, a loop filter 2469, a reference buffer 2470, an image generator 2471, a first selector 2472, and a second selector 2473.

[0311] The subtractor 2461 calculates the difference between the input data (in this case, the local image) input to the generation unit 243A and the predicted image generated by the image generator 2471, and generates a residual (for example, a difference signal between the local image and the predicted image). The subtractor 2461 outputs the generated residual to the DCT 2462.

[0312] The DCT 2462 applies DCT to the residual generated by the subtractor 2461 to generate DCT coefficients. The DCT 2462 outputs the generated DCT coefficients to the quantizer 2463. The quantizer 2463 quantizes the DCT coefficients using the quantized values ​​obtained from the encoding unit 242A, and outputs the quantized DCT coefficients to the encoder 2464 and the inverse quantizer 2466.

[0313] The encoder 2464 encodes the header information and DCT coefficients to generate a codeword. This codeword, or the codeword output by the first generator 2435A, is output as a subcodeword to the subsequent synthesis unit 244. The encoder 2464 outputs the amount of bits generated for the codeword to the bit amount accumulator 2465.

[0314] The bit amount accumulator 2465 calculates the cumulative value of the number of bits generated in one frame. The bit amount accumulator 2465 outputs the calculated cumulative value to the first selector 2472 as the cumulative number of bits generated for the key codeword.

[0315] The inverse quantizer 2466 inversely quantizes the DCT coefficients quantized by the quantizer 2463 and outputs them to the inverse DCT 2467. The inverse DCT 2467 inversely quantizes the DCT coefficients inversely quantized by the inverse quantizer 2466 and generates residual components. The inverse DCT 2467 outputs the residual components to the adder 2468.

[0316] The adder 2468 adds the residual component and the predicted image generated by the image generator 2471 to generate an added image. The adder 2468 outputs the added image to the loop filter 2469. The loop filter 2469 applies loop filtering to the added image. The reference buffer 2470 stores the filtered added image to which the loop filtering has been applied.

[0317] The image generator 2471 generates a predicted image based on the local image and the summed image stored in the reference buffer 2470. For example, the image generator 2471 refers to the local image and the summed image stored in the reference buffer 2470 and generates a predicted image (hereinafter also referred to as an intra-predicted image).

[0318] An example of a locally decoded image (a local image after decoding, corresponding to, for example, an added image) referenced by the image generator 2471 will be explained using Figure 40. Figure 40 is a diagram showing an example of a locally decoded image according to the second embodiment of this disclosure. In Figure 40, the locally decoded image is shown with numbers assigned to the blocks to be encoded (Coding Units: CUs).

[0319] CUs #1 to #6, #7, #12, #13, #18, #19, and #24 to #30 shown in Figure 40 are CUs encoded using interprediction, while CUs #8 to #11, #14 to #17, and #20 to #23 are CUs encoded using intraprediction.

[0320] The intra-predicted image is generated by referencing the pixel values ​​of the spatial neighbors of the block (CU) to be encoded. In this embodiment, the local image of the synthesis target is reconstructed, but other regions are not.

[0321] Therefore, the spatial neighbors that the image generator 2471 can reference are limited to the intra-encoded CUs (local image CUs) contained within the same frame image (composite target). In the example in Figure 40, the CUs that the image generator 2471 can reference are CUs #8 to #11, #14 to #17, and #20 to #23.

[0322] For example, when generating an intra-prediction image for CU#8, the image generator 2471 attempts to reference CU#1 to #3 and #7. However, since CU#1 to #3 and #7 are CUs encoded by inter-prediction (outside the local image), the image generator 2471 cannot reference CU#1 to #3 and #7. In this case, the image generator 2471 determines the reference target using a method predetermined by the standard, etc., and applies intra-prediction to generate an intra-prediction image.

[0323] For example, when generating an intra-prediction image for CU#9, the image generator 2471 attempts to access CU#2 to #4 and #8. Since CU#2 to #4 are CUs encoded using inter-prediction, the image generator 2471 cannot access CU#1 to #3 and #7. On the other hand, since CU#8 is a CU encoded using intra-prediction, the image generator 2471 can access CU#8.

[0324] In this case, the image generator 2471 refers to CU#8, applies intra prediction, and generates an intra prediction image. Alternatively, in addition to this, the image generator 2471 may determine the references corresponding to CU#2 to #4 in a manner predetermined by standards, etc., and apply intra prediction to generate an intra prediction image.

[0325] For example, when generating an intra-prediction image for CU#15, the image generator 2471 attempts to reference CU#8 to #10 and #14. Since CU#8 to #10 and #14 are CUs encoded by inter-prediction, the image generator 2471 can reference CU#8 to #10 and #14. In this case, the image generator 2471 references CU#8 to #10 and #14, applies intra-prediction, and generates an intra-prediction image.

[0326] Thus, the image generator 2471 generates a predicted image by referring to the pixel values ​​of the spatial neighbors of the CU that is the target of intra encoding. Therefore, the reference buffer 2470 referenced by the image generator 2471 is required to store the pixel values ​​of the spatial neighbors of the CU that the image generator 2471 references. On the other hand, for CUs to which inter encoding is applied, the image does not need to be reconstructed and buffering is not required.

[0327] Therefore, for example, CUs #8 to #11, #14 to #17, and #20 to #23 in Figure 40 are stored in the reference buffer 2470. On the other hand, CUs #1 to #6, #7, #12, #13, #18, #19, and #24 to #30, which are not used (not referenced) by the image generator 2471 to generate the predicted image, are not stored in the reference buffer 2470.

[0328] Thus, the reference buffer 2470 stores the local image of the interpolated frame (e.g., the synthesis target) after the loop filter has been applied.

[0329] Returning to Figure 39, the image generator 2471 outputs the predicted image to the subtractor 2461 and the adder 2468.

[0330] The first selector 2472 outputs to the encoding unit 242A either the cumulative generated bit amount calculated by the bit amount accumulator 2465 or the cumulative generated bit amount calculated by the bit amount accumulator 2436. The first selector 2472 selects the cumulative generated bit amount to output based on the reference direction of the motion vector (for example, information indicating whether it is inter prediction or intra prediction).

[0331] For example, if the encoding scheme of the generated codeword is inter predictive coding, the first selector 2472 selects the cumulative generated bit amount calculated by the bit amount accumulator 2436 and outputs it to the encoding unit 242A. On the other hand, if the encoding scheme of the generated codeword is intra predictive coding, the first selector 2472 selects the cumulative generated bit amount calculated by the bit amount accumulator 2465 and outputs it to the encoding unit 242A.

[0332] The second selector 2473 outputs either the codeword encoded by the encoder 2464 or the codeword encoded by the first generator 2435A as a subcodeword to the synthesis unit 244. The second selector 2473 selects the codeword to output based on the reference direction of the motion vector (for example, information indicating whether it is inter-prediction or intra-prediction).

[0333] For example, if the encoding scheme of the generated codeword is inter predictive coding, the second selector 2473 selects the codeword encoded by the first generator 2435A (i.e., the codeword generated by inter coding) and outputs it to the synthesis unit 244. On the other hand, if the encoding scheme of the generated codeword is intra predictive coding, the second selector 2473 selects the codeword encoded by the encoder 2464 (i.e., the codeword generated by intra coding) and outputs it to the synthesis unit 244.

[0334] <2-2. Processing Example> Figure 41 is a flowchart showing an example of the image processing flow according to the second embodiment of the present disclosure. The image processing in Figure 41 is repeatedly performed by the image processing device 200A, for example, while the imaging device 100 is taking images. That is, the image processing is repeatedly performed, for example, while the image processing device 200A is acquiring keyframes and subframes.

[0335] Note that the same reference numerals are used for the same image processing shown in Figure 31, and their explanations are omitted.

[0336] In step S104, the image processing device 200A determines the division method and generates local images according to the keyframe image, subframe image, motion vector, reference direction, and division method (step S201). The image processing device 200A generates local images of the region encompassing the intra-predicted region as the reference direction from the keyframe image and subframe image, etc.

[0337] The image processing device 200A generates subcodes based on codewords generated by intra-coding according to the generated local image and codewords generated by inter-coding according to motion frames, etc. (step S202). For example, the image processing device 200A generates subcodes by performing intra-coding on regions determined to be intra-predicted. Alternatively, for example, the image processing device 200A generates subcodes by performing inter-coding on regions determined to be inter-predicted.

[0338] The subcodeword generated by the image processing device 200A is combined with the key codeword and transmitted to the playback device 300.

[0339] As described above, the image processing apparatus 200A according to this embodiment generates subcodes by performing intra encoding in addition to inter encoding. As a result, the image processing apparatus 200A can generate subcodes without reconstructing the entire synthesis target, even when complex movements such as rotational movements are included in the keyframes.

[0340] As a result, the playback device 300 (an example of a decoding device) can more easily reproduce (decode) image frames with a higher frame rate than the keyframe by decoding the composite codeword, without having to perform processing such as frame interpolation. For example, if the composite codeword is encoded with a standard video codec, the image processing system 10 can reproduce the composite codeword with a standard playback device 300.

[0341] <2-3. Other Configuration Examples> Figure 42 shows another configuration example of the control unit 240A according to the second embodiment of the present disclosure. The control unit 240A in Figure 42 includes, for example, Motion Estimation 241X, Encoder 242AX, Encoder 243AX, Stream Synthesis 244X, and Local region reconstruction 245X.

[0342] Motion Estimation 241X is, for example, an example of the prediction unit 241 in Figure 34. Motion Estimation 241X, similar to Figure 32, uses the motion vector F τ→t F τ→t+1 Generates.

[0343] Local region reconstruction 245X is the local region I of the intermediate frame. loc τ This is an image generation unit that generates the image, and is an example of the image generation unit 245 in Figure 34. Local region I of the intermediate frame loc τ This is an example of the local image described above (corresponding to the reconstructed image in Figure 36).

[0344] Encoder 243X is an encoding unit that encodes intermediate frames, and corresponds, for example, to a part of the generation unit 243A in Figure 34. Encoder 243X, similar to Figure 32, BC t This generates an example of the amount of bits output by the encoder 2424 in Figure 26.

[0345] Encoder 242AX is an encoding unit that encodes local regions of keyframes and intermediate frames, and is, for example, an example of encoding unit 242A in Figure 34. Encoder 242AX has the same configuration as Keyframe Encoder 242X in Figure 32.

[0346] Encoder242AX has keyframe I t or local region I of the intermediate frame loc τ The signal is input. The signal input to Encoder 242AX may be switched, for example, by a switch.

[0347] Keyframe I t If this is entered, Encoder 242AX will input keyframe I t bitstream S t Generates the local region I of the intermediate frame. loc τ If this is input, Encoder 242AX will input local region I of the intermediate frame. loc τbitstream S t Generates the local region I of the intermediate frame. loc τ bitstream S t This corresponds, for example, to the subcode word in Figure 39.

[0348] In other words, the Encoder 242AX in Figure 42 functions as both a Keyframe Encoder 242X (or the encoding unit 242A in Figure 34) and as part of the generation unit 243A in Figure 34. By sharing some of the functions of the generation unit 243A in Figure 34 (the function of generating subcodewords) with the Encoder 242AX in this way, the circuit size of the control unit 240A can be further reduced.

[0349] <<3. Application Examples>> For example, the technology relating to this disclosure may be applied to various types of video processing. Therefore, an example of applying this technology to stream distribution will be described below.

[0350] Figure 43 shows an example of an image processing system 10 related to an application example of the present disclosure. In Figure 43, for example, an imaging system 11 is mounted on an information processing device 30A, such as a smartphone. Also, for example, a display system 12 is mounted on information processing devices 30B to 30D, such as smartphones.

[0351] Furthermore, the information processing device 30, on which at least one of the imaging system 11 and the display system 12 is mounted, is not limited to a smartphone, but may be various devices such as a camera or other sensor device, a tablet terminal, a PC, or a television.

[0352] A videographer captures moving images using, for example, an information processing device 30A equipped with an imaging system 11. The information processing device 30A performs the image processing described above and generates a high-frame-rate composite codeword from low-frame-rate keyframes. This composite codeword corresponds, for example, to the recorded data of the moving images captured by the information processing device 30A.

[0353] The information processing device 30A, for example, distributes the generated composite codeword (recorded data) to the information processing devices 30B to 30D according to the instructions of the videographer. This distribution can, for example, be performed in real time.

[0354] Information processing devices 30B to 30D decode the composite codeword and display it to the viewer. As described above, the composite codeword is decoded according to the encoding method used by information processing device 30A. Therefore, information processing devices 30B to 30D do not need to perform any special processing to decode the composite codeword, such as reconstructing interpolated frames, and can more easily play (decode) high-frame-rate video.

[0355] Furthermore, the information processing device 30A does not reconstruct the interpolated frames, but generates the decoded interpolated frames. Therefore, the information processing device 30A can further reduce the power consumption required to generate the recorded data. Also, because the information processing device 30A does not reconstruct the interpolated frames, it can generate the recorded data at a higher speed. As a result, the information processing device 30A can capture high-frame-rate video in real time and distribute it to the information processing devices 30B to 30D.

[0356] <<4. Hardware Configuration>> The information devices such as the image processing device 200 according to each embodiment described above may be realized by a computer 3000 having the configuration shown in Figure 44. Figure 44 is a hardware configuration diagram showing an example of a computer that realizes the functions of the image processing device 200 according to the proposed technology of this disclosure. The computer 3000 includes a CPU 3100, RAM 3200, ROM 3300, HDD (Hard Disk Drive) 3400, communication interface 3500, and input / output interface 3600. The parts of the computer 3000 are connected by a bus 3050.

[0357] The CPU 3100 operates based on programs stored in the ROM 3300 or HDD 3400 and controls each component. The CPU 3100 loads the programs stored in the ROM 3300 or HDD 3400 into the RAM 3200 and executes processing corresponding to various programs.

[0358] ROM3300 stores boot programs such as the BIOS (Basic Input Output System) that are executed by the CPU3100 when the computer 3000 starts up, as well as programs that depend on the computer 3000's hardware.

[0359] The HDD 3400 is a computer-readable recording medium that non-temporarily records programs executed by the CPU 3100 and data used by such programs. Specifically, the HDD 3400 is a recording medium that records an information processing program according to this disclosure, which is an example of program data 3450.

[0360] The communication interface 3500 is an interface for the computer 3000 to connect to an external network 3550 (such as the Internet). The CPU 3100 receives data from other devices and transmits data it generates to other devices via the communication interface 3500.

[0361] The input / output interface 3600 is an interface for connecting the input / output device 3650 and the computer 3000. The CPU 3100 receives data from input devices such as keyboards and mice via the input / output interface 3600. The CPU 3100 transmits data to output devices such as displays, speakers, and printers via the input / output interface 3600. The input / output interface 3600 can also function as a media interface for reading programs recorded on a predetermined recording medium (media).

[0362] Media include optical recording media such as DVDs (Digital Versatile Discs) and PDs (Phase-change rewritable Disks), magneto-optical recording media such as MOs (Magneto-Optical disks), tape media, magnetic recording media, or semiconductor memory.

[0363] When the computer 3000 functions as an image processing device 200 according to the embodiment, the CPU 3100 of the computer 3000 realizes the functions of the control unit 240 in Figure 13 by executing an image processing program loaded on the RAM 3200. The HDD 3400 stores the image processing program according to this disclosure and data in the storage device.

[0364] The CPU 3100 reads the program data 3450 from the HDD 3400 and executes it. However, as another example, the CPU 3100 can also obtain these programs from other devices via the external network 3550.

[0365] <<5. Other Embodiments>> The embodiments described above are examples only, and various modifications and applications are possible.

[0366] The control device for controlling the image processing device 200 of the embodiment may be implemented by a dedicated computer system or by a general-purpose computer system.

[0367] For example, an image processing program for performing the above-described operations is stored in a computer-readable recording medium such as an optical disc, semiconductor memory, magnetic tape, or flexible disk and distributed. Then, for example, the control device is configured by installing the program on a computer and executing the above-described operations. In this case, the control device may be an external device to the image processing device 200 (for example, a personal computer). Alternatively, the control device may be an internal device to the image processing device 200 (for example, a control unit 240).

[0368] Alternatively, the image processing program may be stored on a disk drive provided by a server device on a network such as the Internet, and made available for download to a computer. Furthermore, the above-mentioned functions may be implemented through the cooperation of an OS (Operating System) and application software. In this case, the parts other than the OS may be stored on a medium and distributed, or the parts other than the OS may be stored on a server device and made available for download to a computer.

[0369] Furthermore, among the processes described in each of the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically by known methods. In addition, the processing procedures, specific names, and information including various data and parameters shown in the above documents and drawings can be changed at will unless otherwise specified. For example, the various information shown in each figure is not limited to the information shown.

[0370] Furthermore, the components of each illustrated device are functionally conceptual and do not necessarily need to be physically configured as shown. In other words, the specific forms of distribution and integration of each device are not limited to those shown, and all or part of them can be functionally or physically distributed and integrated in any unit according to various loads and usage conditions. This distribution and integration configuration may also be performed dynamically.

[0371] Furthermore, the embodiments described above can be combined as appropriate in areas where the processing content is not contradictory. Also, the steps shown in the flowcharts, etc., of the embodiments described above can be changed in order as appropriate.

[0372] Furthermore, each embodiment can also be implemented as any configuration that constitutes the device or system, such as a processor as a system LSI (Large Scale Integration), a module using multiple processors, a unit using multiple modules, or a set with additional functions added to a unit (i.e., a configuration of part of the device).

[0373] In each embodiment, a system refers to a collection of multiple components (devices, modules (parts), etc.), regardless of whether all components are located in the same enclosure. Therefore, multiple devices housed in separate enclosures and connected via a network, and a single device containing multiple modules within a single enclosure, are both considered systems.

[0374] Furthermore, for example, each embodiment can adopt a cloud computing configuration in which a single function is shared and processed collaboratively by multiple devices via a network.

[0375] <<6. Conclusion>> Although embodiments of the present disclosure have been described above, the technical scope of the present disclosure is not limited to the embodiments described above, and various modifications are possible without departing from the gist of the present disclosure. Furthermore, components from different embodiments and modifications may be combined as appropriate.

[0376] Furthermore, the effects described in the embodiments of this specification are merely illustrative and not limiting, and other effects may also occur.

[0377] Note that the present technology can also have the following configurations. (1) An image processing apparatus comprising a control unit that encodes a first frame image to generate a first codeword, generates a motion vector from a second frame image having a higher frame rate than the first frame image, generates a second codeword of an interpolated frame image that interpolates the first frame image according to the motion vector, and synthesizes the first codeword and the second codeword to generate a synthesized codeword. (2) The image processing apparatus according to (1), wherein the control unit calculates a cumulative bit amount of the second codeword and generates the first codeword according to the cumulative bit amount. (3) The image processing apparatus according to (1) or (2), wherein the second codeword includes a codeword by inter-frame predictive coding. (4) The image processing apparatus according to (3), wherein the second codeword includes a codeword of the motion vector. (5) The image processing apparatus according to (4), wherein the second codeword does not include a codeword of a residual signal. (6) The image processing apparatus according to any one of (1) to (5), wherein the control unit generates a local image of the interpolated frame image based on the first frame image and the second frame image, and generates the second codeword according to the local image and the motion vector. (7) The image processing apparatus according to (6), wherein the control unit generates the local image according to a reference direction in the first frame image of the motion vector and a division method of the interpolated frame image. (8) The image processing apparatus according to (6) or (7), wherein the second codeword includes a codeword by intra-frame predictive coding. (9) The image processing apparatus according to (8), wherein the second codeword is generated by referring to a spatial neighborhood of the local image. (10) The image processing apparatus according to (9), wherein the second codeword is generated without referring to areas other than the spatial neighborhood of the local image. (11) The image processing apparatus according to any one of (1) to (10), wherein the first frame image is an RGB image. (12) The image processing apparatus according to any one of (1) to (11), wherein the second frame image is an RGB image having a lower resolution than the first frame image.(13) The image processing apparatus according to any one of (1) to (11), wherein the second frame image is an image including event data indicating the content of an event which is a change in the brightness of light. (14) The image processing apparatus according to any one of (1) to (13), wherein the control unit acquires the first frame image and the second frame image captured by one imaging device. (15) An image processing method comprising: encoding the first frame image to generate a first codeword; generating a motion vector from a second frame image having a higher frame rate than the first frame image; generating a second codeword for an interpolated frame image that interpolates the first frame image according to the motion vector; and combining the first codeword and the second codeword to generate a composite codeword. (16) An imaging system comprising: an imaging device that captures a first frame image and a second frame image having a higher frame rate than the first frame image; and an image processing device that encodes the first frame image to generate a first codeword, generates a motion vector from the second frame image, generates a second codeword for an interpolated frame image that interpolates the first frame image according to the motion vector, and generates a composite codeword by combining the first codeword and the second codeword.

[0378] 10 Image processing system 11 Imaging system 12 Display system 30 Information processing device 100 Imaging device 200 Image processing device 210 External I / F 220 Storage unit 230 Communication unit 240 Control unit 300 Playback device 400 Display device

Claims

1. An image processing apparatus comprising a control unit that encodes a first frame image to generate a first codeword, generates a motion vector from a second frame image having a higher frame rate than the first frame image, generates a second codeword for an interpolated frame image that interpolates the first frame image according to the motion vector, and generates a composite codeword by combining the first codeword and the second codeword.

2. The image processing apparatus according to claim 1, wherein the control unit calculates the cumulative bit amount of the second codeword and generates the first codeword according to the cumulative bit amount.

3. The image processing apparatus according to claim 1, wherein the second codeword includes a codeword obtained by interframe predictive coding.

4. The image processing apparatus according to claim 3, wherein the second codeword includes the codeword for the motion vector.

5. The image processing apparatus according to claim 4, wherein the second codeword does not include the codeword of the residual signal.

6. The image processing apparatus according to claim 1, wherein the control unit generates a local image of the interpolated frame image based on the first frame image and the second frame image, and generates a second codeword according to the local image and the motion vector.

7. The image processing apparatus according to claim 6, wherein the control unit generates the local image according to the reference direction of the motion vector in the first frame image and the method of dividing the interpolated frame image.

8. The image processing apparatus according to claim 6, wherein the second codeword includes a codeword obtained by intraframe predictive coding.

9. The image processing apparatus according to claim 8, wherein the second codeword is generated by referring to the spatial neighborhood of the local image.

10. The image processing apparatus according to claim 9, wherein the second codeword is generated without referring to anything other than the spatial vicinity of the local image.

11. The image processing apparatus according to claim 1, wherein the first frame image is an RGB image.

12. The image processing apparatus according to claim 1, wherein the second frame image is an RGB image with a lower resolution than the first frame image.

13. The image processing apparatus according to claim 1, wherein the second frame image is an image that includes event data indicating the content of an event which is a change in the brightness of light.

14. The image processing apparatus according to claim 1, wherein the control unit acquires the first frame image and the second frame image captured by a single imaging device.

15. An image processing method comprising: encoding a first frame image to generate a first codeword; generating a motion vector from a second frame image having a higher frame rate than the first frame image; generating a second codeword for an interpolated frame image that interpolates the first frame image according to the motion vector; and generating a composite codeword by combining the first codeword and the second codeword.

16. An imaging system comprising: an imaging device that captures a first frame image and a second frame image having a higher frame rate than the first frame image; and an image processing device that encodes the first frame image to generate a first codeword, generates a motion vector from the second frame image, generates a second codeword for an interpolated frame image that interpolates the first frame image according to the motion vector, and generates a composite codeword by combining the first codeword and the second codeword.