Encoding device and encoding method
Patent Information
- Application Number
- PCT/EP2026/058241
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-27
- Filing Date
- 2026-03-24
- Publication Date
- 2026-10-01
Smart Images

Figure EP2026058241_01102026_PF_FP_ABST
Abstract
Description
[0001] 1
[0002] ENCODING DEVICE AND ENCODING METHOD
[0003] FIELD OF THE INVENTION
[0004] The present technology relates to encoding devices and methods for encoding, in particular to encoding that uses affine transformations.
[0005] BACKGROUND
[0006] Usual digital video cameras that are e.g. used as stand-alone devices or integrated into other electronic devices such as mobile phones produce large amounts of data, which are typically reduced by encoding the generated image data / image frames. On the other hand, there is a demand for high-speed videos, i.e. videos having a high frame rate, with high image quality. Such high-speed videos come with an increased amount of data. It is therefore desirable to improve encoding devices such as to allow efficient encoding of videos with high frame rates.
[0007] SUMMARY OF INVENTION
[0008] To this end, an encoding device for encoding a series of frames containing image data for a video is provided. The encoding device is configured to receive and store a first number of frames per second that is smaller than a second number of frames per second of the video, to receive event data indicating a temporal stream of events, where each event indicates intensity changes above a predetermined threshold of light received for one of a plurality of pixels that constitute the frames, and to determine for the stored frames based on the event data and the stored frames whether affine transformations of pixel values of the frames shall be used for reconstructing missing frames that are temporally located between two consecutive, stored frames. If affine transformations shall not be used, the encoding device is configured to determine and store based on the event data and the stored frames motion vectors for reconstructing the missing frames, which motion vectors linearly shift pixel values. If affine transformations shall be used, the encoding device is configured to determine and store based on the event data and the stored frames the affine transformations.
[0009] Further, an encoding method for encoding a series of frames containing image data for a video is provided. The encoding method comprises: receiving and storing a first number of frames per second that is smaller than a second number of frames per second of the video; receiving event data indicating a temporal stream of events, where each event indicates intensity changes above a predetermined threshold of light received for one of a plurality of pixels that constitute the frames; determining for the stored frames based on the event data and the stored frames whether affine transformations of pixel values of the frames shall be used for reconstructing missing frames that are temporally located between two consecutive, stored frames; if affine transformations shall not be used, determining and storing based on the event data and the stored frames motion vectors for reconstructing the missing frames, which motion vectors linearly shift pixel values; and if affine transformations shall be used, determining and storing based on the event data and the stored frames the affine transformations.2
[0010] Event data are used to decide whether motion vectors, i.e. linear transitions, or affine transformations, i.e. rotations, scaling, or skew, should be used for the reconstruction of frames that are added to stored frames (i.e. keyframes) in the final, decoded video. This allows generating interpolating frames based on event data, where usage of motion vectors would lead to qualitative minor outcomes, e.g. since consecutive captured frames show large amounts of rotation, scaling or skew. This helps to enhance the frame rate of low-speed videos by using event data in situations in which reconstruction of traditional motion vectors would be insufficient.
[0011] BRIEF DESCRIPTION OF DRAWINGS
[0012] Fig. 1 is a schematic diagram of an encoding device.
[0013] Fig. 2 is another schematic diagram of an encoding device.
[0014] Fig. 3 is another schematic diagram of an encoding device.
[0015] Fig. 4 is a schematic illustration of motion vectors and affine transformations.
[0016] Fig. 5 is another schematic illustration of motion vectors and affine transformations.
[0017] Fig. 6 is another schematic illustration of affine transformations.
[0018] Fig. 7 is a schematic illustration of an encoding process.
[0019] Fig. 8 is a schematic diagram of a sensor device.
[0020] Fig. 9 is a schematic block diagram of a sensor section.
[0021] Fig. 10 is a schematic block diagram of a pixel array section.
[0022] Fig. 11 is a schematic circuit diagram of a pixel block.
[0023] Fig. 12 is a schematic block diagram illustrating of an event detecting section.
[0024] Fig. 13 is a schematic circuit diagram of a current-voltage converting section.
[0025] Fig. 14 is a schematic circuit diagram of a subtraction section and a quantization section.
[0026] Fig. 15 is a schematic timing chart of an example operation of the sensor section.
[0027] Fig. 16 is a schematic diagram of a frame data generation method based on event data.3
[0028] Fig. 17 is a schematic block diagram of another quantization section.
[0029] Fig. 18 is a schematic diagram of another event detecting section.
[0030] Fig. 19 is a schematic block diagram of another pixel array section.
[0031] Fig. 20 is a schematic circuit diagram of another pixel block.
[0032] Fig. 21 is a schematic block diagram of a scan-type imaging device.
[0033] Fig. 22 is a schematic block diagram of a sensor device comprising an encoding device.
[0034] Fig. 23 shows different schematic block diagrams of a sensor device.
[0035] Fig. 24 is a schematic illustration of a decoding device.
[0036] Fig. 25 illustrates schematically a process flow of an encoding method.
[0037] Fig. 26 is a schematic block diagram of a smart phone.
[0038] Fig. 27 is a schematic block diagram of a vehicle control system.
[0039] Fig. 28 is a diagram of assistance in explaining an example of installation positions of an outside-vehicle information detecting section and an imaging section.
[0040] Figs. 29A and 29B are schematic illustrations of a mobile device and a head mounted display comprising a sensor device.
[0041] DETAILED DESCRIPTION
[0042] Fig. 1 recapitulates typical functions of an encoding device 100. The encoding device 100 receives a series of frames Fl produces by a sensor device 10. The encoding device 100 may comprise a pixel signal processing unit 60 that is configured to receive pixel signals, i.e. intensity signals, from the sensor device 10, and to process the received frames of pixel signals to generate e.g. color frames therefrom. The pixel signal processing unit 60 is for example an image signal processor, ISP, that de-mosaics raw data from the sensor device 10, which are e.g. provided with color filters in a Bayer-pattem, to generate RGB or YUV information for each pixel. But the pixel signal processing unit 60 may also have other functions that preprocess the raw intensity signals such as to improve their encoding capabilities.
[0043] Further, the encoding device 100 also comprises an encoding unit 70 that is configured to generate an encoded video stream that allows reconstructing a video stream of frames F2. The encoding unit 704
[0044] operates in principle as is well known to a skilled person. Its basic functions are schematically illustrated in Fig. 1. A current frame II from the frames Fl is input to the encoding unit 70. Based on the current frame II and a previous frame 10 a prediction unit 702 generates a prediction for the current frame II. This prediction is subtracted from the actual current frame to generate a residual, which is transformed and quantized in a quantization unit 704 and coded in a coder unit 706. This quantized and coded residual becomes part of the encoded frame data, if it is assumed that in decoding the residual can be combined with a previously decoded full frame to also generate a full frame. In the language of HEVC, the residual constitutes the encoded frame data for P- and B-frames, while the full color frame is only encoded for I-frarnes. The quantized residual is dequantized in dequantization unit 708 and combined with the predicted current frame to form the actual current frame II, which is then stored to be used as previous frame in the next time step.
[0045] Fig. 2 shows a more detailed view of the encoding unit 70. As indicated above, the encoding unit 70 comprises a prediction unit 702, a quantization unit 704, a coder unit 706, and a dequantization unit 708. Here, it should be noted that the subtraction step indicated explicitly in Fig. 1 is considered implemented within the prediction unit 702 in Fig. 2, which outputs the predicted image and the residual. Further, Fig.
[0046] 2 explicitly shows the addition of the dequantized residual and the predicted image as well as a fdter unit 709. Fig. 2 also illustrates a buffer unit 703 in which previously encoded images / frames can be stored for the next encoding steps.
[0047] As shown in Fig. 2, the encoding unit 70 may comprise an image-based prediction unit 702a within the prediction unit 702 that performs image prediction as in principle known by a skilled person. Further, the prediction unit 702 may comprise a hybrid-based prediction unit 702b that carries out image prediction mainly based on detected event data. Different possible functions of hybrid-based prediction are possible.
[0048] The image-based prediction unit 702a may comprise an intra-picture estimation unit that decides how to encode incoming frames, an intra-picture prediction unit that generates a predicted image based on intra-picture prediction, a motion estimation unit that measures an optical flow from a previous frame to a current image to determine e.g. motion vectors, and a motion compensation unit that generates a predicted image by applying the optical flow, e.g. the motion vectors, measured by the motion estimation unit to the previous frame. Based on control data it may be decided whether to use the predicted image from intra-picture prediction or from inter-picture prediction, i.e. the predicted image generated by the motion compensation unit. Thus, the image-based prediction unit 702a operates basically as is well known to a skilled person.
[0049] The inputs and outputs of the image-based prediction unit 702a are as follows, a: the frames Fl. b: a control signal, c: the residual, d: intra-prediction data, e: the predicted image, f: motion data, e.g. motion vectors, g: (the) previously encoded frame(s). Optionally also h: event data may be used by the motion estimation unit.
[0050] The encoding unit 70 is also configured to generate the encoded frame data by processing event data as well as frame data. To this end, the hybrid-based prediction unit 702b may be provided. The encoding73961
[0051] 5
[0052] device may then be configured to switch between an encoding mode in which the video data are generated only from the frames Fl, i.e. by using the image-based prediction unit 702a as described above and as commonly known, and an encoding mode in which the video data are generated by processing the event data together with the received frame data.
[0053] This means e.g. that based on the retrieved information like e.g. the amount of motion in the scene obtainable from the event data or the contents of the observed scene (presence of humans, amount of fine details to resolve, etc.) it is decided whether to encode the frame data in a classical manner or whether to base the encoding also on event data. To this end, the encoding device 100 may use an encoding mode, in which it processes event data to generate interpolated encoded frame data located temporally between two of the received frames Fl.
[0054] Fig. 3 shows a schematic illustration of an encoding device 100 and its functions that can carry out the above described encoding based on frame data and event data. The encoding device 100 is used to encode a series of frames F2 that contain image data for a video stream. This means, that the encoded data output by the encoding device 100 can be used to generate the series of frames F2 by decoding the output of the encoding device 100. Thus, in this context “encoding of frames” refer to the frames of the video stream that can be reproduced from the encoded data. This video stream has a certain frame rate that may in general be different from the frame rate of a series of frames that are input into the encoding device 100 and different from the frame rate of frames that are stored (e.g. as keyframes) in the encoded data.
[0055] The encoding device 100 may be constituted by any computing device that is capable to carry out the functions described below. For example, the encoding device 100 may be or may be comprised in a computer, processor, microprocessor, CPU, GPU, FPGA, or ASIC. The encoding device 100 may be constituted by hardware or software or a mixture of hardware and software. In general, the structure, design or layout of the encoding device 100 is arbitrary as long as it is capable to carry out the functions described below.
[0056] As schematically illustrated in Fig. 3, the encoding device 100 receives a video stream containing a first number of frames Fl per second. The first number of frames Fl per second is typically smaller than a second number of frames F2 per second that are present in the video obtainable from the encoded data. The encoding device 100 stores the frames Fl in a first storage area 110. It should be noted that it is not necessary that the encoding device 100 stores all the received frames Fl in the first storage area 110. The encoding device 100 may also drop received frames Fl and store only the remaining frames.
[0057] As also illustrated at the top of Fig. 3, the encoding device 100 receives event data ED that indicate a temporal stream of events.
[0058] Here, as will be described in more detail below with respect to Figs. 8 to 21, each event indicates intensity changes above a predetermined threshold of the light received for one of a plurality of pixels 51 that constitute the frames F 1. In contrast to conventional, frame-based imaging, the generation and output of events is not necessarily frame-based, it can also be asynchronous. This is illustrated in Fig. 3 by showing6
[0059] between each consecutive pair of frames Fl several events occurring at different times. The events are not generated at regular intervals. Every time when a change of light is detected in a single pixel that is larger than a predetermined threshold an event is generated. The generation of events is therefore particularly dependent on motions in the observed scene since these will lead to variations in the observed scene that lead to changes in the observed intensity. Thus, the event data ED constitute information about movements of the objects captured in the frames Fl.
[0060] This information and the information contained in the stored frames Fl is used by the encoding device 100 to determine whether affine transformations AF of pixel values of the frames Fl shall be used for reconstructing missing frames MF that are temporally located between two consecutive, stored frames F 1. That is, the encoding device 100 checks whether transitions between the contents of stored frames Fl can be described by purely linear transitions, i.e. by shifting image contents linearly by a certain amounts, or whether affine transformations have to be used. Here, affine transformations can describe rotations, size changes, and shear of contents shown in the stored frames. The definition of a specific affine transformation requires more parameters, and hence more memory space than the definition of a specific motion vector. However, if only motion vectors are used, the occurrence of considerable rotation, resizing or shear in a video stream requires maintaining a large number of full frames, since such rotations, resizing or shear cannot be encoded by linear motion vectors. When the usage of affine transformations is allowed, the number of full frames can be reduced, since the occurring rotations, resizing or shear can be parameterized by the affine transformations. Hence, affine transformations can help to increase the compression rate. Therefore, affine transformations are used in addition to motion vectors in modem codecs such as AVI or VVC.
[0061] In addition, when interpolating between a given number of frames Fl in order to increase the frame rate of a video, usage of purely motion vector-based approaches will lead to poor results for considerable rotation, resizing or shear between parts of temporally consecutive frames F 1. Thus, in order to obtain high quality videos with increased frame rate usage of affine transformations in addition to motion vectors is advantageous.
[0062] The event data ED allow deducing the optical flow of contents shown in the stored frames Fl, which optical flow indicates the apparent two-dimensional movements of the contents on the frames. Thus, by deducing the optical flow from the event data ED it can be deduced whether portions of temporally consecutive frames Fl are related to a sufficient accuracy by a mere transition, i.e. by a motion vector, or whether an affine transformation relates them. In this manner, the encoding device 100 can decide whether usage of motion vectors MV is sufficient or whether affine transformations AF should be used. The encoding device 100 may determine whether to use motion vectors MV or whether to use affine transformations AF based on respective artificial intelligence models that have been trained for this purpose. However, also rule-based approaches may be used.
[0063] If affine transformations shall not be used, the encoding device 100 determines based on the event data ED and the stored frames Fl motion vectors MV, and stores them in a second storage area 120. Otherwise, i.e. if affine transformations AF shall be used, the encoding device 100 determines based on the event7
[0064] data ED and the stored frames Fl the affine transformations AF, and stores them in the second storage area 120. The encoding device 100 may determine the motion vectors MV and / or the affine transformations AF by respective artificial intelligence models that have been trained for this determination. However, also rule-based approaches may be used.
[0065] Thus, the encoding device determines and stores motion vectors MV and affine transformations AF for each of the stored frames F2, whichever is better for a high-quality compression / high-quality frame interpolation. The second storage area 120 in which the motion vectors MV and the affine transformations 120 are stored may be part of the same or a different memory that contains the first storage area 110.
[0066] The motion vectors MV and / or the affine transformations AF allow, together with the stored frames F2, reconstructing one or more missing frames MF that are temporally located between two consecutive, stored frames Fl. The missing frames MF fill in gaps between the stored frames such as to obtain the higher frame rate of the frames F2 that can be reconstructed from the encoded data. In particular, as will be described below, the stored frames Fl, the motion vectors MV and the affine transformations AF can be used in a decoding device 200 to reconstruct frames that are located between temporally consecutive stored frames F2. This allows either omitting storage of received frames, thereby reducing the data amount of the encoded data compared to the data amount of the received video stream and / or the reconstructed video stream. Or it allows interpolating missing frames MF that are not included in the received frames F 1.
[0067] The motion vectors MV and the affine transformations AF indicate changes in the image data constituting the respective frames. This means that for each region of a frame, a motion vector MV or an affine transformation can be defined that indicates a seeming movement of the contents of said region within the frame. Then, by using the original frame Fl and its motion vectors MV / its affine transformation AF, it is possible to deduce the contents of a missing frame MF. Here, as schematically illustrated in Fig. 3, forward transformations (motion vectors MV or affine transformations AF) may be used that indicate a change in image data from a stored frame Fl to a future missing frame MF. Further, also backward transformations (motion vectors MV or affine transformations AF) may be used that indicate a change in image data from a past missing frame MF to a stored frame Fl .
[0068] Although the present description focuses for ease of description on the case that there is one missing frame MF between two temporally consecutive stored frames Fl, it is also possible to have a plurality of missing frames MF between two stored frames Fl. In this case, transforms can also indicate changes between two temporally consecutive missing frames MF. The number of missing frames MF between two stored frames Fl can vary and may depend e.g. on the change rate of the image data constituting the frames. If there are only small changes, it will be possible to store less frames Fl and interpolate more missing frames MF without generating too many errors. If there are large or rapid changes, only few, one, or even no missing frame MF might be present. It should be noted that the determination of motion vectors MV and / or affine transformations AF and the reconstruction of (one or several) missing frames MF between stored frames F2 is in principle known e.g. from existing codecs such as AVI or VVC.8
[0069] As is illustrated schematically in Fig. 4, for deciding whether to use affine transformations AF or motion vectors MV, the encoding unit can divide the stored frames Fl into blocks Bl, B2. For example, each frame Fl can be divided into blocks Bl, B2 that comprise a certain number of pixels of the frame Fl. As schematically shown in Fig. 4, different block sizes can be used in different areas of the frame in order to take into account spatial or temporal (i.e. motion) details of the shown scene. Examples of such blocks Bl, B2 are e.g. coding tree units, coding tree blocks, coding units, prediction units and the like of existing codecs. In general, any area ofN x M pixels may be used as block Bl, B2.
[0070] Then, the encoding unit 100 determines for each block Bl, B2 whether affine transformations AF shall be used for reconstructing missing frames, and determines and stores the respective motion vector MV or affine transformation AF for each block.
[0071] Fig. 4 shows two temporally consecutive, stored frames Fl on which the movements of a bicycle and of a cloud can be seen. The cloud moves within a block Bl (that may contain sub-blocks). The bicycle moves within blocks B2 (that may also contain sub-blocks). While the bicycle moves essentially linearly from left to right, the cloud rotates and shrinks. Thus, the encoding unit 100 determines that in order to reconstruct missing frames MF between the two stored frames Fl an affine transformation AF has to be used. On the other hand, the movement of the bicycle can be represented by motion vectors MV pointing from left to right.
[0072] The size and direction of the motion vectors MV as well as the parameters of the affine transformations AF are determined on the one hand from the stored frames Fl serving as boundary conditions on the transformations. On the other hand, motion vectors MV and affine transformations AF are dictated by the event data ED that provide precise and temporally highly resolved information about the movements of the objects shown in the frames during the time period between the boundary frames Fl. In this manner it is possible to deduce block wise and with high precision motion vectors MV and affine transformations AF that allow reconstructing of missing frames MF that interpolate between stored frames.
[0073] Fig. 5 shows schematically further details about the parameters to be stored in conjunction with the different transformations. On the one hand, it shows the pure translations of block B2. Here, each subblock, e.g. each pixel, of the block B2 is shifted by the same two-dimensional vector v = (A, B). Thus, only two parameters must be stored for a block B2 that can be encoded via a motion vector.
[0074] On the other hand, Fig. 5 shows three different manners to represent affine transformations AF. Under a) a general affine transformation is shown. A general affine transformation in two dimensions assigns to the coordinates (x,y) of each sub-block of a block B 1 (which may be a single pixel) a translation vector v according to the following formula:
[0075] (vx\ / a b c\ / X\
[0076] [vy = d e f I y )
[0077]
[0078] \ 1 / \0 0 l / 'l'9
[0079] Thus, for a general affine transformation in two dimensions six parameters need to be stored. This general affine transformation parameterizes shifts, rotations, resizing, and shearing as illustrated in a) of Fig. 6.
[0080] However, the amount of parameters to be stored can be reduced, if one divides the affine transformations AF into two different classes. As shown in b) of Fig. 6 on the left side, a translation plus rotation can be represented by two vectors connecting two comers of the transformed block Bl before and after transformation. Also, as shown in b) of Fig. 6 on the right side, a translation plus a resizing in one dimension can be represented by two vectors.
[0081] More general transformations like shearing or a combination of rotation, resizing, and or shearing can be represented by three vectors connecting three comers of the transformed block Bl before and after transformation (see c) in Fig. 6).
[0082] Thus, instead of using a general affine transformation AF as a default, the encoding unit 100 may use affine transformations defined by two control points within one block B 1 and two respective vectors (see also b) in Fig. 5), or affine transformations defined by three control points within one block Bl and three respective vectors (see also c) in Fig. 5).
[0083] Here, the two-vector variant leads to a translation vector v for coordinate (x,y) of each sub-block of block B 1 (which may be a single pixel) according to the following formula:
[0084] „ Vlx ~ VQX
[0085] C = r = - w
[0086] £■ = _£) = -Jy _
[0087]
[0088] h
[0089] (vox, voy) and (vix, viy) are the two vectors at the control points, i.e. the comers of the block Bl, w is the width of the block Bl, and h is its height. Thus, one sees that the affine transformation AF with two control points is parameterized by four parameters, which is less than for the general affine transformation.
[0090] For the affine transformation AF with three control points the above formula also applies, however, the parameters C, D, E, and F are
[0091] r_ Vlx~ Vox
[0092] w
[0093] D _ Vix VQX
[0094] < ”h
[0095] _ Vjy Vpy
[0096] W
[0097] _ V2y— VOy
[0098]
[0099] h10
[0100] and (v2x, V2y) is the third vector. Thus, the affine transformation with three control points is defined by six parameters, just as the general affine transformation.
[0101] Thus, in order to increase data compression, it is preferable that the encoding device 100 distinguishes between affine transformations with two control points and three control points. However, if the encoding speed is more relevant, general affine transformations may be used, since then the additional classification of the used affine transformation can be omitted. Anyhow it is preferably, if the encoding device 100 stores and outputs which transform it uses for which missing frame: pure translations, i.e. motion vectors, or affine transformations, and if affine transformations: which representation - general or control point based. This allows a decoder to use the correct processing for the affine transformation coefficients.
[0102] The encoding device 100 may also be configured to determine residual frames or residual data RD based on the calculated affine transformations AR and motion vectors MV. This is schematically illustrated in Fig. 7.
[0103] In the example of Fig. 7 the encoding device 100 receives a video stream containing less frames Fl per second than should be present in the video stream that is decoded from the encoded data. In addition, as described above and as also illustrated at the top of Fig. 7, the encoding device 100 receives event data ED that indicate a temporal stream of events.
[0104] The received frames Fl and the event data ED are used by the encoding device 100 as described above to determine motion vectors MV and / or affine transformations from the event data ED. This is schematically illustrated as the second step in Fig. 7. Thus, the motion vectors MV and the affine transformations AF are not generated by comparison of frames, but by translating the motion information contained in the event data ED into the transformations, e.g. by extrapolating positions of objects within the frames based on the observed intensity changes that are recorded in the event data ED.
[0105] Then, missing frames MF are reconstructed by the encoding device 100 based on the stored frames Fl and the determined affine transformations AF and / or motion vectors MV as illustrated in the third step of Fig. 7. This is in principle known from existing codecs like AVI or VVC.
[0106] However, since no frames corresponding to the missing frames MF that were reconstructed based on the affine transformation AF and / or motion vectors MV are present in the originally received frames Fl, the encoding device 100 reconstructs a second set of missing frames MF. This is shown in the fourth step of Fig. 7, where the second set of missing frames MF is hatched by dots. The first set of missing frames MF that is reconstructed based on affine transformations AF and / or motion vectors MV and the second set of missing frames MF are generated for the same temporal position. This means that one missing frame MF of the first set and one missing frame MF of the second set are intended to show the same part of the observed scene at the same time. The second set of missing frames MF is generated based on the stored frames Fl by using an artificial intelligence model that has been trained for this reconstruction. Also, if necessary, an artificial intelligence model may be used that can also use the event data ED as an input.Thus, in the example of Fig. 7 two frames for the same time are estimated based on two different methods. First, based on affine transformations AF / motion vectors MV obtained from event data ED. And second, based on an artificial intelligence interpolation of the originally received (and stored) frames Fl that estimates the pixel values of the missing frame MF.
[0107] These two estimations can then be used by the encoding device 100 for residual data RD determination. To this end, the missing frames MF from the first set are compared to the temporally corresponding missing frames MF of the second set as illustrated in the fifth step of Fig. 7. This corresponds in particular to the generation of residuals described above with respect to Fig. 1 except that instead of an originally received frame Fl a missing frame MF that has been estimated by an artificial intelligence model is used. The difference between the respectively temporally corresponding missing frames MF can be determined and constitutes the residual data RD. Thus, in the example of Fig. 7, it is the deviation between the two estimation schemes that constitutes the residual data RD.
[0108] Finally, all reconstructed missing frames MF are dropped to only store the originally received frames Fl, the affine transformations AF, the motion vectors MV and the residual data RD. Thus, a combination of a video stream and event data can be used to encode a video stream that, when decoded, has a higher frame rate than the original video stream. The compression rate can be maintained at a high level, even if the stream contains a high number of affine transformations. Thus, an efficient encoding scheme is provided for image data and event data obtained e.g. from hybrid sensors having conventional imaging pixels and event detection pixels.
[0109] In the following an overview of such hybrid sensors that combine traditional, frame-based sensors, such as active pixel sensors, APS, and event based or dynamic vision sensors, EVS / DVS, will be given. It should be noted that in order to ease the description and also in order to cover an important application example, the following description is focused without prejudice on a specific implementation of a hybrid sensor including both, APS pixels and EVS / DVS pixels. This is of course purely exemplary. It is to be understood that a differently implemented hybrid sensor could be used.
[0110] Fig. 8 is a diagram illustrating a configuration example of a sensor device 10, which is in the example of Fig. 8 constituted by a sensor chip.
[0111] The sensor device 10 is a single-chip semiconductor chip and includes a sensor die (substrate) 11, which serves as a plurality of dies (substrates), and a logic die 12 that are stacked. Note that, the sensor device 10 can also include only a single die or three or more stacked dies.
[0112] In the sensor device 10 of Fig. 8, the sensor die 11 includes (a circuit serving as) a sensor section 21, and the logic die 12 includes a logic section 22. Note that, the sensor section 21 can be partly formed on the logic die 12. Further, the logic section 22 can be partly formed on the sensor die 11.
[0113] The sensor section 21 includes pixels configured to perform photoelectric conversion on incident light to12
[0114] generate electrical signals, and generates event data indicating the occurrence of events that are changes in the electrical signal of the pixels. The sensor section 21 supplies the event data to the logic section 22. That is, the sensor section 21 performs imaging of performing, in the pixels, photoelectric conversion on incident light to generate electrical signals, similarly to a synchronous image sensor, for example. The sensor section 21, however, also generates event data indicating the occurrence of events that are changes in the electrical signal of the pixels in addition to generating image data in a frame format (frame data). The sensor section 21 outputs, to the logic section 22, the event data obtained by the imaging.
[0115] Here, the synchronous image sensor is an image sensor configured to perform imaging in synchronization with a vertical synchronization signal and output frame data that is image data in a frame format. The event detection functions of the sensor section 21 can be regarded as asynchronous (an asynchronous image sensor) in contrast to the synchronous image sensor, since as far as event detection is concerned the sensor section 21 does not operate in synchronization with a vertical synchronization signal when outputting event data.
[0116] Note that, the sensor section 21 generates and outputs, other than event data, frame data, similarly to the synchronous image sensor. In addition, the sensor section 21 can output, together with event data, electrical signals of pixels in which events have occurred, as pixel signals that are pixel values of the pixels in frame data.
[0117] The logic section 22 controls the sensor section 21 as needed. Further, the logic section 22 performs various types of data processing, such as data processing of generating frame data on the basis of event data from the sensor section 21 and image processing on frame data from the sensor section 21 or frame data generated on the basis of the event data from the sensor section 21, and outputs data processing results obtained by performing the various types of data processing on the event data and the frame data.
[0118] Fig. 9 is a block diagram illustrating a configuration example of the sensor section 21 of Fig. 8.
[0119] The sensor section 21 includes a pixel array section 31, a driving section 32, an arbiter 33, an AD (Analog to Digital) conversion section 34, and an output section 35.
[0120] The pixel array section 31 preferably includes a plurality of pixels 51 (Fig. 10) arrayed in a two-dimensional lattice pattern. The pixel array section 31 detects, in a case where a change larger than a predetermined threshold (including a change equal to or larger than the threshold as needed) has occurred in (a voltage corresponding to) a photocurrent that is an electrical signal generated by photoelectric conversion in the pixel 51, the change in the photocurrent as an event. In a case of detecting an event, the pixel array section 31 outputs, to the arbiter 33, a request for requesting the output of event data indicating the occurrence of the event. Then, in a case of receiving a response indicating event data output permission from the arbiter 33, the pixel array section 31 outputs the event data to the driving section 32 and the output section 35. In addition, the pixel array section 31 may output an electrical signal of the pixel 51 in which the event has been detected to the AD conversion section 34, as a pixel signal.13
[0121] The driving section 32 supplies control signals to the pixel array section 31 to drive the pixel array section 31. For example, the driving section 32 drives the pixel 51 regarding which the pixel array section 31 has output event data, so that the pixel 51 in question supplies (outputs) a pixel signal to the AD conversion section 34. But the driving section 32 may also drive the pixels 51 in a synchronous manner to generate frame data.
[0122] The arbiter 33 arbitrates the requests for requesting the output of event data from the pixel array section 31, and returns responses indicating event data output permission or prohibition to the pixel array section 31.
[0123] The AD conversion section 34 includes, for example, a single-slope ADC (AD converter) (not illustrated) in each column of pixel blocks 41 (Fig. 10) described later, for example. The AD conversion section 34 performs, with the ADC in each column, AD conversion on pixel signals of the pixels 51 of the pixel blocks 41 in the column, and supplies the resultant to the output section 35. Note that, the AD conversion section 34 can perform CDS (Correlated Double Sampling) together with pixel signal AD conversion.
[0124] The output section 35 performs necessary processing on the pixel signals from the AD conversion section 34 and the event data from the pixel array section 31 and supplies the resultant to the logic section 22 (Fig.
[0125] 8).
[0126] Here, a change in the photocurrent generated in the pixel 51 can be recognized as a change in the amount of light entering the pixel 51, so that it can also be said that an event is a change in light amount (a change in light amount larger than the threshold) in the pixel 51.
[0127] Event data indicating the occurrence of an event at least includes location information (coordinates or the like) indicating the location of a pixel block in which a change in light amount, which is the event, has occurred. Besides, the event data can also include the polarity (positive or negative) of the change in light amount.
[0128] With regard to the series of event data that is output from the pixel array section 31 at timings at which events have occurred, it can be said that, as long as the event data interval is the same as the event occurrence interval, the event data implicitly includes time point information indicating (relative) time points at which the events have occurred. However, for example, when the event data is stored in a memory and the event data interval is no longer the same as the event occurrence interval, the time point information implicitly included in the event data is lost. Thus, the output section 35 includes, in event data, time point information indicating (relative) time points at which events have occurred, such as timestamps, before the event data interval is changed from the event occurrence interval. The processing of including time point information in event data can be performed in any block other than the output section 35 as long as the processing is performed before time point information implicitly included in event data is lost.
[0129] Fig. 10 is a block diagram illustrating a configuration example of the pixel array section 31 of Fig. 9.73961
[0130] 14
[0131] The pixel array section 31 may include a plurality of pixel blocks 41. Each pixel block 41 may include the IxJ pixels 51 that are one or more pixels arrayed in I rows and J columns (I and J are integers), an event detecting section 52, and a pixel signal generating section 53. The one or more pixels 51 in the pixel block 41 share the event detecting section 52 and the pixel signal generating section 53. Further, in each column of the pixel blocks 41, a VSL (Vertical Signal Line) for connecting the pixel blocks 41 to the ADC of the AD conversion section 34 is wired.
[0132] The pixel 51 receives light incident from an object and performs photoelectric conversion to generate a photocurrent serving as an electrical signal. The pixel 51 supplies the photocurrent to the event detecting section 52 under the control of the driving section 32.
[0133] The event detecting section 52 detects, as an event, a change larger than the predetermined threshold in photocurrent from each of the pixels 51, under the control of the driving section 32. In a case of detecting an event, the event detecting section 52 may supply, to the arbiter 33 (Fig. 9), a request for requesting the output of event data indicating the occurrence of the event. Then, when receiving a response indicating event data output permission to the request from the arbiter 33, the event detecting section 52 outputs the event data to the driving section 32 and the output section 35.
[0134] The pixel signal generating section 53 generates, in the case where the event detecting section 52 has detected an event, a voltage corresponding to a photocurrent from the pixel 51 as a pixel signal, and supplies the voltage to the AD conversion section 34 through the VSL, under the control of the driving section 32.
[0135] Here, detecting a change larger than the predetermined threshold in photocurrent as an event can also be recognized as detecting, as an event, absence of change larger than the predetermined threshold in photocurrent. The pixel signal generating section 53 can generate a pixel signal in the case where absence of change larger than the predetermined threshold in photocurrent has been detected as an event as well as in the case where a change larger than the predetermined threshold in photocurrent has been detected as an event.
[0136] Further, the event detecting section 52 and the pixel signal generating section 53 may operate independently from each other. This means that while event data are generated by the event detecting section 52, pixel signals are generated by the pixel signal generating section 53 in a synchronous manner to generate frame data. The event detecting section 52 and the pixel signal generating section 53 may also operate in a time multiplexed manner such that each of the sections receives the photocurrent from the pixels 51 only during different time intervals.
[0137] Fig. 11 is a circuit diagram illustrating a configuration example of the pixel block 41. The configuration of Fig. 11 may be referred to as “shared pixel configuration”.
[0138] The pixel block 41 includes, as described with reference to Fig. 10, the pixels 51, the event detecting73961
[0139] 15
[0140] section 52, and the pixel signal generating section 53.
[0141] The pixel 51 includes a photoelectric conversion element 61 and transfer transistors 62 and 63.
[0142] The photoelectric conversion element 61 includes, for example, a PD (Photodiode). The photoelectric conversion element 61 receives incident light and performs photoelectric conversion to generate charges.
[0143] The transfer transistor 62 includes, for example, an N (Negative)-type MOS (Metal-Oxide-Semiconductor) FET (Field Effect Transistor). The transfer transistor 62 of the n-th pixel 51 of the IxJ pixels 51 in the pixel block 41 is turned on or off in response to a control signal OFGn supplied from the driving section 32 (Fig. 9). When the transfer transistor 62 is turned on, charges generated in the photoelectric conversion element 61 are transferred (supplied) to the event detecting section 52, as a photocurrent.
[0144] The transfer transistor 63 includes, for example, an N-type MOSFET. The transfer transistor 63 of the n-th pixel 51 of the IxJ pixels 51 in the pixel block 41 is turned on or off in response to a control signal TRGn supplied from the driving section 32. When the transfer transistor 63 is turned on, charges generated in the photoelectric conversion element 61 are transferred to an FD 74 of the pixel signal generating section 53.
[0145] The IxJ pixels 51 in the pixel block 41 are connected to the event detecting section 52 of the pixel block 41 through nodes 60. Thus, photocurrents generated in (the photoelectric conversion elements 61 of) the pixels 51 are supplied to the event detecting section 52 through the nodes 60. As a result, the event detecting section 52 receives the sum of photocurrents from all the pixels 51 in the pixel block 41. Thus, the event detecting section 52 detects, as an event, a change in sum of photocurrents supplied from the IxJ pixels 51 in the pixel block 41.
[0146] The pixel signal generating section 53 includes a reset transistor 71, an amplification transistor 72, a selection transistor 73, and the FD (Floating Diffusion) 74.
[0147] The reset transistor 71, the amplification transistor 72, and the selection transistor 73 include, for example, N-type MOSFETs.
[0148] The reset transistor 71 is turned on or off in response to a control signal RST supplied from the driving section 32 (Fig. 9). When the reset transistor 71 is turned on, the FD 74 is connected to a power supply VDD, and charges accumulated in the FD 74 are thus discharged to the power supply VDD. With this, the FD 74 is reset.
[0149] The amplification transistor 72 has a gate connected to the FD 74, a drain connected to the power supply VDD, and a source connected to the VSL through the selection transistor 73. The amplification transistor 72 is a source follower and outputs a voltage (electrical signal) corresponding to the voltage of the FD 74 supplied to the gate to the VSL through the selection transistor 73.73961
[0150] 16
[0151] The selection transistor 73 is turned on or off in response to a control signal SEL supplied from the driving section 32. When the selection transistor 73 is turned on, a voltage corresponding to the voltage of the FD 74 from the amplification transistor 72 is output to the VSL.
[0152] The FD 74 accumulates charges transferred from the photoelectric conversion elements 61 of the pixels 51 through the transfer transistors 63, and converts the charges to voltages.
[0153] With regard to the pixels 51 and the pixel signal generating section 53, which are configured as described above, the driving section 32 turns on the transfer transistors 62 with control signals OFGn, so that the transfer transistors 62 supply, to the event detecting section 52, photocurrents based on charges generated in the photoelectric conversion elements 61 of the pixels 51. With this, the event detecting section 52 receives a current that is the sum of the photocurrents from all the pixels 51 in the pixel block 41, which might also be only a single pixel.
[0154] When the event detecting section 52 detects, as an event, a change in photocurrent (sum of photocurrents) in the pixel block 41, the driving section 32 may turn off the transfer transistors 62 of all the pixels 51 in the pixel block 41, to thereby stop the supply of the photocurrents to the event detecting section 52. Then, the driving section 32 may sequentially turn on, with the control signals TRGn, the transfer transistors 63 of the pixels 51 in the pixel block 41 in which the event has been detected, so that the transfer transistors 63 transfers charges generated in the photoelectric conversion elements 61 to the FD 74. The FD 74 accumulates the charges transferred from (the photoelectric conversion elements 61 of) the pixels 51. Voltages corresponding to the charges accumulated in the FD 74 are output to the VSL, as pixel signals of the pixels 51, through the amplification transistor 72 and the selection transistor 73.
[0155] As described above, in the sensor section 21 (Fig. 9), in this manner only pixel signals of the pixels 51 in the pixel block 41 in which an event has been detected are sequentially output to the VSL. The pixel signals output to the VSL are supplied to the AD conversion section 34 to be subjected to AD conversion.
[0156] Here, in the pixels 51 in the pixel block 41, the transfer transistors 63 can be turned on not sequentially but simultaneously. In this case, the sum of pixel signals of all the pixels 51 in the pixel block 41 can be output.
[0157] Further, it is also possible to operate the pixel generating section 53 independently from the event detection by controlling the transfer transistors 63 without regard of the event detection or by entirely omitting the transfer transistors 63. In this case frame data can be generated by the pixel generating section 53 in the conventional, synchronous manner, while the event detection can be performed concurrently. If necessary, photocurrent transfer to the event detecting section 52 and the pixel signal generating section 53 may be time multiplexed such that only one of the two transfer transistors 62, 63 per pixel 51 is open at a given point in time.
[0158] In the pixel array section 31 of Fig. 10, the pixel block 41 includes one or more pixels 51, and the one or73961
[0159] 17
[0160] more pixels 51 share the event detecting section 52 and the pixel signal generating section 53. Thus, in the case where the pixel block 41 includes a plurality of pixels 51, the numbers of the event detecting sections 52 and the pixel signal generating sections 53 can be reduced as compared to a case where the event detecting section 52 and the pixel signal generating section 53 are provided for each of the pixels 51, with the result that the scale of the pixel array section 31 can be reduced.
[0161] Note that, in the case where the pixel block 41 includes a plurality of pixels 51, the event detecting section 52 can be provided for each of the pixels 51. In the case where the plurality of pixels 51 in the pixel block 41 share the event detecting section 52, events are detected in units of the pixel blocks 41. In the case where the event detecting section 52 is provided for each of the pixels 51, however, events can be detected in units of the pixels 51.
[0162] Yet, even in the case where the plurality of pixels 51 in the pixel block 41 share the single event detecting section 52, events can be detected in units of the pixels 51 when the transfer transistors 62 of the plurality of pixels 51 are temporarily turned on in a time-division manner.
[0163] Further, in a case where there is no need to output pixel signals, the pixel block 41 can be formed without the pixel signal generating section 53. In the case where the pixel block 41 is formed without the pixel signal generating section 53, the sensor section 21 can be formed without the AD conversion section 34 and the transfer transistors 63. In this case, the scale of the sensor section 21 can be reduced. The sensor will then output the address of the pixel (block) in which the event occurred, if necessary with a time stamp.
[0164] Moreover, additional pixels 51 connected to pixel signal generating sections 53, but not to event detecting sections 52 may be provided on the sensor die 11. These pixels 51 may be interleaved with the pixels 51 connected to the event detecting section 52, but may also form a separate pixel array for generating frame data. The separation of the event detection function and the pixel signal generating function may even by implemented as two pixel arrays on different dies that are arranged such as to observe the same scene. Thus basically any pixel arrangement / pixel circuitry might be used that allows obtaining event data with a first subset of pixels 51 and generating of pixel signals with a second subset of pixels 51.
[0165] Fig. 12 is a block diagram illustrating a configuration example of the event detecting section 52 of Fig. 10.
[0166] The event detecting section 52 includes a current-voltage converting section 81, a buffer 82, a subtraction section 83, a quantization section 84, and a transfer section 85.
[0167] The current-voltage converting section 81 converts (a sum of) photocurrents from the pixels 51 to voltages corresponding to the logarithms of the photocurrents (hereinafter also referred to as a "photovoltage") and supplies the voltages to the buffer 82.
[0168] The buffer 82 buffers photovoltages from the current-voltage converting section 81 and supplies the resultant to the subtraction section 83.73961
[0169] 18
[0170] The subtraction section 83 calculates, at a timing instructed by a row driving signal that is a control signal from the driving section 32, a difference between the current photovoltage and a photovoltage at a timing slightly shifted from the current time, and supplies a difference signal corresponding to the difference to the quantization section 84.
[0171] The quantization section 84 quantizes difference signals from the subtraction section 83 to digital signals and supplies the quantized values of the difference signals to the transfer section 85 as event data.
[0172] The transfer section 85 transfers (outputs), on the basis of event data from the quantization section 84, the event data to the output section 35. That is, the transfer section 85 supplies a request for requesting the output of the event data to the arbiter 33. Then, when receiving a response indicating event data output permission to the request from the arbiter 33, the transfer section 85 outputs the event data to the output section 35.
[0173] Fig. 13 is a circuit diagram illustrating a configuration example of the current-voltage converting section 81 of Fig. 12.
[0174] The current-voltage converting section 81 includes transistors 91 to 93. As the transistors 91 and 93, for example, N-type MOSFETs can be employed. As the transistor 92, for example, a P-type MOSFET can be employed.
[0175] The transistor 91 has a source connected to the gate of the transistor 93, and a photocurrent is supplied from the pixel 51 to the connecting point between the source of the transistor 91 and the gate of the transistor 93. The transistor 91 has a drain connected to the power supply VDD and a gate connected to the drain of the transistor 93.
[0176] The transistor 92 has a source connected to the power supply VDD and a drain connected to the connecting point between the gate of the transistor 91 and the drain of the transistor 93. A predetermined bias voltage Vbias is applied to the gate of the transistor 92. With the bias voltage Vbias, the transistor 92 is turned on or off, and the operation of the current-voltage converting section 81 is turned on or off depending on whether the transistor 92 is turned on or off.
[0177] The source of the transistor 93 is grounded.
[0178] In the current-voltage converting section 81, the transistor 91 has the drain connected on the power supply VDD side and is thus a source follower. The source of the transistor 91, which is the source follower, is connected to the pixels 51 (Fig. 11), so that photocurrents based on charges generated in the photoelectric conversion elements 61 of the pixels 51 flow through the transistor 91 (from the drain to the source). The transistor 91 operates in a subthreshold region, and at the gate of the transistor 91, photovoltages corresponding to the logarithms of the photocurrents flowing through the transistor 91 are generated. As described above, in the current-voltage converting section 81, the transistor 91 converts photocurrents73961
[0179] 19
[0180] from the pixels 51 to photovoltages corresponding to the logarithms of the photocurrents.
[0181] In the current-voltage converting section 81, the transistor 91 has the gate connected to the connecting point between the drain of the transistor 92 and the drain of the transistor 93, and the photovoltages are output from the connecting point in question.
[0182] Fig. 14 is a circuit diagram illustrating configuration examples of the subtraction section 83 and the quantization section 84 of Fig. 12.
[0183] The subtraction section 83 includes a capacitor 101, an operational amplifier 102, a capacitor 103, and a switch 104. The quantization section 84 includes a comparator 111.
[0184] The capacitor 101 has one end connected to the output terminal of the buffer 82 (Fig. 12) and the other end connected to the input terminal (inverting input terminal) of the operational amplifier 102. Thus, photovoltages are input to the input terminal of the operational amplifier 102 through the capacitor 101.
[0185] The operational amplifier 102 has an output terminal connected to the non-inverting input terminal (+) of the comparator 111.
[0186] The capacitor 103 has one end connected to the input terminal of the operational amplifier 102 and the other end connected to the output terminal of the operational amplifier 102.
[0187] The switch 104 is connected to the capacitor 103 to switch the connections between the ends of the capacitor 103. The switch 104 is turned on or off in response to a row driving signal that is a control signal from the driving section 32, to thereby switch the connections between the ends of the capacitor 103.
[0188] A photovoltage on the buffer 82 (Fig. 12) side of the capacitor 101 when the switch 104 is on is denoted by Vinit, and the capacitance (electrostatic capacitance) of the capacitor 101 is denoted by Cl. The input terminal of the operational amplifier 102 serves as a virtual ground terminal, and a charge Qinit that is accumulated in the capacitor 101 in the case where the switch 104 is on is expressed by Expression (1).
[0189] Qinit = Cl x Vinit (1)
[0190] Further, in the case where the switch 104 is on, the connection between the ends of the capacitor 103 is cut (short-circuited), so that no charge is accumulated in the capacitor 103.
[0191] When a photovoltage on the buffer 82 (Fig. 12) side of the capacitor 101 in the case where the switch 104 has thereafter been turned off is denoted by Vafter, a charge Qafter that is accumulated in the capacitor 101 in the case where the switch 104 is off is expressed by Expression (2).
[0192] Qafter = Cl x Vafter (2)73961
[0193] 20
[0194] When the capacitance of the capacitor 103 is denoted by C2 and the output voltage of the operational amplifier 102 is denoted by Vout, a charge Q2 that is accumulated in the capacitor 103 is expressed by Expression (3).
[0195] Q2 = -C2 x Vout (3)
[0196] Since the total amount of charges in the capacitors 101 and 103 does not change before and after the switch 104 is turned off, Expression (4) is established.
[0197] Qinit = Qafter + Q2 (4)
[0198] When Expression (1) to Expression (3) are substituted for Expression (4), Expression (5) is obtained.
[0199] Vout = - (C1 / C2) x (Vafter - Vinit) (5)
[0200] With Expression (5), the subtraction section 83 subtracts the photovoltage Vinit from the photovoltage Vafter, that is, calculates the difference signal (Vout) corresponding to a difference Vafter - Vinit between the photovoltages Vafter and Vinit. With Expression (5), the subtraction gain of the subtraction section 83 is C1 / C2. Since the maximum gain is normally desired, Cl is preferably set to a large value and C2 is preferably set to a small value. Meanwhile, when C2 is too small, kTC noise increases, resulting in a risk of deteriorated noise characteristics. Thus, the capacitance C2 can only be reduced in a range that achieves acceptable noise. Further, since the pixel blocks 41 each have installed therein the event detecting section 52 including the subtraction section 83, the capacitances Cl and C2 have space constraints. In consideration of these matters, the values of the capacitances Cl and C2 are determined.
[0201] The comparator 111 compares a difference signal from the subtraction section 83 with a predetermined threshold (voltage) Vth (>0) applied to the inverting input terminal (-), thereby quantizing the difference signal. The comparator 111 outputs the quantized value obtained by the quantization to the transfer section 85 as event data.
[0202] For example, in a case where a difference signal is larger than the threshold Vth, the comparator 111 outputs an H (High) level indicating 1, as event data indicating the occurrence of an event. In a case where a difference signal is not larger than the threshold Vth, the comparator 111 outputs an L (Low) level indicating 0, as event data indicating that no event has occurred.
[0203] The transfer section 85 supplies a request to the arbiter 33 in a case where it is confirmed on the basis of event data from the quantization section 84 that a change in light amount that is an event has occurred, that is, in the case where the difference signal (Vout) is larger than the threshold Vth. When receiving a response indicating event data output permission, the transfer section 85 outputs the event data indicating the occurrence of the event (for example, H level) to the output section 35.73961
[0204] 21
[0205] The output section 35 includes, in event data from the transfer section 85, location / address information regarding (the pixel block 41 including) the pixel 51 in which an event indicated by the event data has occurred and time point information indicating a time point at which the event has occurred, and further, as needed, the polarity of a change in light amount that is the event, i.e. whether the intensity did increase or decrease. The output section 35 outputs the event data.
[0206] As the data format of event data including location information regarding the pixel 51 in which an event has occurred, time point information indicating a time point at which the event has occurred, and the polarity of a change in light amount that is the event, for example, the data format called "AER (Address Event Representation)" can be employed.
[0207] Note that, a gain A of the entire event detecting section 52 is expressed by the following expression where the gain of the current-voltage converting section 81 is denoted by CGiogand the gain of the buffer 82 is 1.
[0208] A = CGiogC 1 / C2 (ZiPhoto_n) (6)
[0209] Here, iPhoto_n denotes a photocurrent of the n-th pixel 51 of the IxJ pixels 51 in the pixel block 41. In Expression (6), E denotes the summation of n that takes integers ranging from 1 to IxJ.
[0210] Note that, the pixel 51 can receive any light as incident light with an optical fdter through which predetermined light passes, such as a color fdter. For example, in a case where the pixel 51 receives visible light as incident light, event data indicates the occurrence of changes in pixel value in images including visible objects. Further, for example, in a case where the pixel 51 receives, as incident light, infrared light, millimeter waves, or the like for ranging, event data indicates the occurrence of changes in distances to objects. In addition, for example, in a case where the pixel 51 receives infrared light for temperature measurement, as incident light, event data indicates the occurrence of changes in temperature of objects. In the present embodiment, the pixel 51 is assumed to receive visible light as incident light.
[0211] Fig. 15 is a timing chart illustrating an example of the operation of the sensor section 21 of Fig. 9.
[0212] At Timing TO, the driving section 32 changes all the control signals OFGn from the L level to the H level, thereby turning on the transfer transistors 62 of all the pixels 51 in the pixel block 41. With this, the sum of photocurrents from all the pixels 51 in the pixel block 41 is supplied to the event detecting section 52. Here, the control signals TRGn are all at the L level, and hence the transfer transistors 63 of all the pixels 51 are off.
[0213] For example, at Timing Tl, when detecting an event, the event detecting section 52 outputs event data at the H level in response to the detection of the event.
[0214] At Timing T2, the driving section 32 sets all the control signals OFGn to the L level on the basis of the event data at the H level, to stop the supply of the photocurrents from the pixels 51 to the event detecting section 52. Further, the driving section 32 sets the control signal SEL to the H level, and sets the control73961
[0215] 22
[0216] signal RST to the H level over a certain period of time, to control the FD 74 to discharge the charges to the power supply VDD, thereby resetting the FD 74. The pixel signal generating section 53 outputs, as a reset level, a pixel signal corresponding to the voltage of the FD 74 when the FD 74 has been reset, and the AD conversion section 34 performs AD conversion on the reset level.
[0217] At Timing T3 after the reset level AD conversion, the driving section 32 sets a control signal TRG1 to the H level over a certain period to control the first pixel 51 in the pixel block 41 in which the event has been detected (or which is triggered for other reasons, as e.g. time multiplexed output and / or synchronous readout of frame data) to transfer, to the FD 74, charges generated by photoelectric conversion in (the photoelectric conversion element 61 of) the first pixel 51. The pixel signal generating section 53 outputs, as a signal level, a pixel signal corresponding to the voltage of the FD 74 to which the charges have been transferred from the pixel 51, and the AD conversion section 34 performs AD conversion on the signal level.
[0218] The AD conversion section 34 outputs, to the output section 35, a difference between the signal level and the reset level obtained after the AD conversion, as a pixel signal serving as a pixel value of the image (frame data).
[0219] Here, the processing of obtaining a difference between a signal level and a reset level as a pixel signal serving as a pixel value of an image is called "CDS." CDS can be performed after the AD conversion of a signal level and a reset level, or can be simultaneously performed with the AD conversion of a signal level and a reset level in a case where the AD conversion section 34 performs single-slope AD conversion. In the latter case, AD conversion is performed on the signal level by using the AD conversion result of the reset level as an initial value.
[0220] At Timing T4 after the AD conversion of the pixel signal of the first pixel 51 in the pixel block 41, the driving section 32 sets a control signal TRG2 to the H level over a certain period of time to control the second pixel 51 in the pixel block 41 in which the event has been detected to output a pixel signal.
[0221] In the sensor section 21, similar processing is executed thereafter, so that pixel signals of the pixels 51 in the pixel block 41 in which the event has been detected are sequentially output.
[0222] When the pixel signals of all the pixels 51 in the pixel block 41 are output, the driving section 32 sets all the control signals OFGn to the H level to turn on the transfer transistors 62 of all the pixels 51 in the pixel block 41.
[0223] Fig. 16 is a diagram illustrating an example of a frame data generation method based on event data.
[0224] The logic section 22 sets a frame interval and a frame width on the basis of an externally input command, for example. Here, the frame interval represents the interval of frames of frame data that is generated on the basis of event data. The frame width represents the time width of event data that is used for generating frame data on a single frame. A frame interval and a frame width that are set by the logic section 22 are73961
[0225] 23
[0226] also referred to as a "set frame interval" and a "set frame width," respectively.
[0227] The logic section 22 generates, on the basis of the set frame interval, the set frame width, and event data from the sensor section 21, frame data that is image data in a frame format, to thereby convert the event data to the frame data.
[0228] That is, the logic section 22 generates, in each set frame interval, frame data on the basis of event data in the set frame width from the beginning of the set frame interval.
[0229] Here, it is assumed that event data includes time point information ti indicating a time point at which an event has occurred (hereinafter also referred to as an "event time point") and coordinates (x, y) serving as location information regarding (the pixel block 41 including) the pixel 51 in which the event has occurred (hereinafter also referred to as an "event location").
[0230] In Fig. 16, in a three-dimensional space (time and space) with the x axis, the y axis, and the time axis t, points representing event data are plotted on the basis of the event time point t and the event location (coordinates) (x, y) included in the event data.
[0231] That is, when a location (x, y, t) on the three-dimensional space indicated by the event time point t and the event location (x, y) included in event data is regarded as the space-time location of an event, in Fig. 16, the points representing the event data are plotted on the space-time locations (x, y, t) of the events.
[0232] The logic section 22 starts to generate frame data on the basis of event data by using, as a generation start time point at which frame data generation starts, a predetermined time point, for example, a time point at which frame data generation is externally instructed or a time point at which the sensor device 10 is powered on.
[0233] Here, cuboids each having the set frame width in the direction of the time axis t in the set frame intervals, which appear from the generation start time point, are referred to as a "frame volume." The size of the frame volume in the x-axis direction or the y-axis direction is equal to the number of the pixel blocks 41 or the pixels 51 in the x-axis direction or the y-axis direction, for example.
[0234] The logic section 22 generates, in each set frame interval, frame data on a single frame on the basis of event data in the frame volume having the set frame width from the beginning of the set frame interval.
[0235] Frame data can be generated by, for example, setting white to a pixel (pixel value) in a frame at the event location (x, y) included in event data and setting a predetermined color such as gray to pixels at other locations in the frame.
[0236] Besides, in a case where event data includes the polarity of a change in light amount that is an event, frame data can be generated in consideration of the polarity included in the event data. For example, white can be set to pixels in the case a positive polarity, while black can be set to pixels in the case of a73961
[0237] 24
[0238] negative polarity.
[0239] In addition, in the case where pixel signals of the pixels 51 are also output when event data is output as described with reference to Fig. 10 and Fig. 11, frame data can be generated on the basis of the event data by using the pixel signals of the pixels 51. That is, frame data can be generated by setting, in a frame, a pixel at the event location (x, y) (in a block corresponding to the pixel block 41) included in event data to a pixel signal of the pixel 51 at the location (x, y) and setting a predetermined color such as gray to pixels at other locations.
[0240] Note that, in the frame volume, there are a plurality of pieces of event data that are different in the event time point t but the same in the event location (x, y) in some cases. In this case, for example, event data at the latest or oldest event time point t can be prioritized. Further, in the case where event data includes polarities, the polarities of a plurality of pieces of event data that are different in the event time point t but the same in the event location (x, y) can be added together, and a pixel value based on the added value obtained by the addition can be set to a pixel at the event location (x, y).
[0241] Here, in a case where the frame width and the frame interval are the same, the frame volumes are adjacent to each other without any gap. Further, in a case where the frame interval is larger than the frame width, the frame volumes are arranged with gaps. In a case where the frame width is larger than the frame interval, the frame volumes are arranged to be partly overlapped with each other.
[0242] As explained above, the pixel signal generating section 53 may also generate frame data in the conventional, synchronous manner.
[0243] Fig. 17 is a block diagram illustrating another configuration example of the quantization section 84 of Fig.
[0244] 12.
[0245] Note that, in Fig. 17, parts corresponding to those in the case of Fig. 14 are denoted by the same reference signs, and the description thereof is omitted as appropriate below.
[0246] In Fig. 17, the quantization section 84 includes comparators 111 and 112 and an output section 113.
[0247] Thus, the quantization section 84 of Fig. 17 is similar to the case of Fig. 14 in including the comparator 111. However, the quantization section 84 of Fig. 17 is different from the case of Fig. 14 in newly including the comparator 112 and the output section 113.
[0248] The event detecting section 52 (Fig. 12) including the quantization section 84 of Fig. 17 detects, in addition to events, the polarities of changes in light amount that are events.
[0249] In the quantization section 84 of Fig. 17, the comparator 111 outputs, in the case where a difference signal is larger than the threshold Vth, the H level indicating 1, as event data indicating the occurrence of an event having the positive polarity. The comparator 111 outputs, in the case where a difference signal is73961
[0250] 25
[0251] not larger than the threshold Vth, the L level indicating 0, as event data indicating that no event having the positive polarity has occurred.
[0252] Further, in the quantization section 84 of Fig. 17, a threshold Vth' (<Vth) is supplied to the non-inverting input terminal (+) of the comparator 112, and difference signals are supplied to the inverting input terminal (-) of the comparator 112 from the subtraction section 83. Here, for the sake of simple description, it is assumed that the threshold Vth' is equal to -Vth, for example, which needs however not to be the case.
[0253] The comparator 112 compares a difference signal from the subtraction section 83 with the threshold Vth' applied to the inverting input terminal (-), thereby quantizing the difference signal. The comparator 112 outputs, as event data, the quantized value obtained by the quantization.
[0254] For example, in a case where a difference signal is smaller than the threshold Vth' (the absolute value of the difference signal having a negative value is larger than the threshold Vth), the comparator 112 outputs the H level indicating 1, as event data indicating the occurrence of an event having the negative polarity. Further, in a case where a difference signal is not smaller than the threshold Vth' (the absolute value of the difference signal having a negative value is not larger than the threshold Vth), the comparator 112 outputs the L level indicating 0, as event data indicating that no event having the negative polarity has occurred.
[0255] The output section 113 outputs, on the basis of event data output from the comparators 111 and 112, event data indicating the occurrence of an event having the positive polarity, event data indicating the occurrence of an event having the negative polarity, or event data indicating that no event has occurred to the transfer section 85.
[0256] For example, the output section 113 outputs, in a case where event data from the comparator 111 is the H level indicating 1, +V volts indicating +1, as event data indicating the occurrence of an event having the positive polarity, to the transfer section 85. Further, the output section 113 outputs, in a case where event data from the comparator 112 is the H level indicating 1, -V volts indicating -1, as event data indicating the occurrence of an event having the negative polarity, to the transfer section 85. In addition, the output section 113 outputs, in a case where each event data from the comparators 111 and 112 is the L level indicating 0, 0 volts (GND level) indicating 0, as event data indicating that no event has occurred, to the transfer section 85.
[0257] The transfer section 85 supplies a request to the arbiter 33 in the case where it is confirmed on the basis of event data from the output section 113 of the quantization section 84 that a change in light amount that is an event having the positive polarity or the negative polarity has occurred. After receiving a response indicating event data output permission, the transfer section 85 outputs event data indicating the occurrence of the event having the positive polarity or the negative polarity (+V volts indicating 1 or -V volts indicating -1) to the output section 35.73961
[0258] 26
[0259] Preferably, the quantization section 84 has a configuration as illustrated in Fig. 17.
[0260] Fig. 18 is a diagram illustrating another configuration example of the event detecting section 52.
[0261] In Fig. 18, the event detecting section 52 includes a subtractor 430, a quantizer 440, a memory 451, and a controller 452. The subtractor 430 and the quantizer 440 correspond to the subtraction section 83 and the quantization section 84, respectively.
[0262] Note that, in Fig. 18, the event detecting section 52 further includes blocks corresponding to the currentvoltage converting section 81 and the buffer 82, but the illustrations of the blocks are omitted in Fig. 18.
[0263] The subtractor 430 includes a capacitor 431, an operational amplifier 432, a capacitor 433, and a switch 434. The capacitor 431, the operational amplifier 432, the capacitor 433, and the switch 434 correspond to the capacitor 101, the operational amplifier 102, the capacitor 103, and the switch 104, respectively.
[0264] The quantizer 440 includes a comparator 441. The comparator 441 corresponds to the comparator 111.
[0265] The comparator 441 compares a voltage signal (difference signal) from the subtractor 430 with the predetermined threshold voltage Vth applied to the inverting input terminal (-). The comparator 441 outputs a signal indicating the comparison result, as a detection signal (quantized value).
[0266] The voltage signal from the subtractor 430 may be input to the input terminal (-) of the comparator 441, and the predetermined threshold voltage Vth may be input to the input terminal (+) of the comparator 441.
[0267] The controller 452 supplies the predetermined threshold voltage Vth applied to the inverting input terminal (-) of the comparator 441. The threshold voltage Vth which is supplied may be changed in a time-division manner. For example, the controller 452 supplies a threshold voltage Vthl corresponding to ON events (for example, positive changes in photocurrent) and a threshold voltage Vth2 corresponding to OFF events (for example, negative changes in photocurrent) at different timings to allow the single comparator to detect a plurality of types of address events (events).
[0268] The memory 451 accumulates output from the comparator 441 on the basis of Sample signals supplied from the controller 452. The memory 451 may be a sampling circuit, such as a switch, plastic, or capacitor, or a digital memory circuit, such as a latch or flip-flop. For example, the memory 451 may hold, in a period in which the threshold voltage Vth2 corresponding to OFF events is supplied to the inverting input terminal (-) of the comparator 441, the result of comparison by the comparator 441 using the threshold voltage Vthl corresponding to ON events. Note that, the memory 451 may be omitted, may be provided inside the pixel (pixel block 41), or may be provided outside the pixel.
[0269] Fig. 19 is a block diagram illustrating another configuration example of the pixel array section 31 of Fig.
[0270] 9.73961
[0271] 27
[0272] Note that, in Fig. 19, parts corresponding to those in the case of Fig. 10 are denoted by the same reference signs, and the description thereof is omitted as appropriate below.
[0273] In Fig. 19, the pixel array section 31 includes the plurality of pixel blocks 41. The pixel block 41 includes the lx J pixels 51 that are one or more pixels and the event detecting section 52.
[0274] Thus, the pixel array section 31 of Fig. 19 is similar to the case of Fig. 10 in that the pixel array section 31 includes the plurality of pixel blocks 41 and that the pixel block 41 includes one or more pixels 51 and the event detecting section 52. However, the pixel array section 31 of Fig. 19 is different from the case of Fig.
[0275] 10 in that the pixel block 41 does not include the pixel signal generating section 53.
[0276] As described above, in the pixel array section 31 of Fig. 19, the pixel block 41 does not include the pixel signal generating section 53, so that the sensor section 21 (Fig. 9) can be formed without the AD conversion section 34. As further described above, the pixel signal generating section 53 and pixels 51 feeding into it may be placed on a different part of the die 11 or even on another sensor chip.
[0277] Fig. 20 is a circuit diagram illustrating a configuration example of the pixel block 41 of Fig. 19.
[0278] As described with reference to Fig. 19, the pixel block 41 includes the pixels 51 and the event detecting section 52, but does not include the pixel signal generating section 53.
[0279] In this case, the pixel 51 can only include the photoelectric conversion element 61 without the transfer transistors 62 and 63.
[0280] Note that, in the case where the pixel 51 has the configuration illustrated in Fig. 20, the event detecting section 52 can output a voltage corresponding to a photocurrent from the pixel 51, as a pixel signal.
[0281] Above, the sensor device 10 was described to be an asynchronous imaging device configured to read out events by the asynchronous readout system. However, the event readout system is not limited to the asynchronous readout system and may be the synchronous readout system. An imaging device to which the synchronous readout system is applied is a scan type imaging device that is the same as a general imaging device configured to perform imaging at a predetermined frame rate. Further, the event data detection may be performed asynchronously, while the pixel signal generation may be performed synchronously.
[0282] Fig. 21 is a block diagram illustrating a configuration example of a scan type imaging device.
[0283] As illustrated in Fig. 21, an imaging device 510 includes a pixel array section 521, a driving section 522, a signal processing section 525, a read-out region selecting section 527, and a signal generating section 528.
[0284] The pixel array section 521 includes a plurality of pixels 530. The plurality of pixels 530 each output an73961
[0285] 28
[0286] output signal in response to a selection signal from the read-out region selecting section 527. The plurality of pixels 530 can each include an in-pixel quantizer as illustrated in Fig. 18, for example. The plurality of pixels 530 output output signals corresponding to the amounts of change in light intensity. The plurality of pixels 530 may be two-dimensionally disposed in a matrix as illustrated in Fig. 21.
[0287] The driving section 522 drives the plurality of pixels 530, so that the pixels 530 output pixel signals generated in the pixels 530 to the signal processing section 525 through an output line 514. Note that, the driving section 522 and the signal processing section 525 are circuit sections for acquiring grayscale information. Thus, in a case where only event information (event data) is acquired, the driving section 522 and the signal processing section 525 may be omitted.
[0288] The read-out region selecting section 527 selects some of the plurality of pixels 530 included in the pixel array section 521. For example, the read-out region selecting section 527 selects one or a plurality of rows included in the two-dimensional matrix structure corresponding to the pixel array section 521. The readout region selecting section 527 sequentially selects one or a plurality of rows on the basis of a cycle set in advance. Further, the read-out region selecting section 527 may determine a selection region on the basis of requests from the pixels 530 in the pixel array section 521.
[0289] The signal generating section 528 generates, on the basis of output signals of the pixels 530 selected by the read-out region selecting section 527, event signals corresponding to active pixels in which events have been detected of the selected pixels 530. The events mean an event that the intensity of light changes. The active pixels mean the pixel 530 in which the amount of change in light intensity corresponding to an output signal exceeds or falls below a threshold set in advance. For example, the signal generating section 528 compares output signals from the pixels 530 with a reference signal, and detects, as an active pixel, a pixel that outputs an output signal larger or smaller than the reference signal. The signal generating section 528 generates an event signal (event data) corresponding to the active pixel.
[0290] The signal generating section 528 can include, for example, a column selecting circuit configured to arbitrate signals input to the signal generating section 528. Further, the signal generating section 528 can output not only information regarding active pixels in which events have been detected, but also information regarding non-active pixels in which no event has been detected, i.e. it can operate as a conventional synchronous image sensorthat generates a series of consecutive image frames.
[0291] The signal generating section 528 outputs, through an output line 515, address information and timestamp information (for example, (X, Y, T)) regarding the active pixels in which the events have been detected. However, the data that is output from the signal generating section 528 may not only be the address information and the timestamp information, but also information in a frame format (for example, (0, 0, 1, o, -)).
[0292] In the above description a sensor device 10 has been described in which event data generation and pixel signal generation may depend on each other or may be independent of each other. Moreover, pixels 51 may be shared between event detection circuitry and pixel signal generating circuitry either by respective73961
[0293] 29
[0294] circuitry or by time multiplexing. But pixels 51 may also be divided such that there are event detection pixels and pixel signal generating pixels. These pixels 51 may be interleaved in the same pixel array or may be part of different pixel arrays on the same die or even be arranged on different dies.
[0295] A sensor device 10 that comprises the above-described encoding device 100 and a hybrid image sensor as described above is schematically illustrated in Fig. 22. Besides the encoding device 100 the sensor device 10 comprises a plurality of pixels 51 that are each configured to receive light from a scene and perform photoelectric conversion to generate an electrical signal. Here, the pixels 51 might be arranged in pixel blocks 41 as described with respect to Fig. 10. For the ease of description, it is assumed in the following without limitation that there is only one pixel 51 per pixel block 41, i.e. that pixels 51 and pixel blocks 41 refer to the same structural elements. However, the following description could be easily generalized by understanding any reference to a pixel 51 as a reference to a pixel block 41 comprising several pixels 51. In Fig. 22 a two-dimensional array of pixels 51 is shown. However, also a different arrangement of pixels 51 is possible, such as e.g. a one-dimensional line or a scattered, arbitrary arrangement. The following reference to a two-dimensional pixel array is only chosen to ease the description.
[0296] The sensor device 10 comprises event detection circuitry 20 that is configured to detect as event data intensity changes above a predetermined threshold of the light received by each of a first subset SI of the pixels 51. The event detection circuitry 20 is basically constituted by the event detecting sections 52 as described above that are capable to receive the photocurrent of the pixels 51 and to detect events due to intensity changes of the received light that are larger than predetermined (but potentially dynamically adjustable) event thresholds. Although the event detection circuitry 20 has been depicted in Fig. 22 as a block adjacent to the pixels 51 this has to be understood as merely symbolic. As discussed above, the event detecting sections 52 of the event detection circuitry 20 may be arranged close or next to the pixels 51. However, the event detecting sections 52 could also be grouped together in a dedicated area on the sensor chip. In particular, the event detection circuitry 20 could be arranged on a different wafer / die than the pixels 51 and be vertically connected with the pixels 51.
[0297] The pixels 51 comprise the first subset SI. In Fig. 22 the pixels 51 of the first subset SI are indicated by vertical hatching. These pixels 51 are connected to the event detection circuitry 20, e.g. as described above, and allow event detection on the light received by them.
[0298] The pixels 51 further comprise a second subset S2 that is formed in Fig. 22 by the unhatched pixels 51. The pixels 51 of the second subset S2 are used to generate full intensity information, i.e. to generate data representing the (absolute) intensity of the light received by the pixels 51. Here, it is understood that appropriate color filters may be provided to each of the pixels 51 of the second subset S2 to allow generation of a color image. Color filters may also be provided to the pixels 51 of the first subset S 1.
[0299] The pixels 51 of the first subset S 1 and the pixels 51 of the second subset S2 are arranged such that they observe the same scene. This means that pixels 51 of the first subset S 1 and pixels 51 of the second subset 52 receive light from the same location of the scene. A pixel 51 of the first subset S 1 and a pixel 51 of the second subset S2 are associated with each other when they receive light from the same location.73961
[0300] 30
[0301] The sensor device 10 comprises pixel signal generating circuitry 30 that is configured to generate a pixel signal indicating intensity values of the received light for each pixel 51 of the second subset S2 of the pixels 51. For example, the pixel signal generating circuitry 30 may be constituted by the pixel signal generating sections 53 as described above that can produce pixel signals indicating the received light intensity in an asynchronous (event triggered) or conventional, synchronous manner. In particular, the pixels 51 of the second subset S2 may operate according to the known principles of an active pixel sensor, APS. As is commonly known, the pixel signal generating circuitry 30 is configured to generate the pixel signals based on the amount of light received at the respective pixels 51 of the second subset S2 during an exposure period of said pixel 51.
[0302] Just as the event detection circuitry 20 also the pixel signal generating circuitry 30 may be formed either in a distributed manner, where e.g. one pixel signal generating section 53 is arranged next to or close to one pixel 51 of the second subset S2, or in a dedicated region of the sensor chip. The pixel signal generating circuitry 30 may also be considered to include all control, reset, and / or selection lines and the like that are necessary to operate the sensor device 10 as an APS that is capable to generate a consecutive stream of pixel signals indicating the intensity of the received light over a given time period.
[0303] The sensor device comprises further a control unit 40 that is configured to control the event detection circuitry 20 and the pixel signal generating circuitry 30. The control unit 40 may be any arrangement of circuitry that is capable to carry out the functions necessary to control the event detection circuitry 20 and the pixel signal generating circuitry 30. For example, the control unit 40 may be constituted by a processor. The control unit 40 may be part of the pixel section of the sensor device 10 and may be placed on the same die(s) as the other components of the sensor device 10. But the control unit 10 may also be arranged separately, e.g. on a separate die. The functions of the control unit 10 may be fully implemented in hardware, in software or may be implemented as a mixture of hardware and software functions. The control unit 40 may further be configured to provide the video stream generated by the pixels 51 of the second subset S2 and the event data generated by the pixels 51 of the first subset S 1 to the encoding device 100 to enable it to carry out the encoding process as described above with respect to Fig. 4.
[0304] In the schematic illustration of the sensor device 10 given in Fig. 22, which is also shown in Fig. 23 a) the pixels 51 in the first subset SI differ from the pixels 51 in the second subset S2. Thus, while the event detection circuitry 20 receives signals only from the first subset SI in this example, the pixel signal generating unit 30 receives signals only from the second subset S2.
[0305] As shown in Figs. 22 and 23 a) one implementation for such a division of pixels 51 is to arrange all pixels 51 in a single pixel array. Then, a specific number of pixels 51 can be designed for event detection, such as by providing circuitry as described above with respect to Fig. 20, for example for one pixel of a N x M pixel group, e.g. of a 2 x 2 pixel group as shown in Figs. 22 and 23 a). These pixels 51 form then the first subset SI of pixels 51. The remaining pixels 51 form the second subset S2 of pixels 51, and can be provided with conventional circuitry for a synchronous readout.73961
[0306] 31
[0307] However, at least a part of the pixels 51 may belong to the first subset of pixels 51 as well as to the second subset of pixels 51. For example, all pixels 51 of a pixel array may function as pixels 51 of the first subset SI and the second subset S2 as schematically illustrated in Fig. 23 b). The pixels 51 may then be connected to the event detection circuitry 20 and the pixel signal circuitry 30 e.g. as described above with respect to Fig. 11, i.e. each pixel 51 can be selectively chosen to transfer its photocurrent to the event detection circuitry 20 or the pixel signal generation circuitry 30, e.g. in a time multiplexed manner. This shared pixel architecture allows to have the same spatial resolution for event detection and intensity information generation.
[0308] In the above description it was assumed that pixels 51 of both pixel subsets SI, S2 were part of the same pixel array. However, as schematically illustrated in Fig. 23 c) it is also conceivable that two different pixel arrays are used that each contain only pixels 51 of one pixel subset, either on different parts of a single die or even on different sensor chips. As long as both pixel arrays observe the same scene it will nevertheless be possible to reconstruct intensity information for the pixels 51 of the second subset S2 based on the events detected by the pixels 51 of the first subset S 1.
[0309] In general, any geometrical arrangement of pixels 51 with shared or divided functionality will be sufficient as long as event data and intensity information of the same scene can be obtained. If there is a mismatch in spatial resolution or if it is not possible to map the solid angles observed by pixels 51 in different subsets in a one to one manner, this will be solvable in principle by using interpolation techniques to adapt / align the two different data sets.
[0310] In Fig. 24 a decoding device 200 and its functions are schematically illustrated. The decoding device 200 can be used to decode encoded data that has been generated as described above with respect to Figs. 3 to 7. That is, the decoding device 200 generates a series of frames F2 containing image data with the second number of frames F2 per second. As the encoding device 100, the decoding device 200 may be constituted by any computing device that is capable to carry out the functions described below. For example, the decoding device 200 may be or may be comprised in a computer, processor, microprocessor, CPU, GPU, FPGA, or ASIC. The decoding device 200 may be constituted by hardware or software or a mixture of hardware and software. In general, the structure, design or layout of the decoding device 200 is arbitrary as long as it is capable to carry out the functions described below.
[0311] As illustrated at the top of Fig. 24, the decoding device receives encoded data that contain a first number of frames Fl per second that is smaller than the second number of frames F2 per second, affine transformations AF, and motion vectors MV for each of the received frames Fl that allow, based on the motion vectors MV and the received frames Fl, reconstructing missing frames MF that are temporally located between two consecutive, received frames Fl, and for at least one (but generally for some or all) of the missing frames MF that are to be reconstructed, residual data RD that allow full reconstruction of missing frames MF, where the reconstruction by affine transformations AF and motion vectors MV is insufficient. That is, the decoding device 200 receives encoded data that has been generated as described above with respect to Figs. 3 to 7.73961
[0312] 32
[0313] Then, the decoding device reconstructs missing frames MF based on the received frames Fl and the affine transformations and / or the motion vectors MV, as also schematically illustrated in Fig. 24. As described above the reconstruction process is in principle known, e.g. from AVI or VVC codecs.
[0314] Based on the reconstructed missing frames MF and the residual data RD, the decoding device 200 fills in gaps between the received frames Fl with one or more corrected, reconstructed frames. Together, these frames constitute the frames F2 of the decoded video stream. The decoding device 200 can then output this video stream.
[0315] Besides being used in a codec, the information about affine transformations can be used for various other different purposes, in which precise knowledge of the optical flow between full frames is necessary. For example, in image-based 3D reconstruction algorithms, such as structure from motion, the optical flow is needed to measure distances to captured objects in order to determine their three-dimensional shape. Using event data ED to determine affine transformations can improve the results of such algorithms since the optical flow can be described more precisely.
[0316] Above, various implementations of the basic idea to determine affine transformations based on event data have been described. The basic methods underlying these implementations are summarized below with respect to Fig. 25.
[0317] Fig. 25 shows schematically a process flow of an encoding method for encoding a series of frames containing image data for a video.
[0318] At S 101 a first number of frames per second is received and stored that is smaller than a second number of frames per second of the video that is to be reconstructed from the encoded data.
[0319] At SI 02 event data are received that indicate a temporal stream of events, where each event indicates intensity changes above a predetermined threshold of light received for one of a plurality of pixels that constitute the frames.
[0320] At SI 03 it is determined for the stored frames based on the event data and the stored frames whether affine transformations of pixel values of the frames shall be used for reconstructing missing frames that are temporally located between two consecutive, stored frames.
[0321] If affine transformations shall not be used (No), then at SI 04 motion vectors are determined and stored based on the event data and the stored frames for reconstructing the missing frames, which motion vectors linearly shift pixel values.
[0322] If affine transformations shall be used (Yes), then at SI 05 these affine transformations are determined and stored based on the event data and the stored frames.
[0323] In this manner it is possible to generate high frame rate videos even for video streams containing a highnumber of affine transformations while at the same time ensuring that the quality of the decoded data remains high.
[0324] Fig. 26 is a block diagram illustrating an example of a schematic configuration of a smartphone 900 that may constitute or comprise a sensor device 10 as described above. The smartphone 900 is equipped with a processor 901, memory 902, storage 903, an external connection interface 904, a camera 906, a sensor 907, a microphone 908, an input device 909, a display device 910, a speaker 911, a radio communication interface 912, one or more antenna switches 915, one or more antennas 916, a bus 917, a battery 918, and an auxiliary controller 919.
[0325] The processor 901 may be a CPU or system -on-a-chip (SoC), for example, and controls functions in the application layer and other layers of the smartphone 900. The memory 902 includes RAM and ROM, and stores programs executed by the processor 901 as well as data. The storage 903 may include a storage medium such as semiconductor memory or a hard disk. The external connection interface 904 is an interface for connecting an externally attached device, such as a memory card or Universal Serial Bus (USB) device, to the smartphone 900. The processor may function as control unit 40.
[0326] The camera 906 includes an image sensor as described above. The sensor 907 may include a sensor group such as a positioning sensor, a gyro sensor, a geomagnetic sensor, and an acceleration sensor, for example. The microphone 908 converts audio input into the smartphone 900 into an audio signal. The input device 909 includes devices such as a touch sensorthat detects touches on a screen of the display device 910, a keypad, a keyboard, buttons, or switches, and receives operations or information input from a user. The display device 910 includes a screen such as a liquid crystal display (UCD) or an organic light-emitting diode (OUED) display, and displays an output image of the smartphone 900. The speaker 911 converts an audio signal output from the smartphone 900 into audio.
[0327] The radio communication interface 912 supports a cellular communication scheme such as UTE or LTE-Advanced, and executes radio communication. Typically, the radio communication interface 912 may include a BB processor 913, an RF circuit 914, and the like. The BB processor 913 may conduct processes such as encoding / decoding, modulation / demodulation, and multiplexing / demultiplexing, for example, and executes various signal processing for radio communication. Meanwhile, the RF circuit 914 may include components such as a mixer, a filter, and an amp, and transmits or receives a radio signal via an antenna 916. The radio communication interface 912 may also be a one-chip module integrating the BB processor 913 and the RF circuit 914. The radio communication interface 912 may also include multiple BB processors 913 and multiple RF circuits 91. Note that although Fig. 26 illustrates an example of the radio communication interface 912 including multiple BB processors 913 and multiple RF circuits 914, the radio communication interface 912 may also include a single BB processor 913 or a single RF circuit 914.
[0328] Furthermore, in addition to a cellular communication scheme, the radio communication interface 912 may also support other types of radio communication schemes such as a short-range wireless communication scheme, a near field wireless communication scheme, or a wireless local area network (LAN) scheme. In34
[0329] this case, a BB processor 913 and an RF circuit 914 may be included for each radio communication scheme.
[0330] Each antenna switch 915 switches the destination of an antenna 916 among multiple circuits included in the radio communication interface 912 (for example, circuits for different radio communication schemes).
[0331] Each antenna 916 includes a single or multiple antenna elements (for example, multiple antenna elements constituting a MIMO antenna), and is used by the radio communication interface 912 to transmit and receive radio signals. The smartphone 900 may also include multiple antennas 916 as illustrated in Fig. 26. Note that although Fig. 26 illustrates an example of the smartphone 900 including multiple antennas 916, the smartphone 900 may also include a single antenna 916.
[0332] Furthermore, the smartphone 900 may also be equipped with an antenna 916 for each radio communication scheme. In this case, the antenna switch 915 may be omitted from the configuration of the smartphone 900.
[0333] The bus 917 interconnects the processor 901, the memory 902, the storage 903, the external connection interface 904, the camera 906, the sensor 907, the microphone 908, the input device 909, the display device 910, the speaker 911, the radio communication interface 912, and the auxiliary controller 919. The battery 918 supplies electric power to the respective blocks of the smartphone 900 illustrated in Fig. 26 via power supply lines partially illustrated with dashed lines in the drawing. The auxiliary controller 919 causes minimal functions of the smartphone 900 to operate while in a sleep mode, for example.
[0334] Fig. 27 is a block diagram depicting an example of schematic configuration of a vehicle control system as an example of a mobile body control system to which the technology according to an embodiment of the present disclosure can be applied.
[0335] The vehicle control system 12000 includes a plurality of electronic control units connected to each other via a communication network 12001. In the example depicted in Fig. 27, the vehicle control system 12000 includes a driving system control unit 12010, a body system control unit 12020, an outside -vehicle information detecting unit 12030, an in-vehicle information detecting unit 12040, and an integrated control unit 12050. In addition, a microcomputer 12051, a sound / image output section 12052, and a vehicle-mounted network interface (I / F) 12053 are illustrated as a functional configuration of the integrated control unit 12050.
[0336] The driving system control unit 12010 controls the operation of devices related to the driving system of the vehicle in accordance with various kinds of programs. For example, the driving system control unit 12010 functions as a control device for a driving force generating device for generating the driving force of the vehicle, such as an internal combustion engine, a driving motor, or the like, a driving force transmitting mechanism for transmitting the driving force to wheels, a steering mechanism for adjusting the steering angle of the vehicle, a braking device for generating the braking force of the vehicle, and the like.35
[0337] The body system control unit 12020 controls the operation of various kinds of devices provided to a vehicle body in accordance with various kinds of programs. For example, the body system control unit 12020 functions as a control device for a keyless entry system, a smart key system, a power window device, or various kinds of lamps such as a headlamp, a backup lamp, a brake lamp, a turn signal, a fog lamp, or the like. In this case, radio waves transmitted from a mobile device as an alternative to a key or signals of various kinds of switches can be input to the body system control unit 12020. The body system control unit 12020 receives these input radio waves or signals, and controls a door lock device, the power window device, the lamps, or the like of the vehicle.
[0338] The outside-vehicle information detecting unit 12030 detects information about the outside of the vehicle including the vehicle control system 12000. For example, the outside-vehicle information detecting unit 12030 is connected with an imaging section 12031. The outside-vehicle information detecting unit 12030 makes the imaging section 12031 image an image of the outside of the vehicle, and receives the imaged image. On the basis of the received image, the outside-vehicle information detecting unit 12030 may perform processing of detecting an object such as a human, a vehicle, an obstacle, a sign, a character on a road surface, or the like, or processing of detecting a distance thereto.
[0339] The imaging section 12031 is an optical sensor that receives light, and which outputs an electric signal corresponding to a received light amount of the light. The imaging section 12031 can output the electric signal as an image, or can output the electric signal as information about a measured distance. In addition, the light received by the imaging section 12031 may be visible light, or may be invisible light such as infrared rays or the like.
[0340] The in-vehicle information detecting unit 12040 detects information about the inside of the vehicle. The in-vehicle information detecting unit 12040 is, for example, connected with a driver state detecting section 12041 that detects the state of a driver. The driver state detecting section 12041, for example, includes a camera that images the driver. On the basis of detection information input from the driver state detecting section 12041, the in-vehicle information detecting unit 12040 may calculate a degree of fatigue of the driver or a degree of concentration of the driver, or may determine whether the driver is dozing.
[0341] The microcomputer 12051 can calculate a control target value for the driving force generating device, the steering mechanism, or the braking device on the basis of the information about the inside or outside of the vehicle which information is obtained by the outside-vehicle information detecting unit 12030 or the in-vehicle information detecting unit 12040, and output a control command to the driving system control unit 12010. For example, the microcomputer 12051 can perform cooperative control intended to implement functions of an advanced driver assistance system (ADAS) which functions include collision avoidance or shock mitigation for the vehicle, following driving based on a following distance, vehicle speed maintaining driving, a warning of collision of the vehicle, a warning of deviation of the vehicle from a lane, or the like.
[0342] In addition, the microcomputer 12051 can perform cooperative control intended for automatic driving,73961
[0343] 36
[0344] which makes the vehicle to travel autonomously without depending on the operation of the driver, or the like, by controlling the driving force generating device, the steering mechanism, the braking device, or the like on the basis of the information about the outside or inside of the vehicle which information is obtained by the outside-vehicle information detecting unit 12030 or the in-vehicle information detecting unit 12040.
[0345] In addition, the microcomputer 12051 can output a control command to the body system control unit 12020 on the basis of the information about the outside of the vehicle which information is obtained by the outside-vehicle information detecting unit 12030. For example, the microcomputer 12051 can perform cooperative control intended to prevent a glare by controlling the headlamp so as to change from a high beam to a low beam, for example, in accordance with the position of a preceding vehicle or an oncoming vehicle detected by the outside-vehicle information detecting unit 12030.
[0346] The sound / image output section 12052 transmits an output signal of at least one of a sound and an image to an output device capable of visually or auditorily notifying information to an occupant of the vehicle or the outside of the vehicle. In the example ofFig. 27, an audio speaker 12061, a display section 12062, and an instrument panel 12063 are illustrated as the output device. The display section 12062 may, for example, include at least one of an on-board display and a head-up display.
[0347] Fig. 28 is a diagram depicting an example of the installation position of the imaging section 12031.
[0348] In Fig. 28, the imaging section 12031 includes imaging sections 12101, 12102, 12103, 12104, and 12105.
[0349] The imaging sections 12101, 12102, 12103, 12104, and 12105 are, for example, disposed at positions on a front nose, sideview mirrors, a rear bumper, and a back door of the vehicle 12100 as well as a position on an upper portion of a windshield within the interior of the vehicle. The imaging section 12101 provided to the front nose and the imaging section 12105 provided to the upper portion of the windshield within the interior of the vehicle obtain mainly an image of the front of the vehicle 12100. The imaging sections 12102 and 12103 provided to the sideview mirrors obtain mainly an image of the sides of the vehicle 12100. The imaging section 12104 provided to the rear bumper or the back door obtains mainly an image of the rear of the vehicle 12100. The imaging section 12105 provided to the upper portion of the windshield within the interior of the vehicle is used mainly to detect a preceding vehicle, a pedestrian, an obstacle, a signal, a traffic sign, a lane, or the like.
[0350] Incidentally, Fig. 28 depicts an example of photographing ranges of the imaging sections 12101 to 12104. An imaging range 12111 represents the imaging range of the imaging section 12101 provided to the front nose. Imaging ranges 12112 and 12113 respectively represent the imaging ranges of the imaging sections 12102 and 12103 provided to the sideview mirrors. An imaging range 12114 represents the imaging range of the imaging section 12104 provided to the rear bumper or the back door. A bird’s-eye image of the vehicle 12100 as viewed from above is obtained by superimposing image data imaged by the imaging sections 12101 to 12104, for example.73961
[0351] 37
[0352] At least one of the imaging sections 12101 to 12104 may have a function of obtaining distance information. For example, at least one of the imaging sections 12101 to 12104 may be a stereo camera constituted of a plurality of imaging elements, or may be an imaging element having pixels for phase difference detection.
[0353] For example, the microcomputer 12051 can determine a distance to each three-dimensional object within the imaging ranges 12111 to 12114 and a temporal change in the distance (relative speed with respect to the vehicle 12100) on the basis of the distance information obtained from the imaging sections 12101 to 12104, and thereby extract, as a preceding vehicle, a nearest three-dimensional object in particular that is present on a traveling path of the vehicle 12100 and which travels in substantially the same direction as the vehicle 12100 at a predetermined speed (for example, equal to or more than 0 km / hour). Further, the microcomputer 12051 can set a following distance to be maintained in front of a preceding vehicle in advance, and perform automatic brake control (including following stop control), automatic acceleration control (including following start control), or the like. It is thus possible to perform cooperative control intended for automatic driving that makes the vehicle travel autonomously without depending on the operation of the driver or the like.
[0354] For example, the microcomputer 12051 can classify three-dimensional object data on three-dimensional objects into three-dimensional object data of a two-wheeled vehicle, a standard-sized vehicle, a largesized vehicle, a pedestrian, a utility pole, and other three-dimensional objects on the basis of the distance information obtained from the imaging sections 12101 to 12104, extract the classified three-dimensional object data, and use the extracted three-dimensional object data for automatic avoidance of an obstacle. For example, the microcomputer 12051 identifies obstacles around the vehicle 12100 as obstacles that the driver of the vehicle 12100 can recognize visually and obstacles that are difficult for the driver of the vehicle 12100 to recognize visually. Then, the microcomputer 12051 determines a collision risk indicating a risk of collision with each obstacle. In a situation in which the collision risk is equal to or higher than a set value and there is thus a possibility of collision, the microcomputer 12051 outputs a warning to the driver via the audio speaker 12061 or the display section 12062, and performs forced deceleration or avoidance steering via the driving system control unit 12010. The microcomputer 12051 can thereby assist in driving to avoid collision.
[0355] At least one of the imaging sections 12101 to 12104 may be an infrared camera that detects infrared rays. The microcomputer 12051 can, for example, recognize a pedestrian by determining whether or not there is a pedestrian in imaged images of the imaging sections 12101 to 12104. Such recognition of a pedestrian is, for example, performed by a procedure of extracting characteristic points in the imaged images of the imaging sections 12101 to 12104 as infrared cameras and a procedure of determining whether or not it is the pedestrian by performing pattern matching processing on a series of characteristic points representing the contour of the object. When the microcomputer 12051 determines that there is a pedestrian in the imaged images of the imaging sections 12101 to 12104, and thus recognizes the pedestrian, the sound / image output section 12052 controls the display section 12062 so that a square contour line for emphasis is displayed so as to be superimposed on the recognized pedestrian. The sound / image output section 12052 may also control the display section 12062 so that an icon or the like73961
[0356] 38
[0357] representing the pedestrian is displayed at a desired position.
[0358] An example of the vehicle control system to which the technology according to the present disclosure is applicable has been described above. The technology according to the present disclosure is applicable to the imaging section 12031 among the above-mentioned configurations. Specifically, the sensor device 10 is applicable to the imaging section 12031. The imaging section 12031 to which the technology according to the present disclosure has been applied flexibly acquires event data and performs data processing on the event data, thereby being capable of providing appropriate driving assistance.
[0359] Further possible implementations of the sensor device 10 are mobile devices 3000 such as cell phones, tablets, smart watches and the like as shown in Fig. 29A or head-mounted displays 4000 as shown in Fig.
[0360] 29B. Further, the sensor device 10 is useable in augmented and / or virtual reality applications / cameras or in surveillance systems like 360° cameras.
[0361] Note that, the embodiments of the present technology are not limited to the above-mentioned embodiment, and various modifications can be made without departing from the gist of the present technology.
[0362] Further, the effects described herein are only exemplary and not limited, and other effects may be provided.
[0363] Note that, the present technology can also take the following configurations.
[0364] [1] An encoding device (100) for encoding a series of frames (F2) containing image data for a video, the encoding device (100) being configured to
[0365] receive and store a first number of frames (Fl) per second that is smaller than a second number of frames (F2) per second of the video;
[0366] receive event data (ED) indicating a temporal stream of events, where each event indicates intensity changes above a predetermined threshold of light received for one of a plurality of pixels (51) that constitute the frames;
[0367] determine for the stored frames (Fl) based on the event data (ED) and the stored frames (Fl) whether affine transformations (AF) of pixel values of the frames (Fl) shall be used for reconstructing missing frames (MF) that are temporally located between two consecutive, stored frames (Fl);
[0368] if affine transformations shall not be used, determine and store based on the event data (ED) and the stored frames (Fl) motion vectors (MV) for reconstructing the missing frames (MF), which motion vectors (MV) linearly shift pixel values; and
[0369] if affine transformations (AF) shall be used, determine and store based on the event data (ED) and the stored frames (Fl) the affine transformations (AF).
[0370] [2] The encoding device (100) according to [1], wherein
[0371] the stored frames (Fl) are divided into blocks (Bl, B2);
[0372] whether affine transformations (AF) shall be used for reconstructing missing frames is determined for each block (Bl, B2); and39
[0373] the respective motion vector (MV) or affine transformation (AF) for each block are determined and stored.
[0374] [3] The encoding device (100) according to [2], wherein
[0375] the determined affine transformations (AF) are either general affine transformations, affine transformations defined by two control points within one block (Bl) and two respective vectors, or affine transformations defined by three control points within one block (Bl) and three respective vectors.
[0376] [4] The encoding device (100) according to any one of [1] to [3], wherein
[0377] the stored motion vectors (MV) and / or affine transformations (AF) contain forward transformations that indicate a change in image data from a stored frame (Fl) to a future missing frame (MF), and backward transformations that indicate a change in image data from a past missing frame (MF) to a stored frame (Fl).
[0378] [5] The encoding device (100) according to any one of [1] to [4], wherein the encoding device (100) is further configured to
[0379] reconstruct missing frames (MF) based on the stored frames and the determined affine transformations (AF) and / or motion vectors (MV);
[0380] reconstruct missing frames (MF) based on the stored frames (Fl) by using an artificial intelligence model that has been trained for this reconstruction;
[0381] compare missing frames (MF) for the same temporal position in the video stream that have been reconstructed based on the affine transformations (AF) and / or the motion vectors (MV) and based on the artificial intelligence model to determine residual data (RD); and
[0382] store the residual data (RD) together with the stored frames (Fl) and the affine transformations (AF) and / or motion vectors (MV).
[0383] [6] The encoding device (100) according to any one of [1] to [5], wherein
[0384] the determination of motion vectors (MV) and affine transformations (AF) is performed by respective artificial intelligence models that have been trained for this determination.
[0385] [7] A sensor device (10) comprising:
[0386] the encoding device (100) according to any one of [1] to [6];
[0387] a plurality of pixels (51) each configured to receive light from a scene and to perform photoelectric conversion to generate an electrical signal;
[0388] event detection circuitry (20) that is configured to detect as event data intensity changes above a predetermined threshold of the light received by each of a first subset (S 1) of the pixels (51);
[0389] pixel signal generating circuitry (30) that is configured to generate a pixel signal indicating intensity values of the received light for each pixel (51) of a second subset (S2) of the pixels (51); and a control unit (40) that is configured to provide the pixel signals as the first number of frames (Fl) per second and the event data (ED) to the encoding device (100).
[0390] [8] The sensor device (10) according to [7], wherein40
[0391] the pixels (51) are arranged in a two-dimensional pixel array; and
[0392] at least a part of the pixels (51) belongs to both, the first subset (SI) of pixels (51) and the second subset (S2) of pixels (51).
[0393] [9] The sensor device (10) according to [7], wherein
[0394] the pixels (51) are arranged in a two-dimensional pixel array; and
[0395] pixels in the first subset (SI) of pixels (51) are different from pixels (51) in the second subset (S2) of pixels (51).
[0396]
[0010] The sensor device (10) according to [7], wherein
[0397] the pixels (51) in the first subset (SI) of pixels (51) are arranged in a first two-dimensional pixel array; and
[0398] the pixels (51) in the second subset (S2) of pixels (51) are arranged in a different, second two-dimensional pixel array.
[0399]
[0011] An encoding method for encoding a series of frames containing image data for a video, the encoding method comprising:
[0400] receiving and storing a first number of frames per second that is smaller than a second number of frames per second of the video;
[0401] receiving event data indicating a temporal stream of events, where each event indicates intensity changes above a predetermined threshold of light received for one of a plurality of pixels that constitute the frames;
[0402] determining for the stored frames based on the event data and the stored frames whether affine transformations of pixel values of the frames shall be used for reconstructing missing frames that are temporally located between two consecutive, stored frames;
[0403] if affine transformations shall not be used, determining and storing based on the event data and the stored frames motion vectors for reconstructing the missing frames, which motion vectors linearly shift pixel values; and
[0404] if affine transformations shall be used, determining and storing based on the event data and the stored frames the affine transformations.
Claims
41CLAIMS1. An encoding device for encoding a series of frames containing image data for a video, the encoding device being configured toreceive and store a first number of frames per second that is smaller than a second number of frames per second of the video;receive event data indicating a temporal stream of events, where each event indicates intensity changes above a predetermined threshold of light received for one of a plurality of pixels that constitute the frames;determine for the stored frames based on the event data and the stored frames whether affine transformations of pixel values of the frames shall be used for reconstructing missing frames that are temporally located between two consecutive, stored frames;if affine transformations shall not be used, determine and store based on the event data and the stored frames motion vectors for reconstructing the missing frames, which motion vectors linearly shift pixel values; andif affine transformations shall be used, determine and store based on the event data and the stored frames the affine transformations.
2. The encoding device according to claim 1, whereinthe stored frames are divided into blocks;whether affine transformations shall be used for reconstructing missing frames is determined for each block; andthe respective motion vector or affine transformation for each block are determined and stored.
3. The encoding device according to claim 2, whereinthe determined affine transformations are either general affine transformations, affine transformations defined by two control points within one block and two respective vectors, or affine transformations defined by three control points within one block and three respective vectors.
4. The encoding device according to claim 1, whereinthe stored motion vectors and / or affine transformations contain forward transformations that indicate a change in image data from a stored frame to a future missing frame, and backward transformations that indicate a change in image data from a past missing frame to a stored frame.
5. The encoding device according to claim 1, wherein the encoding device is further configured toreconstruct missing frames based on the stored frames and the determined affine transformations and / or motion vectors;reconstruct missing frames based on the stored frames by using an artificial intelligence model that has been trained for this reconstruction;compare missing frames for the same temporal position in the video stream that have been42reconstructed based on the affine transformation and / or motion vectors and based on the artificial intelligence model to determine residual data; andstore the residual data together with the stored frames and the affine transformations and / or the motion vectors.
6. The encoding device according to claim 1, whereinthe determination of motion vectors and affine transformations is performed by respective artificial intelligence models that have been trained for this determination.
7. A sensor device comprising:the encoding device according to claim 1 ;a plurality of pixels each configured to receive light from a scene and to perform photoelectric conversion to generate an electrical signal;event detection circuitry that is configured to detect as event data intensity changes above a predetermined threshold of the light received by each of a first subset of the pixels;pixel signal generating circuitry that is configured to generate a pixel signal indicating intensity values of the received light for each pixel of a second subset of the pixels; anda control unit that is configured to provide the pixel signals as the first number of frames per second and the event data to the encoding device.
8. The sensor device according to claim 7, whereinthe pixels are arranged in a two-dimensional pixel array; andat least a part of the pixels belongs to both, the first subset of pixels and the second subset of pixels.
9. The sensor device according to claim 7, whereinthe pixels are arranged in a two-dimensional pixel array; andpixels in the first subset of pixels are different from pixels in the second subset of pixels.
10. The sensor device according to claim 7, whereinthe pixels in the first subset of pixels are arranged in a first two-dimensional pixel array; and the pixels in the second subset of pixels are arranged in a different, second two-dimensional pixel array.
11. An encoding method for encoding a series of frames containing image data for a video, the encoding method comprising:receiving and storing a first number of frames per second that is smaller than a second number of frames per second of the video;receiving event data indicating a temporal stream of events, where each event indicates intensity changes above a predetermined threshold of light received for one of a plurality of pixels that constitute the frames;determining for the stored frames based on the event data and the stored frames whether43affine transformations of pixel values of the frames shall be used for reconstructing missing frames that are temporally located between two consecutive, stored frames;if affine transformations shall not be used, determining and storing based on the event data and the stored frames motion vectors for reconstructing the missing frames, which motion vectors linearly shift pixel values; andif affine transformations shall be used, determining and storing based on the event data and the stored frames the affine transformations.