Encoding device, decoding device, encoding method and decoding method
Patent Information
- Application Number
- PCT/EP2026/056697
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-13
- Filing Date
- 2026-03-11
- Publication Date
- 2026-09-17
Smart Images

Figure EP2026056697_17092026_PF_FP_ABST
Abstract
Description
[0001] 1
[0002] ENCODING DEVICE, DECODING DEVICE, ENCODING METHOD AND DECODING METHOD
[0003] FIELD OF THE INVENTION
[0004] The present technology relates to encoding and decoding devices and methods for encoding and decoding, in particular for zero-residual mode encoding and decoding.
[0005] BACKGROUND
[0006] Usual digital video cameras that are e.g. used as stand-alone devices or integrated into other electronic devices such as mobile phones produce large amounts of data, which are typically reduced by encoding the generated image data / image frames. However, even after encoding high data rates may be present. Further, encoding power consumption and computational costs are also an issue for large amounts of data.
[0007] It is therefore desirable to improve the encoding and decoding devices such as to allow a further reduction of data.
[0008] SUMMARY OF INVENTION
[0009] To this end, an encoding device is provided for encoding a series of frames containing image data for a video stream having a first number of frames per second. The encoding device is configured to store a second number of frames per second that is smaller than the first number of frames per second, determine and store motion vectors for each of the stored frames that allow, based on the motion vectors and the stored frames, reconstructing missing frames that are temporally located between two consecutive, stored frames, and store, for at least one of the missing frames that are to be reconstructed, error data that indicate areas in a reconstruction of said at least one missing frame that need correction.
[0010] Further, a decoding device is provided for decoding a series of frames containing image data for a video stream having a first number of frames per second. The decoding device is configured to receive a second number of frames per second that is smaller than the first number of frames per second, motion vectors for each of the received frames that allow, based on the motion vectors and the received frames, reconstructing missing frames that are temporally located between two consecutive, received frames, and for at least one of the missing frames that are to be reconstructed, error data that indicate areas in a reconstruction of said at least one missing frame that need correction, reconstruct missing frames based on the received frames and the motion vectors, and correct in the at least one reconstructed missing frame areas indicated by the error data.
[0011] Further, an encoding method is provided for encoding a series of frames containing image data for a video stream having a first number of frames per second. The encoding method comprises: storing a second number of frames per second that is smaller than the first number of frames per second; determining and storing motion vectors for each of the stored frames that allow, based on the motion vectors and the stored frames, reconstructing missing frames that are temporally located between two73960
[0012] 2
[0013] consecutive, stored frames; and storing, for at least one of the missing frames that are to be reconstructed, error data that indicate areas in a reconstruction of said at least one missing frame that need correction.
[0014] Further, a decoding method is provided for decoding a series of frames containing image data for a video stream having a first number of frames per second. The decoding method comprises: receiving a second number of frames per second that is smaller than the first number of frames per second, motion vectors for each of the received frames that allow, based on the motion vectors and the received frames, reconstructing missing frames that are temporally located between two consecutive, received frames, and, for at least one of the missing frames that are to be reconstructed, error data that indicate areas in a reconstruction of said at least one missing frame that need correction; reconstructing missing frames based on the received frames and the motion vectors; and correcting in the at least one reconstructed missing frame areas indicated by the error data.
[0015] Thus, encoding is performed in zero-residual mode in which no residual that indicates differences between a reconstructed (or predicted) frame and the original frame (that has been omitted to reduce data) is included in the stored, encoded data. This reduces the amount of memory necessary to store the encoded data. To avoid a deterioration of the image data after decoding, error data are stored, too. These error data allow identifying regions in reconstructed frames that need correction in order to resemble the originally captured scene (i.e. the omitted or missing frame). In this manner, it is possible to reduce the size of encoded data, while at the same time ensuring that the quality of the decoded image data is sufficiently high.
[0016] BRIEF DESCRIPTION OF DRAWINGS
[0017] Fig. 1 is a schematic diagram of an encoding device.
[0018] Fig. 2 is a schematic illustration of error data used in encoding.
[0019] Fig. 3 is a schematic illustration of the generation of error data.
[0020] Fig. 4 is another schematic illustration of the generation of error data.
[0021] Fig. 5 is a schematic diagram of a sensor device.
[0022] Fig. 6 is a schematic block diagram of a sensor section.
[0023] Fig. 7 is a schematic block diagram of a pixel array section.
[0024] Fig. 8 is a schematic circuit diagram of a pixel block.
[0025] Fig. 9 is a schematic block diagram illustrating of an event detecting section.3
[0026] Fig. 10 is a schematic circuit diagram of a current-voltage converting section.
[0027] Fig. 11 is a schematic circuit diagram of a subtraction section and a quantization section.
[0028] Fig. 12 is a schematic timing chart of an example operation of the sensor section.
[0029] Fig. 13 is a schematic diagram of a frame data generation method based on event data.
[0030] Fig. 14 is a schematic block diagram of another quantization section.
[0031] Fig. 15 is a schematic diagram of another event detecting section.
[0032] Fig. 16 is a schematic block diagram of another pixel array section.
[0033] Fig. 17 is a schematic circuit diagram of another pixel block.
[0034] Fig. 18 is a schematic block diagram of a scan-type imaging device.
[0035] Fig. 19 is a schematic block diagram of a sensor device comprising an encoding device.
[0036] Fig. 20 shows different schematic block diagrams of a sensor device.
[0037] Fig. 21 is a schematic illustration of a decoding device.
[0038] Fig. 22 is a schematic illustration of the usage of error data for decoding.
[0039] Fig. 23 illustrates schematically a process flow of an encoding method.
[0040] Fig. 24 illustrates schematically a process flow of a decoding method.
[0041] Fig. 25 is a schematic block diagram of a smart phone.
[0042] Fig. 26 is a schematic block diagram of a vehicle control system.
[0043] Fig. 27 is a diagram of assistance in explaining an example of installation positions of an outside-vehicle information detecting section and an imaging section.
[0044] Figs. 28A and 28B are schematic illustrations of a mobile device and a head mounted display comprising a sensor device
[0045] DETAILED DESCRIPTION4
[0046] Fig. 1 shows a schematic illustration of an encoding device 100 and its functions. The encoding device 100 is used to encoding a series of frames F3 that contain image data for a video stream. This means, that the encoded data output by the encoding device 100 can be used to generate the series of frames F3 by decoding the output of the encoding device 100. Thus, in this context “encoding of frames” refer to the frames of the video stream that can be reproduced from the encoded data. This video stream has a certain frame rate, i.e. a first number of frames per second, that may in general be different from the frame rate of a series of frame Fl that is input into the encoding device 100 and that is typically different from the frame rate of frames F2 that are stored (e.g. as keyframes) in the encoded data.
[0047] The encoding device 100 may be constituted by any computing device that is capable to carry out the functions described below. For example, the encoding device 100 may be or may be comprised in a computer, processor, microprocessor, CPU, GPU, FPGA, or ASIC. The encoding device 100 may be constituted by hardware or software or a mixture of hardware and software. In general, the structure, design or layout of the encoding device 100 is arbitrary as long as it is capable to carry out the functions described below.
[0048] As schematically illustrated in Fig. 1, the encoding device 100 receives a video stream containing a certain number of frames Fl. These frames Fl are transformed by the encoding device to frames F2 that are stored as part of the encoded data, e.g. in a first storage area 110. These frames come in a second number of frames F2 per second that is smaller than the first number of frames per second, i.e. the encoded data contain less frames F2 per second than contained in the video stream that is to be reproduced based on the encoded data. As described later in detail, the stored frames F2 may be all the frames Fl that are received by the encoding device 100. In this case, the received video stream has less frames Fl than the video stream obtainable after decoding the encoded data. Or the stored frames F2 are obtained from the received frames Fl by omitting frames or by combining frames. Then, the encoded data are typically used to reconstruct the original received frames Fl.
[0049] The encoding device determines and stores motion vectors MV for each of the stored frames F2. The motion vectors MV may be stored in a second storage area 120 that may be part of the same or a different memory that contains the first storage area 110. The motion vectors MV allow, based on the motion vectors MV and the stored frames F2, reconstructing missing frames MF that are temporally located between two consecutive, stored frames F2. The motion vectors MV are generated from the received frames Fl. The missing frames MF fill in gaps between the stored frames such as to obtain the higher frame rate of the frames F3 that can be reconstructed from the encoded data. In particular, as will be described below, the stored frames F2 and the motion vectors MV are used in a decoding device 200 to reconstruct frames that are located between temporally consecutive stored frames F2. This allows omitting frames of the received frames Fl, thereby reducing the data amount of the encoded data compared to the data amount of the received video stream and / or the reconstructed video stream. Further, as will be described later, it is also possible to provide interpolating missing frames MF that are not included in the received frames F 1.
[0050] The motion vectors MV indicate changes in the image data constituting the respective frames. This means73960
[0051] 5
[0052] that for each region of a frame, a motion vector MV can be defined that indicates a seeming movement of the contents of said region within the frame. Then, by using the original frame F2 and its motion vectors MV, it is possible to deduce the contents of a missing frame MF. Here, as schematically illustrated in Fig.
[0053] 1, forward motion vectors may be used that indicate a change in image data from a stored frame F2 to a future missing frame MF. Further, also backward motion vectors may be used that indicate a change in image data from a past missing frame MF to a stored frame F2. In addition, although the present description focuses for ease of description on the case that there is one missing frame MF between two temporally consecutive stored frames F2, it is also possible to have a plurality of missing frames MF between two stored frames F2. In this case, motion vectors MV can also indicate changes between two temporally consecutive missing frames MF. The number of missing frames MF between two stored frames F2 can vary and may depend e.g. on the change rate of the image data constituting the frames. If there are only small changes, it will be possible to store less frames F2 and interpolate more missing frames MF without generating too many errors. If there are large or rapid changes, only few, one, or even no missing frame MF might be present. It should be noted that the determination of motion vectors MV and the reconstruction of (one or several) missing frames MF between stored frames F2 is in principle known e.g. from the MPEG and / or HEVC standards such as H.264 or H.265. Thus, further details can be omitted here.
[0054] In addition to the stored frames F2 and the motion vectors MV the encoding device 100 stores, e.g. in a third storage area 130 that may be the same or different from the memories that contains the first and second storage areas 110, 120, for at least one (but generally for some or all) of the missing frames MF that are to be reconstructed, error data FL. The error data FL indicate areas in a reconstruction of said missing frames MF that need correction. That is, for areas in frames MF that need to be reconstructed in order to reproduce the decoded frames F3, in which it is to be expected that the reconstructed image data do not correctly resemble a truthful interpolation between the stored frames F2, error data FL are added to the encoded data. Based on these error data a decoding device 200 can identify the incorrect areas and provide measures for error correction, as will be described in more detail below.
[0055] Thus, the encoded data consist only of the stored frames F2, the motion vectors MV, and the error data FL. In particular, the encoding device 100 operates in a zero- or non-residual mode and does not store residuals between reconstructed missing frames MF and original frames Fl. Such residuals are typically obtained when from an original series of frames Fl frames are to be dropped for the encoding. It is then possible to compare the estimation / prediction / calculation of the missing frames MF, which is based on the stored, i.e. maintained, frames F2 and the motion vectors MV, and the original frames Fl to be dropped. Deviations between these two are then stored as so-called residuals. While the residuals allow a high-quality reproduction of the original frames Fl, they also have a comparably large data size.
[0056] Therefore, as described above, to reduce the data amount of the encoded data, the encoding device 100 does not store residuals. It only stores error data FL that indicate where deviations have become too large, e.g. larger than a predetermined threshold. For example, pixel values in residuals can be analyzed to determine whether the deviations are too large, e.g. by summing the squares of the pixel values in a certain area and comparing the sum (or its square root) to a threshold. Thus, instead of storing all the pixel6
[0057] values for a certain area, only a single error indication needs to be stored per area. As will be discussed below, this will be sufficient during a decoding process to generate the decoded frames F3 with high quality, i.e. with a high resemblance of originally received frames Fl or the originally captured scene.
[0058] The process of error data generation is further schematically illustrated in Fig. 2. For example, each frame can be divided into blocks B that comprise a certain number of pixels of the frame. As schematically shown in Fig. 2, different block sizes can be used in different areas of the frame in order to take into account spatial or temporal (i.e. motion) details of the shown scene. Examples of such blocks B are e.g. coding tree units, coding tree blocks, coding units, prediction units and the like of the HEVC standard. In general, any area of N x M pixels may be used as block B.
[0059] Then, also the reconstruction of each missing frame MF can be divided into blocks B corresponding to the blocks B of the stored frame F2 from which the missing frame MF was reconstructed. The error data FL may then be flags that indicate for each block B whether the respective block B needs correction or not.
[0060] In the example of Fig. 2 two temporally consecutive stored frames F2 are shown at the top of Fig. 2. The frames F2 are divided into blocks B, e.g. according to the interesting features shown in the frames F2. The frames show a bicycle and a cloud. The bicycle moves to the right and the cloud moves basically upwards. From these stored frames F2 it is possible to deduce motion vectors MV that show the rightward movement of the bicycle and the upward movement of the cloud. Further, it is possible to reconstruct a missing frame MF from the motion vectors MV and the stored frames F2 that is temporally located between the two shown stored frames F2. In this simple example, the missing frame MF shows the bicycle and the cloud in the middle of their respective positions in the two stored frames F2.
[0061] Then, error data FL for each block can be generated. In Fig. 2, each of the back blocks B in the lower left image is classified as erroneous. This can be translated into a flag representation, in which for each block B a “1” indicates an error, while “0” indicates no necessity for correction.
[0062] Here, the generation of error data FL may be based on additional information. For example, the error data FL may be calculated based on a residual as described above. Also, additional motion information might be used such as event as will be described in more detail below. Further, the encoding device 100 may also use a database or an artificial intelligence model to determine, based on the stored frames F2 used for reconstructing the missing frame MF, whether the objects shown in the missing frame MF look correct. For example, if in the example of Fig. 2 the reconstructed missing frame F2 would show a car instead of a bicycle, or a bicycle with distorted shape or wrong color this could be recognized by the encoding device 100, e.g. by a respectively trained artificial intelligence model. In this sense, erroneous blocks or areas in a frame could be identified based on a “consistency check” with the temporally adjacent frames without reference to a residual and without reference to additional sensor data.
[0063] In the following, two approaches to generate the error data FL will be described in more detail. First, with respect to Fig. 3 the generation of error data FL based on a residual will be described. Second, with7
[0064] respect to Fig. 4 the generation of error data FL based on event data will be described.
[0065] In the example of Fig. 3 the encoding device 100 is configured to receive a video stream of frames Fl that has the same frame rate as the video stream that is to be reproduced by decoding the encoded data, i.e. it contains the first number of frames per second. From this video stream the encoding device 100 designates frames D for dropping and designating the remaining frames F2 for storing. This is illustrated at the top of Fig. 3, where hatched frames D are to be dropped. Again, although one frame D to be dropped between two frames F2 to be stored is indicated in Fig.3, there may also be more than one frame D that is to be dropped. The frames to be dropped constitute missing frames MF as discussed above.
[0066] As schematically indicated in Fig. 3 by the arrows A, the encoding device 100 determines the motion vectors MV from the received frames Fl, in particular by comparing frames F2 for storing with frames to be dropped (or if more than one frame D is dropped between two frames F2 to be stored, by comparing frames D to be dropped). Thus, since the frame rate of the received frames Fl is the same as the frame rate of the frames F3 generated by decoding the encoded data, motion vectors MV can be determined based on in principle known techniques (as e.g. indicated in MPEG or HEVC standards) from the received frames Fl.
[0067] Then, the encoding device 100 reconstructs missing frames MF as would be done by a decoding device 200 based on the frames F2 that are to be stored and based on the determined motion vectors MV. Again, this is in principle known. This step is shown as the second step in Fig. 3, where the missing frames MF are indicated by an oblique hatching.
[0068] As illustrated in the third step of Fig. 3, the encoding device 100 compares then the reconstructed missing frames MF (oblique hatching) with the respective frames D to be dropped (hatched with points) to determine the error data FL. In particular, each missing frame MF has been generated to replace a respective frame D after decoding. Thus, a residual can be calculated by comparing the respective pairs of missing frame MF and dropped frame D, e.g. by subtracting the two frames. Blocks in the residual can be compared to a threshold, e.g. by summing squares of pixel values of pixels in the block. If the sum (or its square root) is above the threshold, the block is marked as erroneous in the error data FL, e.g. by a flag.
[0069] Then, as illustrated at the bottom of Fig. 3, the frames D designated for dropping are dropped, i.e. not stored, and only the remaining frames F2 are stored. Further, also the comparison between the reconstructed missing frames MF and the dropped frames D are dropped, i.e. no residuals are stored. Only the error data FL are stored.
[0070] Accordingly, the data amount of the originally received frames Fl is reduced by the data amount of all dropped frames D. In addition, in comparison to conventional encoding schemes that use residuals, only error data FL are stored, such as e.g. one flag for each block in a reconstructed missing frame MF. This reduces the size of the encoded data further in comparison to the conventional encoding schemes. However, due to the presence of the error data FL, a decoding device 200 will be informed about erroneous areas in the reconstructed frames and can apply measures to correct these errors. Accordingly,8
[0071] no (significant) deterioration of the image quality due to encoding and decoding will occur.
[0072] In the example of Fig. 4 the encoding device 100 receives a video stream containing less frames per second than should be present in the video stream that is decoded from the encoded data. In particular, the encoding device 100 only receive the frames F2 to be stored with the second number of frames per second. In addition, as also illustrated at the top of Fig. 4, the encoding device 100 receives event data ED that indicate a temporal stream of events.
[0073] Here, as will be described in more detail below with respect to Figs. 5 to 18, each event indicates intensity changes above a predetermined threshold of the light received for one of a plurality of pixels 51 that constitute the frames. In contrast to conventional, frame-based imaging, the generation and output of events is not necessarily frame-based, it can also be asynchronous. This is illustrated in Fig. 4 by showing between each consecutive pair of frames F2 several events occurring at different times. The events are not generated at regular intervals. Every time when a change of light is detected in a single pixel that is larger than a predetermined threshold an event is generated. The generation of events is therefore particularly dependent on motions in the observed scene since these will lead to variations in the observed scene that lead to changes in the observed intensity. Thus, the event data ED constitute information about movements of the objects captured in the frames F2.
[0074] This information is used by the encoding device 100 to determine motion vectors MV from the event data ED and the frames F2 of the received video stream. This is schematically illustrated as the second step in Fig. 4. Thus, in the present example the motion vectors MV are not generated by comparison of frames, but by translating the motion information contained in the event data ED into motion vectors MV, e.g. by extrapolating positions of objects within the frames based on the observed intensity changes that are recorded in the event data ED. Here, the determination of motion vectors based on the event data may be performed by an artificial intelligence model that has been trained for this determination.
[0075] Then, just as in the example described with respect to Fig. 3, missing frames MF are reconstructed by the encoding device 100 based on the stored frames F2 and the determined motion vectors MV as illustrated in the third step of Fig. 4.
[0076] However, since no frames corresponding to the missing frames MF that were reconstructed based on the motion vectors MV are present in the originally received frames Fl, the encoding device 100 reconstructs a second set of missing frames MF. This is shown in the fourth step of Fig. 4, where the second set of missing frames MF is hatched by dots. The first set of missing frames MF that is reconstructed based on motion vectors MV and the second set of missing frames MF are generated for the same temporal position. This means that one missing frame MF of the first set and one missing frame MF of the second set are intended to show the same part of the observed scene at the same time. The second set of missing frames MF is generated based on the stored frames F2 by using a further artificial intelligence model that has been trained for this reconstruction. Also, if necessary, an artificial intelligence model may be used that can also use the event data ED as an input.9
[0077] Thus, in the example of Fig. 4 two frames for the same time are estimated based on two different methods. First, based on motion vectors MV obtained from event data ED. And second, based on an artificial intelligence interpolation of the originally received (and stored) frames F2 that estimates the pixel values of the missing frame MF.
[0078] These two estimations can then be used by the encoding device 100 for error data FL determination. To this end, the missing frames MF from the first set are compared to the temporally corresponding missing frames MF of the second set as illustrated in the fifth step of Fig. 4. This corresponds in particular to the generation of residuals described above with respect to Fig. 3 except that instead of an originally received frame Fl a missing frame MF that has been estimated by an artificial intelligence model is used. Again, the difference between the respectively temporally corresponding missing frames MF can be determined and pixel values in blocks of this residual can be compared to a threshold in order to decide whether or not the block contains an error. Thus, in the example of Fig. 4, it is the deviation between the two estimation schemes that indicates whether an error in an area of a reconstructed missing frame MF exists.
[0079] Finally, similar to the description of Fig. 3, all the reconstructed missing frames MF and the comparison between the reconstructed missing frames are dropped to only store the originally received frames F2 and the error data FL. Thus, a combination of a video stream and event data can be used to encode a video stream that, when decoded, has a higher frame rate than the original video stream. However, in this process, the data amount of the original data is not increased, since no additional frame information but only (comparably small) error data are stored together with the original frames Fl. Thus, an efficient encoding scheme is provided for image data and event data obtained e.g. from hybrid sensors having conventional imaging pixels and event detection pixels. Of course, the operation modes of Figs. 3 and 4 could be combined, e.g. by additionally dropping frames in Fig. 4 or by additionally providing event data in Fig. 3.
[0080] In the following an overview of such hybrid sensors that combine traditional, frame-based sensors, such as active pixel sensors, APS, and event based or dynamic vision sensors, EVS / DVS, will be given. It should be noted that in order to ease the description and also in order to cover an important application example, the following description is focused without prejudice on a specific implementation of a hybrid sensor including both, APS pixels and EVS / DVS pixels. This is of course purely exemplary. It is to be understood that a differently implemented hybrid sensor could be used.
[0081] Fig. 5 is a diagram illustrating a configuration example of a sensor device 10, which is in the example of Fig. 5 constituted by a sensor chip.
[0082] The sensor device 10 is a single-chip semiconductor chip and includes a sensor die (substrate) 11, which serves as a plurality of dies (substrates), and a logic die 12 that are stacked. Note that, the sensor device 10 can also include only a single die or three or more stacked dies.
[0083] In the sensor device 10 of Fig. 5, the sensor die 11 includes (a circuit serving as) a sensor section 21, and the logic die 12 includes a logic section 22. Note that, the sensor section 21 can be partly formed on the10
[0084] logic die 12. Further, the logic section 22 can be partly formed on the sensor die 11.
[0085] The sensor section 21 includes pixels configured to perform photoelectric conversion on incident light to generate electrical signals, and generates event data indicating the occurrence of events that are changes in the electrical signal of the pixels. The sensor section 21 supplies the event data to the logic section 22. That is, the sensor section 21 performs imaging of performing, in the pixels, photoelectric conversion on incident light to generate electrical signals, similarly to a synchronous image sensor, for example. The sensor section 21, however, also generates event data indicating the occurrence of events that are changes in the electrical signal of the pixels in addition to generating image data in a frame format (frame data). The sensor section 21 outputs, to the logic section 22, the event data obtained by the imaging.
[0086] Here, the synchronous image sensor is an image sensor configured to perform imaging in synchronization with a vertical synchronization signal and output frame data that is image data in a frame format. The event detection functions of the sensor section 21 can be regarded as asynchronous (an asynchronous image sensor) in contrast to the synchronous image sensor, since as far as event detection is concerned the sensor section 21 does not operate in synchronization with a vertical synchronization signal when outputting event data.
[0087] Note that, the sensor section 21 generates and outputs, other than event data, frame data, similarly to the synchronous image sensor. In addition, the sensor section 21 can output, together with event data, electrical signals of pixels in which events have occurred, as pixel signals that are pixel values of the pixels in frame data.
[0088] The logic section 22 controls the sensor section 21 as needed. Further, the logic section 22 performs various types of data processing, such as data processing of generating frame data on the basis of event data from the sensor section 21 and image processing on frame data from the sensor section 21 or frame data generated on the basis of the event data from the sensor section 21, and outputs data processing results obtained by performing the various types of data processing on the event data and the frame data.
[0089] Fig. 6 is a block diagram illustrating a configuration example of the sensor section 21 of Fig. 5.
[0090] The sensor section 21 includes a pixel array section 31, a driving section 32, an arbiter 33, an AD (Analog to Digital) conversion section 34, and an output section 35.
[0091] The pixel array section 31 preferably includes a plurality of pixels 51 (Fig. 7) arrayed in a two-dimensional lattice pattern. The pixel array section 31 detects, in a case where a change larger than a predetermined threshold (including a change equal to or larger than the threshold as needed) has occurred in (a voltage corresponding to) a photocurrent that is an electrical signal generated by photoelectric conversion in the pixel 51, the change in the photocurrent as an event. In a case of detecting an event, the pixel array section 31 outputs, to the arbiter 33, a request for requesting the output of event data indicating the occurrence of the event. Then, in a case of receiving a response indicating event data output permission from the arbiter 33, the pixel array section 31 outputs the event data to the driving section 3273960
[0092] 11
[0093] and the output section 35. In addition, the pixel array section 31 may output an electrical signal of the pixel 51 in which the event has been detected to the AD conversion section 34, as a pixel signal.
[0094] The driving section 32 supplies control signals to the pixel array section 31 to drive the pixel array section 31. For example, the driving section 32 drives the pixel 51 regarding which the pixel array section 31 has output event data, so that the pixel 51 in question supplies (outputs) a pixel signal to the AD conversion section 34. But the driving section 32 may also drive the pixels 51 in a synchronous manner to generate frame data.
[0095] The arbiter 33 arbitrates the requests for requesting the output of event data from the pixel array section 31, and returns responses indicating event data output permission or prohibition to the pixel array section 31.
[0096] The AD conversion section 34 includes, for example, a single-slope ADC (AD converter) (not illustrated) in each column of pixel blocks 41 (Fig. 7) described later, for example. The AD conversion section 34 performs, with the ADC in each column, AD conversion on pixel signals of the pixels 51 of the pixel blocks 41 in the column, and supplies the resultant to the output section 35. Note that, the AD conversion section 34 can perform CDS (Correlated Double Sampling) together with pixel signal AD conversion.
[0097] The output section 35 performs necessary processing on the pixel signals from the AD conversion section 34 and the event data from the pixel array section 31 and supplies the resultant to the logic section 22 (Fig.
[0098] 5).
[0099] Here, a change in the photocurrent generated in the pixel 51 can be recognized as a change in the amount of light entering the pixel 51, so that it can also be said that an event is a change in light amount (a change in light amount larger than the threshold) in the pixel 51.
[0100] Event data indicating the occurrence of an event at least includes location information (coordinates or the like) indicating the location of a pixel block in which a change in light amount, which is the event, has occurred. Besides, the event data can also include the polarity (positive or negative) of the change in light amount.
[0101] With regard to the series of event data that is output from the pixel array section 31 at timings at which events have occurred, it can be said that, as long as the event data interval is the same as the event occurrence interval, the event data implicitly includes time point information indicating (relative) time points at which the events have occurred. However, for example, when the event data is stored in a memory and the event data interval is no longer the same as the event occurrence interval, the time point information implicitly included in the event data is lost. Thus, the output section 35 includes, in event data, time point information indicating (relative) time points at which events have occurred, such as timestamps, before the event data interval is changed from the event occurrence interval. The processing of including time point information in event data can be performed in any block other than the output section 35 as long as the processing is performed before time point information implicitly included inevent data is lost.
[0102] Fig. 7 is a block diagram illustrating a configuration example of the pixel array section 31 of Fig. 6.
[0103] The pixel array section 31 may include a plurality of pixel blocks 41. Each pixel block 41 may include the IxJ pixels 51 that are one or more pixels arrayed in I rows and J columns (I and J are integers), an event detecting section 52, and a pixel signal generating section 53. The one or more pixels 51 in the pixel block 41 share the event detecting section 52 and the pixel signal generating section 53. Further, in each column of the pixel blocks 41, a VSL (Vertical Signal Line) for connecting the pixel blocks 41 to the ADC of the AD conversion section 34 is wired.
[0104] The pixel 51 receives light incident from an object and performs photoelectric conversion to generate a photocurrent serving as an electrical signal. The pixel 51 supplies the photocurrent to the event detecting section 52 under the control of the driving section 32.
[0105] The event detecting section 52 detects, as an event, a change larger than the predetermined threshold in photocurrent from each of the pixels 51, under the control of the driving section 32. In a case of detecting an event, the event detecting section 52 may supply, to the arbiter 33 (Fig. 6), a request for requesting the output of event data indicating the occurrence of the event. Then, when receiving a response indicating event data output permission to the request from the arbiter 33, the event detecting section 52 outputs the event data to the driving section 32 and the output section 35.
[0106] The pixel signal generating section 53 generates, in the case where the event detecting section 52 has detected an event, a voltage corresponding to a photocurrent from the pixel 51 as a pixel signal, and supplies the voltage to the AD conversion section 34 through the VSL, under the control of the driving section 32.
[0107] Here, detecting a change larger than the predetermined threshold in photocurrent as an event can also be recognized as detecting, as an event, absence of change larger than the predetermined threshold in photocurrent. The pixel signal generating section 53 can generate a pixel signal in the case where absence of change larger than the predetermined threshold in photocurrent has been detected as an event as well as in the case where a change larger than the predetermined threshold in photocurrent has been detected as an event.
[0108] Further, the event detecting section 52 and the pixel signal generating section 53 may operate independently from each other. This means that while event data are generated by the event detecting section 52, pixel signals are generated by the pixel signal generating section 53 in a synchronous manner to generate frame data. The event detecting section 52 and the pixel signal generating section 53 may also operate in a time multiplexed manner such that each of the sections receives the photocurrent from the pixels 51 only during different time intervals.
[0109] Fig. 8 is a circuit diagram illustrating a configuration example of the pixel block 41. The configuration of13
[0110] Fig. 8 may be referred to as “shared pixel configuration”.
[0111] The pixel block 41 includes, as described with reference to Fig. 7, the pixels 51, the event detecting section 52, and the pixel signal generating section 53.
[0112] The pixel 51 includes a photoelectric conversion element 61 and transfer transistors 62 and 63.
[0113] The photoelectric conversion element 61 includes, for example, a PD (Photodiode). The photoelectric conversion element 61 receives incident light and performs photoelectric conversion to generate charges.
[0114] The transfer transistor 62 includes, for example, an N (Negative)-type MOS (Metal-Oxide-Semiconductor) FET (Field Effect Transistor). The transfer transistor 62 of the n-th pixel 51 of the IxJ pixels 51 in the pixel block 41 is turned on or off in response to a control signal OFGn supplied from the driving section 32 (Fig. 6). When the transfer transistor 62 is turned on, charges generated in the photoelectric conversion element 61 are transferred (supplied) to the event detecting section 52, as a photocurrent.
[0115] The transfer transistor 63 includes, for example, an N-type MOSFET. The transfer transistor 63 of the n-th pixel 51 of the IxJ pixels 51 in the pixel block 41 is turned on or off in response to a control signal TRGn supplied from the driving section 32. When the transfer transistor 63 is turned on, charges generated in the photoelectric conversion element 61 are transferred to an FD 74 of the pixel signal generating section 53.
[0116] The IxJ pixels 51 in the pixel block 41 are connected to the event detecting section 52 of the pixel block 41 through nodes 60. Thus, photocurrents generated in (the photoelectric conversion elements 61 of) the pixels 51 are supplied to the event detecting section 52 through the nodes 60. As a result, the event detecting section 52 receives the sum of photocurrents from all the pixels 51 in the pixel block 41. Thus, the event detecting section 52 detects, as an event, a change in sum of photocurrents supplied from the IxJ pixels 51 in the pixel block 41.
[0117] The pixel signal generating section 53 includes a reset transistor 71, an amplification transistor 72, a selection transistor 73, and the FD (Floating Diffusion) 74.
[0118] The reset transistor 71, the amplification transistor 72, and the selection transistor 73 include, for example, N-type MOSFETs.
[0119] The reset transistor 71 is turned on or off in response to a control signal RST supplied from the driving section 32 (Fig. 6). When the reset transistor 71 is turned on, the FD 74 is connected to a power supply VDD, and charges accumulated in the FD 74 are thus discharged to the power supply VDD. With this, the FD 74 is reset.
[0120] The amplification transistor 72 has a gate connected to the FD 74, a drain connected to the power supply14
[0121] VDD, and a source connected to the VSL through the selection transistor 73. The amplification transistor 72 is a source follower and outputs a voltage (electrical signal) corresponding to the voltage of the FD 74 supplied to the gate to the VSL through the selection transistor 73.
[0122] The selection transistor 73 is turned on or off in response to a control signal SEL supplied from the driving section 32. When the selection transistor 73 is turned on, a voltage corresponding to the voltage of the FD 74 from the amplification transistor 72 is output to the VSL.
[0123] The FD 74 accumulates charges transferred from the photoelectric conversion elements 61 of the pixels 51 through the transfer transistors 63, and converts the charges to voltages.
[0124] With regard to the pixels 51 and the pixel signal generating section 53, which are configured as described above, the driving section 32 turns on the transfer transistors 62 with control signals OFGn, so that the transfer transistors 62 supply, to the event detecting section 52, photocurrents based on charges generated in the photoelectric conversion elements 61 of the pixels 51. With this, the event detecting section 52 receives a current that is the sum of the photocurrents from all the pixels 51 in the pixel block 41, which might also be only a single pixel.
[0125] When the event detecting section 52 detects, as an event, a change in photocurrent (sum of photocurrents) in the pixel block 41, the driving section 32 may turn off the transfer transistors 62 of all the pixels 51 in the pixel block 41, to thereby stop the supply of the photocurrents to the event detecting section 52. Then, the driving section 32 may sequentially turn on, with the control signals TRGn, the transfer transistors 63 of the pixels 51 in the pixel block 41 in which the event has been detected, so that the transfer transistors 63 transfers charges generated in the photoelectric conversion elements 61 to the FD 74. The FD 74 accumulates the charges transferred from (the photoelectric conversion elements 61 of) the pixels 51. Voltages corresponding to the charges accumulated in the FD 74 are output to the VSL, as pixel signals of the pixels 51, through the amplification transistor 72 and the selection transistor 73.
[0126] As described above, in the sensor section 21 (Fig. 6), in this manner only pixel signals of the pixels 51 in the pixel block 41 in which an event has been detected are sequentially output to the VSL. The pixel signals output to the VSL are supplied to the AD conversion section 34 to be subjected to AD conversion.
[0127] Here, in the pixels 51 in the pixel block 41, the transfer transistors 63 can be turned on not sequentially but simultaneously. In this case, the sum of pixel signals of all the pixels 51 in the pixel block 41 can be output.
[0128] Further, it is also possible to operate the pixel generating section 53 independently from the event detection by controlling the transfer transistors 63 without regard of the event detection or by entirely omitting the transfer transistors 63. In this case frame data can be generated by the pixel generating section 53 in the conventional, synchronous manner, while the event detection can be performed concurrently. If necessary, photocurrent transfer to the event detecting section 52 and the pixel signal generating section 53 may be time multiplexed such that only one of the two transfer transistors 62, 6373960
[0129] 15
[0130] per pixel 51 is open at a given point in time.
[0131] In the pixel array section 31 of Fig. 7, the pixel block 41 includes one or more pixels 51, and the one or more pixels 51 share the event detecting section 52 and the pixel signal generating section 53. Thus, in the case where the pixel block 41 includes a plurality of pixels 51, the numbers of the event detecting sections 52 and the pixel signal generating sections 53 can be reduced as compared to a case where the event detecting section 52 and the pixel signal generating section 53 are provided for each of the pixels 51, with the result that the scale of the pixel array section 31 can be reduced.
[0132] Note that, in the case where the pixel block 41 includes a plurality of pixels 51, the event detecting section 52 can be provided for each of the pixels 51. In the case where the plurality of pixels 51 in the pixel block 41 share the event detecting section 52, events are detected in units of the pixel blocks 41. In the case where the event detecting section 52 is provided for each of the pixels 51, however, events can be detected in units of the pixels 51.
[0133] Yet, even in the case where the plurality of pixels 51 in the pixel block 41 share the single event detecting section 52, events can be detected in units of the pixels 51 when the transfer transistors 62 of the plurality of pixels 51 are temporarily turned on in a time-division manner.
[0134] Further, in a case where there is no need to output pixel signals, the pixel block 41 can be formed without the pixel signal generating section 53. In the case where the pixel block 41 is formed without the pixel signal generating section 53, the sensor section 21 can be formed without the AD conversion section 34 and the transfer transistors 63. In this case, the scale of the sensor section 21 can be reduced. The sensor will then output the address of the pixel (block) in which the event occurred, if necessary with a time stamp.
[0135] Moreover, additional pixels 51 connected to pixel signal generating sections 53, but not to event detecting sections 52 may be provided on the sensor die 11. These pixels 51 may be interleaved with the pixels 51 connected to the event detecting section 52, but may also form a separate pixel array for generating frame data. The separation of the event detection function and the pixel signal generating function may even by implemented as two pixel arrays on different dies that are arranged such as to observe the same scene. Thus basically any pixel arrangement / pixel circuitry might be used that allows obtaining event data with a first subset of pixels 51 and generating of pixel signals with a second subset of pixels 51.
[0136] Fig. 9 is a block diagram illustrating a configuration example of the event detecting section 52 of Fig. 7.
[0137] The event detecting section 52 includes a current-voltage converting section 81, a buffer 82, a subtraction section 83, a quantization section 84, and a transfer section 85.
[0138] The current-voltage converting section 81 converts (a sum of) photocurrents from the pixels 51 to voltages corresponding to the logarithms of the photocurrents (hereinafter also referred to as a "photovoltage") and supplies the voltages to the buffer 82.73960
[0139] 16
[0140] The buffer 82 buffers photovoltages from the current-voltage converting section 81 and supplies the resultant to the subtraction section 83.
[0141] The subtraction section 83 calculates, at a timing instructed by a row driving signal that is a control signal from the driving section 32, a difference between the current photovoltage and a photovoltage at a timing slightly shifted from the current time, and supplies a difference signal corresponding to the difference to the quantization section 84.
[0142] The quantization section 84 quantizes difference signals from the subtraction section 83 to digital signals and supplies the quantized values of the difference signals to the transfer section 85 as event data.
[0143] The transfer section 85 transfers (outputs), on the basis of event data from the quantization section 84, the event data to the output section 35. That is, the transfer section 85 supplies a request for requesting the output of the event data to the arbiter 33. Then, when receiving a response indicating event data output permission to the request from the arbiter 33, the transfer section 85 outputs the event data to the output section 35.
[0144] Fig. 10 is a circuit diagram illustrating a configuration example of the current-voltage converting section 81 of Fig. 9.
[0145] The current-voltage converting section 81 includes transistors 91 to 93. As the transistors 91 and 93, for example, N-type MOSFETs can be employed. As the transistor 92, for example, a P-type MOSFET can be employed.
[0146] The transistor 91 has a source connected to the gate of the transistor 93, and a photocurrent is supplied from the pixel 51 to the connecting point between the source of the transistor 91 and the gate of the transistor 93. The transistor 91 has a drain connected to the power supply VDD and a gate connected to the drain of the transistor 93.
[0147] The transistor 92 has a source connected to the power supply VDD and a drain connected to the connecting point between the gate of the transistor 91 and the drain of the transistor 93. A predetermined bias voltage Vbias is applied to the gate of the transistor 92. With the bias voltage Vbias, the transistor 92 is turned on or off, and the operation of the current-voltage converting section 81 is turned on or off depending on whether the transistor 92 is turned on or off.
[0148] The source of the transistor 93 is grounded.
[0149] In the current-voltage converting section 81, the transistor 91 has the drain connected on the power supply VDD side and is thus a source follower. The source of the transistor 91, which is the source follower, is connected to the pixels 51 (Fig. 8), so that photocurrents based on charges generated in the photoelectric conversion elements 61 of the pixels 51 flow through the transistor 91 (from the drain to the source). The73960
[0150] 17
[0151] transistor 91 operates in a subthreshold region, and at the gate of the transistor 91, photovoltages corresponding to the logarithms of the photocurrents flowing through the transistor 91 are generated. As described above, in the current-voltage converting section 81, the transistor 91 converts photocurrents from the pixels 51 to photovoltages corresponding to the logarithms of the photocurrents.
[0152] In the current-voltage converting section 81, the transistor 91 has the gate connected to the connecting point between the drain of the transistor 92 and the drain of the transistor 93, and the photovoltages are output from the connecting point in question.
[0153] Fig. 11 is a circuit diagram illustrating configuration examples of the subtraction section 83 and the quantization section 84 of Fig. 9.
[0154] The subtraction section 83 includes a capacitor 101, an operational amplifier 102, a capacitor 103, and a switch 104. The quantization section 84 includes a comparator 111.
[0155] The capacitor 101 has one end connected to the output terminal of the buffer 82 (Fig. 9) and the other end connected to the input terminal (inverting input terminal) of the operational amplifier 102. Thus, photovoltages are input to the input terminal of the operational amplifier 102 through the capacitor 101.
[0156] The operational amplifier 102 has an output terminal connected to the non-inverting input terminal (+) of the comparator 111.
[0157] The capacitor 103 has one end connected to the input terminal of the operational amplifier 102 and the other end connected to the output terminal of the operational amplifier 102.
[0158] The switch 104 is connected to the capacitor 103 to switch the connections between the ends of the capacitor 103. The switch 104 is turned on or off in response to a row driving signal that is a control signal from the driving section 32, to thereby switch the connections between the ends of the capacitor 103.
[0159] A photovoltage on the buffer 82 (Fig. 9) side of the capacitor 101 when the switch 104 is on is denoted by Vinit, and the capacitance (electrostatic capacitance) of the capacitor 101 is denoted by Cl. The input terminal of the operational amplifier 102 serves as a virtual ground terminal, and a charge Qinit that is accumulated in the capacitor 101 in the case where the switch 104 is on is expressed by Expression (1).
[0160] Qinit = Cl x Vinit (1)
[0161] Further, in the case where the switch 104 is on, the connection between the ends of the capacitor 103 is cut (short-circuited), so that no charge is accumulated in the capacitor 103.
[0162] When a photovoltage on the buffer 82 (Fig. 9) side of the capacitor 101 in the case where the switch 104 has thereafter been turned off is denoted by Vafter, a charge Qafter that is accumulated in the capacitor73960
[0163] 18
[0164] 101 in the case where the switch 104 is off is expressed by Expression (2).
[0165] Qafter = Cl x Vafter (2)
[0166] When the capacitance of the capacitor 103 is denoted by C2 and the output voltage of the operational amplifier 102 is denoted by Vout, a charge Q2 that is accumulated in the capacitor 103 is expressed by Expression (3).
[0167] Q2 = -C2 x Vout (3)
[0168] Since the total amount of charges in the capacitors 101 and 103 does not change before and after the switch 104 is turned off, Expression (4) is established.
[0169] Qinit = Qafter + Q2 (4)
[0170] When Expression (1) to Expression (3) are substituted for Expression (4), Expression (5) is obtained.
[0171] Vout = - (C1 / C2) x (Vafter - Vinit) (5)
[0172] With Expression (5), the subtraction section 83 subtracts the photovoltage Vinit from the photovoltage Vafter, that is, calculates the difference signal (Vout) corresponding to a difference Vafter - Vinit between the photovoltages Vafter and Vinit. With Expression (5), the subtraction gain of the subtraction section 83 is C1 / C2. Since the maximum gain is normally desired, Cl is preferably set to a large value and C2 is preferably set to a small value. Meanwhile, when C2 is too small, kTC noise increases, resulting in a risk of deteriorated noise characteristics. Thus, the capacitance C2 can only be reduced in a range that achieves acceptable noise. Further, since the pixel blocks 41 each have installed therein the event detecting section 52 including the subtraction section 83, the capacitances Cl and C2 have space constraints. In consideration of these matters, the values of the capacitances Cl and C2 are determined.
[0173] The comparator 111 compares a difference signal from the subtraction section 83 with a predetermined threshold (voltage) Vth (>0) applied to the inverting input terminal (-), thereby quantizing the difference signal. The comparator 111 outputs the quantized value obtained by the quantization to the transfer section 85 as event data.
[0174] For example, in a case where a difference signal is larger than the threshold Vth, the comparator 111 outputs an H (High) level indicating 1, as event data indicating the occurrence of an event. In a case where a difference signal is not larger than the threshold Vth, the comparator 111 outputs an L (Low) level indicating 0, as event data indicating that no event has occurred.
[0175] The transfer section 85 supplies a request to the arbiter 33 in a case where it is confirmed on the basis of event data from the quantization section 84 that a change in light amount that is an event has occurred, that is, in the case where the difference signal (Vout) is larger than the threshold Vth. When receiving a73960
[0176] 19
[0177] response indicating event data output permission, the transfer section 85 outputs the event data indicating the occurrence of the event (for example, H level) to the output section 35.
[0178] The output section 35 includes, in event data from the transfer section 85, location / address information regarding (the pixel block 41 including) the pixel 51 in which an event indicated by the event data has occurred and time point information indicating a time point at which the event has occurred, and further, as needed, the polarity of a change in light amount that is the event, i.e. whether the intensity did increase or decrease. The output section 35 outputs the event data.
[0179] As the data format of event data including location information regarding the pixel 51 in which an event has occurred, time point information indicating a time point at which the event has occurred, and the polarity of a change in light amount that is the event, for example, the data format called "AER (Address Event Representation)" can be employed.
[0180] Note that, a gain A of the entire event detecting section 52 is expressed by the following expression where the gain of the current-voltage converting section 81 is denoted by CGiogand the gain of the buffer 82 is 1.
[0181] A = CGiogC 1 / C2 (ZiPhoto_n) (6)
[0182] Here, iPhoto_n denotes a photocurrent of the n-th pixel 51 of the IxJ pixels 51 in the pixel block 41. In Expression (6), E denotes the summation of n that takes integers ranging from 1 to IxJ.
[0183] Note that, the pixel 51 can receive any light as incident light with an optical fdter through which predetermined light passes, such as a color fdter. For example, in a case where the pixel 51 receives visible light as incident light, event data indicates the occurrence of changes in pixel value in images including visible objects. Further, for example, in a case where the pixel 51 receives, as incident light, infrared light, millimeter waves, or the like for ranging, event data indicates the occurrence of changes in distances to objects. In addition, for example, in a case where the pixel 51 receives infrared light for temperature measurement, as incident light, event data indicates the occurrence of changes in temperature of objects. In the present embodiment, the pixel 51 is assumed to receive visible light as incident light.
[0184] Fig. 12 is a timing chart illustrating an example of the operation of the sensor section 21 of Fig. 6.
[0185] At Timing TO, the driving section 32 changes all the control signals OFGn from the L level to the H level, thereby turning on the transfer transistors 62 of all the pixels 51 in the pixel block 41. With this, the sum of photocurrents from all the pixels 51 in the pixel block 41 is supplied to the event detecting section 52. Here, the control signals TRGn are all at the L level, and hence the transfer transistors 63 of all the pixels 51 are off.
[0186] For example, at Timing Tl, when detecting an event, the event detecting section 52 outputs event data at the H level in response to the detection of the event.73960
[0187] 20
[0188] At Timing T2, the driving section 32 sets all the control signals OFGn to the L level on the basis of the event data at the H level, to stop the supply of the photocurrents from the pixels 51 to the event detecting section 52. Further, the driving section 32 sets the control signal SEL to the H level, and sets the control signal RST to the H level over a certain period of time, to control the FD 74 to discharge the charges to the power supply VDD, thereby resetting the FD 74. The pixel signal generating section 53 outputs, as a reset level, a pixel signal corresponding to the voltage of the FD 74 when the FD 74 has been reset, and the AD conversion section 34 performs AD conversion on the reset level.
[0189] At Timing T3 after the reset level AD conversion, the driving section 32 sets a control signal TRG1 to the H level over a certain period to control the first pixel 51 in the pixel block 41 in which the event has been detected (or which is triggered for other reasons, as e.g. time multiplexed output and / or synchronous readout of frame data) to transfer, to the FD 74, charges generated by photoelectric conversion in (the photoelectric conversion element 61 of) the first pixel 51. The pixel signal generating section 53 outputs, as a signal level, a pixel signal corresponding to the voltage of the FD 74 to which the charges have been transferred from the pixel 51, and the AD conversion section 34 performs AD conversion on the signal level.
[0190] The AD conversion section 34 outputs, to the output section 35, a difference between the signal level and the reset level obtained after the AD conversion, as a pixel signal serving as a pixel value of the image (frame data).
[0191] Here, the processing of obtaining a difference between a signal level and a reset level as a pixel signal serving as a pixel value of an image is called "CDS." CDS can be performed after the AD conversion of a signal level and a reset level, or can be simultaneously performed with the AD conversion of a signal level and a reset level in a case where the AD conversion section 34 performs single-slope AD conversion. In the latter case, AD conversion is performed on the signal level by using the AD conversion result of the reset level as an initial value.
[0192] At Timing T4 after the AD conversion of the pixel signal of the first pixel 51 in the pixel block 41, the driving section 32 sets a control signal TRG2 to the H level over a certain period of time to control the second pixel 51 in the pixel block 41 in which the event has been detected to output a pixel signal.
[0193] In the sensor section 21, similar processing is executed thereafter, so that pixel signals of the pixels 51 in the pixel block 41 in which the event has been detected are sequentially output.
[0194] When the pixel signals of all the pixels 51 in the pixel block 41 are output, the driving section 32 sets all the control signals OFGn to the H level to turn on the transfer transistors 62 of all the pixels 51 in the pixel block 41.
[0195] Fig. 13 is a diagram illustrating an example of a frame data generation method based on event data.
[0196] The logic section 22 sets a frame interval and a frame width on the basis of an externally input command,73960
[0197] 21
[0198] for example. Here, the frame interval represents the interval of frames of frame data that is generated on the basis of event data. The frame width represents the time width of event data that is used for generating frame data on a single frame. A frame interval and a frame width that are set by the logic section 22 are also referred to as a "set frame interval" and a "set frame width," respectively.
[0199] The logic section 22 generates, on the basis of the set frame interval, the set frame width, and event data from the sensor section 21, frame data that is image data in a frame format, to thereby convert the event data to the frame data.
[0200] That is, the logic section 22 generates, in each set frame interval, frame data on the basis of event data in the set frame width from the beginning of the set frame interval.
[0201] Here, it is assumed that event data includes time point information ti indicating a time point at which an event has occurred (hereinafter also referred to as an "event time point") and coordinates (x, y) serving as location information regarding (the pixel block 41 including) the pixel 51 in which the event has occurred (hereinafter also referred to as an "event location").
[0202] In Fig. 13, in a three-dimensional space (time and space) with the x axis, the y axis, and the time axis t, points representing event data are plotted on the basis of the event time point t and the event location (coordinates) (x, y) included in the event data.
[0203] That is, when a location (x, y, t) on the three-dimensional space indicated by the event time point t and the event location (x, y) included in event data is regarded as the space-time location of an event, in Fig. 13, the points representing the event data are plotted on the space-time locations (x, y, t) of the events.
[0204] The logic section 22 starts to generate frame data on the basis of event data by using, as a generation start time point at which frame data generation starts, a predetermined time point, for example, a time point at which frame data generation is externally instructed or a time point at which the sensor device 10 is powered on.
[0205] Here, cuboids each having the set frame width in the direction of the time axis t in the set frame intervals, which appear from the generation start time point, are referred to as a "frame volume." The size of the frame volume in the x-axis direction or the y-axis direction is equal to the number of the pixel blocks 41 or the pixels 51 in the x-axis direction or the y-axis direction, for example.
[0206] The logic section 22 generates, in each set frame interval, frame data on a single frame on the basis of event data in the frame volume having the set frame width from the beginning of the set frame interval.
[0207] Frame data can be generated by, for example, setting white to a pixel (pixel value) in a frame at the event location (x, y) included in event data and setting a predetermined color such as gray to pixels at other locations in the frame.73960
[0208] 22
[0209] Besides, in a case where event data includes the polarity of a change in light amount that is an event, frame data can be generated in consideration of the polarity included in the event data. For example, white can be set to pixels in the case a positive polarity, while black can be set to pixels in the case of a negative polarity.
[0210] In addition, in the case where pixel signals of the pixels 51 are also output when event data is output as described with reference to Fig. 7 and Fig. 8, frame data can be generated on the basis of the event data by using the pixel signals of the pixels 51. That is, frame data can be generated by setting, in a frame, a pixel at the event location (x, y) (in a block corresponding to the pixel block 41) included in event data to a pixel signal of the pixel 51 at the location (x, y) and setting a predetermined color such as gray to pixels at other locations.
[0211] Note that, in the frame volume, there are a plurality of pieces of event data that are different in the event time point t but the same in the event location (x, y) in some cases. In this case, for example, event data at the latest or oldest event time point t can be prioritized. Further, in the case where event data includes polarities, the polarities of a plurality of pieces of event data that are different in the event time point t but the same in the event location (x, y) can be added together, and a pixel value based on the added value obtained by the addition can be set to a pixel at the event location (x, y).
[0212] Here, in a case where the frame width and the frame interval are the same, the frame volumes are adjacent to each other without any gap. Further, in a case where the frame interval is larger than the frame width, the frame volumes are arranged with gaps. In a case where the frame width is larger than the frame interval, the frame volumes are arranged to be partly overlapped with each other.
[0213] As explained above, the pixel signal generating section 53 may also generate frame data in the conventional, synchronous manner.
[0214] Fig. 14 is a block diagram illustrating another configuration example of the quantization section 84 of Fig.
[0215] 9.
[0216] Note that, in Fig. 14, parts corresponding to those in the case of Fig. 11 are denoted by the same reference signs, and the description thereof is omitted as appropriate below.
[0217] In Fig. 14, the quantization section 84 includes comparators 111 and 112 and an output section 113.
[0218] Thus, the quantization section 84 of Fig. 14 is similar to the case of Fig. 11 in including the comparator 111. However, the quantization section 84 of Fig. 14 is different from the case of Fig. 11 in newly including the comparator 112 and the output section 113.
[0219] The event detecting section 52 (Fig. 9) including the quantization section 84 of Fig. 14 detects, in addition to events, the polarities of changes in light amount that are events.73960
[0220] 23
[0221] In the quantization section 84 of Fig. 14, the comparator 111 outputs, in the case where a difference signal is larger than the threshold Vth, the H level indicating 1, as event data indicating the occurrence of an event having the positive polarity. The comparator 111 outputs, in the case where a difference signal is not larger than the threshold Vth, the L level indicating 0, as event data indicating that no event having the positive polarity has occurred.
[0222] Further, in the quantization section 84 of Fig. 14, a threshold Vth' (<Vth) is supplied to the non-inverting input terminal (+) of the comparator 112, and difference signals are supplied to the inverting input terminal (-) of the comparator 112 from the subtraction section 83. Here, for the sake of simple description, it is assumed that the threshold Vth' is equal to -Vth, for example, which needs however not to be the case.
[0223] The comparator 112 compares a difference signal from the subtraction section 83 with the threshold Vth' applied to the inverting input terminal (-), thereby quantizing the difference signal. The comparator 112 outputs, as event data, the quantized value obtained by the quantization.
[0224] For example, in a case where a difference signal is smaller than the threshold Vth' (the absolute value of the difference signal having a negative value is larger than the threshold Vth), the comparator 112 outputs the H level indicating 1, as event data indicating the occurrence of an event having the negative polarity. Further, in a case where a difference signal is not smaller than the threshold Vth' (the absolute value of the difference signal having a negative value is not larger than the threshold Vth), the comparator 112 outputs the L level indicating 0, as event data indicating that no event having the negative polarity has occurred.
[0225] The output section 113 outputs, on the basis of event data output from the comparators 111 and 112, event data indicating the occurrence of an event having the positive polarity, event data indicating the occurrence of an event having the negative polarity, or event data indicating that no event has occurred to the transfer section 85.
[0226] For example, the output section 113 outputs, in a case where event data from the comparator 111 is the H level indicating 1, +V volts indicating +1, as event data indicating the occurrence of an event having the positive polarity, to the transfer section 85. Further, the output section 113 outputs, in a case where event data from the comparator 112 is the H level indicating 1, -V volts indicating -1, as event data indicating the occurrence of an event having the negative polarity, to the transfer section 85. In addition, the output section 113 outputs, in a case where each event data from the comparators 111 and 112 is the L level indicating 0, 0 volts (GND level) indicating 0, as event data indicating that no event has occurred, to the transfer section 85.
[0227] The transfer section 85 supplies a request to the arbiter 33 in the case where it is confirmed on the basis of event data from the output section 113 of the quantization section 84 that a change in light amount that is an event having the positive polarity or the negative polarity has occurred. After receiving a response indicating event data output permission, the transfer section 85 outputs event data indicating the73960
[0228] 24
[0229] occurrence of the event having the positive polarity or the negative polarity (+V volts indicating 1 or -V volts indicating -1) to the output section 35.
[0230] Preferably, the quantization section 84 has a configuration as illustrated in Fig. 14.
[0231] Fig. 15 is a diagram illustrating another configuration example of the event detecting section 52.
[0232] In Fig. 15, the event detecting section 52 includes a subtractor 430, a quantizer 440, a memory 451, and a controller 452. The subtractor 430 and the quantizer 440 correspond to the subtraction section 83 and the quantization section 84, respectively.
[0233] Note that, in Fig. 15, the event detecting section 52 further includes blocks corresponding to the currentvoltage converting section 81 and the buffer 82, but the illustrations of the blocks are omitted in Fig. 15.
[0234] The subtractor 430 includes a capacitor 431, an operational amplifier 432, a capacitor 433, and a switch 434. The capacitor 431, the operational amplifier 432, the capacitor 433, and the switch 434 correspond to the capacitor 101, the operational amplifier 102, the capacitor 103, and the switch 104, respectively.
[0235] The quantizer 440 includes a comparator 441. The comparator 441 corresponds to the comparator 111.
[0236] The comparator 441 compares a voltage signal (difference signal) from the subtractor 430 with the predetermined threshold voltage Vth applied to the inverting input terminal (-). The comparator 441 outputs a signal indicating the comparison result, as a detection signal (quantized value).
[0237] The voltage signal from the subtractor 430 may be input to the input terminal (-) of the comparator 441, and the predetermined threshold voltage Vth may be input to the input terminal (+) of the comparator 441.
[0238] The controller 452 supplies the predetermined threshold voltage Vth applied to the inverting input terminal (-) of the comparator 441. The threshold voltage Vth which is supplied may be changed in a time-division manner. For example, the controller 452 supplies a threshold voltage Vthl corresponding to ON events (for example, positive changes in photocurrent) and a threshold voltage Vth2 corresponding to OFF events (for example, negative changes in photocurrent) at different timings to allow the single comparator to detect a plurality of types of address events (events).
[0239] The memory 451 accumulates output from the comparator 441 on the basis of Sample signals supplied from the controller 452. The memory 451 may be a sampling circuit, such as a switch, plastic, or capacitor, or a digital memory circuit, such as a latch or flip-flop. For example, the memory 451 may hold, in a period in which the threshold voltage Vth2 corresponding to OFF events is supplied to the inverting input terminal (-) of the comparator 441, the result of comparison by the comparator 441 using the threshold voltage Vthl corresponding to ON events. Note that, the memory 451 may be omitted, may be provided inside the pixel (pixel block 41), or may be provided outside the pixel.73960
[0240] 25
[0241] Fig. 16 is a block diagram illustrating another configuration example of the pixel array section 31 of Fig.
[0242] 6.
[0243] Note that, in Fig. 16, parts corresponding to those in the case of Fig. 7 are denoted by the same reference signs, and the description thereof is omitted as appropriate below.
[0244] In Fig. 16, the pixel array section 31 includes the plurality of pixel blocks 41. The pixel block 41 includes the lx J pixels 51 that are one or more pixels and the event detecting section 52.
[0245] Thus, the pixel array section 31 of Fig. 16 is similar to the case of Fig. 7 in that the pixel array section 31 includes the plurality of pixel blocks 41 and that the pixel block 41 includes one or more pixels 51 and the event detecting section 52. However, the pixel array section 31 of Fig. 16 is different from the case of Fig.
[0246] 7 in that the pixel block 41 does not include the pixel signal generating section 53.
[0247] As described above, in the pixel array section 31 of Fig. 16, the pixel block 41 does not include the pixel signal generating section 53, so that the sensor section 21 (Fig. 6) can be formed without the AD conversion section 34. As further described above, the pixel signal generating section 53 and pixels 51 feeding into it may be placed on a different part of the die 11 or even on another sensor chip.
[0248] Fig. 17 is a circuit diagram illustrating a configuration example of the pixel block 41 of Fig. 16.
[0249] As described with reference to Fig. 16, the pixel block 41 includes the pixels 51 and the event detecting section 52, but does not include the pixel signal generating section 53.
[0250] In this case, the pixel 51 can only include the photoelectric conversion element 61 without the transfer transistors 62 and 63.
[0251] Note that, in the case where the pixel 51 has the configuration illustrated in Fig. 17, the event detecting section 52 can output a voltage corresponding to a photocurrent from the pixel 51, as a pixel signal.
[0252] Above, the sensor device 10 was described to be an asynchronous imaging device configured to read out events by the asynchronous readout system. However, the event readout system is not limited to the asynchronous readout system and may be the synchronous readout system. An imaging device to which the synchronous readout system is applied is a scan type imaging device that is the same as a general imaging device configured to perform imaging at a predetermined frame rate. Further, the event data detection may be performed asynchronously, while the pixel signal generation may be performed synchronously.
[0253] Fig. 18 is a block diagram illustrating a configuration example of a scan type imaging device.
[0254] As illustrated in Fig. 18, an imaging device 510 includes a pixel array section 521, a driving section 522, a signal processing section 525, a read-out region selecting section 527, and a signal generating section73960
[0255] 26
[0256] 528.
[0257] The pixel array section 521 includes a plurality of pixels 530. The plurality of pixels 530 each output an output signal in response to a selection signal from the read-out region selecting section 527. The plurality of pixels 530 can each include an in-pixel quantizer as illustrated in Fig. 15, for example. The plurality of pixels 530 output output signals corresponding to the amounts of change in light intensity. The plurality of pixels 530 may be two-dimensionally disposed in a matrix as illustrated in Fig. 18.
[0258] The driving section 522 drives the plurality of pixels 530, so that the pixels 530 output pixel signals generated in the pixels 530 to the signal processing section 525 through an output line 514. Note that, the driving section 522 and the signal processing section 525 are circuit sections for acquiring grayscale information. Thus, in a case where only event information (event data) is acquired, the driving section 522 and the signal processing section 525 may be omitted.
[0259] The read-out region selecting section 527 selects some of the plurality of pixels 530 included in the pixel array section 521. For example, the read-out region selecting section 527 selects one or a plurality of rows included in the two-dimensional matrix structure corresponding to the pixel array section 521. The readout region selecting section 527 sequentially selects one or a plurality of rows on the basis of a cycle set in advance. Further, the read-out region selecting section 527 may determine a selection region on the basis of requests from the pixels 530 in the pixel array section 521.
[0260] The signal generating section 528 generates, on the basis of output signals of the pixels 530 selected by the read-out region selecting section 527, event signals corresponding to active pixels in which events have been detected of the selected pixels 530. The events mean an event that the intensity of light changes. The active pixels mean the pixel 530 in which the amount of change in light intensity corresponding to an output signal exceeds or falls below a threshold set in advance. For example, the signal generating section 528 compares output signals from the pixels 530 with a reference signal, and detects, as an active pixel, a pixel that outputs an output signal larger or smaller than the reference signal. The signal generating section 528 generates an event signal (event data) corresponding to the active pixel.
[0261] The signal generating section 528 can include, for example, a column selecting circuit configured to arbitrate signals input to the signal generating section 528. Further, the signal generating section 528 can output not only information regarding active pixels in which events have been detected, but also information regarding non-active pixels in which no event has been detected, i.e. it can operate as a conventional synchronous image sensorthat generates a series of consecutive image frames.
[0262] The signal generating section 528 outputs, through an output line 515, address information and timestamp information (for example, (X, Y, T)) regarding the active pixels in which the events have been detected. However, the data that is output from the signal generating section 528 may not only be the address information and the timestamp information, but also information in a frame format (for example, (0, 0, 1, o, -)).73960
[0263] 27
[0264] In the above description a sensor device 10 has been described in which event data generation and pixel signal generation may depend on each other or may be independent of each other. Moreover, pixels 51 may be shared between event detection circuitry and pixel signal generating circuitry either by respective circuitry or by time multiplexing. But pixels 51 may also be divided such that there are event detection pixels and pixel signal generating pixels. These pixels 51 may be interleaved in the same pixel array or may be part of different pixel arrays on the same die or even be arranged on different dies.
[0265] A sensor device 10 that comprises the above described encoding device 100 and a hybrid image sensor as described above is schematically illustrated in Fig. 19. Besides the encoding device 100 the sensor device 10 comprises a plurality of pixels 51 that are each configured to receive light from a scene and perform photoelectric conversion to generate an electrical signal. Here, the pixels 51 might be arranged in pixel blocks 41 as described with respect to Fig. 7. For the ease of description it is assumed in the following without limitation that there is only one pixel 51 per pixel block 41, i.e. that pixels 51 and pixel blocks 41 refer to the same structural elements. However, the following description could be easily generalized by understanding any reference to a pixel 51 as a reference to a pixel block 41 comprising several pixels 51. In Fig. 19 a two-dimensional array of pixels 51 is shown. However, also a different arrangement of pixels 51 is possible, such as e.g. a one-dimensional line or a scattered, arbitrary arrangement. The following reference to a two-dimensional pixel array is only chosen to ease the description.
[0266] The sensor device 10 comprises event detection circuitry 20 that is configured to detect as event data intensity changes above a predetermined threshold of the light received by each of a first subset S 1 of the pixels 51. The event detection circuitry 20 is basically constituted by the event detecting sections 52 as described above that are capable to receive the photocurrent of the pixels 51 and to detect events due to intensity changes of the received light that are larger than predetermined (but potentially dynamically adjustable) event thresholds. Although the event detection circuitry 20 has been depicted in Fig. 19 as a block adjacent to the pixels 51 this has to be understood as merely symbolic. As discussed above, the event detecting sections 52 of the event detection circuitry 20 may be arranged close or next to the pixels 51. However, the event detecting sections 52 could also be grouped together in a dedicated area on the sensor chip. In particular, the event detection circuitry 20 could be arranged on a different wafer / die than the pixels 51 and be vertically connected with the pixels 51.
[0267] The pixels 51 comprise the first subset SI. In Fig. 15 the pixels 51 of the first subset SI are indicated by vertical hatching. These pixels 51 are connected to the event detection circuitry 20, e.g. as described above, and allow event detection on the light received by them.
[0268] The pixels 51 further comprise a second subset S2 that is formed in Fig. 15 by the unhatched pixels 51. The pixels 51 of the second subset S2 are used to generate full intensity information, i.e. to generate data representing the (absolute) intensity of the light received by the pixels 51. Here, it is understood that appropriate color filters may be provided to each of the pixels 51 of the second subset S2 to allow generation of a color image. Color filters may also be provided to the pixels 51 of the first subset S 1.
[0269] The pixels 51 of the first subset S 1 and the pixels 51 of the second subset S2 are arranged such that they73960
[0270] 28
[0271] observe the same scene. This means that pixels 51 of the first subset S 1 and pixels 51 of the second subset S2 receive light from the same location of the scene. A pixel 51 of the first subset S 1 and a pixel 51 of the second subset S2 are associated with each other when they receive light from the same location.
[0272] The sensor device 10 comprises pixel signal generating circuitry 30 that is configured to generate a pixel signal indicating intensity values of the received light for each pixel 51 of the second subset S2 of the pixels 51. For example, the pixel signal generating circuitry 30 may be constituted by the pixel signal generating sections 53 as described above that can produce pixel signals indicating the received light intensity in an asynchronous (event triggered) or conventional, synchronous manner. In particular, the pixels 51 of the second subset S2 may operate according to the known principles of an active pixel sensor, APS. As is commonly known, the pixel signal generating circuitry 30 is configured to generate the pixel signals based on the amount of light received at the respective pixels 51 of the second subset S2 during an exposure period of said pixel 51.
[0273] Just as the event detection circuitry 20 also the pixel signal generating circuitry 30 may be formed either in a distributed manner, where e.g. one pixel signal generating section 53 is arranged next to or close to one pixel 51 of the second subset S2, or in a dedicated region of the sensor chip. The pixel signal generating circuitry 30 may also be considered to include all control, reset, and / or selection lines and the like that are necessary to operate the sensor device 10 as an APS that is capable to generate a consecutive stream of pixel signals indicating the intensity of the received light over a given time period.
[0274] The sensor device comprises further a control unit 40 that is configured to control the event detection circuitry 20 and the pixel signal generating circuitry 30. The control unit 40 may be any arrangement of circuitry that is capable to carry out the functions necessary to control the event detection circuitry 20 and the pixel signal generating circuitry 30. For example, the control unit 40 may be constituted by a processor. The control unit 40 may be part of the pixel section of the sensor device 10 and may be placed on the same die(s) as the other components of the sensor device 10. But the control unit 10 may also be arranged separately, e.g. on a separate die. The functions of the control unit 10 may be fully implemented in hardware, in software or may be implemented as a mixture of hardware and software functions. The control unit 40 may further be configured to provide the video stream generated by the pixels 51 of the second subset S2 and the event data generated by the pixels 51 of the first subset S 1 to the encoding device 100 to enable it to carry out the encoding process as described above with respect to Fig. 4.
[0275] In the schematic illustration of the sensor device 10 given in Fig. 19, which is also shown in Fig. 20 a) the pixels 51 in the first subset SI differ from the pixels 51 in the second subset S2. Thus, while the event detection circuitry 20 receives signals only from the first subset SI in this example, the pixel signal generating unit 30 receives signals only from the second subset S2.
[0276] As shown in Figs. 19 and 20 a) one implementation for such a division of pixels 51 is to arrange all pixels 51 in a single pixel array. Then, a specific number of pixels 51 can be designed for event detection, such as by providing circuitry as described above with respect to Fig. 17, for example for one pixel of a N x M pixel group, e.g. of a 2 x 2 pixel group as shown in Figs. 19 and 20 a). These pixels 51 form then the first73960
[0277] 29
[0278] subset SI of pixels 51. The remaining pixels 51 form the second subset S2 of pixels 51, and can be provided with conventional circuitry for a synchronous readout.
[0279] However, at least a part of the pixels 51 may belong to the first subset of pixels 51 as well as to the second subset of pixels 51. For example, all pixels 51 of a pixel array may function as pixels 51 of the first subset SI and the second subset S2 as schematically illustrated in Fig. 20 b). The pixels 51 may then be connected to the event detection circuitry 20 and the pixel signal circuitry 30 e.g. as described above with respect to Fig. 8, i.e. each pixel 51 can be selectively chosen to transfer its photocurrent to the event detection circuitry 20 or the pixel signal generation circuitry 30, e.g. in a time multiplexed manner. This shared pixel architecture allows to have the same spatial resolution for event detection and intensity information generation.
[0280] In the above description it was assumed that pixels 51 of both pixel subsets SI, S2 were part of the same pixel array. However, as schematically illustrated in Fig. 20 c) it is also conceivable that two different pixel arrays are used that each contain only pixels 51 of one pixel subset, either on different parts of a single die or even on different sensor chips. As long as both pixel arrays observe the same scene it will nevertheless be possible to reconstruct intensity information for the pixels 51 of the second subset S2 based on the events detected by the pixels 51 of the first subset S 1.
[0281] In general, any geometrical arrangement of pixels 51 with shared or divided functionality will be sufficient as long as event data and intensity information of the same scene can be obtained. If there is a mismatch in spatial resolution or if it is not possible to map the solid angles observed by pixels 51 in different subsets in a one to one manner, this will be solvable in principle by using interpolation techniques to adapt / align the two different data sets.
[0282] In Fig. 21 a decoding device 200 and its functions are schematically illustrated. The decoding device 200 can be used to decode encoded data that has been generated as described above with respect to Figs. 1 to 4. That is, the decoding device 200 generates a series of frames F3 containing image data with the first number of frames per second. As the encoding device 100, the decoding device 200 may be constituted by any computing device that is capable to carry out the functions described below. For example, the decoding device 200 may be or may be comprised in a computer, processor, microprocessor, CPU, GPU, FPGA, or ASIC. The decoding device 200 may be constituted by hardware or software or a mixture of hardware and software. In general, the structure, design or layout of the decoding device 200 is arbitrary as long as it is capable to carry out the functions described below.
[0283] As illustrated at the top of Fig. 21, the decoding device receives encoded data that contain a second number of frames F2 per second that is smaller than the first number of frames per second, motion vectors MV for each of the received frames F2 that allow, based on the motion vectors MV and the received frames F2, reconstructing missing frames MF that are temporally located between two consecutive, received frames F2, and for at least one (but generally for some or all) of the missing frames MF that are to be reconstructed, error data FU that indicate areas in a reconstruction of said at least one missing frame MF that need correction. That is, the decoding device 200 receives encoded data that has been generated73960
[0284] 30
[0285] as described above with respect to Figs. 1 to 4.
[0286] Then, the decoding device reconstructs missing frames MF based on the received frames F2 and the motion vectors MV, as also schematically illustrated in Fig. 21. As described above the reconstruction process is in principle known, e.g. from the MPEG or HEVC standards.
[0287] Based on the reconstructed missing frames MF and the received error data FL, the decoding device 200 corrects areas in missing frame MF that were indicated as erroneous in the error data FL. In this manner the decoding device 200 fdls in gaps between the received frame F2 with corrected, reconstructed frames. Together, these frames constitute the frames F3 of the decoded video stream having the first number of frames per second. The decoding device 200 can then output this video stream.
[0288] As for example illustrated schematically in Fig. 22, the areas indicated by the error data FL may be corrected based on surrounding areas that are not indicated by the error data FL as areas that need correction. Fig. 22 shows the reconstructed missing frame MF with erroneous blocks as indicated by the error data FL. Then, the decoding device may extrapolate the contents of all non-erroneous blocks such as to fill in or correct the erroneous blocks.
[0289] For example, as illustrated on the right-hand side of Fig. 22 by horizontal hatching, the encoding device 200 may totally discard or drop all data of the reconstructed frame in the areas indicated by the error data FL. The image data in these areas are then newly generated, e.g. by inpainting, based on surrounding areas that are not indicated by the error data FL as areas that need correction. As such interpolation processes, e.g. the process of inpainting, are in principle known a detailed description thereof can be omitted. Basically, the data of the discarded areas are recreated by interpolating between non-erroneous areas. Of course, instead of totally discarding the erroneous areas, the image data of these areas could be used to support the reconstruction process, e.g. in cases where an inpainting process gives no clear solution for specific pixel values. Reconstruction, and in particular inpainting, may be based on respectively trained artificial intelligence models.
[0290] Above various implementations of the basic idea to encode and decode image data with error data instead of residuals have been described. The basic methods underlying these implementations are summarized below with respect to Figs. 23 and 24.
[0291] Fig. 23 shows schematically a process flow of an encoding method for encoding a series of frames containing image data for a video stream having a first number of frames per second.
[0292] At S101 a second number of frames per second is stored that is smaller than the first number of frames per second. At SI 02 motion vectors MV are determined and stored for each of the stored frames F2, which motion vectors MV allow, based on the motion vectors MV and the stored frames F2, reconstructing missing frames MF that are temporally located between two consecutive, stored frames F2. At SI 03 for at least one (or some or all) of the missing frames MF that are to be reconstructed, error data FL are stored that indicate areas in a reconstruction of said at least one missing frame MF that need73960
[0293] 31
[0294] correction.
[0295] Fig. 24 shows schematically a process flow of a decoding method for decoding a series of frames containing image data for a video stream having a first number of frames per second.
[0296] At S201 a second number of frames per second is received that is smaller than the first number of frames per second. At S202 motion vectors MV for each of the received frames F2 are received that allow, based on the motion vectors MV and the received frames F2, reconstructing missing frames MF that are temporally located between two consecutive, received frames F2. And at S203 for at least one (or some or all) of the missing frames MF that are to be reconstructed, error data FL are received that indicate areas in a reconstruction of said at least one missing frame MF that need correction. At S204 missing frames MF are reconstructed based on the received frames F2 and the motion vectors MV. At S205 in the at least one reconstructed missing frame MF areas indicated by the error data FL are corrected.
[0297] In this manner it is possible to reduce the data amount of encoded data while at the same time ensuring that the quality of the decoded data remains high.
[0298] Fig. 25 is a block diagram illustrating an example of a schematic configuration of a smartphone 900 that may constitute or comprise a sensor device 10 as described above. The smartphone 900 is equipped with a processor 901, memory 902, storage 903, an external connection interface 904, a camera 906, a sensor 907, a microphone 908, an input device 909, a display device 910, a speaker 911, a radio communication interface 912, one or more antenna switches 915, one or more antennas 916, a bus 917, a battery 918, and an auxiliary controller 919.
[0299] The processor 901 may be a CPU or system -on-a-chip (SoC), for example, and controls functions in the application layer and other layers of the smartphone 900. The memory 902 includes RAM and ROM, and stores programs executed by the processor 901 as well as data. The storage 903 may include a storage medium such as semiconductor memory or a hard disk. The external connection interface 904 is an interface for connecting an externally attached device, such as a memory card or Universal Serial Bus (USB) device, to the smartphone 900. The processor may function as control unit 40.
[0300] The camera 906 includes an image sensor as described above. The sensor 907 may include a sensor group such as a positioning sensor, a gyro sensor, a geomagnetic sensor, and an acceleration sensor, for example. The microphone 908 converts audio input into the smartphone 900 into an audio signal. The input device 909 includes devices such as a touch sensorthat detects touches on a screen of the display device 910, a keypad, a keyboard, buttons, or switches, and receives operations or information input from a user. The display device 910 includes a screen such as a liquid crystal display (LCD) or an organic light-emitting diode (OLED) display, and displays an output image of the smartphone 900. The speaker 911 converts an audio signal output from the smartphone 900 into audio.
[0301] The radio communication interface 912 supports a cellular communication scheme such as LTE or LTE-Advanced, and executes radio communication. Typically, the radio communication interface 912 may73960
[0302] 32
[0303] include a BB processor 913, an RF circuit 914, and the like. The BB processor 913 may conduct processes such as encoding / decoding, modulation / demodulation, and multiplexing / demultiplexing, for example, and executes various signal processing for radio communication. Meanwhile, the RF circuit 914 may include components such as a mixer, a fdter, and an amp, and transmits or receives a radio signal via an antenna 916. The radio communication interface 912 may also be a one-chip module integrating the BB processor 913 and the RF circuit 914. The radio communication interface 912 may also include multiple BB processors 913 and multiple RF circuits 91. Note that although Fig. 25 illustrates an example of the radio communication interface 912 including multiple BB processors 913 and multiple RF circuits 914, the radio communication interface 912 may also include a single BB processor 913 or a single RF circuit 914.
[0304] Furthermore, in addition to a cellular communication scheme, the radio communication interface 912 may also support other types of radio communication schemes such as a short-range wireless communication scheme, a near field wireless communication scheme, or a wireless local area network (LAN) scheme. In this case, a BB processor 913 and an RF circuit 914 may be included for each radio communication scheme.
[0305] Each antenna switch 915 switches the destination of an antenna 916 among multiple circuits included in the radio communication interface 912 (for example, circuits for different radio communication schemes).
[0306] Each antenna 916 includes a single or multiple antenna elements (for example, multiple antenna elements constituting a MIMO antenna), and is used by the radio communication interface 912 to transmit and receive radio signals. The smartphone 900 may also include multiple antennas 916 as illustrated in Fig. 25. Note that although Fig. 25 illustrates an example of the smartphone 900 including multiple antennas 916, the smartphone 900 may also include a single antenna 916.
[0307] Furthermore, the smartphone 900 may also be equipped with an antenna 916 for each radio communication scheme. In this case, the antenna switch 915 may be omitted from the configuration of the smartphone 900.
[0308] The bus 917 interconnects the processor 901, the memory 902, the storage 903, the external connection interface 904, the camera 906, the sensor 907, the microphone 908, the input device 909, the display device 910, the speaker 911, the radio communication interface 912, and the auxiliary controller 919. The battery 918 supplies electric power to the respective blocks of the smartphone 900 illustrated in Fig. 25 via power supply lines partially illustrated with dashed lines in the drawing. The auxiliary controller 919 causes minimal functions of the smartphone 900 to operate while in a sleep mode, for example.
[0309] Fig. 26 is a block diagram depicting an example of schematic configuration of a vehicle control system as an example of a mobile body control system to which the technology according to an embodiment of the present disclosure can be applied.
[0310] The vehicle control system 12000 includes a plurality of electronic control units connected to each othervia a communication network 12001. In the example depicted in Fig. 26, the vehicle control system 12000 includes a driving system control unit 12010, a body system control unit 12020, an outside -vehicle information detecting unit 12030, an in-vehicle information detecting unit 12040, and an integrated control unit 12050. In addition, a microcomputer 12051, a sound / image output section 12052, and a vehicle-mounted network interface (I / F) 12053 are illustrated as a functional configuration of the integrated control unit 12050.
[0311] The driving system control unit 12010 controls the operation of devices related to the driving system of the vehicle in accordance with various kinds of programs. For example, the driving system control unit 12010 functions as a control device for a driving force generating device for generating the driving force of the vehicle, such as an internal combustion engine, a driving motor, or the like, a driving force transmitting mechanism for transmitting the driving force to wheels, a steering mechanism for adjusting the steering angle of the vehicle, a braking device for generating the braking force of the vehicle, and the like.
[0312] The body system control unit 12020 controls the operation of various kinds of devices provided to a vehicle body in accordance with various kinds of programs. For example, the body system control unit 12020 functions as a control device for a keyless entry system, a smart key system, a power window device, or various kinds of lamps such as a headlamp, a backup lamp, a brake lamp, a turn signal, a fog lamp, or the like. In this case, radio waves transmitted from a mobile device as an alternative to a key or signals of various kinds of switches can be input to the body system control unit 12020. The body system control unit 12020 receives these input radio waves or signals, and controls a door lock device, the power window device, the lamps, or the like of the vehicle.
[0313] The outside-vehicle information detecting unit 12030 detects information about the outside of the vehicle including the vehicle control system 12000. For example, the outside-vehicle information detecting unit 12030 is connected with an imaging section 12031. The outside-vehicle information detecting unit 12030 makes the imaging section 12031 image an image of the outside of the vehicle, and receives the imaged image. On the basis of the received image, the outside-vehicle information detecting unit 12030 may perform processing of detecting an object such as a human, a vehicle, an obstacle, a sign, a character on a road surface, or the like, or processing of detecting a distance thereto.
[0314] The imaging section 12031 is an optical sensor that receives light, and which outputs an electric signal corresponding to a received light amount of the light. The imaging section 12031 can output the electric signal as an image, or can output the electric signal as information about a measured distance. In addition, the light received by the imaging section 12031 may be visible light, or may be invisible light such as infrared rays or the like.
[0315] The in-vehicle information detecting unit 12040 detects information about the inside of the vehicle. The in-vehicle information detecting unit 12040 is, for example, connected with a driver state detecting section 12041 that detects the state of a driver. The driver state detecting section 12041, for example, includes a camera that images the driver. On the basis of detection information input from the driver state34
[0316] detecting section 12041, the in-vehicle information detecting unit 12040 may calculate a degree of fatigue of the driver or a degree of concentration of the driver, or may determine whether the driver is dozing.
[0317] The microcomputer 12051 can calculate a control target value for the driving force generating device, the steering mechanism, or the braking device on the basis of the information about the inside or outside of the vehicle which information is obtained by the outside-vehicle information detecting unit 12030 or the in-vehicle information detecting unit 12040, and output a control command to the driving system control unit 12010. For example, the microcomputer 12051 can perform cooperative control intended to implement functions of an advanced driver assistance system (ADAS) which functions include collision avoidance or shock mitigation for the vehicle, following driving based on a following distance, vehicle speed maintaining driving, a warning of collision of the vehicle, a warning of deviation of the vehicle from a lane, or the like.
[0318] In addition, the microcomputer 12051 can perform cooperative control intended for automatic driving, which makes the vehicle to travel autonomously without depending on the operation of the driver, or the like, by controlling the driving force generating device, the steering mechanism, the braking device, or the like on the basis of the information about the outside or inside of the vehicle which information is obtained by the outside-vehicle information detecting unit 12030 or the in-vehicle information detecting unit 12040.
[0319] In addition, the microcomputer 12051 can output a control command to the body system control unit 12020 on the basis of the information about the outside of the vehicle which information is obtained by the outside-vehicle information detecting unit 12030. For example, the microcomputer 12051 can perform cooperative control intended to prevent a glare by controlling the headlamp so as to change from a high beam to a low beam, for example, in accordance with the position of a preceding vehicle or an oncoming vehicle detected by the outside-vehicle information detecting unit 12030.
[0320] The sound / image output section 12052 transmits an output signal of at least one of a sound and an image to an output device capable of visually or auditorily notifying information to an occupant of the vehicle or the outside of the vehicle. In the example ofFig. 26, an audio speaker 12061, a display section 12062, and an instrument panel 12063 are illustrated as the output device. The display section 12062 may, for example, include at least one of an on-board display and a head-up display.
[0321] Fig. 27 is a diagram depicting an example of the installation position of the imaging section 12031.
[0322] In Fig. 27, the imaging section 12031 includes imaging sections 12101, 12102, 12103, 12104, and 12105.
[0323] The imaging sections 12101, 12102, 12103, 12104, and 12105 are, for example, disposed at positions on a front nose, sideview mirrors, a rear bumper, and a back door of the vehicle 12100 as well as a position on an upper portion of a windshield within the interior of the vehicle. The imaging section 12101 provided to the front nose and the imaging section 12105 provided to the upper portion of the windshield within the interior of the vehicle obtain mainly an image of the front of the vehicle 12100. The imaging35
[0324] sections 12102 and 12103 provided to the sideview mirrors obtain mainly an image of the sides of the vehicle 12100. The imaging section 12104 provided to the rear bumper or the back door obtains mainly an image of the rear of the vehicle 12100. The imaging section 12105 provided to the upper portion of the windshield within the interior of the vehicle is used mainly to detect a preceding vehicle, a pedestrian, an obstacle, a signal, a traffic sign, a lane, or the like.
[0325] Incidentally, Fig. 27 depicts an example of photographing ranges of the imaging sections 12101 to 12104. An imaging range 12111 represents the imaging range of the imaging section 12101 provided to the front nose. Imaging ranges 12112 and 12113 respectively represent the imaging ranges of the imaging sections 12102 and 12103 provided to the sideview mirrors. An imaging range 12114 represents the imaging range of the imaging section 12104 provided to the rear bumper or the back door. A bird’s-eye image of the vehicle 12100 as viewed from above is obtained by superimposing image data imaged by the imaging sections 12101 to 12104, for example.
[0326] At least one of the imaging sections 12101 to 12104 may have a function of obtaining distance information. For example, at least one of the imaging sections 12101 to 12104 may be a stereo camera constituted of a plurality of imaging elements, or may be an imaging element having pixels for phase difference detection.
[0327] For example, the microcomputer 12051 can determine a distance to each three-dimensional object within the imaging ranges 12111 to 12114 and a temporal change in the distance (relative speed with respect to the vehicle 12100) on the basis of the distance information obtained from the imaging sections 12101 to 12104, and thereby extract, as a preceding vehicle, a nearest three-dimensional object in particular that is present on a traveling path of the vehicle 12100 and which travels in substantially the same direction as the vehicle 12100 at a predetermined speed (for example, equal to or more than 0 km / hour). Further, the microcomputer 12051 can set a following distance to be maintained in front of a preceding vehicle in advance, and perform automatic brake control (including following stop control), automatic acceleration control (including following start control), or the like. It is thus possible to perform cooperative control intended for automatic driving that makes the vehicle travel autonomously without depending on the operation of the driver or the like.
[0328] For example, the microcomputer 12051 can classify three-dimensional object data on three-dimensional objects into three-dimensional object data of a two-wheeled vehicle, a standard-sized vehicle, a largesized vehicle, a pedestrian, a utility pole, and other three-dimensional objects on the basis of the distance information obtained from the imaging sections 12101 to 12104, extract the classified three-dimensional object data, and use the extracted three-dimensional object data for automatic avoidance of an obstacle. For example, the microcomputer 12051 identifies obstacles around the vehicle 12100 as obstacles that the driver of the vehicle 12100 can recognize visually and obstacles that are difficult for the driver of the vehicle 12100 to recognize visually. Then, the microcomputer 12051 determines a collision risk indicating a risk of collision with each obstacle. In a situation in which the collision risk is equal to or higher than a set value and there is thus a possibility of collision, the microcomputer 12051 outputs a warning to the driver via the audio speaker 12061 or the display section 12062, and performs forced73960
[0329] 36
[0330] deceleration or avoidance steering via the driving system control unit 12010. The microcomputer 12051 can thereby assist in driving to avoid collision.
[0331] At least one of the imaging sections 12101 to 12104 may be an infrared camera that detects infrared rays. The microcomputer 12051 can, for example, recognize a pedestrian by determining whether or not there is a pedestrian in imaged images of the imaging sections 12101 to 12104. Such recognition of a pedestrian is, for example, performed by a procedure of extracting characteristic points in the imaged images of the imaging sections 12101 to 12104 as infrared cameras and a procedure of determining whether or not it is the pedestrian by performing pattern matching processing on a series of characteristic points representing the contour of the object. When the microcomputer 12051 determines that there is a pedestrian in the imaged images of the imaging sections 12101 to 12104, and thus recognizes the pedestrian, the sound / image output section 12052 controls the display section 12062 so that a square contour line for emphasis is displayed so as to be superimposed on the recognized pedestrian. The sound / image output section 12052 may also control the display section 12062 so that an icon or the like representing the pedestrian is displayed at a desired position.
[0332] An example of the vehicle control system to which the technology according to the present disclosure is applicable has been described above. The technology according to the present disclosure is applicable to the imaging section 12031 among the above-mentioned configurations. Specifically, the sensor device 10 is applicable to the imaging section 12031. The imaging section 12031 to which the technology according to the present disclosure has been applied flexibly acquires event data and performs data processing on the event data, thereby being capable of providing appropriate driving assistance.
[0333] Further possible implementations of the sensor device 10 are mobile devices 3000 such as cell phones, tablets, smart watches and the like as shown in Fig. 28A or head-mounted displays 4000 as shown in Fig.
[0334] 28B. Further, the sensor device 10 is useable in augmented and / or virtual reality applications / cameras or in surveillance systems like 360° cameras.
[0335] Note that, the embodiments of the present technology are not limited to the above-mentioned embodiment, and various modifications can be made without departing from the gist of the present technology.
[0336] Further, the effects described herein are only exemplary and not limited, and other effects may be provided.
[0337] Note that, the present technology can also take the following configurations.
[0338] [1] An encoding device (100) for encoding a series of frames (F3) containing image data for a video stream having a first number of frames per second, the encoding device (100) being configured to store a second number of frames (F2) per second that is smaller than the first number of frames per second;
[0339] determine and store motion vectors (MV) for each of the stored frames (F2) that allow, based on the motion vectors (MV) and the stored frames (F2), reconstructing missing frames (MF) that are73960
[0340] 37
[0341] temporally located between two consecutive, stored frames (F2); and
[0342] store, for at least one of the missing frames (MF) that are to be reconstructed, error data (FL) that indicate areas in a reconstruction of said at least one missing frame (MF) that need correction.
[0343] [2] The encoding device (100) according to [1], wherein
[0344] the reconstruction of said at least one missing frame (MF) is divided into blocks (B) and the error data (FL) are flags that indicate for each block (B) whether the respective block (B) needs correction.
[0345] [3] The encoding device (100) according to [1] or [2], wherein
[0346] the stored motion vectors (MV) contain forward motion vectors that indicate a change in image data from a stored frame (F2) to a future missing frame (MF), and backward motion vectors that indicate a change in image data from a past missing frame (MF) to a stored frame (F2).
[0347] [4] The encoding device (100) according to any one of [1] to [3], wherein the encoding device (100) is further configured to
[0348] receive a video stream (Fl) having the first number of frames per second;
[0349] designate frames of the received video stream for dropping and designating the remaining frames (F2) for storing as the stored frames (F2), while the frames to be dropped constitute the missing frames (MF);
[0350] determine the motion vectors (MV) by comparing frames (F2) for storing with frames to be dropped;
[0351] reconstruct missing frames (MF) based on the frames for storing (F2) and the determined motion vectors (MV);
[0352] compare the reconstructed missing frames (MF) with the respective frames to be dropped to determine the error data (FL);
[0353] drop the respectively designated frames from the received video stream to store only the remaining frames as the stored frames (F2);
[0354] drop the comparison between the reconstructed missing frames (MF) and the respective frames to be dropped to store only the error data (FL).
[0355] [5] The encoding device (100) according to any one of [1] to [4], wherein the encoding device (100) is further configured to
[0356] receive a video stream (Fl) having the second number of frames per second as the frames (F2) to be stored;
[0357] receive event data (ED) indicating a temporal stream of events, where each event indicates intensity changes above a predetermined threshold of the light received for one of a plurality of pixels (51) that constitute the frames;
[0358] determine from the event data (ED) and the frames (F2) of the received video stream (Fl) the motion vectors (MV);
[0359] reconstruct missing frames (MF) based on the stored frames (F2) and the determined motion vectors (MV);
[0360] reconstruct missing frames (MF) based on the stored frames (F2) by using an artificial73960
[0361] 38
[0362] intelligence model that has been trained for this reconstruction;
[0363] compare missing frames (MF) for the same temporal position in the video stream that have been reconstructed based on the motion vectors (MV) and based on the artificial intelligence model to determine the error data (FL);
[0364] drop the reconstructed missing frames (MF) and the comparison between the reconstructed missing frames to only store the error data (FL).
[0365] [6] The encoding device according to [5], wherein
[0366] the determination of motion vectors is performed by an artificial intelligence model that has been trained for this determination.
[0367] [7] A sensor device (10) comprising:
[0368] the encoding device (100) according to claim [5] or [6];
[0369] a plurality of pixels (51) each configured to receive light from a scene and to perform photoelectric conversion to generate an electrical signal;
[0370] event detection circuitry (20) that is configured to detect as event data intensity changes above a predetermined threshold of the light received by each of a first subset (S 1) of the pixels (51);
[0371] pixel signal generating circuitry (30) that is configured to generate a pixel signal indicating intensity values of the received light for each pixel (51) of a second subset (S2) of the pixels (51); and a control unit (40) that is configured to provide the pixel signals as the video stream and the event data to the encoding device (100).
[0372] [8] The sensor device (10) according to [7], wherein
[0373] the pixels (51) are arranged in a two-dimensional pixel array; and
[0374] at least a part of the pixels (51) belongs to both, the first subset (SI) of pixels (51) and the second subset (S2) of pixels (51).
[0375] [9] The sensor device (10) according to [7], wherein
[0376] the pixels (51) are arranged in a two-dimensional pixel array; and
[0377] pixels (51) in the first subset (SI) of pixels (51) are different from pixels (51) in the second subset (S2) of pixels (51).
[0378]
[0010] The sensor device (10) according to [7], wherein
[0379] the pixels (51) in the first subset (SI) of pixels (51) are arranged in a first two-dimensional pixel array; and
[0380] the pixels (51) in the second subset (S2) of pixels (51) are arranged in a different, second two-dimensional pixel array.
[0381]
[0011] A decoding device (200) for decoding a series of frames (F3) containing image data for a video stream having a first number of frames per second, the decoding device (200) being configured to receive a second number of frames (F2) per second that is smaller than the first number of frames per second, motion vectors (MV) for each of the received frames (F2) that allow, based on the39
[0382] motion vectors (MV) and the received frames (F2), reconstructing missing frames (MF) that are temporally located between two consecutive, received frames (F2), and for at least one of the missing frames (MF) that are to be reconstructed, error data (FL) that indicate areas in a reconstruction of said at least one missing frame (MF) that need correction;
[0383] reconstruct missing frames (MF) based on the received frames (F2) and the motion vectors (MV); and
[0384] correct in the at least one reconstructed missing frame (MF) areas indicated by the error data (FL).
[0385]
[0012] The decoding device (200) according to
[0011] , wherein
[0386] the areas indicated by the error data (FL) are corrected based on surrounding areas that are not indicated by the error data (FL) as areas that need correction.
[0387]
[0013] The decoding device (200) according to
[0012] , wherein
[0388] the areas indicated by the error data (FL) are corrected by dropping image data of these areas and newly generating image data of these areas by inpainting based on surrounding areas that are not indicated by the error data (FL) as areas that need correction.
[0389]
[0014] An encoding method for encoding a series of frames (F3) containing image data for a video stream having a first number of frames per second, the encoding method comprising:
[0390] storing a second number of frames (F2) per second that is smaller than the first number of frames per second;
[0391] determining and storing motion vectors (MV) for each of the stored frames (F2) that allow, based on the motion vectors (MV) and the stored frames (F2), reconstructing missing frames (MF) that are temporally located between two consecutive, stored frames (F2); and
[0392] storing, for at least one of the missing frames (MF) that are to be reconstructed, error data (FL) that indicate areas in a reconstruction of said at least one missing frame (MF) that need correction.
[0393]
[0015] A decoding method for decoding a series of frames (F3) containing image data for a video stream having a first number of frames per second, the decoding method comprising:
[0394] receiving a second number of frames (F2) per second that is smaller than the first number of frames per second, motion vectors (MV) for each of the received frames (F2) that allow, based on the motion vectors (MV) and the received frames (F2), reconstructing missing frames (MF) that are temporally located between two consecutive, received frames (F2), and, for at least one of the missing frames (MF) that are to be reconstructed, error data (FL) that indicate areas in a reconstruction of said at least one missing frame (MF) that need correction;
[0395] reconstructing missing frames (MF) based on the received frames (F2) and the motion vectors (MV); and
[0396] correcting in the at least one reconstructed missing frame (MF) areas indicated by the error data (FL).
Claims
40CLAIMS1. An encoding device for encoding a series of frames containing image data for a video stream having a first number of frames per second, the encoding device being configured tostore a second number of frames per second that is smaller than the first number of frames per second;determine and store motion vectors for each of the stored frames that allow, based on the motion vectors and the stored frames, reconstructing missing frames that are temporally located between two consecutive, stored frames; andstore, for at least one of the missing frames that are to be reconstructed, error data that indicate areas in a reconstruction of said at least one missing frame that need correction.
2. The encoding device according to claim 1, whereinthe reconstruction of said at least one missing frame is divided into blocks and the error data are flags that indicate for each block whether the respective block needs correction.
3. The encoding device according to claim 1, whereinthe stored motion vectors contain forward motion vectors that indicate a change in image data from a stored frame to a future missing frame, and backward motion vectors that indicate a change in image data from a past missing frame to a stored frame.
4. The encoding device according to claim 1, wherein the encoding device is further configured toreceive a video stream having the first number of frames per second;designate frames of the received video stream for dropping and designating the remaining frames for storing as the stored frames, while the frames to be dropped constitute the missing frames;determine the motion vectors by comparing frames for storing with frames to be dropped; reconstruct missing frames based on the frames for storing and the determined motion vectors;compare the reconstructed missing frames with the respective frames to be dropped to determine the error data;drop the respectively designated frames from the received video stream to store only the remaining frames as the stored frames;drop the comparison between the reconstructed missing frames and the respective frames to be dropped to store only the error data.
5. The encoding device according to claim 1, wherein the encoding device is further configured toreceive a video stream having the second number of frames per second as the frames to be stored;receive event data indicating a temporal stream of events, where each event indicates intensity changes above a predetermined threshold of the light received for one of a plurality of pixels that7396041constitute the frames;determine from the event data and the frames of the received video stream the motion vectors;reconstruct missing frames based on the stored frames and the determined motion vectors; reconstruct missing frames based on the stored frames by using an artificial intelligence model that has been trained for this reconstruction;compare missing frames for the same temporal position in the video stream that have been reconstructed based on the motion vectors and based on the artificial intelligence model to determine the error data;drop the reconstructed missing frames and the comparison between the reconstructed missing frames to only store the error data.
6. The encoding device according to claim 5, whereinthe determination of motion vectors is performed by an artificial intelligence model that has been trained for this determination.
7. A sensor device comprising:the encoding device according to claim 5;a plurality of pixels each configured to receive light from a scene and to perform photoelectric conversion to generate an electrical signal;event detection circuitry that is configured to detect as event data intensity changes above a predetermined threshold of the light received by each of a first subset of the pixels;pixel signal generating circuitry that is configured to generate a pixel signal indicating intensity values of the received light for each pixel of a second subset of the pixels; anda control unit that is configured to provide the pixel signals as the video stream and the event data to the encoding device.
8. The sensor device according to claim 7, whereinthe pixels are arranged in a two-dimensional pixel array; andat least a part of the pixels belongs to both, the first subset of pixels and the second subset of pixels.
9. The sensor device according to claim 7, whereinthe pixels are arranged in a two-dimensional pixel array; andpixels in the first subset of pixels are different from pixels in the second subset of pixels.
10. The sensor device according to claim 7, whereinthe pixels in the first subset of pixels are arranged in a first two-dimensional pixel array; and the pixels in the second subset of pixels are arranged in a different, second two-dimensional pixel array.
11. A decoding device for decoding a series of frames containing image data for a video streamhaving a first number of frames per second, the decoding device being configured toreceive a second number of frames per second that is smaller than the first number of frames per second, motion vectors for each of the received frames that allow, based on the motion vectors and the received frames, reconstructing missing frames that are temporally located between two consecutive, received frames, and for at least one of the missing frames that are to be reconstructed, error data that indicate areas in a reconstruction of said at least one missing frame that need correction;reconstruct missing frames based on the received frames and the motion vectors; and correct in the at least one reconstructed missing frame areas indicated by the error data.
12. The decoding device according to claim 11, whereinthe areas indicated by the error data are corrected based on surrounding areas that are not indicated by the error data as areas that need correction.
13. The decoding device according to claim 12, whereinthe areas indicated by the error data are corrected by dropping image data of these areas and newly generating image data of these areas by inpainting based on surrounding areas that are not indicated by the error data as areas that need correction.
14. An encoding method for encoding a series of frames containing image data for a video stream having a first number of frames per second, the encoding method comprising:storing a second number of frames per second that is smaller than the first number of frames per second;determining and storing motion vectors for each of the stored frames that allow, based on the motion vectors and the stored frames, reconstructing missing frames that are temporally located between two consecutive, stored frames; andstoring, for at least one of the missing frames that are to be reconstructed, error data that indicate areas in a reconstruction of said at least one missing frame that need correction.
15. A decoding method for decoding a series of frames containing image data for a video stream having a first number of frames per second, the decoding method comprising:receiving a second number of frames per second that is smaller than the first number of frames per second, motion vectors for each of the received frames that allow, based on the motion vectors and the received frames, reconstructing missing frames that are temporally located between two consecutive, received frames, and, for at least one of the missing frames that are to be reconstructed, error data that indicate areas in a reconstruction of said at least one missing frame that need correction;reconstructing missing frames based on the received frames and the motion vectors; and correcting in the at least one reconstructed missing frame areas indicated by the error data.