Dynamic Vision Sensor Color Camera

By independently operating each pixel in a DVS color camera and transmitting a multi-time variation characterized light pattern, the problem of traditional cameras generating redundant data and insensitive to light colors in a stationary scene is solved, and efficient color image capture is achieved.

CN118302718BActive Publication Date: 2025-05-06RAMOT AT TEL AVIV UNIVERSITY LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202280075909.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-11-15
Filing Date
2022-11-14
Publication Date
2025-05-06
Estimated Expiration
2042-11-14

AI Technical Summary

Technical Problem

Traditional cameras generate a large amount of redundant data when recording stationary and unchanged scenes, and are insensitive to the color of light, making it difficult to provide efficient color images.

Method used

Using a dynamic vision sensor (DVS) color camera (DVS-CCam), reflected light is collected and processed to provide a color image by operating each pixel independently in a DVS photoelectric sensor and emitting light patterns characterized by multiple temporal changes in intensity and color.

Benefits of technology

It significantly reduces the redundancy of image data, improves the utilization efficiency of energy, bandwidth and memory resources, and can provide color images efficiently.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118302718B_ABST
    Figure CN118302718B_ABST
Patent Text Reader

Abstract

A dynamic vision sensor (DVS) color camera (DVS‑CCam) is operable to acquire a color image of a scene, the DVS‑CCam comprising a DVS photosensor having DVS pixels, an illuminator operable to emit a light pattern characterized by temporal variations in intensity and color to illuminate the scene, an optical system configured to collect light reflected from features in the scene by the light pattern emitted by the illuminator and to focus the reflected light onto a plurality of pixels of the DVS photosensor, and a processor configured to process DVS signals generated by the DVS pixels in response to the temporal variations in the reflected light to provide a color image of the scene.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 279,255, filed on November 15, 2021, which is expressly incorporated herein by reference in its entirety. Technical Field

[0003] The subject matter disclosed herein relates generally to imaging systems and, more particularly, to methods and apparatus for providing a dynamic vision sensor (DVS) camera operable to provide a color image of a scene. Background Art

[0004] Conventional cameras commonly encountered in almost all everyday appliances and devices (e.g., mobile phones, laptops, and tablets) employ photoelectric sensors in which all pixels are typically controlled to record light intensities from a scene imaged by the camera at substantially the same time. All pixels are read out from the photoelectric sensor to the same frame to provide a black-and-white or color contrast image of the scene. The camera operates to provide a video of the scene from which feature changes and motion in the scene can be determined by acquiring a sequence of contrast images of the scene at a sufficiently large frame rate per second (fps). Since data from all pixels are used to provide each image in a sequence of images in the video, pixels that image image features of a static and unchanged scene repeatedly contribute the same data to each image of the video. Therefore, although current conventional cameras can be considered to provide video of unprecedented quality, they generate large amounts of redundant data and are relatively inefficient in the use of energy, bandwidth, and memory resources.

[0005] On the other hand, a dynamic vision system (DVS) camera, instead of controlling all pixels in a DVS photoelectric sensor simultaneously to record light from a scene and then reading out the signals of all pixels into the same frame, operates each pixel in a DVS photoelectric sensor independently. Each pixel in a DVS generates and transmits an image data signal (hereinafter referred to as a DVS image signal) to a controller for processing to provide an image only when the change in the intensity of the incident light from the scene is greater than a predetermined threshold intensity. The image data signal includes the address of the pixel and optionally includes a sign indicating whether the change in the sensed intensity is positive or negative. DVS pixels that image static and unchanged features of a scene do not generate an image data signal. Therefore, DVS cameras have a higher sensitivity to changes in the scene, significantly reduce the redundancy of the image data generated by the camera, and more efficiently utilize the energy, bandwidth, and memory resources of the camera. However, although DVS cameras have a high sensitivity to light intensity, they are relatively insensitive to the color of light and are relatively unsuitable for providing color data of images of the scenes they image. Summary of the invention

[0006] In various embodiments, a dynamic vision sensor (DVS) color camera (DVS-CCam) is provided, which is operable to acquire a color image of a scene, the DVS-CCam comprising a DVS photosensor having a plurality of DVS pixels, an illuminator operable to emit a light pattern characterized by a plurality of temporal variations in intensity and color to illuminate the scene, an optical system configured to collect light reflected by the light pattern emitted by the illuminator from a plurality of features in the scene and to focus the reflected light onto the plurality of pixels of the DVS photosensor, and a processor configured to process a plurality of DVS signals generated by the plurality of DVS pixels in response to the plurality of temporal variations in the reflected light to provide a color image of the scene.

[0007] In various embodiments, a method of acquiring a color image of a scene using a dynamic vision sensor (DVS) color camera (DVS-CCam) including a DVS photosensor is provided, the DVS photosensor including a plurality of DVS pixels, the method comprising: emitting a light pattern characterized by a plurality of temporal variations in intensity and color to illuminate a scene, collecting light reflected by the emitted light pattern from a plurality of features in the scene and focusing the reflected light onto the plurality of DVS pixels, and processing a plurality of DVS signals generated by the plurality of DVS pixels in response to the plurality of temporal variations in the reflected light to provide a color image of the scene.

[0008] In some embodiments, the emitted light pattern comprises a plurality of light pulses that are temporally consecutive.

[0009] In some embodiments, the emitted light pattern comprises a plurality of light pulses that are discrete and temporally separated.

[0010] In some embodiments, the emitted light pattern includes red light, green light, and blue light.

[0011] In some embodiments, each pixel of the plurality of DVS pixels generates a DVS image signal stream responsive to a variation in light intensity during the plurality of temporal variations in the reflected light.

[0012] In some embodiments, a pixel response value is determined for each pixel of the plurality of DVS pixels, the pixel response value being a function of a color of a feature in the scene imaged by the pixel.

[0013] In some embodiments, the pixel response value is a sum of a number of DVS signals in at least a portion of a DVS image signal stream generated by the pixel in response to the temporal variation of the reflected light.

[0014] In some embodiments, the pixel response values ​​are used to determine a best fit color that minimizes a linear mean square estimate of the pixel response values.

[0015] In some embodiments, a neural network (NN) is also included.

[0016] In some embodiments, the neural network comprises a convolutional neural network (CNN).

[0017] In some embodiments, the neural network provides a color responsive to a feature of the feature vector based on a pixel response value associated with the pixel.

[0018] In some embodiments, the neural network includes a contraction portion, an expansion portion, and a plurality of layers connecting the contraction portion and the expansion portion.

[0019] In some embodiments, the constriction reduces the spatial dimension and increases the number of channels.

[0020] In some embodiments, the expansion portion increases the spatial dimension and reduces the number of channels.

[0021] In some embodiments, the reduced number of channels includes three channels.

[0022] In some embodiments, the plurality of layers are configured to add a plurality of weights to the NN.

[0023] In some embodiments, the loss function includes a weighted average of L1 and an SSIM index.

[0024] In some embodiments, the DVS photosensor comprises an event-based photosensor.

[0025] In some embodiments, the neural network is configured to extract spectral components from visible light other than RGB light for generating a plurality of hyperspectral images.

[0026] In some embodiments, information associated with an acquired color image is used to reconstruct a next color image. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Non-limiting examples of the embodiments disclosed herein are described below with reference to the drawings listed after this paragraph. The drawings and descriptions are intended to explain and illustrate the embodiments disclosed herein and should not be considered limiting in any way. The same elements in different drawings may be represented by the same numbers. The elements in the drawings are not necessarily drawn to scale. In the drawings:

[0028] Figure 1 A DVS camera suitable for reproducing color images is schematically shown according to an embodiment of the present invention;

[0029] Figure 2 A DVS camera for reproducing a color image and configured to emit a plurality of discrete and temporally separated light pulses is schematically shown according to an embodiment of the present invention;

[0030] Figure 3 A DVS camera for reproducing a color image and comprising a neural network is schematically shown according to an embodiment of the present invention;

[0031] Figure 4 The results of linear reconstruction of a color image using a DVS camera system in a test setup are shown according to an embodiment of the present invention. The color image is converted into a black and white image, where different shades of gray represent different colors. This conversion is also applied to the following figures;

[0032] Figure 5 Several results using a DVS system including a convolutional neural network (CNN) in a test setup for color reconstruction of a three-dimensional (3D) scene are shown according to an embodiment of the present invention; and

[0033] Figure 6 An embodiment according to the present invention shows that when exposed to the same ambient brightness as used in the training of the CNN, Figure 5 Results on the robustness of the DVS camera system to distance variations and its impact on color reconstruction. DETAILED DESCRIPTION

[0034] One aspect of embodiments of the present disclosure is directed to providing a DVS camera, also referred to as a DVS color camera (DVS-CCam), operable to provide a color image of a scene imaged by the camera.

[0035] According to one embodiment, DVS-CCam optionally includes a DVS photoelectric sensor, an illuminator operable to emit a temporal light pattern, wherein the intensity and / or color of the light pattern of the temporal light pattern varies over time, a processor, and a controller that controls the illuminator and the processor. The temporal light pattern can be a light pattern characterized by substantially continuously changing intensity and / or color, or a temporally continuous or temporally separated sequence of light pulses of different intensities and / or different colors. Optionally, DVS-CCam can include any type of event-based camera sensor, including neuromorphic camera sensors, silicon retina camera sensors, and other sensors suitable for performing the functions described herein, and may not necessarily be limited to dynamic vision sensors.

[0036] When imaging a scene, the controller may, for example, control an illuminator (also referred to hereinafter as a "flicker illuminator" or "flicker") to emit a temporal light pattern comprising a sequence of light pulses of different intensities and different colors to illuminate the scene. At the rising and falling edges of each emitted light pulse, or at the transitions between substantially temporally consecutive emitted light pulses of different colors and / or amplitudes, each pixel in the DVS photosensor senses changes in the intensity and / or color of the incident light as changes in intensity. In response to the sensed intensity changes, each pixel generates a stream of DVS image signals, each DVS image signal including the address of the pixel for processing by the processor. The rate at which a given pixel generates DVS image signals and the total number of DVS images generated by a given pixel in response to the sensed changes are a function of the color, amplitude, and / or shape of the emitted light pulses or the transitions between consecutive light pulses that generate the sensed changes, the impulse response function of the pixel, and the color of the features in the scene imaged by the given pixel.

[0037] In one embodiment, a processor processes a DVS image signal stream generated by a pixel in response to a change generated by an emitted light pulse or a transition between consecutive light pulses to determine a value (hereinafter referred to as a "DVS pixel response") that is a function of the color of a feature imaged by the pixel. Optionally, for example, the DVS pixel response determined by the processor is the sum of the number of DVS signals in at least a portion of the DVS image signal stream generated by the pixel in response to the change. The processor processes the DVS pixel responses of the pixel in response to a plurality of sensed intensity changes generated by a plurality of light pulses emitted by a scintillator to illuminate a scene to determine a color of the feature that best fits the pixel DVS pixel response. Optionally, the best fit color is a function of the color that minimizes a linear mean square estimator of the DVS pixel response and a prediction of the DVS pixel response based on the feature color.

[0038] In one embodiment, the best fit color is determined by a neural network that provides the color of the feature in response to a feature vector based on the DVS response for a given pixel. Any of a variety of color representations, such as CIE color space, HSL (hue, saturation, brightness), HSV (hue, saturation, value), or RGB (red, blue, green) can be used to characterize the color of the feature. In one embodiment, the processor uses the best fit color determined for each of the plurality of pixels in the DVS photosensor to construct a color image of the scene.

[0039] In one embodiment, the constructed color image of the scene (hereinafter also referred to as a "reconstructed" color image) can be used to reconstruct subsequent frames in the video. Some (or optionally all) of the information associated with the colors in the image at a given time t0 can be used to reconstruct a new image at a later time t1.

[0040] Figure 1An embodiment of a DVS-CCam 20 is schematically shown. The DVS-CCam 20 optionally includes a scintillator 22 operable to illuminate a scene with a light pattern exhibiting variations in intensity and color, a DVS photosensor 23 including a plurality of light-sensitive pixels 24, and an optical system represented by a lens 26 that focuses light reflected by the light pattern emitted by the scintillator 22 from a plurality of features in the scene onto the DVS photosensor. The light-sensitive pixels 24 are configured to generate DVS image signals in response to sensing variations in the intensity of light incident on the pixels. A controller 28 controls the operation of the components of the DVS-CCam 20. According to an embodiment of the present disclosure, the controller may include a processor and a memory for processing the DVS image signals generated by the pixels 24 in response to variations in the intensity of the incident light sensed by the pixels.

[0041] exist Figure 1 In the figure, according to one embodiment of the present disclosure, the DVS-CCam 20 schematically shows imaging a scene represented by an ellipse 40, and the controller 28 is shown controlling the scintillator 22 to optionally emit a sequence of multiple light pulses 50 that are substantially continuous in time to illuminate the scene. The emitted light pulses 50 include light pulses characterized by red (R), green (G), and blue (B) light of different intensities and different (optionally relatively narrow) frequency bands. Optionally, the light pulses 50 are emitted in three light pulse groups 51, 52, and 53. Each light pulse group optionally includes emitted R, G, and B light pulses 50, which are respectively distinguished by unique hatching and marked by a group number and a subscript r, g, or b, the subscript r, g, or b identifying the color R, G, or B of the light in the light pulse. For example, the emitted R, G, and B light pulses in group 51 are represented by 51r, 51g, and 51b, respectively. Light pulses in the same group, optionally as Figure 1 The light pulses in the groups shown have substantially the same intensity, while the light pulses in different groups have different intensities.

[0042] A plurality of features in the scene 40 reflect the light of the emitted light pulse 50, and in the reflected light pulse, the reflected light is imaged by the optical system 26 on a plurality of DVS pixels 24 of the photosensor 23 to image the features respectively. The light intensity in the reflected light pulse reflected by the features in the scene 40 from the emitted light pulse 50 is a function of the light intensity of the emitted light pulse, the light color of the emitted light pulse, and the color of the feature. Assuming that the frequency band of the emitted light pulse 50 is sufficiently narrow, the color of the light reflected by a given feature in the scene 40 from a given emitted light pulse 50 is substantially the same as the color of the given emitted light pulse.

[0043] For example, features in scene 40 indicated by circles 41R and 42G are assumed to be substantially red and substantially green features respectively. These features are schematically shown as reflected light reflected from emitted light pulses 50 in a sequence of light pulses corresponding to reflections of the emitted sequence of light pulses.

[0044] The reflected light pulses formed by the characteristic 41R for the emission light pulse groups 51, 52 and 53 in the emission light pulse 50 are marked by the reflected light pulse group numbers 41 / 51, 41 / 52 and 41 / 53, respectively, and the subscripts r, g or b are used to indicate the corresponding colors of the reflected light pulses and the corresponding colors of the emission light pulses reflected by the reflected light pulses. For example, the reflected light pulses formed by the characteristic 41R for the emission light pulses 51r, 51g and 51b in the emission light pulse group 51 are marked as reflected light pulses 41 / 51r, 41 / 51g and 41 / 51b, respectively. The reflected light pulses formed by the characteristic 41R for the emission light pulses 52r, 52g and 52b in the emission light pulse group 52 are marked as reflected light pulses 41 / 52r, 41 / 52g and 41 / 52b, respectively. The reflected light pulses formed by the characteristic 41R for the reflection can be collectively referred to as reflected light pulses 41 / 50. The reflected light pulses reflected by the feature 41R can be generally referred to by the reference numerals 41 / 51, 41 / 52 or 41 / 53 of the reflected light pulse group to which the reference light pulse belongs in the same light pulse group. Figure 1 , the light intensity in each reflected light pulse reflected by feature 41R is schematically represented by the height of the light pulse.

[0045] Because feature 41R is assumed to be substantially red, the feature exhibits a relatively high reflectivity to R light relative to the reflectivities of G light and B light. As a result, although the transmitted R, G, and B light pulses 51r, 51g, and 51b in the transmitted light pulse group 51 have the same intensity, the reflected light pulse 41 / 51r in the corresponding reflected light pulse group 41 / 51 has a much greater intensity than the reflected G light pulse 41 / 51g or the reflected B light pulse 41 / 51b. The intensity of the reflected light pulse 41 / 51g is schematically shown as being greater than the intensity of the reflected light pulse 41 / 51b because, for example, it is assumed that the reflectivity of feature 41R to G light is greater than the reflectivity to B light. The intensity of the reflected light pulses in the reflected light pulse group 41 / 52 is less than the intensity of the corresponding reflected light pulses in the reflected light pulse group 41 / 51 because the intensity of the light pulses in the transmitted light pulse group 52 is less than the intensity of the transmitted light pulses in the transmitted light pulse group 51. The intensity of the reflected light pulses in group 41 / 53 is less than the intensity of the corresponding reflected light pulses in group 41 / 52.

[0046] Similarly, the reflected light pulses formed by the green feature 42G for the emitted light pulses in the emitted light pulse groups 51, 52 and 53 may be collectively referred to as reflected light pulses 42 / 50 and individually referred to by the group number to which they belong with the appropriate subscript r, g or b. Reflected R light pulses 41 / 51r, 41 / 52r and 41 / 52r exhibit enhanced intensity compared to the light pulse 41 / 50 formed by the red feature 41R, and reflected G light pulses 41 / 51g, 41 / 52g and 41 / 52g exhibit enhanced intensity for the reflected light pulse 42 / 50 formed by the green feature 42G.

[0047] Light from reflected light pulse 41 / 50 reflected by feature 41R is collected by optical system 26 and imaged onto DVS pixel 24-41. In response to incident light, DVS pixel 24-41 generates a DVS image signal stream of changes in intensity and / or color of light sensed by the pixel, the DVS image signal stream being generated at each transition between temporally adjacent reflected light pulses 41 / 50. The characteristics of the DVS image signal stream generated for a given transition between reflected light pulses 41 / 50 are a function of the intensity and color of the given reflected light pulse at the transition, that is, a function of the color of feature 41R.

[0048] Similarly, light from reflected light pulse 42 / 50 reflected by feature 42G is collected by optical system 26 and imaged onto DVS pixel 24-42. DVS pixel 24-42 generates a DVS image signal stream of changes in intensity and / or color of light sensed by the pixel, the DVS image signal stream being generated at each transition between temporally adjacent reflected light pulses 42 / 50. The characteristics of the DVS image signal stream generated for a given transition between reflected light pulses 42 / 50 are a function of the intensity and color of the given reflected light pulse at the time of the transition, that is, a function of the color of feature 41G.

[0049] Because reflected light pulses 41 / 50 reflected from feature 41R and reflected light pulses 42 / 50 reflected from feature 42G exhibit different intensity patterns, the DVS image signal stream generated by DVS pixel 24-41 in response to transitions between reflected light pulses 41 / 50 from feature 41R is different from the DVS image signal stream generated by DVS pixel 24-42 in response to transitions between reflected light pulses 42 / 50 from feature 41G. The DVS image signal streams generated by DVS pixel 24-41 and DVS pixel 24-42 distinguish the color of feature 41R from the color of feature 42G.

[0050] According to an embodiment of the present disclosure, DVS-CCam 20 processes DVS image signal streams from DVS pixels 24-41 and 24-42, which image features 41R and 41G, and processes a DVS image signal stream generated by DVS pixel 24, which images other features of scene 40, thereby determining the colors of the respective features, and determining a color image of scene 40 based on the determined colors.

[0051] In one embodiment, for a plurality of DVS image signal streams generated by a given DVS pixel 24 in photosensor 23 in response to changes in sensed light intensity generated by transitions between reflected light pulses 50 , DVS-CCam 20 determines a DVS pixel response for each DVS image signal stream.

[0052] Optionally, the DVS pixel response is equal to the sum of the number of DVS image signals generated by a given pixel in response to transitions between the first and second boundary times. For abrupt changes in the color and / or intensity of light at a transition, the signal rate at which the DVS pixel 24 generates the DVS image signal can be modeled by a very fast rise time to a maximum signal rate followed by an exponentially decaying fall time to a background signal rate. Optionally, for example, the first boundary time can be the time at which the signal rate reaches a maximum value, and the second boundary time can be the time at which the signal rate decays to a rate that is considered equal to the background signal rate. Assume that there are N light pulses emitted to illuminate the scene 40, and there are N transitions for which the pixel generates a stream of DVS image signals. The DVS pixel response generated by a given DVS pixel 24 for the nth transition between the nth pulse emitted by the scintillator 22 and the previous (n-1)th pulse is represented as DVS-PR(x,y)(n-1),n,2≤n≤N, where x and y are the row and column coordinates of a given pixel in the photosensor 23. DVS-PR(x,y)(n-1),n can be expected to be equal to the pixel response function Provided value. Pixel response function is expected to be a function of the color of the feature in scene 40 imaged at DVS pixel 24 and the color and intensity of the transmit pulse 50 used to illuminate the feature imaged at the pixel and whereby the feature forms the transition of reflected light to the pixel. In notation:

[0053]

[0054] In expression (1), F(x,y,rf,gf,bf) is a feature imaged on a DVS pixel 24 at row and column coordinates x,y and having an unknown color represented by RGB color components represented by rf, gf and bf, respectively. TP50n(rn,gn,bn,In) represents the nth emitted light pulse 50 having known color coordinates rn, gn and bn and intensity In.

[0055] In one embodiment, by determining the actual value of DVS-PR(x,y)n and the predicted value The unknown characteristic colors rf, gf and bf can be determined by minimizing the loss function of the difference between (F(x, y, rf, gf, bf), TP50n(r, g, b, I)). f , g^ f , b^ f This can be determined by minimizing the least squares cost function with respect to (rf, gf, bf):

[0056] r^ f , g^ f , b^ f = arg min

[0057]

[0058] It should be noted that although Figure 1 It is assumed that the emitted light pulses 50 are continuous in time, but the implementation of the embodiments of the present disclosure is not limited to continuous light pulses. DVS-CCam according to the embodiments can illuminate the scene with continuously changing light patterns or temporally separated light pulses. For example, Figure 2 An embodiment of a DVS-CCam 20 is schematically shown which emits discrete and temporally separated pulses of light to illuminate a scene 40. The DVS image signals generated by the pixels in response to the rise and fall times of the emission pulses can be used to determine the color of features in the scene imaged at the pixels.

[0059] It should also be noted that although Figure 1 and Figure 2 While schematically showing DVS-CCam 20 illuminating scene 40 with RGB visible light, DVS-CCam according to one embodiment may illuminate the scene with light in multiple wavelength bands in addition to the RGB or visible wavelength bandwidth and generate a "color" image of the scene in the other wavelength bands.

[0060] Figure 3 Schematically shows Figure 1 and Figure 2 , and includes a neural network (NN) 29. The NN 29 may optionally be a convolutional neural network (CNN), which may include nonlinear estimation for color, and may additionally take into account inherent spatial correlation in the output of the DVS-CCam 20 to reduce noise. The NN 29 may determine a characteristic color by receiving a feature vector for each DVS pixel 24 having a component DVS-PR(x,y)n (1≤n≤N).

[0061] NN 29 may include a contraction path and an expansion path, and may include multiple layers connecting the paths to optionally increase weights to improve the NN model. An exemplary architecture of NN 29 may be a CNN based on U-Net and Xception. Each layer in the contraction path may (optionally) use repeated Xception layers based on separable convolutions to reduce the spatial dimension and may increase the number of channels. Each layer in the expansion path may use separable transposed convolutions to increase the spatial dimension and may reduce the channel. The end of the expansion path may provide a desired output size, which may optionally be the same as the input size, and the channel may be reduced to 3 channels, 1 channel for each RGB color. Optionally, the path connecting the contraction layer and the expansion layer may retain the size of the data.

[0062] A loss function may be used, which may optionally be a weighted average of the L1 norm and the structural similarity index measure (SSIM) index, and may be given by the following equation, where MS-SSIM is multiscale SSIM:

[0063]

[0064] in and Y represent the reconstructed and true images.

[0065] It should be noted that although the NN 29 is described as an expansion path to reduce the channels to 3 channels for RGB colors, a neural network according to an embodiment may allow for the extraction of more spectral components from light sources of different colors other than RGB light, thereby allowing DVS-CCam to generate "color" images of scenes in other wavelength bands. Optionally, from invisible light sources, such as infrared light sources. This aspect may be particularly advantageous because hyperspectral images can be generated for possible use in applications involving image segmentation, classification and recognition, as well as other suitable applications. It should be further noted that although the loss function is described as including L1 and SSIM index (MM-SSIM), other loss functions may also be used.

[0066] The applicant conducted a number of tests to evaluate the efficacy of the DVS camera of the present invention. In the first round of tests, the DVS camera used linear estimation to reconstruct the color image, and in the second round of tests, the DVS including the CNN performed nonlinear estimation and reconstructed the color image. The DVS system used to conduct the tests is described in the test setup section below.

[0067] A. Test Setup

[0068] The test setup consisted of using a commercial DVS (Samsung DVS Gen3, 640 x 480 resolution) facing a static scene 14 inches (35.5 cm) away. For the scintillator, a screen capable of producing light of varying wavelengths and intensities was used. The scintillator covered a larger surface area than the scene area captured by the DVS in order to provide relatively uniform illumination, and it was placed directly behind the camera facing the scene. The scintillator changed the color of its emission at a frequency of 3 Hz. For the calibration process (in order to train the CNN), an RGB camera (Point Grey Grasshopper3 U3, 2.3MP) adjacent to the DVS was used. The scene was static to prevent the camera from detecting an image without the scintillator changing color or intensity. The test setup was designed to allow capturing the bitstream generated by the DVS in response to the scintillator, thereby producing a single RGB frame. This frame is an RGB image of the scene at the same resolution as the original DVS video.

[0069] To evaluate the DVS including the CNN, the CNN was first trained. A labeled dataset was created using a stereo system combining a DVS and an RGB camera. The camera had a resolution of 1920x1200, a frame rate of 163fps, and a depth of 8-bit color for each of the three color channels. The calibration process produced a set of matching points in each sensor using the Harris Corner Detector algorithm, which was then used to compute a holography that transforms the view of the RGB sensor to the view of the DVS sensor. The calibration process assumed that the captured scene was in a dark room on a flat surface, 14” away from the DVS sensor. Therefore, training data was acquired on 2D scenes to maintain calibration accuracy. Each training sample contained a sequence of frames, where most frames maintained the response of the scene to the scintillator change and a few frames were background noise frames before the scintillator change. For example, in the case of an RGB scintillator with three intensities, 32 frames were used for each color and intensity, for a total of 288 frames.

[0070] B. Test Results

[0071] Figure 4 The results of linear reconstruction of a color image are shown. Image 400 is a static scene, which is a color matrix from X-Rite. Image 402 is a color reconstruction. In order to evaluate DVS using linear estimation, the color of each pixel is reconstructed based on the sum of DVS events for each specific pixel in order to reconstruct a complete RGB image. Although the color reconstruction is successful, the quality of the reconstruction is noisy. This is due to ignoring the spatial correlation between the colors of adjacent pixels.

[0072] Figure 5 Several results of color reconstruction of a three-dimensional (3D) scene using a CNN are shown at a distance of 14 inches from the DVS sensor (the distance is measured from the DVS lens to the center of the 3D scene). Images 500 and 504 are original images, respectively, and images 502 and 506 are reconstructed images, respectively. The root mean square error (RMSE) of image 502 is 45, while the RMSE of image 506 is 47. The resulting reconstruction shows high fidelity and low noise. Therefore, this solution is feasible for color reconstruction of both 3D and 2D scenes.

[0073] Figure 6Results are shown for the robustness of the system to distance variations and their effect on color reconstruction when exposed to the same ambient brightness as used in the training of the CNN. The scene (color matrix) was initially placed at a distance of 14 inches (35.5 cm) from the DVS camera, and the color image was reconstructed at increasing distances at 2 inch intervals. Image 600 is a color reconstruction of increasing distances starting from an initial 14 inches (35.5 cm) between the scene and the DVS camera at 5.08 cm, image 602 is an image of increasing distances of 11.176 cm, image 604 is an image of increasing distances of 17.272 cm, image 606 is an image of increasing distances of 23.368 cm, image 608 is an image of increasing distances of 29.454 cm, and image 610 is an image of increasing distances of 35.56 cm. As can be seen from these images, color reconstruction using CNNs is not limited to fixed training distances, but can achieve distance generalization.

[0074] Some stages (steps) of the above method may also be implemented in a computer program for running on a computer system, including at least a code portion for executing the steps of the relevant method or enabling the programmable device to perform the functions of the device or system according to the present disclosure when running on a programmable device such as a computer system. Such a method may also be implemented in a computer program for running on a computer system, including at least a code portion for enabling a computer to perform the steps of the method according to the present disclosure.

[0075] A computer program is a list of instructions such as a specific application and / or operating system. For example, a computer program may include one or more of a subroutine, a function, a procedure, a method, an implementation, an executable application, an applet, a servlet, source code, code, a shared library / dynamically loaded library, and / or other instruction sequences designed for execution on a computer system.

[0076] The computer program may be stored internally on a non-transitory computer-readable medium. All or some of the computer programs may be provided on a computer-readable medium that is permanently, removably, or remotely coupled to an information processing system. The computer-readable medium may include, for example but not limited to, any number of the following media: magnetic storage media including magnetic disk and tape storage media; optical storage media such as optical disk media (e.g., CD-ROM, CD-R, etc.) and digital video disk storage media; non-volatile memory storage media including semiconductor-based storage cells such as flash memory, electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), read-only memory; ferromagnetic digital memory; magnetoresistive random access memory (MRAM); volatile storage media including registers, buffers or caches, main memory, random access memory, etc.

[0077] A computer process typically consists of the program or portion of a program being executed (running), current program values ​​and state information, and the resources used by the operating system to manage the execution of the process. An operating system (OS) is software that manages the sharing of computer resources and provides programmers with interfaces for accessing those resources. An operating system processes system data and user input and responds by allocating and managing tasks and internal system resources as services to the system's users and programs.

[0078] A computer system may, for example, include at least one processing unit, associated memory, and a plurality of input / output (I / O) devices. When executing a computer program, the computer system processes information according to the computer program and generates resultant output information through the I / O devices.

[0079] Unless otherwise indicated, the use of "and / or" between the last two members of a list of options for selection indicates that selection of one or more of the listed options is appropriate and can be made.

[0080] It should be understood that where the claims or specification refer to "a" or "an" element, such reference should not be interpreted as there being only one of the element.

[0081] All references mentioned in this specification are incorporated by reference in their entirety into this specification to the same extent as each individual reference is expressly and individually indicated to be incorporated by reference into this specification. In addition, the citation or identification of any reference in this application should not be construed as an admission that the reference is available as prior art to the present disclosure.

[0082] Although the present disclosure is described in terms of certain embodiments and generally related methods, variations and permutations of the embodiments and methods will be apparent to those skilled in the art. The present disclosure should be understood not to be limited to the specific embodiments described herein, but only to the scope of the appended claims.

Claims

1. A dynamic vision sensor DVS color camera DVS-CCam, operable to acquire a color image of a scene, the DVS-CCam comprising: A DVS photoelectric sensor including a plurality of DVS pixels; an illuminator operable to emit a light pattern characterized by a plurality of temporal variations of intensity and color to illuminate a scene; an optical system configured to collect light reflected from the light pattern emitted by the illuminator off a plurality of features in the scene and to focus the reflected light onto the plurality of DVS pixels of a DVS photosensor; as well as a processor configured to process a plurality of DVS signals generated by the plurality of DVS pixels in response to a plurality of temporal variations in the reflected light to provide a color image of the scene, wherein a pixel response value is determined for each pixel of the plurality of DVS pixels, the pixel response value being a function of a color of a feature in the scene imaged by the pixel, The pixel response value is a sum of the number of DVS signals in at least a portion of a DVS image signal stream generated by the pixel in response to the temporal variation of the reflected light.

2. DVS-CCam according to claim 1, wherein: The emitted light pattern comprises a plurality of light pulses that are sequential in time.

3. The DVS-CCam according to claim 1, wherein: The emitted light pattern includes a plurality of discrete and temporally separated light pulses.

4. The DVS-CCam according to claim 1, wherein: The emitted light pattern includes red light, green light and blue light.

5. The DVS-CCam according to claim 1, wherein: Each pixel of the plurality of DVS pixels generates the DVS image signal stream in response to a change in light intensity during the plurality of temporal changes in the reflected light.

6. The DVS-CCam according to claim 1, wherein: The pixel response values ​​are used to determine a best fit color that minimizes a linear mean square estimate of the pixel response values.

7. The DVS-CCam according to claim 1 further comprises a neural network NN.

8. The DVS-CCam according to claim 1, wherein: Information associated with the acquired color image is used to reconstruct the next image.

9. The DVS-CCam according to claim 1, wherein: The DVS photoelectric sensor comprises an event-based photoelectric sensor.

10. A method for acquiring a color image of a scene using a dynamic vision sensor DVS color camera DVS-CCam including a DVS photoelectric sensor, wherein the DVS photoelectric sensor includes a plurality of DVS pixels, the method comprising: emitting a light pattern characterized by a plurality of temporal variations of intensity and color to illuminate a scene; collecting light reflected by the emitted light pattern from a plurality of features in the scene and focusing the reflected light onto the plurality of DVS pixels; as well as processing a plurality of DVS signals generated by the plurality of DVS pixels in response to a plurality of temporal variations in the reflected light to provide a color image of the scene, wherein a pixel response value is determined for each pixel of the plurality of DVS pixels, the pixel response value being a function of a color of a feature in the scene imaged by the pixel, The pixel response value is a sum of the number of DVS signals in at least a portion of the DVS image signal stream generated by the plurality of pixels in response to the temporal variation of the reflected light.

11. The method according to claim 10, wherein: The emitted light pattern comprises a plurality of light pulses that are sequential in time.

12. The method according to claim 10, wherein: The emitted light pattern includes a plurality of discrete and temporally separated light pulses.

13. The method according to claim 10, wherein: The emitted light pattern includes red light, green light and blue light.

14. The method according to claim 10, wherein: Each pixel of the plurality of DVS pixels generates the DVS image signal stream in response to a change in light intensity during the plurality of temporal changes in the reflected light.

15. The method according to claim 10, wherein: The pixel response values ​​are used to determine a best fit color that minimizes a linear mean square estimate of the pixel response values.

16. The method according to claim 10, further comprising a neural network (NN).

17. A dynamic vision sensor DVS color camera DVS-CCam, operable to acquire a color image of a scene, the DVS-CCam comprising: A DVS photoelectric sensor including a plurality of DVS pixels; an illuminator operable to emit a light pattern characterized by a plurality of temporal variations of intensity and color to illuminate a scene; an optical system configured to collect light reflected from the light pattern emitted by the illuminator off a plurality of features in the scene and to focus the reflected light onto the plurality of DVS pixels of a DVS photosensor; as well as a processor configured to process a plurality of DVS signals generated by the plurality of DVS pixels in response to a plurality of temporal variations in the reflected light to provide a color image of the scene, wherein a pixel response value is determined for each pixel of the plurality of DVS pixels, the pixel response value being a function of a color of a feature in the scene imaged by the pixel, The pixel response value is used to determine a best-fit color, and the best-fit color minimizes a linear mean square estimate of the pixel response value.

18. A method for acquiring a color image of a scene using a dynamic vision sensor DVS color camera DVS-CCam including a DVS photoelectric sensor, wherein the DVS photoelectric sensor includes a plurality of DVS pixels, the method comprising: emitting a light pattern characterized by a plurality of temporal variations of intensity and color to illuminate a scene; collecting light reflected by the emitted light pattern from a plurality of features in the scene and focusing the reflected light onto the plurality of DVS pixels; as well as processing a plurality of DVS signals generated by the plurality of DVS pixels in response to a plurality of temporal variations in the reflected light to provide a color image of the scene, wherein a pixel response value is determined for each pixel of the plurality of DVS pixels, the pixel response value being a function of a color of a feature in the scene imaged by the pixel, The pixel response value is used to determine a best-fit color, and the best-fit color minimizes a linear mean square estimate of the pixel response value.

Citation Information

Patent Citations

  • Dynamic vision sensor and projector for depth imaging

    US20190045173A1

  • Image sensor having combined responses from linear and logarithmic pixel circuits

    US9894296B1