Depth sensor device and method for operating the depth sensor device
The sensor device uses event-based sensors to project light patches in alternating directions, correlating reflections to generate accurate depth maps by averaging projection angles, addressing storage and latency issues in conventional systems.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-05
- Publication Date
- 2026-03-19
AI Technical Summary
Conventional depth estimation systems require large storage space and high latency due to the need to store intensity values from all camera pixels, limiting real-time applications and increasing processing time, especially when dealing with high pixel counts.
A sensor device using event-based sensors projects time-series light patches in alternating directions, correlating projected light with reflections to generate depth maps, averaging projection angles from different sweeps to compensate for brightness-dependent delays in event detection.
This approach reduces latency and storage requirements while improving depth map accuracy by compensating for brightness-induced delays, enhancing real-time performance without increasing system complexity or cost.
Smart Images

Figure 2026509449000001_ABST
Abstract
Description
Technical Field
[0005]
[0001] The present disclosure relates to a sensor device and a method for operating the sensor device. The present disclosure is particularly related to the generation of a depth map of a scene.
Background Art
[0002] In recent years, technologies for automatically measuring distances by transmitting and receiving light have attracted attention. Such technologies include the use of structured light, that is, a method of irradiating an object with static or temporally changing sparse light patterns at various solid angles to generate, for example, line, bar, or checkerboard patterns, or an active stereo method. When the directions of a light source and a camera are known, the shape and distance of an object can be calculated from triangulation based on the known positions of the light source and the camera, the direction of the emitted light in space, and the position of the corresponding optical signal on the camera.
[0003] Further improvements to the above structured light system are desirable.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Patent Document 3
Patent Document 4
Patent Document 5
Patent Document 6
Summary of the Invention
Problems to be Solved by the Invention
発明の概要
発明が解決しようとする課題
Summary of the Invention
Problems to be Solved by the Invention
[0005] unchanged as per the instruction. If there are any specific requirements regarding these tags, please let me know. And there seems to be a typo in "
発明が解決しようとする課題
発明が解決しようとする課題
発明の概要
[0006] These shortcomings of conventional depth estimation techniques can be mitigated by using event-based sensors, i.e., sensors that are sensitive only to changes in the received signal. This disclosure aims to improve depth sensing based on event-based sensors. [Means for solving the problem]
[0007] For this purpose, a sensor device for generating a depth map of a scene is provided. The sensor device comprises a projection unit configured to project a plurality of time-series light patches onto the scene, in which a sequence of light patches in one series is shifted along a first direction or a second direction opposite to the first direction; a light-receiving unit positioned at a predetermined distance from the projection unit and configured to observe the reflection of the light patches from the scene; and a control unit configured to temporally correlate the projected light patches with the observed reflections and generate the depth map of the scene using the correlation and the predetermined distance. Here, the projection unit is configured to alternately project a first sweep light and a second sweep light, each first sweep light being a sequence of light patches shifted along the first direction, and each second sweep light being a sequence of light patches shifted along the second direction. The light-receiving unit comprises a plurality of event detection pixels that indicate the occurrence of an event when the light intensity measured by each event detection pixel changes by a predetermined amount or more over a given time period. Furthermore, the time it takes for one event detection pixel to indicate the occurrence of the event depends on the absolute value of the light intensity before the change in light intensity, each event detection pixel is configured to receive light only from a predetermined solid angle, the solid angles observed by adjacent event detection pixels are adjacent, and the observed solid angles of all event detection pixels cover the field of view of the light receiving unit. The control unit is configured to calculate, for each event occurrence during the first sweep light projection and the second sweep light projection, the predetermined solid angle of each event detection pixel and the projection angle at which the projection unit projected a single light fragment at the time the event occurred, and for each pair of event occurrences, the control unit is configured to average the respective calculated projection angles for each consecutive first sweep light and second sweep light, and for each consecutive second sweep light and first sweep light, and to calculate the depth map of the scene based on the pair of predetermined solid angle and average projection angle calculated for each event occurrence.
[0008] Furthermore, a method for operating the sensor device described above, comprising projecting a plurality of time-series light fragments onto the scene, wherein a series of light fragments in one sequence are shifted along a first direction or a second direction opposite to the first direction, the projection being performed by alternately projecting a first sweep light and a second sweep light, each first sweep light being a series of consecutive light fragments shifted along the first direction, each second sweep light being a series of consecutive light fragments shifted along the second direction, observing the reflection of the light fragments from the scene, and temporally correlating the projected light fragments with the observed reflections. A method for calculating the depth map of the scene using the correlation and the predetermined distance; during the first sweep light projection and during the second sweep light projection, for each event occurrence, calculating the predetermined solid angle of each event detection pixel and the projection angle of one light fragment projected at the time the event occurred; for each pair of event occurrences, averaging the respective calculated projection angles for each consecutive first sweep light and second sweep light, and for each consecutive second sweep light and first sweep light, and calculating the depth map of the scene based on the pair of predetermined solid angle and average projection angle calculated for each event occurrence.
[0009] By averaging the projection angles detected in different directions of light strip sweep, the adverse effects of delay differences between event detection pixels caused by brightness differences in the scene can be suppressed. [Brief explanation of the drawing]
[0010] [Figure 1A] This is a simplified block diagram of an event detection circuit for a solid-state imaging device that includes a pixel array. [Figure 1B] Figure 1A is a simplified block diagram of the pixel array. [Figure 1C] Figure 1A is a simplified block diagram of the imaging signal readout circuit of the solid-state imaging device. [Figure 2] A schematic diagram of the sensor device is shown below. [Figure 3] A schematic diagram of another sensor device is shown below. [Figure 4] Schematically shows the dependence of the delay of the event detection pixel on the luminance. [Figure 5A] Schematically shows the operating principle of the sensor device. [Figure 5B] Schematically shows the operating principle of the sensor device. [Figure 5C] Schematically shows the operating principle of the sensor device. [Figure 6] Schematically shows the operating principle of the sensor device. [Figure 7] Schematically shows the operating principle of the sensor device. [Figure 8] Schematically shows the generation of light patches of different colors. [Figure 9] Schematically shows the color filter array used in the sensor device. [Figure 10A] Schematically shows different application examples of a camera equipped with a depth sensor device. [Figure 10B] Schematically shows different application examples of a camera equipped with a depth sensor device. [Figure 11] Schematically shows a head-mounted display equipped with a depth sensor device. [Figure 12] Schematically shows an industrial manufacturing device equipped with a depth sensor device. [Figure 13] Shows a schematic process flow of a method for operating the sensor device. [Figure 14A] Schematically shows the alternation of the operating principle of the sensor device. [Figure 14B] Schematically shows the alternation of the operating principle of the sensor device. [Figure 15A] Schematically shows the readout frequencies in different operating modes of the sensor device. [Figure 15B] Schematically shows the readout frequencies in different operating modes of the sensor device. [Figure 15C] Schematically shows the readout frequencies in different operating modes of the sensor device. [Figure 16]This is a simplified perspective view of a solid-state imaging device having a stacked structure according to one embodiment of the present disclosure. [Figure 17] This is a simplified diagram of an example configuration of a multilayer solid-state imaging device to which the technology disclosed herein may be applied. [Figure 18] This block diagram shows an example of a schematic configuration of a vehicle control system. [Figure 19] Figure 18 is an explanatory diagram showing an example of the installation locations of the external information detection unit and imaging unit of the vehicle control system. [Modes for carrying out the invention]
[0011] This disclosure relies on event detection using event vision sensors / dynamic vision sensors. These sensors are generally well known to those skilled in the art, but their outlines will be briefly described with reference to Figures 1A to 1C.
[0012] Figure 1A is a block diagram of a solid-state imaging device 100 employing event-based change detection. The solid-state imaging device 100 includes a pixel array 110 having one or more imaging pixels 111, each pixel 111 including a photoelectric conversion element PD. The pixel array 110 can be a one-dimensional pixel array (line sensor) in which the photoelectric conversion elements PD of all pixels are arranged along straight or meandering lines. In particular, the pixel array 110 can be a two-dimensional array, in which the photoelectric conversion elements PD of the pixels 111 can be arranged along straight or meandering rows and along straight or meandering columns.
[0013] A two-dimensional array of pixels 111 is shown. The pixels 111 are arranged along straight rows and along straight columns orthogonal to these rows. Each pixel 111 converts incident light into an imaging signal representing the intensity of the incident light and an event signal indicating a change in light intensity (e.g., an increase (positive polarity) of at least an upper threshold amount and / or a decrease (negative polarity) of at least a lower threshold amount). If necessary, the intensity detection and event detection functions of each pixel 111 may be separated, and different pixels observing the same solid angle may perform their respective functions. These different pixels may be sub-pixels and may be implemented to share parts of the circuit. Different pixels may be part of different image sensors. In this disclosure, when referring to a pixel capable of generating imaging signals and event signals, this should be understood to include combinations of pixels that perform these functions separately as described above. Preferably, in this disclosure, a pixel 111 may generate only event signals.
[0014] The controller 120 controls the flow of each process within the pixel array 110. For example, the controller 120 can control a threshold generation circuit 130 that calculates and supplies a threshold for each individual pixel 111 within the pixel array 110. The readout circuit 140 supplies control signals corresponding to each individual pixel 111 and outputs information regarding the position of the pixel 111 that indicates an event. Since the solid-state imaging device 100 employs event-based change detection, the readout circuit 140 can output a variable amount of data at time intervals.
[0015] Figure 1B illustrates details regarding the event detection capabilities of the imaging pixels 111 in Figure 1A. Of course, any other arbitrary embodiment that enables event detection can also be employed. Each pixel 111 includes a photoreceptor module PR and is assigned to a pixel backend 300. In this case, each complete pixel backend 300 can be assigned to a single photoreceptor module PR. Alternatively, a pixel backend 300 or a portion thereof may be assigned to two or more photoreceptor modules PR. In this case, the shared portion of the pixel backend 300 can be sequentially connected to the assigned photoreceptor modules PR via multiplexing.
[0016] The photoreceptor module PR includes a photoelectric conversion element PD, such as a photodiode or other type of light sensor. The photoelectric conversion element PD converts incident light 9 into a photocurrent Iphoto. In this case, the amount of photocurrent Iphoto is a function of the light intensity of the incident light 9.
[0017] The photoreceptor circuit PRC converts the photocurrent Iphoto into a photoreceptor signal Vpr. The voltage of the photoreceptor signal Vpr is a function of the photocurrent Iphoto.
[0018] The memory capacitor 310 stores charge and holds a memory voltage that depends on past photoreceptor signals Vpr. In particular, the memory capacitor 310 receives the photoreceptor signal Vpr and causes the first electrode of the memory capacitor 310 to carry a charge in response to the photoreceptor signal Vpr, i.e., the light received by the photoelectric conversion element PD. The second electrode of the memory capacitor C1 is connected to the comparator node (inverting input) of the comparator circuit 340. Therefore, the voltage Vdiff at the comparator node fluctuates in accordance with the change in the photoreceptor signal Vpr.
[0019] The comparator circuit 340 compares the difference between the current photoreceptor signal Vpr and past photoreceptor signals with a threshold. The comparator circuit 340 can be located within each pixel backend 300 or shared among subsets of pixels (e.g., columns). For example, each pixel 111 has a pixel backend 300 including the comparator circuit 340, the comparator circuit 340 is essential to the imaging pixel 111, and each imaging pixel 111 has its own dedicated comparator circuit 340.
[0020] The memory element 350 stores the output of the comparator in response to a sample signal from the controller 120. The memory element 350 may include a sampling circuit (e.g., a switch and a parasitic or explicit capacitor) and / or a digital memory circuit such as a latch or a flip-flop. The memory element 350 can be a sampling circuit. The memory element 350 can be configured to store 1 bit, 2 bits, or more binary bits.
[0021] The output signal of the reset circuit 380 sets the inverting input terminal of the comparator circuit 340 to a predetermined potential. The output signal of the reset circuit 380 can be controlled in response to the contents of the memory element 350 and / or in response to a global reset signal received from the controller 120.
[0022] The solid-state imaging device 100 operates as follows: Changes in the light intensity of the incident radiation 9 are converted into changes in the photoreceptor signal Vpr. At a timing specified by the controller 120, the comparator circuit 340 compares Vdfiff at the inverting input terminal (comparator node) with a threshold Vb applied to its non-inverting input terminal. Simultaneously, the controller 120 operates the memory element 350 to store the comparator output signal Vcomp. The memory element 350 may be located in the pixel circuit 111 or the readout circuit 140 shown in Figure 1A.
[0023] If the state of the stored comparator output signal indicates a change in light intensity and the global reset signal GlobalReset (controlled by controller 120) is active, the conditional reset circuit 380 outputs a reset output signal and resets Vdiff to a known level.
[0024] The memory element 350 may include information indicating that the change in light intensity detected by the pixel 111 has exceeded a threshold.
[0025] The solid-state imaging device 120 can output the address of the pixel 111 where a change in light intensity was detected (the address of pixel 111 corresponds to its row number and column number). A change in light intensity detected at a given pixel is called an event. More specifically, the term "event" means that the photoreceptor signal, which represents the light intensity of the pixel and is a function of it, has changed by an amount greater than or equal to a threshold applied by the controller via the threshold generation circuit 130. To transmit an event, the address of the corresponding pixel 111 is transmitted along with (optionally) data indicating whether the change in light intensity is positive or negative. The data indicating whether the change in light intensity is positive or negative may contain a single bit.
[0026] To detect changes in light intensity between the present and past points in time, each pixel 111 stores a value (representation) that represents the light intensity at a past point in time.
[0027] More specifically, each pixel 111 stores a voltage Vdiff that represents the difference between the photoreception signal from the last registered event at that pixel 111 and the current photoreception signal at that pixel 111.
[0028] To detect an event, the Vdiff of the comparator node is first compared to a first threshold for detecting an increase in light intensity (ON event). The comparator output is sampled in a (explicit or parasitic) capacitor or stored in a flip-flop. Next, the Vdiff of the comparator node is compared to a second threshold for detecting a decrease in light intensity (OFF event). The comparator output is sampled in a (explicit or parasitic) capacitor or stored in a flip-flop.
[0029] A global reset signal is transmitted to all pixels 111. At each pixel 111, this global reset signal is logically ANDed with the sampled comparator output, and only the pixels where an event was detected are reset. The sampled comparator output voltage is then read, and the corresponding pixel address is transmitted to the data receiving device. The time precision of event detection can be in the microsecond range.
[0030] Figure 1C shows an example configuration of a solid-state imaging device 100, including an image sensor assembly 10 used for reading out an intensity imaging signal in the form of an active pixel sensor (APS). Herein, Figure 1C is merely an example. The reading out of the imaging signal can be performed by any other known method. As described above, the image sensor assembly 10 may use identical pixels 111, or these pixels 111 may be complemented by additional pixels that each observe the same solid angle. In the following description, we select an exemplary case in which identical pixel arrays 110 are used.
[0031] The image sensor assembly 10 includes a pixel array 110, an address decoder 12, a pixel timing drive unit 13, an ADC (analog-to-digital converter) 14, and a sensor controller 15.
[0032] The pixel array 110 includes a plurality of pixel circuits 11P arranged in a matrix of rows and columns. Each pixel circuit 11P includes a photosensitive element and a FET (field-effect transistor) that controls the signal output from the photosensitive element.
[0033] The address decoder 12 and the pixel timing drive unit 13 control the driving of each pixel circuit 11P arranged in the pixel array 110. Specifically, the address decoder 12 supplies the pixel timing drive unit 13 with control signals that specify the pixel circuit 11P to be driven, based on the address, latch signal, etc., supplied from the sensor controller 15. The pixel timing drive unit 13 drives the FETs of the pixel circuit 11P based on the drive timing signal supplied from the sensor controller 15 and the control signal supplied from the address decoder 12. The electrical signals (pixel output signal, imaging signal) of the pixel circuit 11P are supplied to the ADC 14 via the vertical signal line VSL. In this case, each ADC 14 is connected to one of the vertical signal lines VSL, and each vertical signal line VSL is connected to all pixel circuits 11P in one row of the pixel array unit 11. Each ADC 14 performs analog-to-digital conversion on the pixel output signals that are sequentially output from the row of the pixel array unit 11, and outputs digital pixel data DPXS to the signal processing unit. For this purpose, each ADC14 is equipped with a comparator 23, a digital-to-analog converter (DAC) 22, and a counter 24.
[0034] The sensor controller 15 controls the image sensor assembly 10. For example, the sensor controller 15 supplies address and latch signals to the address decoder 12 and drives timing signals to the pixel timing drive unit 13. Additionally, the sensor controller 15 may supply control signals to control the ADC 14.
[0035] The pixel circuit 11P includes a photoelectric conversion element PD as a photosensitive element. The photoelectric conversion element PD may include, for example, a photodiode or be composed of a photodiode. For one photoelectric conversion element PD, the pixel circuit 11P may have four FETs that function as active elements: a transfer transistor TG, a reset transistor RST, an amplification transistor AMP, and a selection transistor SEL.
[0036] A photoelectric converter (PD) converts incident light into electric charge (in this case, electrons) through photoelectric conversion. The amount of charge generated by the PD corresponds to the amount of incident light.
[0037] The transfer transistor TG is connected between the photoelectric conversion element PD and the floating diffusion region FD. The transfer transistor TG functions as a transfer element that transfers charge from the photoelectric conversion element PD to the floating diffusion region FD. The floating diffusion region FD functions as a temporary local charge storage unit. A transfer signal, which functions as a control signal, is supplied to the gate (transfer gate) of the transfer transistor TG via a transfer control line.
[0038] Therefore, the transfer transistor TG can transfer electrons photoelectrically converted by the photoelectric conversion element PD to the floating diffusion FD.
[0039] The reset transistor RST is connected between the floating diffusion FD and the power line to which the positive power supply voltage VDD is supplied. The reset signal, which functions as a control signal, is supplied to the gate of the reset transistor RST via the reset control line.
[0040] Therefore, the reset transistor RST, which functions as a reset element, resets the potential of the floating diffusion FD to the potential of the power line.
[0041] The floating diffusion FD is connected to the gate of the amplification transistor AMP, which functions as an amplification element. In other words, the floating diffusion FD functions as the input node of the amplification transistor AMP.
[0042] The amplification transistor AMP and the selection transistor SEL are connected in series between the power line VDD and the vertical signal line VSL.
[0043] Therefore, the amplification transistor AMP is connected to the signal line VSL via the selection transistor SEL and forms a source follower circuit with the constant current source 21, which is shown as part of the ADC 14.
[0044] Next, a selection signal, which functions as a control signal corresponding to the address signal, is supplied to the gate of the selection transistor SEL via the selection control line, causing the selection transistor SEL to turn on.
[0045] When the selector transistor SEL is turned on, the amplification transistor AMP amplifies the potential of the floating diffuser FD and outputs a voltage corresponding to the potential of the floating diffuser FD to the signal line VSL. The signal line VSL transfers the pixel output signal from the pixel circuit 11P to the ADC 14.
[0046] The gates of the transfer transistor TG, reset transistor RST, and selection transistor SEL are connected, for example, on a row-by-row basis, so their operations are performed simultaneously for each row of pixel circuits 11P. Furthermore, it is also possible to selectively read out a single pixel or a group of pixels.
[0047] The ADC14 may include a DAC22, a constant current source21 connected to the vertical signal line VSL, a comparator23, and a counter24.
[0048] The vertical signal line VSL, the constant current source 21, and the amplification transistor AMP of the pixel circuit 11P constitute a source follower circuit.
[0049] DAC22 generates and outputs a reference signal. By performing a digital-to-analog conversion on a digital signal that increases by, for example, 1 at regular intervals, DAC22 can generate a reference signal that includes a reference voltage ramp. Within this voltage ramp, the reference signal increases steadily per unit of time. This increase can be linear or nonlinear.
[0050] The comparator 23 has two input terminals. The reference signal output from the DAC 22 is supplied to the first input terminal of the comparator 23 via the first capacitor C1. The pixel output signal transmitted via the vertical signal line VSL is supplied to the second input terminal of the comparator 23 via the second capacitor C2.
[0051] Comparator 23 compares the pixel output signals supplied to the two input terminals with a reference signal and outputs a comparator output signal representing the comparison result. In other words, comparator 23 outputs a comparator output signal that represents the relative magnitudes of the pixel output signal and the reference signal. For example, the comparator output signal may be high level when the pixel output signal is greater than the reference signal, and low level otherwise. Or, the opposite may also be true. The comparator output signal VCO is supplied to counter 24.
[0052] The counter 24 counts the count value in synchronization with a predetermined clock. Specifically, the counter 24 starts counting the count value at the beginning of the P phase or D phase when the DAC 22 begins to decrease the reference signal, and continues counting the count value until the relationship between the pixel output signal and the reference signal changes and the comparator output signal inverts. When the comparator output signal inverts, the counter 24 stops counting the count value and outputs the count value at that time as the AD conversion result of the pixel output signal (digital pixel data DPXS).
[0053] The event sensors described above may be used in the following contexts when referring to event detection. However, any other embodiment of event detection is also applicable.
[0054] Figure 2 schematically shows a sensor device 1000 that generates a depth map of a scene including an object O. That is, it is a device that can estimate the distance from the surface elements of the object O to the sensor device 1000. The sensor device 1000 may generate depth information itself, or it may generate only data for establishing depth information in subsequent processing steps. This sensor device comprises a projection unit 1010, a light receiving unit 1020, and a control unit 1030.
[0055] The projection unit 1010 is configured to illuminate different positions of the object O using illumination patterns over different time periods. In particular, the projection unit 1010 is configured to project multiple time-series light patches onto the scene, and consecutive light patches in one series are shifted along a first direction x1 or along a second direction x2 opposite to the first direction x1.
[0056] In the example in Figure 2 and the following description, a lighting pattern consisting of lines L is used. The position of the lines L changes over time, and different parts of the scene are illuminated by these lines L at different time intervals. Preferably, the solid angles illuminated by the lines L are adjacent to each other and completely satisfy the projected solid angle PS of the projection section 1010. However, the lines L may be separated by a certain distance from each other.
[0057] Figure 2 shows an example where only one line is projected onto the scene, but as schematically shown in Figure 3, multiple lines can also be projected simultaneously. Note that the equally spaced arrangement of lines in Figure 3 is merely a simplified example; the lines can be placed at any position. Furthermore, the number of lines, i.e., the number of solid angles illuminated, can change over time.
[0058] The following explanation focuses on the line examples shown in Figures 2 and 3, but those skilled in the art will readily understand that other sparse lighting patterns, such as checkerboard patterns or pixel-level illumination, can also be used. However, any series of light fragments projected onto the scene must move across the entire scene in a predefined, reversible manner. In the example of line L shown in Figure 2, this means that for a series of lines sweeping the scene along a first direction x1 (right to left in Figure 2), there exists a series of lines sweeping the scene in the opposite direction, i.e., a second direction x2 (left to right in Figure 2). Here, a single sweep range does not necessarily have to cover the entire projected solid angle PS, but may cover only a portion of this projected solid angle. For example, a sweep in one direction may include 10, 50, or 100 different light projection positions, which can then be demonstrated in reverse order.
[0059] As mentioned above, the example of projection lines was chosen for the sake of simplicity of explanation. The projection unit 1010 can project any kind of light fragment that can be reversibly swept across the scene. For example, dot-like light can be moved in rows across the observed scene, and the scanning operation can be reversed after (or during) scanning the entire scene. Similarly, multiple dot-like light fragments can be used. Furthermore, line segments can be used instead of lines that cross the entire projection solid angle PS.
[0060] Illumination variations can be achieved, for example, by using a fixed light source and deflecting its light at different angles at different times. For instance, a tilting mirror in a MEMS (Micro-Electro-Mechanical System) can be used to deflect the illumination pattern, and / or a refractive diffraction grating can be used to generate multiple lines. Alternatively, a vertical-cavity surface-emitting laser (VCSEL) or any other array of laser LEDs can be used to illuminate different parts of the scene / object at different times. Furthermore, it is conceivable to generate time-varying illumination patterns using shielding optics such as slit plates or liquid crystal panels.
[0061] The projection unit 1010 may be arranged such that some points within its field of view are not illuminated by the illumination pattern.
[0062] Alternatively, the illumination pattern emitted from the projection unit 1010 can be fixed, and the object O can be moved across the illumination pattern first in one direction, and then in the opposite direction. In principle, the method of generating the illumination pattern and the method of moving the object across it are arbitrary, as long as different positions of the object O are reversibly illuminated over different time periods.
[0063] The sensor device 1000 includes a light-receiving unit 1020 positioned at a predetermined distance d from the projection unit 1010 and configured to observe the reflection of light fragments from the scene. Due to the surface structure of the scene / object O, the light fragments projected from the projection unit 1010 are reflected by the object O in a distorted form, forming an image I on the light-receiving unit 1020.
[0064] The light-receiving unit 1020 includes a plurality of event detection pixels 1025. These event detection pixels 1025 indicate the occurrence of an event when the light intensity measured by the event detection pixels 1025 changes by a predetermined amount or more within a predetermined time. Therefore, as explained in Figures 1A to 1C, the light-receiving unit 1020 can function as an event sensor that can detect changes in light intensity exceeding a given threshold. Here, both positive and negative changes can be detected, and so-called positive or negative events can occur. Furthermore, the event detection threshold is dynamically adaptable and may differ between positive and negative events.
[0065] The time it takes for an event detection pixel 1025 to indicate the occurrence of an event depends on the absolute value of the light intensity, i.e., the brightness of the observed portion of the scene before (or after) the change in light intensity. Here, the received brightness depends on several factors, such as the distance to the object, surface reflectance, type of reflection (diffuse reflection, directional reflection), ambient light conditions, and characteristics of the projected light. Furthermore, each event detection pixel 1025 is configured to receive only light from a predetermined solid angle. In this case, the solid angles observed by adjacent event detection pixels 102 are adjacent, and the solid angles observed by all event detection pixels 1025 cover the field of view of the light receiving unit 1020. In other words, the brighter the solid angle observable by an event detection pixel 1025, the faster the event detection speed of that event detection pixel 1025. A brightly illuminated event detection pixel 1025 detects the occurrence of an event almost immediately, but the delay in event detection increases as the brightness decreases.
[0066] This is schematically illustrated in Figure 4. Figure 4 shows that, compared to brightly lit areas (represented by white rectangles), the event detection delay in this example increases by up to 20 microseconds on average in areas with minimal illumination (longer delays may also occur). In Figure 4, the error bars reflect the effect of jitter.
[0067] This inherent dependence of the exact time of event detection on detectable brightness causes problems with depth measurement using event detection pixels 1025 when the scene consists of bright and dark regions (e.g., highly reflective or white regions and low-reflective or black regions).
[0068] This problem is schematically illustrated in Figures 5A to 5C. Figure 5A shows a portion of an object O in an observation scene that includes a bright region A1 (e.g., a white or light-colored surface) and a dark region A2 (e.g., a black or dark-colored surface). A line L crossing both regions A1 and A2 is swept across the object O first in a first direction x1 and then in a second direction x2.
[0069] Figure 5B illustrates event detection for the reflection of line L from a bright region A1. An event is generated by a change in brightness when the reflection R of line L becomes visible to the event detection pixel 1025. In the bright region A1, this occurs almost instantaneously. By utilizing the time of event detection and the fact that the projection angle of line L is always known, the projection angle that triggered the event can be identified. Since each event detection pixel 1025 observes only a small solid angle that can be assumed as a line of sight, and the distance d between the projection unit 1010 and the light receiving unit 1020 is known, the distance between the reflection position of line L on the object O and the sensor device 1000 can be calculated.
[0070] As explained in Figure 5C, when observing the dark region A2, this operating principle of the sensor device 1000 deteriorates.
[0071] Similar to the case where a bright region A1 is illuminated, the illumination line L becomes a reflected light R, which is received by the event detection pixel 1025. However, since this reflection occurs in a dark region A2, the luminance of the reflected light R, i.e., the absolute value of the intensity of the reflected light, is lower than that of the reflection in the bright region A1. Therefore, the occurrence of the event is not immediately detected by the event detection pixel 1025, but is detected with a delay Δt. Consequently, the event timestamp will indicate a time t+Δt later than the time t when the line L was projected. When calculating the projection angle of the line L, it is assumed that the line L was projected at the later time t+Δt. Depending on the sweep direction of the line, this may result in the projection angle being estimated to be smaller or larger than it actually is.
[0072] In particular, as schematically shown in Figure 5C, when line L moves in a first direction x1, i.e., from right to left in Figure 5C, it is assumed that the reflection was caused not by line L, but by line L' which is projected later at time t+Δt. The projection angle of this line L' is Δα smaller than the projection angle α of line L projected at time t. Therefore, when sweeping the dark region A2 in the first direction x1, it is assumed that all reflections R were caused by a line with a projection angle Δα smaller than the true projection angle α. In triangulation, this leads to an underestimation of the distance between the dark region A2 and the sensor device 1000. Therefore, when sweeping in the first direction, i.e., towards the light-receiving unit 1020, the dark region is estimated to be closer than it actually is.
[0073] When sweeping in the opposite direction, i.e., the second direction x2, i.e., away from the light-receiving unit 1020, the delay Δt suggests that the reflection R was caused by a line L'' having a projection angle Δα larger than the projection angle of line L projected at time t. In triangulation, this large projection angle causes the distance between the dark region A2 and the sensor device 1000 to be estimated to be longer than it actually is. Therefore, when sweeping in the second direction, i.e., away from the light-receiving unit 1020, the dark region appears to be further away.
[0074] Therefore, without countermeasures, depth maps generated based on events detected by event detection pixels with brightness-dependent delays will not accurately represent the shape of the observed object. Possible solutions to this problem include, for example, increasing the light intensity of the projection ray, estimating absolute intensity to compensate for the delay, or adjusting the sensor circuit. However, increasing the light intensity only alleviates the problem; it does not solve it. This is because dark surfaces always appear darker than bright surfaces. Furthermore, while it is theoretically possible to estimate absolute brightness, for example, by event counting per pixel, measuring jitter / noise dependent on absolute brightness, or using an additional camera system, these approaches require additional processing and / or additional structural elements. This increases the complexity and cost of the system, as well as its processing time. The same applies to adjusting the sensor circuit.
[0075] Therefore, an alternative method is proposed to solve, or at least mitigate, the delay problem of event detection pixels in the structured light system described above. This solution utilizes the fact that the sign of the change in projection angle is determined by the shift direction of the light fragment projected across the scene. Since the direction of change differs for different shift / scan / sweep directions, combining measurements taken in different sweep directions should at least reduce, if not eliminate, the error.
[0076] For this purpose, a control unit 1030 is provided. This control unit 1030 is configured to temporally correlate the projected light fragment with the observed reflection and generate a depth map of the scene using this correlation and a predetermined distance. In particular, the control unit 1030 is configured to calculate a predetermined solid angle of each event detection pixel 1025 and the projection angle at which the projection unit 1010 projected the light fragment at the time of the event, each time an event occurs while the light fragment is shifting in a first direction x1 and while the light fragment is shifting in a second direction x2. That is, for each event, the control unit 1030 calculates, for each sweep in one direction, the line of sight of the event detection pixel 1025 that detected the event, the time the event was detected, and the angle at which the light fragment that may have passed through the line of sight of the event detection pixel 1025 was projected. Then, the same process is performed for sweeps in the opposite direction. Thus, at least a pair of projection angles are calculated for each event detection pixel 1025 / line of sight. The difference in these angles depends on the brightness of the received light, i.e., the brightness of the surface that reflected the projected light fragment. If this surface is very bright, no delay occurs, and therefore these two projection angles are essentially equal.
[0077] The control unit 1030 averages the projection angles calculated for each given event occurrence pair (i.e., each projection angle pair) for the shift of the optical element in the first direction x1 and the shift in the second direction x2. Then, it calculates a depth map of the scene based on the pair of a predetermined solid angle and the average projection angle calculated for each event occurrence. In other words, triangulation is performed based on the average value of projection angles measured during sweeping in different directions, rather than the apparent projection angle, which may be estimated to be smaller or larger than the actual value. Since the deviations from the true projection angle compensate for each other, the true distance to the observed object can be obtained by triangulation. Therefore, a more reliable depth map can be generated.
[0078] In the above, we assumed that we would use the average of two apparent projection angles obtained from a single sweep in each of the two directions. Of course, it is also possible to perform multiple sweeps (but the same number) in each direction and average all the projection angle values detected before triangulation.
[0079] Furthermore, the above assumes that the shift velocity of the optical elements in both directions is equal and constant. Therefore, a delay of Δt always results in the same change in projection angle. Of course, different sweep velocities, i.e., the amount of change in projection angle per unit of time, can be used both within a sweep in one direction and between sweeps in different directions. However, since the projector pattern is known, the velocity change is also known. In this case, the velocity difference can be compensated for by weighting or normalizing the calculated projection angles in the averaging process. The term "average" used above should be understood to include such compensation.
[0080] In this way, the depth map errors caused by the different delay times of the event detection pixels 1025 used can be compensated for, or at least reduced, by simple arithmetic operations and using standard components typical of structured optical systems. Therefore, the reliability of the generated depth map can be increased with relatively little increase in system complexity and processing load. Thus, depth map generation can be made more reliable without unnecessarily increasing system costs or energy consumption.
[0081] Here, the control unit 1030 can be any configuration of circuitry capable of performing the functions described herein. For example, the control unit 1030 can be a processor. The control unit 1030 is part of the pixel section of the sensor device 1000 and can be mounted on the same die(s) as the other components of the sensor device 1000. However, the control unit 1030 can also be located, for example, on a separate die, on a separate processor, or on a separate computer. The functions of the control unit 1030 may be implemented entirely in hardware or software, or as a mixture of hardware and software functions.
[0082] As described above, the time-series light fragments can be time-series rays L that sweep across the entire scene. In the above description, the rays L are straight lines that intersect the first direction x1 at 90°, but in principle, these rays may intersect the first direction x1 at any predetermined angle sufficiently different from 0°, as long as it can completely cover the entire scene. Alternatively, these rays may be curves. As described above, the light fragments in the above series may also include point-like or dot-like rays, line segments, or combinations thereof, as long as a particular movement of a light fragment in one time series can be reproduced in the reverse direction in another time series.
[0083] As described above, the apparent projection angle measured during sweeping in different directions is used to compensate for the deviation Δα from the true projection angle α. However, additionally or alternatively, the control unit 1030 may be configured to calculate, for each event occurrence pair, the absolute value of the light intensity before the change in light intensity by subtracting the projection angles calculated for the shift of the light element in the first direction x1 and the shift of the light element in the second direction x2, respectively. In other words, the delay in event detection due to dark areas, and the resulting overestimated / underestimated projection angles, can also be used to calculate the absolute light intensity received by the event detection pixel 1025. In this way, an image of the scene can be obtained.
[0084] This is schematically shown in Figures 6 and 7. Here, Figure 6 again shows how the deviation of the projection angle is compensated for. In both figures, a flat region with a bright region A1 and a dark region A2 is observed. When sweeping in the first direction x1, the projection angles of the lines projected onto the bright region A1 and the dark region A2 are estimated to be (sufficiently) accurate in the bright region A1, while in the dark region A2 they are estimated to be smaller than the actual values by Δα. When the sweep direction is reversed, the projection angle in the dark region A2 is estimated to be larger than the actual values. Averaging the two measurements of the projection angle gives the true projection angle α in the bright region A1 and the dark region A2.
[0085] On the other hand, subtracting the two measured values yields a measure of the "darkness" of the observation area, as shown in Figure 7. Subtracting these values yields a projection angle shift Δα. This shift can be associated with a time delay in the corresponding event detection pixel 1025. From the manufacturing calibration, it is known how this time delay relates to the observed absolute brightness before (or after) the event. Therefore, starting from an intensity level that does not cause a delay, i.e., a projection angle shift Δα = 0, the intensity level can be estimated based on the projection angle shift Δα. The larger the shift Δα, the smaller the absolute light intensity in the event detection pixel 1025.
[0086] Therefore, in addition to compensating for errors in the depth map due to time delay, the above procedure can also be used to obtain additional information about the observed absolute brightness values.
[0087] The above assumes that all light fragments emitted from the projection unit are of the same color, i.e., generated by a light source having a given wavelength (spectrum), such as a laser. However, multiple time-series light fragments may also include series of light fragments of multiple different colors. That is, for each series of light fragments of one color shifted in the first direction x1, there exists a series of light fragments of the same color shifted in the second direction x2.
[0088] This means that the apparent projection angle measurements described above are performed for different colors, as schematically shown in Figure 8. Here, it is shown that a sweep is first performed in the first direction x1 using red light R. Next, a reverse sweep is performed in the second direction x2 using red light R. Then, a sweep is performed in the first and second directions using green light G, and then a sweep is performed using blue light B. Of course, other colors / wavelengths and other color sweep sequences can be used, as long as sweeps are performed in both directions for each color.
[0089] By using projection light of different colors, the effects of delay in event detection pixels can be further reduced. In fact, surfaces in a scene may have different brightness levels depending on their color. Therefore, the delay in event detection will differ for each color. Consequently, depth maps generated with different colors will have varying degrees of error. In particular, the error on the observed surface may be large for some colors and small for others. For example, a sweep with green light will make green surfaces appear brighter, while other colors will make them appear darker. Therefore, by combining the sweep results of different colors, the accuracy of the resulting depth map can be further improved, even if delay compensation is not completely successful. Thus, using different colors for the light segments in the above sequence further enhances the reliability of the depth map.
[0090] Additionally or alternatively, the control unit 1030 can be configured to calculate the absolute values of the light intensity of each different color used in multiple time-series optical segments. Thus, as described above, it becomes possible to calculate the absolute light intensity for each color. This can be used to supplement color information in the generated depth map or to generate a color image of the scene.
[0091] Additionally or alternatively, the light-receiving unit 1020 may include a color filter array 1027 that provides different colored color filters 1028 to the event detection pixels 1025. An example of such a color filter array 1027 is shown in Figure 9 as an RGB color filter 1028. Of course, any other arbitrary colors or arrangement of color filters 1028 can be used. The control unit 1030 is configured to calculate the absolute value of the light intensity of each different color used in the color filters 1028. Furthermore, the use of a color filter array allows for the acquisition of luminance information of different colors, which can be used to supplement color information in the generated depth map or to generate a color image of the scene. Of course, combining the use of projected light of different colors with the use of color filters can further improve the reliability of the color information.
[0092] In this way, a highly reliable depth map, with brightness or color information supplemented as needed, can be generated using a method with high temporal resolution obtained through event detection.
[0093] The following briefly describes some example application areas for the sensor device 1000 described above.
[0094] Figures 10A and 10B schematically show a camera device 2000 including the sensor device 1000 described above. Here, the camera device 2000 is configured to generate depth information of a shooting scene including the object O in the manner described above.
[0095] Figure 10A shows a smartphone used to acquire depth information, such as a depth map, of an object O. This could be used to improve the augmented reality capabilities of smartphones or to enhance the gaming experience available on smartphones. Figure 10B shows a face capture sensor that can be used, for example, for face recognition at airports or border control, viewpoint correction or virtual makeup in web conferences, or for animating chat avatars for web conferences or games. Furthermore, filmmakers / animators can use such EVS-enhanced face capture sensors to adapt animated characters to real people.
[0096] Figure 11 shows a head-mounted display 3000 equipped with the sensor device 1000 described above as a further example. This head-mounted display 3000 is configured to generate depth information of an object O that is viewed through the head-mounted display 3000 as described above. This example could be used for accurate hand tracking or gesture recognition in augmented reality or virtual reality applications, for example, to assist in complex medical tasks.
[0097] Figure 12 schematically shows an industrial manufacturing apparatus 4000 including the sensor device 1000 described above. This industrial manufacturing apparatus 4000 includes means 4010 for moving an object O in front of the projection unit 1010, thereby (partially) realizing the projection of illumination patterns to different positions on the object O. The industrial manufacturing apparatus 4000 is also configured to generate depth information of the object O based on the positions of the images of these illumination patterns. For example, since the conveyor belt constituting the means 4010 for moving the object O moves at high speed, depth information can only be generated if the light receiving unit 1020 has a sufficiently high temporal resolution, so this application is particularly suitable for EVS-enhanced depth sensors. The EVS-enhanced sensor device 1000 described above satisfies this condition, making it possible to obtain accurate and high-speed depth maps of industrial manufactured objects O, thereby realizing fully automated, high-precision, and high-speed quality control of manufactured objects O. Depth information may include a depth map and / or object position classification on the conveyor belt, deviation information from a desired manufacturing standard, error classification, etc. Here, the means 4010 for moving the object O must pass the object in front of the sensor device 1000 at least twice in opposite directions. Alternatively, the projection unit 1010 must sweep in both directions. During this time, the conveyor belt must be stopped (or the conveyor belt may be moving if the resulting motion blur is compensated for). In this example, synchronization is performed after knowing the conveyor speed that is fed back to the sensor device 1000.
[0098] Figure 13 summarizes the steps for generating a scene depth map using the sensor device 1000 described above. The method for operating the sensor device 1000 is as follows.
[0099] In S110, multiple time-series light fragments are projected onto the scene. Consecutive light fragments in one series are shifted along a first direction x1 or a second direction x2 opposite to the first direction x1.
[0100] In S120, the reflection of light fragments from the scene is observed.
[0101] In S130, the projected light fragments and the observed reflections are correlated in time, and a depth map of the scene is generated using this correlation and a predetermined distance.
[0102] In S140, while the optical element is shifting in the first direction x1 and while the optical element is shifting in the second direction x2, for each event occurrence, a predetermined solid angle of each event detection pixel 1025 and the projection angle onto which the optical element was projected at the time of the event are calculated.
[0103] Then, in S150, for each given pair of event occurrences, the projection angles calculated for the shift of the light fragment in the first direction x1 and the shift of the light fragment in the second direction x2 are averaged, and a depth map of the scene is calculated based on the pair of a predetermined solid angle and the average projection angle calculated for each event occurrence.
[0104] In this way, the accuracy and reliability of the event-based depth map generation described above can be improved.
[0105] As described above, calculating an accurate depth map requires the projection of each series of light fragments in two directions. Typically, such projection is performed by sweeping, for example, using light fragments such as lines, to project the entire projection range of the projection unit 1010 in a continuous manner, i.e., without jumps. In this way, the projection unit 1010 can project a first sweep light and a second sweep light moving in opposite directions onto the object. In each first sweep light, the continuous series of light fragments is shifted along the first direction x1, and in each second sweep light, the continuous series of light fragments is shifted along the second direction x2.
[0106] This is schematically illustrated in Figures 14A and 14B. Figure 14A shows the projection unit 1010 and light-receiving unit 1020 of the sensor device 1000 under two different sweeps. On the left side of Figure 14A, the projection angle α of the light fragment (e.g., a line) is increasing. That is, the continuous position of the light fragment is continuously shifting along the first direction x1. On the right side of Figure 14A, the projection angle α is decreasing. As a result, the light fragment is continuously moving along the second direction x2. In other words, the left side of Figure 14A shows the first sweep light, and the right side of Figure 14A shows the second sweep light.
[0107] The alternating operation of the first and second sweeping beams can be produced, for example, by periodically increasing or decreasing the projection angle α, as shown in Figure 14B. Each pair of the first and second sweeping beams can, in principle, be used to calculate an accurate depth map, as described above. Note that, as shown in Figure 14B, the speed at which the angle changes, i.e., the speed at which the light strip traverses the illumination scene, does not necessarily have to be constant; it can be any speed as long as the scene is scanned along the first direction x1 and the second direction x2. Preferably, the speed during the sweep along the second direction x2 is the same as during the sweep along the first direction, but changes in the opposite direction. That is, the change in projection angle α is temporally symmetric with respect to the turning point shown by the dotted line in Figure 14B.
[0108] As shown in Figures 15A and 15B, this can lead to a decrease in the readout frequency compared to the typical case where a depth map is established after each sweep. Figure 15A illustrates this typical case. Here, at the end of each sweep, the correlation between each projection angle α and the solid angle of the pixel that detects light from the object at the corresponding angle is evaluated to construct a depth map of the entire object. Therefore, the readout frequency in this case is f1 = 1 / t, where t is the sweep time.
[0109] In contrast, the method described above requires two sweeps to calculate the depth map. This means that the integration time becomes twice as long as 2t, and the readout frequency becomes f2 = 1 / 2t = f1 / 2. Therefore, the readout frequency f2 is only half of the readout frequency f1 in the general case.
[0110] To mitigate this problem, the control unit 1030 is configured to calculate a depth map of the scene based on a predetermined solid angle and an average projection angle calculated for each event occurrence, for each consecutive first sweep light and second sweep light, and for each consecutive second sweep light and first sweep light, for the shift of the light element in the first direction x1 and the shift in the second direction x2, respectively.
[0111] This is shown in Figure 15C. In contrast to Figure 15B, we do not wait for a pair of sweeps to be completely finished before performing a single readout. Instead, each sweep is used for two readout processes. For example, in Figure 15C, the sweep numbered "2" is combined with the sweep numbered "1" to construct a depth map. Subsequently, the sweep numbered "2" is also combined with the next sweep numbered "3" to construct a depth map. In this way, a depth map can be provided after each sweep. That is, two readouts are possible during an integration time of 2t, and the readout frequency is again f1 = 2·1 / 2t = 1 / t.
[0112] It should be noted that, in principle, readouts can be performed at the end of any time window of length 2t. Therefore, the number of readouts can be increased during a single integration time of 2t. However, this does not mean that more information will be obtained than when readouts are performed after time t, i.e., the time it takes for one sweep to complete. Furthermore, the readout times can be shifted by the same amount arbitrarily. In other words, readouts are not necessarily tied to the end of a single sweep. However, each readout must take into account the information collected in the current sweep and the information obtained in the previous sweep performed in the opposite direction over a time interval of (2n+1)t, where n is a natural number or 0 (preferably n=0).
[0113] This type of interleaved readout makes it possible to generate accurate and reliable event-based depth maps without reducing the readout frequency.
[0114] Figure 16 is a perspective view showing an example of a stacked structure of a solid-state imaging device 23020 in which multiple pixels are arranged in a matrix, capable of implementing the functions described above. Each pixel includes at least one photoelectric conversion element.
[0115] The solid-state imaging device 23020 has a stacked structure consisting of a first chip (upper chip) 910 and a second chip (lower chip) 920.
[0116] The stacked first chip 910 and second chip 920 can be electrically connected to each other via through-contact (silicon) vias (TC(S)V) formed in the first chip 910.
[0117] The solid-state imaging device 23020 can be formed to have a stacked structure by bonding a first chip 910 and a second chip 920 at the wafer level and cutting them out by dicing.
[0118] In a stacked structure of two chips, the first chip 910 can be an analog chip (sensor chip) that includes at least one analog component for each pixel, for example, photoelectric conversion elements arranged in an array. For example, the first chip 910 may include only photoelectric conversion elements.
[0119] Alternatively, the first chip 910 may include further elements of each photoreceptor module. For example, the first chip 910 may include at least some or all of the n-channel MOSFETs of the photoreceptor module in addition to the photoelectric conversion element. Alternatively, the first chip 910 may include each element of the photoreceptor module.
[0120] The first chip 910 may also include a portion of the pixel backend 300. For example, the first chip 910 may include a memory capacitor, or in addition to the memory capacitor, a sampling / hold circuit and / or buffer circuit electrically connected between the memory capacitor and the event detection comparator circuit. Alternatively, the first chip 910 may include the entire pixel backend. Referring to Figure 13A, the first chip 910 may also include the readout circuit 140, the threshold generation circuit 130, and / or the controller 120, or at least a portion of the entire control unit.
[0121] The second chip 920 may be a logic chip (digital chip) that mainly includes elements that complement the circuitry on the first chip 910 for the solid-state imaging device 23020. The second chip 920 may also include analog circuitry, such as circuitry that quantizes the analog signals transferred from the first chip 910 via the TCV.
[0122] The second tip 920 may have one or more bonding pads BPD, and the first tip 910 may have an opening OPN used for wire bonding to the second tip 920.
[0123] A solid-state imaging device 23020 having a stacked structure of two chips 910 and 920 may have the following characteristic configuration.
[0124] The electrical connection between the first chip 910 and the second chip 920 is made, for example, via a TCV. The TCV can be located at the chip edge or between the pad area and the circuit area. The TCVs for transmitting control signals and supplying power can be concentrated, for example, mainly at the four corners of the solid-state imaging device 23020, thereby reducing the signal wiring area of the first chip 910.
[0125] Typically, the first chip 910 includes a p-type substrate, and the formation of a p-channel MOSFET usually implies the formation of an n-type doped well that separates the p-type source and drain regions of the p-channel MOSFET from each other and further separates them from other p-type regions. Therefore, by avoiding the formation of a p-channel MOSFET, the manufacturing process of the first chip 910 can be simplified.
[0126] Figure 17 schematically shows an example of the configuration of solid-state imaging devices 23010 and 23020.
[0127] The single-layer solid-state imaging device 23010 shown in part A of Figure 17 includes a single die (semiconductor substrate) 23011. The single die 23011 has a pixel region 23012 (photoelectric conversion element), a control circuit 23013 (readout circuit, threshold generation circuit, controller, control unit), and a logic circuit 23014 (pixel backend) mounted and / or formed on it. Pixels are arranged in an array within the pixel region 23012. The control circuit 23013 performs various controls, including pixel drive control. The logic circuit 23014 performs signal processing.
[0128] Sections B and C of Figure 17 show a schematic configuration example of a multilayer solid-state imaging device 23020 having a stacked structure. As shown in sections B and C of Figure 17, the solid-state imaging device 23020 has two dies (chips) stacked on top of each other: a sensor die 23021 (first chip) and a logic die 23024 (second chip). These dies are electrically connected to form a single semiconductor chip.
[0129] Referring to section B of Figure 17, the pixel region 23012 and the control circuit 23013 are formed or mounted on the sensor die 23021, and the logic circuit 23014 is formed or mounted on the logic die 23024. The logic circuit 23014 may include at least a portion of the pixel backend. The pixel region 23012 includes at least a photoelectric conversion element.
[0130] Referring to section C in Figure 17, the pixel region 23012 is formed or mounted on the sensor die 23021. On the other hand, the control circuit 23013 and the logic circuit 23014 are formed or mounted on the logic die 23024.
[0131] In another example (not shown), the pixel region 23012 and the logic circuit 23014, or a portion of the pixel region 23012 and the logic circuit 23014, may be formed or mounted on the sensor die 23021, and the control circuit 23013 may be formed or mounted on the logic die 23024.
[0132] In a solid-state imaging device having multiple photoreceptor modules PR, all photoreceptor modules PR can operate in the same mode. Alternatively, a first subset of photoreceptor modules PR may operate in a low SNR and high temporal resolution mode, while a second complementary subset operates in a high SNR and low temporal resolution mode. The control signal may be a function of user settings, for example, rather than a function of illumination conditions.
[0133] <Examples of applications to mobile devices> The technology disclosed herein (the Technology) can be applied to a variety of products. For example, the Technology disclosed herein may be implemented as a device mounted on any type of mobile vehicle, such as an automobile, electric vehicle, hybrid electric vehicle, motorcycle, bicycle, personal mobility device, airplane, drone, ship, or robot.
[0134] Figure 18 is a block diagram showing a schematic configuration example of a vehicle control system, which is an example of a mobile control system to which the technology described herein may be applied.
[0135] The vehicle control system 12000 comprises multiple electronic control units connected via a communication network 12001. In the example shown in Figure 18, the vehicle control system 12000 includes a drive system control unit 12010, a body system control unit 12020, an external information detection unit 12030, an internal information detection unit 12040, and an integrated control unit 12050. The functional configuration of the integrated control unit 12050 is shown in the figure, which includes a microcomputer 12051, an audio / image output unit 12052, and an in-vehicle network interface 12053.
[0136] The drivetrain control unit 12010 controls the operation of devices related to the vehicle's drivetrain according to various programs. For example, the drivetrain control unit 12010 functions as a control device for a drivetrain generating device that generates driving force for the vehicle, such as an internal combustion engine or a drive motor; a drivetrain transmission mechanism that transmits driving force to the wheels; a steering mechanism that adjusts the steering angle of the vehicle; and a braking device that generates braking force for the vehicle.
[0137] The body system control unit 12020 controls the operation of various devices mounted on the vehicle body according to various programs. For example, the body system control unit 12020 functions as a control device for a keyless entry system, a smart key system, a power window system, or various lamps such as headlights, reverse lights, brake lights, turn signals, or fog lights. In this case, the body system control unit 12020 may receive radio waves transmitted from a portable device that replaces a key or signals from various switches. The body system control unit 12020 receives these radio waves or signals and controls the vehicle's door lock system, power window system, lamps, etc.
[0138] The external information detection unit 12030 detects information from outside the vehicle equipped with the vehicle control system 12000. For example, an imaging unit 12031 is connected to the external information detection unit 12030. The external information detection unit 12030 causes the imaging unit 12031 to capture images of the outside of the vehicle and receives the captured images. Based on the received images, the external information detection unit 12030 may perform object detection processing such as detecting people, cars, obstacles, signs, or characters on the road surface, or distance detection processing.
[0139] The imaging unit 12031 is a solid-state image sensor equipped with the event detection function and photoreceiving module according to this disclosure, or may include such a sensor. The imaging unit 12031 can output an electrical signal as positional information to identify the pixel that detected the event. The light received by the imaging unit 12031 may be visible light or invisible light such as infrared light.
[0140] The in-vehicle information detection unit 12040 detects information inside the vehicle. The in-vehicle information detection unit 12040 is or may include an event detection function and a photoreceptor module as described herein. The in-vehicle information detection unit 12040 is connected to, for example, a driver state detection unit 12041 that detects the driver's state. The driver state detection unit 12041 includes, for example, a camera pointed at the driver, and the in-vehicle information detection unit 12040 may calculate the driver's level of fatigue or concentration, or determine whether the driver is drowsy, based on the detection information input from the driver state detection unit 12041.
[0141] The microcomputer 12051 can calculate control target values for the drive force generator, steering mechanism, or braking system based on information from inside and outside the vehicle acquired by the external information detection unit 12030 or the internal information detection unit 12040, and output control commands to the drive system control unit 12010. For example, the microcomputer 12051 can perform cooperative control aimed at realizing ADAS (Advanced Driver Assistance System) functions, including collision avoidance or impact mitigation, following based on distance between vehicles, maintaining vehicle speed, vehicle collision warning, or vehicle lane departure warning.
[0142] Furthermore, the microcomputer 12051 can perform cooperative control for purposes such as autonomous driving, where the vehicle drives autonomously without driver intervention, by controlling the drive force generating device, steering mechanism, or braking device, etc., based on information about the vehicle's surroundings acquired by the external information detection unit 12030 or the internal information detection unit 12040.
[0143] Furthermore, the microcomputer 12051 can output control commands to the body system control unit 12030 based on external information acquired by the external information detection unit 12030. For example, the microcomputer 12051 can control the headlights according to the position of a preceding or oncoming vehicle detected by the external information detection unit 12030, and perform coordinated control aimed at reducing glare, such as switching from high beams to low beams.
[0144] The audio-image output unit 12052 transmits at least one output signal, either audio or image, to an output device capable of visually or audibly notifying information to the vehicle's occupants or to those outside the vehicle. In the example shown in Figure 18, the output devices include an audio speaker 12061, a display unit 12062, and an instrument panel 12063. The display unit 12062 may include, for example, at least one of an onboard display or a head-up display.
[0145] Figure 19 shows an example of the installation position of the imaging unit 12031. In this figure, the imaging unit 12031 can have imaging units 12101, 12102, 12103, 12104, and 12105.
[0146] The imaging units 12101, 12102, 12103, 12104, and 12105 are installed, for example, on the front nose, side mirrors, rear bumper, back door, and the upper part of the windshield inside the vehicle 12100. The imaging unit 12101 installed on the front nose and the imaging unit 12105 installed on the upper part of the windshield inside the vehicle mainly acquire images of the front of the vehicle 12100. The imaging units 12102 and 12103 installed on the side mirrors mainly acquire images of the sides of the vehicle 12100. The imaging unit 12104 installed on the rear bumper or back door mainly acquires images of the rear of the vehicle 12100. The imaging unit 12105 installed on the upper part of the windshield inside the vehicle is mainly used for detecting preceding vehicles, pedestrians, obstacles, traffic lights, traffic signs, or lanes.
[0147] Figure 19 shows an example of the imaging range of imaging units 12101 to 12104. Imaging range 12111 indicates the imaging range of imaging unit 12101 located on the front nose, imaging ranges 12112 and 12113 indicate the imaging ranges of imaging units 12102 and 12103 located on the side mirrors, respectively, and imaging range 12114 indicates the imaging range of imaging unit 12104 located on the rear bumper or back door. For example, by superimposing the image data captured by imaging units 12101 to 12104, an overhead view image of the vehicle 12100 can be obtained.
[0148] At least one of the imaging units 12101 to 12104 may have a function for acquiring distance information. For example, at least one of the imaging units 12101 to 12104 may be a stereo camera consisting of multiple image sensors, or an image sensor having pixels for phase difference detection.
[0149] For example, the microcomputer 12051, based on distance information obtained from imaging units 12101 to 12104, can determine the distance to each object within the imaging range 12111 to 12114 and the temporal change of this distance (relative speed to vehicle 12100). In particular, it can extract the nearest object on the vehicle 12100's path that is traveling in approximately the same direction as vehicle 12100 at a predetermined speed (e.g., 0 km / h or more) as the preceding vehicle. Furthermore, the microcomputer 12051 can set a predetermined distance to be maintained before the preceding vehicle and perform automatic braking control (including follow-and-stop control) and automatic acceleration control (including follow-and-start control), etc. In this way, cooperative control aimed at autonomous driving, where the vehicle drives autonomously without driver intervention, can be performed.
[0150] For example, the microcomputer 12051 can use distance information obtained from imaging units 12101 to 12104 to classify and extract three-dimensional object data related to three-dimensional objects, such as motorcycles, passenger cars, heavy vehicles, pedestrians, utility poles, and other three-dimensional objects, and use this data for automatic obstacle avoidance. For example, the microcomputer 12051 identifies obstacles around the vehicle 12100 into obstacles that are visible to the driver of the vehicle 12100 and obstacles that are difficult to see. The microcomputer 12051 then determines the collision risk, which indicates the degree of risk of collision with each obstacle. If the collision risk is above a set value and there is a possibility of collision, the microcomputer 12051 can provide driving assistance to avoid collisions by outputting a warning to the driver via the audio speaker 12061 or display unit 12062, or by performing forced deceleration or evasive steering via the drive system control unit 12010.
[0151] At least one of the imaging units 12101 to 12104 may be an infrared camera that detects infrared light. For example, the microcomputer 12051 can recognize pedestrians by determining whether or not pedestrians are present in the images captured by the imaging units 12101 to 12104. Such pedestrian recognition is performed, for example, by a procedure to extract feature points from the images captured by the imaging units 12101 to 12104 as infrared cameras, and a procedure to perform pattern matching on a series of feature points that indicate the contour of an object to determine whether or not it is a pedestrian. When the microcomputer 12051 determines that a pedestrian is present in the images captured by the imaging units 12101 to 12104 and recognizes a pedestrian, the audio-image output unit 12052 controls the display unit 12062 to superimpose a rectangular contour line for emphasis on the recognized pedestrian. The audio-image output unit 12052 may also control the display unit 12062 to display an icon indicating a pedestrian at a desired position.
[0152] The above describes an example of a vehicle control system to which the system described herein can be applied. By applying an optical reception module that acquires event-triggered image information, it may be possible to reduce the amount of image data transmitted over the communication network, thereby reducing power consumption without adversely affecting driver assistance.
[0153] In addition, the embodiments of this technology are not limited to those described above, and various modifications are possible within the scope of this technology without departing from the gist of it.
[0154] The solid-state imaging device according to this disclosure may be any device used to analyze and / or process radiation such as visible light, infrared light, ultraviolet light, and X-rays. For example, the solid-state imaging device may be any electronic device in fields such as transportation, consumer electronics, medical and healthcare, security, beauty, sports, agriculture, and image reproduction.
[0155] Specifically, in the field of image reproduction, solid-state imaging devices can be devices that capture images provided for viewing, such as digital cameras, smartphones, and mobile phone devices with camera functions. In the field of transportation, solid-state imaging devices can be incorporated, for example, into on-board sensors that capture the front, rear, surroundings, and interior of a vehicle for safe driving purposes such as automatic stopping or driver status recognition, into surveillance cameras that monitor moving vehicles and roads, or into distance measuring sensors that measure the distance between vehicles.
[0156] In the consumer electronics sector, solid-state imaging devices can be used in devices mounted on home appliances such as television receivers, refrigerators, and air conditioners, and can be incorporated into all kinds of sensors that can capture user gestures and operate the device accordingly. Thus, solid-state imaging devices can be incorporated into home appliances such as television receivers, refrigerators, and air conditioners, and / or devices that control these appliances. Furthermore, in the medical and healthcare sector, solid-state imaging devices can be incorporated into all kinds of sensors (e.g., solid-state image sensors) used for medical and healthcare applications, such as endoscopes or devices that receive infrared light for angiography.
[0157] In the security field, solid-state imaging devices can be incorporated into devices used for security purposes, such as surveillance cameras or personal authentication cameras. Furthermore, in the beauty field, solid-state imaging devices can be used in devices used for beauty purposes, such as skin measuring instruments for photographing skin or microscopes for photographing probes. In the sports field, solid-state imaging devices can be incorporated into devices used for sports purposes, such as action cameras or wearable cameras for sports. Furthermore, in the agriculture field, solid-state imaging devices can be used in devices used for agricultural purposes, such as cameras for monitoring the condition of farmland and crops.
[0158] This technology can also be configured as follows:
[0159] [1] A sensor device (1000) that generates a depth map of a scene, A projection unit (1010) is configured to project multiple time-series light patches onto the scene, in which a series of light patches in one sequence are shifted along a first direction (x1) or a second direction (x2) opposite to the first direction (x1), A light receiving unit (1020) is positioned at a predetermined distance (d) from the projection unit (1010) and is configured to observe the reflection of the light fragments from the scene, A control unit (1030) is configured to temporally correlate the projected light fragment and the observed reflection, and to generate the depth map of the scene using the correlation and the predetermined distance (d). It is equipped with, The projection unit (1010) is configured to alternately project a first sweep light and a second sweep light, each first sweep light being a series of continuous light fragments shifted along the first direction (x1), and each second sweep light being a series of continuous light fragments shifted along the second direction (x2). The light receiving unit (1020) includes a plurality of event detection pixels (1025) that indicate the occurrence of an event when the light intensity measured by each event detection pixel (1025) changes by a predetermined amount or more during a given time period. The time it takes for one event detection pixel (1025) to indicate the occurrence of the event depends on the absolute value of the light intensity before the change in light intensity. Each event detection pixel (1025) is configured to receive light only from a predetermined solid angle, the solid angles observed by adjacent event detection pixels (1025) are adjacent, and the observed solid angles of all event detection pixels (1025) cover the field of view of the light receiving unit (1020). The control unit (1030) is configured to calculate, during the first sweep light projection and during the second sweep light projection, for each event that occurs, the predetermined solid angle of each event detection pixel and the projection angle at which the projection unit (1010) projected one light fragment at the time the event occurred, and The control unit (1030) is configured to calculate a depth map of the scene based on the pair of a predetermined solid angle and the average projection angle calculated for each pair of event occurrences, for each consecutive first sweep light and second sweep light, and for each consecutive second sweep light and first sweep light. Sensor device (1000). [2] The sensor device (1000) described in [1], The aforementioned time-series optical fragments are time-series rays (L). Sensor device (1000). [3] [2] The sensor device (1000) described above, The aforementioned ray (L) is a straight line that intersects the first direction (x1) at a predetermined angle, or The aforementioned ray is curved. Sensor device (1000). [4] A sensor device (1000) described in any one of items [1] to [3], The control unit (1030) is configured to calculate the absolute value of the light intensity before the change in light intensity by subtracting the projection angles calculated for the shift of the light element in the first direction (x1) and the shift of the light element in the second direction (x2) for each pair of event occurrences. Sensor device (1000). [5] A sensor device (1000) described in any one of items [1] to [4], The plurality of time-series light fragments include light fragments of different colors in each series, and For each series of light fragments of one color that is shifted in the first direction (x1), there exist light fragments of the same color that are shifted in the second direction (x2). Sensor device (1000). The sensor device (1000) described in [6] [5], The control unit (1030) is configured to calculate the absolute value of the light intensity for each of the different colors used in the plurality of time-series light strips. Sensor device (1000). A sensor device (1000) described in any one of the items [7] [1] to [6], The light-receiving unit (1020) includes a color filter array (1027) that provides a color filter (1028) of a different color to the event detection pixel (1025), and The control unit (1030) is configured to calculate the absolute value of the light intensity for each of the different colors used in the color filter (1028). Sensor device (1000). [8] A method for operating the sensor device (1000) described in [1], Projecting onto the scene a plurality of time-series light fragments in which a series of light fragments in one sequence are shifted along a first direction (x1) or a second direction (x2) opposite to the first direction (x1), wherein the projection is performed by alternately projecting a first sweep and a second sweep, each first sweep being a series of consecutive light fragments shifted along the first direction, and each second sweep being a series of consecutive light fragments shifted along the second direction. Observe the reflection of the light element from the aforementioned scene, The projected light fragment and the observed reflection are correlated in time, and the depth map of the scene is generated using the correlation and the predetermined distance (d). During the first sweep light projection and during the second sweep light projection, for each event that occurs, the predetermined solid angle of each event detection pixel (1025) and the projection angle of one light fragment projected at the time the event occurred are calculated, and For each pair of event occurrences, the projection angles calculated for each consecutive first and second sweep light, and for each consecutive second and first sweep light, are averaged, and the depth map of the scene is calculated based on the pair of predetermined solid angle and average projection angle calculated for each event occurrence. method.
Claims
1. A sensor device that generates a depth map of a scene, A projection unit configured to project multiple time-series light patches onto the scene, in which a series of light patches in one sequence are shifted along a first direction or a second direction opposite to the first direction, A light receiving unit is positioned at a predetermined distance from the projection unit and configured to observe the reflection of the light fragments from the scene, A control unit configured to temporally correlate the projected light fragment with the observed reflection, and to generate the depth map of the scene using the correlation and the predetermined distance. It is equipped with, The projection unit is configured to alternately project a first sweep light and a second sweep light, each first sweep light being a series of continuous light fragments shifted along the first direction, and each second sweep light being a series of continuous light fragments shifted along the second direction. The light receiving unit includes a plurality of event detection pixels that indicate the occurrence of an event when the light intensity measured by each event detection pixel changes by a predetermined amount or more during a given time period. The time it takes for one event detection pixel to indicate the occurrence of the event depends on the absolute value of the light intensity before the change in light intensity. Each event detection pixel is configured to receive light only from a predetermined solid angle, the solid angles observed by adjacent event detection pixels are adjacent, and the observed solid angles of all event detection pixels cover the field of view of the light receiving unit. The control unit is configured to calculate, during the first sweep light projection and during the second sweep light projection, for each event that occurs, the predetermined solid angle of each event detection pixel and the projection angle at which the projection unit projected one light fragment at the time the event occurred, and The control unit is configured to calculate a depth map of the scene based on the pair of a predetermined solid angle and the average projection angle calculated for each pair of event occurrences, for each consecutive first sweep light and second sweep light, and for each consecutive second sweep light and first sweep light. Sensor device.
2. A sensor device according to claim 1, The aforementioned time-series light fragments are time-series light rays. Sensor device.
3. A sensor device according to claim 2, The aforementioned ray is a straight line that intersects the first direction at a predetermined angle, or The aforementioned ray is curved. Sensor device.
4. A sensor device according to claim 1, The control unit is configured to calculate the absolute value of the light intensity before the change in light intensity by subtracting the projection angles calculated for the shift of the light element in the first direction and the shift of the light element in the second direction, for each pair of event occurrences. Sensor device.
5. A sensor device according to claim 1, The plurality of time-series light fragments include light fragments of different colors in each series, and For each series of light particles of a single color that is shifted in the first direction, there exist light particles of the same color that are shifted in the second direction. Sensor device.
6. A sensor device according to claim 5, The control unit is configured to calculate the absolute value of the light intensity for each of the different colors used in the plurality of time-series light pieces. Sensor device.
7. A sensor device according to claim 1, The light receiving unit includes a color filter array that provides a color filter of a different color to the event detection pixel, and The control unit is configured to calculate the absolute value of the light intensity for each of the different colors used in the color filter. Sensor device.
8. A method for operating the sensor device described in claim 1, Projecting onto the scene a plurality of time-series light fragments in which a series of light fragments in one sequence are shifted along a first direction or a second direction opposite to the first direction, wherein the projection is performed by alternately projecting a first sweep and a second sweep, each first sweep being a series of consecutive light fragments shifted along the first direction, and each second sweep being a series of consecutive light fragments shifted along the second direction. Observe the reflection of the light element from the aforementioned scene, The projected light fragment and the observed reflection are correlated in time, and the depth map of the scene is generated using the correlation and the predetermined distance. During the first sweep light projection and during the second sweep light projection, for each event that occurs, the predetermined solid angle of each event detection pixel and the projection angle of one light fragment projected at the time the event occurred are calculated, and For each pair of event occurrences, the projection angles calculated for each consecutive first and second sweep light, and for each consecutive second and first sweep light, are averaged, and the depth map of the scene is calculated based on the pair of predetermined solid angle and average projection angle calculated for each event occurrence. method.
Citation Information
Patent Citations
Depth data measuring head, measurement device and measuring method
US20220155059A1
Multiple channel locating
US9696137B2
Motion compensation in phase-shifted structured light illumination for measuring dimensions of freely moving objects
WO2019008330A1
Sensor device and method for operating a sensor device
WO2023001916A1
Solid-state imaging device and method for operating a solid-state imaging device
WO2023001943A1