Method for depth measurement using a time-of-flight camera using amplitude-modulated continuous light

By adopting a method of acquiring a sample sequence and determining a confidence value at a high sampling frequency in an amplitude-modulated continuous light time-of-flight camera, the problem of inaccurate depth measurement caused by object motion is solved, and more accurate depth measurement is achieved.

CN113474673BActive Publication Date: 2025-09-30IEE INT ELECTRONICS & ENG SA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080014378.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-02-15
Filing Date
2020-02-14
Publication Date
2025-09-30
Estimated Expiration
2040-02-14

AI Technical Summary

Technical Problem

Existing amplitude-modulated continuous light time-of-flight cameras are easily disturbed by object motion during the measurement process, resulting in inaccurate depth measurements, especially erroneous and unexpected depth values ​​near the edges of objects.

Method used

A sample sequence is acquired at a sampling frequency higher than the amplitude modulation frequency, a confidence value is determined for each pixel, and a contribution is made to the depth value of the binned area based on the confidence value, thereby reducing the influence of motion artifacts through an intelligent binning method.

Benefits of technology

The accuracy and reliability of depth measurement are improved, the error caused by object motion is reduced, and the reliability and accuracy of binned depth values ​​are ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113474673B_ABST
    Figure CN113474673B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for depth measurement using a time-of-flight camera (1) using amplitude-modulated continuous light, the method comprising the following steps: for each pixel of a plurality of pixels (3) of a sensor array (2) of the camera (1), acquiring (120) at least one sample sequence at a sampling frequency higher than the modulation frequency of the amplitude-modulated continuous light, the sample sequence comprising at least four amplitude samples (A0, A1, A2, A3). In order to enable accurate and effective depth measurement using the time-of-flight camera, the present invention provides that the method further comprises: - determining (130) a confidence value (C) for each sample sequence of each pixel (3), the confidence value indicating the degree of correspondence between the amplitude sample (A0, A1, A2, A3) and the sinusoidal time evolution of the amplitude; and - determining (260) a binned depth value (D) for each binning area of ​​a plurality of binning areas (4) based on the amplitude samples (A0, A1, A2, A3) of the sample sequence of the pixels (3) of the binning area (4) b ), each bin area includes a plurality of pixels (3), wherein the depth value (D b ) depends on its confidence value (C).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] In general terms, the present disclosure relates to a method for depth measurement using a time-of-flight camera using amplitude-modulated continuous light. Background Art

[0002] Time-of-flight cameras are used to provide pixel-by-pixel depth information in images of three-dimensional objects or landscapes. The camera comprises a (typically two-dimensional) sensor array with multiple pixels. Each pixel provides information from which the depth of the recorded point in space (i.e., its distance from the camera) can be derived. In addition to TOF cameras that use light pulses, another type of TOF camera uses amplitude-modulated continuous light. In other words, the camera emits a continuous field of amplitude-modulated light, which is reflected from objects in the camera's field of view. The reflected light is received by each pixel. Due to the amplitude modulation, the phase of the received light can be deduced from the amplitude, and the time of flight can be derived from the relative phase difference, thereby determining the distance to the reflecting object. According to a known method, locked pixels are used, in which the readout of each pixel is synchronized with the modulation frequency of the light. Specifically, the readout frequency of each pixel can be four times the modulation frequency. This is also known as the 4-tap method. It is based on receiving and evaluating four consecutive amplitude samples at four consecutive time points, with each time interval corresponding to a 90° phase shift. Each amplitude measurement can be referred to as a tap.

[0003] A potential problem with the 4-tap approach is that the object may move during the measurement. Consequently, the amplitude samples detected for consecutive taps may correspond to different actual depths, such as the depth of a moving object in the foreground and the depth of the background. For example, if the object's motion has a component perpendicular to the camera's optical axis, a given pixel may correspond to part of the object at one tap and part of the background at the next, or vice versa. This problem can occur with pixels near the edge of an object in the image and often results in erroneous and undesirable depths. This can also occur if the object merely moves away from or towards the camera, due to the stretching or shrinking perceived by the camera, so pixels near the edge may also change between object and background. This effect is also referred to in the literature as "flying pixels."

[0004] EP 2 966 475 A1 discloses a method for binning time-of-flight (TOF) data from a scene to improve the accuracy of TOF measurements and reduce noise therein, wherein the TOF data includes phase data and confidence data. According to the method, multiple TOF data are acquired by illuminating the scene with multiple modulated signals, and each modulated signal is associated with a vector defined by phase and confidence data. The multiple vectors are summed to obtain binned vectors, and the phase and confidence data of the binned vectors are processed to obtain depth data of the scene. According to the description, the "confidence" corresponds to the amplitude of the reflected signal.

[0005] DE 10 2015 195 161 A1 discloses a device for detecting the motion of an object in a target space, wherein the object is located at a certain distance from an image capture device, the image capture device being configured to measure the distance and provide a sensor signal indicative of the distance, wherein if the object is stationary, the sensor signal can be decomposed into a decomposition including odd harmonics. The device comprises a determination circuit configured to receive the sensor signal and generate at least one motion signal based on at least one even harmonic of the decomposition of the sensor signal; and a detection circuit configured to detect the motion of the object based on the at least one motion signal and provide a detection signal indicative of the motion of the object. The image capture device is configured to capture an image comprising a plurality of pixels, and the determination circuit is configured to receive the sensor signal and generate a motion signal for each of the plurality of pixels without relying on neighboring pixels of the plurality of pixels. Summary of the Invention

[0006] Purpose of the Invention

[0007] Therefore, an object of the present invention is to achieve accurate and efficient depth measurement using a time-of-flight camera.

[0008] According to one aspect of the present invention, there is provided a method for performing depth measurement using a time-of-flight camera using amplitude-modulated continuous light, the method comprising:

[0009] - acquiring, for each pixel of a plurality of pixels of a sensor array of the camera, at least one sample sequence at a sampling frequency higher than a modulation frequency of the amplitude-modulated continuous light, the at least one sample sequence comprising at least four amplitude samples;

[0010] Characterized in that the method further comprises:

[0011] - determining for each sequence of samples for each pixel a confidence value, said confidence value indicating how well said amplitude sample corresponds to the sinusoidal time evolution of the amplitude; and

[0012] —Determining a depth value for a bin for each of a plurality of bin regions based on amplitude samples of a sample sequence of pixels from the bin region, each of the plurality of bin regions comprising a plurality of pixels, wherein a contribution of the sample sequence to the depth value of the bin depends on a confidence value thereof.

[0013] The present invention provides a method for depth measurement using a time-of-flight camera using amplitude-modulated continuous light. Depth measurement in this context naturally refers to measuring the distance from the camera in order to obtain a 3D image. The principles of a time-of-flight (TOF) camera using amplitude-modulated continuous light are well known and have been explained above. While the term "light" may refer to visible light, it should be understood that infrared or ultraviolet light may also be used.

[0014] In a first step, the method comprises acquiring, for each of a plurality of pixels of a sensor array of a camera, at least one sample sequence at a sampling frequency higher than the modulation frequency of the amplitude-modulated continuous light, the sample sequence comprising at least four amplitude samples. The pixels may in particular be lock pixels. The sensor array comprises a plurality (typically between several hundred and several thousand) of pixels, typically arranged in a two-dimensional pattern, although a one-dimensional arrangement is also conceivable. The sample sequence comprises at least four amplitude samples, which are sampled at a sampling frequency higher than the modulation frequency of the amplitude-modulated continuous light. In particular, the sampling frequency may be an integer multiple of the modulation frequency. The amplitude samples typically correspond to the amplitude of a correlation function between the transmitted signal and the received signal. For pixels corresponding to the surface of a stationary object, the correlation function should be sinusoidal, i.e. it should correspond to a sine function (or a cosine function, respectively) with a normal non-zero phase shift.

[0015] In another step of the method, a confidence value is determined for each sample sequence for each pixel. This confidence value indicates the degree to which the amplitude samples correspond to the sinusoidal time evolution of the amplitude. There are various methods for determining the confidence value, some of which are discussed below. Typically, the confidence value is defined so that a high degree of correspondence has a high confidence value, while a low degree of correspondence has a low confidence value. This method step is based on the assumption that for stationary objects, the amplitude samples should correspond to a sinusoidal function. Due to various factors such as measurement error, the amplitude samples will typically not correspond perfectly to the sinusoidal time evolution, but only to a certain extent. However, if the object is in motion and some amplitude samples actually correspond to signals received from the object's surface while others correspond to signals received from the background, there will typically be a significant discrepancy between the amplitude samples and any sinusoidal function. In other words, the amplitude samples may not correspond to the sinusoidal time evolution at all, which will affect the confidence value. It should be noted that at least four amplitude samples are generally required to determine the degree of correspondence. A sinusoidal function can be described by four parameters: amplitude, offset, phase, and frequency. Since the frequency of the sinusoidal function is known in this case, three parameters remain. Therefore, it is always possible to find a "fitting" sine function for 3 (or fewer) amplitude samples. On the other hand, if there are 4 or more amplitude samples, any deviation from the sinusoidal time evolution can be determined.

[0016] In another step of the method, for each of a plurality of binning regions, each binning region comprising a plurality of pixels, a depth value for the bin is determined based on amplitude samples of a sample sequence of pixels from the binning region, wherein the contribution of the sample sequence to the depth value of the bin depends on its confidence value. This method step may be described as a "binning step." Such binning is known in the art, but the method of the present invention applies a previously unknown variant that may be referred to as "smart binning," etc. A plurality of binning regions are defined, each of which comprises a plurality of pixels. Each binning region may be rectangular, for example, comprising mxn pixels or nxn pixels. The binning regions may also be referred to as pixel groups. Typically, each binning region is coherent, i.e., it corresponds to a coherent region of the sensor array. It is conceivable that two binning regions overlap, such that a given pixel belongs to more than one binning region. However, typically, the different binning regions are separate. In this case, the binning regions may collectively be considered as units or "pixels" of the low-resolution image, while the pixels of the sensor array correspond to the high-resolution image. The magnitude samples of the pixel sample sequences from a binned region contribute to the depth value of the bin for that binned region. Alternatively, the information from the magnitude samples of the pixel sample sequences from the binned region is combined to obtain the depth value for the bin. There are several possibilities for calculating the depth value for a bin, some of which are discussed further below.

[0017] If a depth value were determined for a single pixel (and a single sample sequence), this would be accomplished by determining the relative phase of a sine function, which would correspond to the depth value. However, as mentioned above, since the phase of the sine function is determined based on the amplitude samples, if some amplitude samples correspond to the object and some correspond to its background, the phase information will be affected. Any depth information derived from a sample sequence with such amplitude samples is largely useless. However, according to the method of the present invention, a confidence value is determined for each sample sequence for each pixel, thereby allowing the reliability of the phase information to be assessed. Then, when using information from the amplitude samples of a given pixel to determine the depth value of a bin, not all sample sequences of pixels in the binned area are treated equally (as in binning methods known in the art). Instead, the contribution of a sample sequence to the depth value of the bin depends on its confidence value. Qualitatively, the contribution of a sample sequence with a high confidence value is generally greater than that of a sample sequence with a low confidence value. The latter's contribution may even be zero, meaning that the sample sequence can be completely ignored.

[0018] The concept of the present invention allows for increased reliability of the binned depth values, since the "corrupted" sequence of amplitude samples is less or not considered at all. Furthermore, the confidence value can be determined based solely on the information of each pixel, i.e., without taking other pixels into account. As will become apparent below, the confidence value can also be calculated with minimal processing power and memory. This also means that the calculation can be performed in real time, even with simple, low-cost processing units.

[0019] Preferably, the method includes acquiring four amplitude samples at a sampling frequency four times higher than the modulation frequency of the amplitude modulated continuous light. This corresponds to the well-known 4-tap method. The sine wave modulated received signal with amplitude A and phase φ can be represented by a two-dimensional vector:

[0020] r(A,φ)=(A·cosφ,A·sinφ)=(d 13 ,d 02 ) (Formula 1)

[0021] Among them, d 02 =A0-A2, hereinafter referred to as the first difference, and d 13 =A1-A3, hereinafter referred to as the second difference, which are two amplitude samples A k , k = 0…3. The first amplitude sample A0 corresponds to a phase angle of 0°, the second amplitude sample A1 corresponds to a phase angle of 90°, the third amplitude sample A2 corresponds to a phase angle of 180°, and the fourth amplitude sample A3 corresponds to a phase angle of 270°. Therefore, the amplitude and phase of the received signal can be calculated as

[0022] φ=atan2(d 02 ,d13 ) (Formula 2)

[0023]

[0024] While the amplitude A of the signal is proportional to the number of received photons, the phase φ is proportional to the depth D of the object seen by the corresponding pixel.

[0025]

[0026] Where D is the measured depth of the object from the camera, c is the speed of light, and f mod is the modulation frequency of the signal.

[0027] Therefore, the depth can be calculated as

[0028]

[0029] There are various options for how the confidence value can be defined. One possible definition of the confidence value C is as follows:

[0030]

[0031] The amplitude A is calculated according to formula 3, but can be approximated as

[0032]

[0033] In a simple embodiment, only one sample sequence is acquired for each pixel, so that for the confidence value and for the depth value of the bin, only the amplitude samples of this sample sequence can be considered. In such an embodiment, it can also be said that a (single) confidence value is determined for each pixel, because each pixel corresponds to a single sample sequence. In another embodiment, the method comprises: acquiring multiple sample sequences for at least one pixel. These can also be referred to as multiple exposures. On the one hand, the multiple exposures that have to be performed in sequence increase the likelihood that some pixels are affected by the movement of the object. On the other hand, considering multiple sample sequences can help to reduce the influence of noise or other measurement errors. In particular, tap measurements can be performed for different sample sequences for acquiring amplitude samples with different integration times, i.e. each sample sequence corresponds to a different integration time. In addition to changing the exposure time, other parameters can also be changed, such as the modulation frequency.

[0034] Even if multiple sample sequences are acquired for a given pixel, each confidence value can be determined by the relationship between the amplitude samples of the individual sample sequences. In other words, the amplitude samples of a sample sequence are considered without considering the amplitude samples of other sample sequences (if any). Equation 6 is an example of such an embodiment. In most cases, this approach is reliable because the amplitude samples of a single sample sequence are not affected by, for example, different exposure times. However, it should be noted that if multiple sample sequences have been determined, a separate confidence value should be determined for each sample sequence, for example because one sample sequence for a given pixel may not be affected by object motion, while another sample sequence is affected, thus producing unreliable data.

[0035] There are several possibilities for how different sample sequences can contribute to the depth value of a bin based on their respective confidence values. For example, there can be a continuous range of contributions or weighting factors. Another possibility can be called a "binary" classification. In this case, the method includes classifying each sample sequence as valid if the confidence value meets a predefined criterion and as invalid otherwise. In other words, each sample sequence can only be valid or invalid. In this case, the method also includes using the amplitude samples of the sample sequence to determine the depth value of the bin only if the pixel is valid. In other words, if a sample sequence is considered to be invalid, the amplitude samples of this sample sequence will be completely ignored and will not affect the depth value of the bin. If only one sample sequence is considered for each pixel, it can also be said that each pixel is classified as valid or invalid separately. The same is true if only the sample sequence of a specific exposure is considered separately.

[0036] In particular, the sample sequence can be classified based on the relationship between the confidence value and a first threshold. In other words, it is determined whether the confidence value is above or below the first threshold, and the sample sequence is classified as valid or invalid depending on the result. Generally, for a high degree of confidence, the confidence value is high, and if the confidence value is above the first threshold, the sample sequence is classified as valid. The first threshold (which may also be referred to as a motion parameter) is usually predefined. It can be estimated, calculated or can be determined by calibration, for example using a still landscape in the absence of moving objects. Graphically, this approach can be described by a confidence mask that masks all pixels having a confidence value below the first threshold.

[0037] According to one embodiment, the depth value of a bin is determined based on a linear combination of amplitude samples from sample sequences of pixels in the binned region, where the contribution of each sample sequence to the linear combination depends on the confidence value of the respective sample sequence. In other words, in the linear combination, each amplitude sample is assigned a coefficient that depends on the confidence value of the corresponding sample sequence. In general, of course, the coefficient will be higher for high confidence values ​​and lower for low confidence values. Specifically, if the confidence value is above a first threshold, the coefficient may be 1, while if the confidence value is below the first threshold, the coefficient may be zero. In this case, the linear combination corresponds to the sum of all valid sample sequences, while ignoring all invalid sample sequences. For example, the first and second differences described above may be defined for each valid sample sequence, and then a "binned" first and second difference corresponding to the sum of the individual difference values ​​for all valid sample sequences may be defined. It will be appreciated that this is equivalent to first summing the amplitude samples to determine the "binned" amplitude sample, and then calculating the first and second difference values ​​for the bin. Once the binned interpolation values ​​are determined, the "binned" phase can be calculated according to Equation 2. Summing the differences can also be viewed as vector addition, where the first and second differences are components of a vector, and this vector addition yields the binned vector. Therefore, this approach can also be referred to as "vector binning" or "weighted vector binning."

[0038] According to another embodiment, the depth value D of the bin b is determined by averaging pixel depth values ​​of sample sequences of pixels from the binned region, wherein the weight of each pixel depth value depends on the confidence value of the corresponding sample sequence of the corresponding pixel, and wherein the pixel depth value is determined based on the amplitude samples of the sample sequence of the pixel. In other words, the pixel depth value D is determined separately for each sample sequence of pixels in the binned region, or may be determined only for valid sample sequences of pixels in the binned region. These pixel depth values ​​D may be determined, for example, using Equation 2 and Equation 5. The pixel depth values ​​are then averaged to determine the depth value D for the bin. b , but in a weighted manner such that the weight depends on the confidence value of the corresponding sample sequence of the corresponding pixel. This approach may also be referred to as "depth binning" or "weighted depth binning". Similarly, for high confidence values, the weight of the pixel depth value is higher, while for low confidence values, the weight of the pixel depth value is lower. Specifically, if the confidence value is below a first threshold, i.e., only the pixel depth value for a valid sample sequence is considered, the weight of the pixel depth value may be zero. If there is only one sample sequence for each pixel, the calculation may be performed as follows:

[0039]

[0040] In such an embodiment, the depth value of the bin corresponds to the (arithmetic) mean of the pixel depth values ​​of all valid pixels (i.e., pixels with a valid sample sequence). If multiple sample sequences are acquired for each pixel, the pixel depth value is determined separately for each sample sequence, and Equation 8 must be modified to average all valid sample sequences for all pixels (or all valid pixels for all integration times).

[0041] In one embodiment, the method comprises determining a first difference between a first amplitude sample and a third amplitude sample of a sample sequence of a pixel, and assigning sample sequences having a positive first difference to a first group, and assigning sample sequences having a negative first difference to a second group. Specifically, this may refer only to valid sample sequences, while invalid sample sequences are not included in either of the two groups. The first difference d has already been mentioned above. 02 =A0-A2. When the first difference and the second difference are considered as components of a vector, the phase difference between any two vectors with a positive first difference is less than 180°, and the phase difference between any two vectors with a negative first difference is less than 180°. When two vectors from the first group (or the second group, respectively) are added, the phase of the resulting vector is between the phases of the added vectors. Accordingly, according to Formula 5, the depth corresponding to the result vector is also between the depths corresponding to the two added vectors. It should be noted that, in general, it is possible that all sample sequences have positive first differences, or it is also possible that all sample sequences have negative first differences. If this is the case, there is of course no need to divide the sample sequences into groups.

[0042] Furthermore, the method may include defining a vector having a second difference between the second amplitude sample and the fourth amplitude sample as a first component and having the first difference as a second component. The second difference d has been explained above. 13 =A1-A3. In addition, the method may include: defining a first set of vectors r P =[x P, y P ], which is a linear combination of the confidence values ​​of the vectors in the first group based on the corresponding sample sequence, and the second group of vectors r M =[x M, y M], which is a linear combination of the vectors in the second group based on the confidence values ​​of the corresponding sample sequences. In the first group of vectors, the vectors in the first group are linearly combined based on the confidence values ​​of the respective sample sequences. In other words, the coefficients or weights of the vectors depend on the confidence values. Specifically, if the confidence value is below a first threshold, the weight may be zero, and if the confidence value is above the first threshold, the weight may be one, in which case only valid sample sequences are added. The same applies to the linear combination of the vectors in the second group. More specifically, each group vector may be the sum of all vectors in the corresponding group, in which case the first and second group vectors r P ,r M The components of are calculated as follows, where the summation of multiple integration times is optional:

[0043]

[0044] Here, each sample sequence corresponds to a separate integration time. In Formulas 9a-9d, it is assumed that there are the same number of sample sequences (or integration times, respectively) for each pixel. The sum of "valid pixels" should be understood as the sum of all pixels with valid It-th (It=0...n) sample sequences. Alternatively, all pixels and all valid sample sequences of the corresponding pixels can be summed. Similar to Formula 2, the phase of the group vector is calculated as follows:

[0045] as well as

[0046] As described above, the phases (and therefore the depths) corresponding to the first set of vectors are within the interval of the respective vectors of the first set. Furthermore, the phases corresponding to the second set of vectors are within the interval of the respective vectors of the second set. In another step, the method includes determining the depth values ​​for the bins based on the phase difference between the second set of vectors and the first set of vectors. In other words, the phase difference (or an amount dependent on the phase difference) between the second set of vectors and the first set of vectors is determined, and the determination (or calculation) of the depth values ​​for the bins is dependent on the phase difference. Typically, it is assumed that both phases are between 0° and 360°.

[0047] According to one embodiment, the method further comprises: determining the depth value of the bin based on both the first set of vectors and the second set of vectors if the phase difference is below a second threshold; and determining the depth value of the bin based on only one of the first set of vectors and the second set of vectors if the phase difference is above the second threshold. As with the first threshold, the second threshold may also be estimated, calculated or determined by calibration using a stationary scene. If the phase difference is below the second threshold, this typically corresponds to a situation where all (or most) of the first difference values ​​have a relatively low absolute value, where some of the first difference values ​​are positive and some are negative, while the second difference values ​​are mostly negative. In this case, it may be assumed that the first set of vectors and the second set of vectors differ only to a small extent, and the first set of vectors may be added to the second set of vectors,

[0048] r b =r P +r M (Formula 10)

[0049] Afterwards, Formula 5 can be used again to determine the depth value of the bin based on the resulting binned vectors. On the other hand, if the phase difference is above the second threshold, this indicates that the first set of vectors and the second set of vectors may correspond to different depths, such as the depth of a foreground object and the background. In this case, it is more appropriate to completely discard one set of vectors and determine the depth value of the bin based only on the other vector.

[0050] According to one embodiment, the second threshold is 180°. This can be seen as a minimum condition for avoiding unrealistic depth values ​​outside the interval given by the first and second sets of vectors. However, the second threshold can be smaller, for example less than 90° or less than 60°. If the second threshold is 180°, then it is not necessary to explicitly calculate the phase difference. Instead, it can be shown that the phase difference is below 180° under the following conditions:

[0051] x M y P <x P y M

[0052] Testing this condition requires only two multiplications and processing power can therefore be saved.Alternatively, the phase difference can also be calculated explicitly based on the phases of the first and second set of vectors.

[0053] If the phase difference is above a second threshold, there are a number of possible criteria for determining which of the first and second group vectors is considered more reliable. In general, it is reasonable to assume that a group vector is more reliable if it is based on a greater number of valid sample sequences. Therefore, if the phase difference is above a second threshold, the depth value of the bin can be determined based on the group vector of the group with more valid sample sequences. In other words, if the first group has more valid sample sequences than the second group, the depth value of the bin is determined based on the first group vector, and vice versa:

[0054]

[0055] Among them, N P ,N M are the number of valid sample sequences in the first and second groups, respectively.

[0056] According to another embodiment, the method comprises: if the first components x of the two group vectors P ,x M If both the first component and the second component are negative, the depth value of the bin is determined based on both the first and second set of vectors; and if at least one first component is positive, the depth value of the bin is determined based on only one of the first and second set of vectors. If both the first component of the first and second set of vectors are negative, it means that the group vectors are located in the second and third quadrants, corresponding to a phase close to 180° and an object depth close to half the blur depth. However, if at least one first component is positive, the vectors are either located in opposite quadrants (the first and third quadrants or the second and fourth quadrants, respectively), or in the first and fourth quadrants, where one vector corresponds to a phase close to 360° or a depth close to the blur depth. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Further details and advantages of the present invention will become apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings, in which:

[0058] Figure 1 is a schematic diagram of a TOF camera and an object that can be used in the method of the present invention;

[0059] Figure 2 is a graph showing the time evolution of the function and four amplitude samples;

[0060] Figure 3 It is a vector graph;

[0061] Figure 4 is a graph illustrating amplitude values ​​on a sensor array;

[0062] Figure 5 is another diagram showing the time evolution of a function and a number of amplitude samples;

[0063] FIG6 is a first diagram illustrating a result of depth measurement according to the prior art;

[0064] FIG7 is a second diagram illustrating the result of depth measurement according to the prior art;

[0065] FIG8 is a vector diagram illustrating vector addition;

[0066] FIG9 is a third diagram showing the result of depth measurement according to the prior art;

[0067] Figure 10 is a flow chart illustrating a first embodiment of the method of the present invention.

[0068] Figure 11 is a diagram illustrating a binary confidence mask;

[0069] Figure 12 is a diagram illustrating the construction of a binary confidence mask and the application of the confidence mask;

[0070] Figure 13 It is another vector diagram that illustrates vector addition;

[0071] Figure 14 is a vector diagram illustrating the positions of two sets of vectors;

[0072] Figure 15 is a first diagram showing the result of depth measurement according to the present invention;

[0073] Figure 16 is a second diagram showing the result of depth measurement according to the present invention;

[0074] Figure 17 is a third diagram showing the result of depth measurement according to the present invention;

[0075] Figure 18 is a flow chart illustrating a second embodiment of the method of the present invention; and

[0076] Figure 19 is a fourth diagram showing the result of depth measurement according to the present invention. DETAILED DESCRIPTION

[0077] Figure 1A TOF camera 1 suitable for depth measurement using amplitude-modulated continuous light is schematically shown. It comprises a rectangular sensor array 2 having a plurality (e.g. thousands or tens of thousands) of pixels 3. In addition, it may comprise a memory and a processing unit, which are not shown for simplicity. The camera 1 is configured to emit amplitude-modulated continuous light 10 using one or more light emitters 5. The light 10 is reflected by a 3D object 20 or scene in the field of view of the camera 1 and the reflected light 11 is received by the pixels 3 of the sensor array 2. The original modulation function s(t) with a phase delay τ is correlated with the reception function q(t) to produce a correlation function c(τ). The amplitude of the reception function q(t) is proportional to the modulation frequency f of the light 10. mod In other words, four amplitude samples A 0…3 , also known as a tap, is used to retrieve the phase φ of the modulated light, as Figure 2 As shown. Every four amplitude samples A 0…3 is part of the sample sequence for the corresponding pixel 3.

[0078] A sinusoidal signal with amplitude A and phase φ can be represented by a 2D vector that can be determined from the 4 amplitude samples determined from the tap measurements, i.e.

[0079] r(A,φ)=(A·cosφ,A·sinφ)=(d 13 ,d 02 ) (Formula 1)

[0080] Among them, d 02 =A0-A2, hereinafter referred to as the first difference, and d 13 =A1-A3, hereinafter referred to as the second difference, which are two amplitude samples A k , k = 0…3. Therefore, the amplitude and phase of the received signal can be calculated as

[0081]

[0082] While the amplitude A of the signal is proportional to the number of received photons, the phase φ is proportional to the depth D of the object seen by the corresponding pixel.

[0083]

[0084] Where D is the pixel depth value, that is, the distance of the pixel from the camera, c is the speed of light, and f mod is the modulation frequency of the signal. Therefore, the depth can be calculated as

[0085]

[0086] Figure 3 When the modulation frequency is f mod= 20 MHz from an object at depth D = 2 m.

[0087] However, motion artifacts may occur along the edges of an object 20 moving in the scene of camera 1. Since tap measurements are performed subsequently, a pixel 3 close to the object edge may "see" the object surface during one acquisition of amplitude samples, while in a subsequent acquisition it may see the background. Figure 4 The occurrence of motion artifacts is illustrated by way of example. During the acquisition of each amplitude sample, an object 20 having the shape of an 'O' moves 1 pixel upward and 1 pixel to the left. The grayscale value represents the number of tap measurements taken when the object was present at the corresponding pixel. Black pixels represent zero tap measurements when the object was present, while white pixels represent four (out of four) tap measurements when the object was present. The full image is on the right, and a zoom around the outer edge of the upper left corner of the 'O' is on the left.

[0088] Figure 5 The error introduced by the motion of the object 20 is illustrated. Depending on what the pixel 3 sees during acquisition, the measured amplitude sample A k Located on the sinusoidal cross-correlation curve of the foreground or background signal. If the 4 subsequent amplitude samples A of a pixel 3 are k Adding together, it is possible to identify blurring effects along the edges of the object 20. In a neighborhood of 4x4 pixels 3, it is possible to find pixels 3 that "see" the foreground object 20 in all taps, as well as pixels 3 that partially see the foreground object 20 or the background. If an acquisition with a different integration time is subsequently performed, this corresponds to an additional sequence of samples, which may also correspond to different depths.

[0089] Since the depth and reflectivity of foreground and background objects are different, the amplitude sample A k can vary significantly. Therefore, the phase and depth calculated according to Equations 2 and 5 may be erroneous. The results are illustrated in Figure 6, which is a high-resolution depth image of an "O"-shaped target at a depth of 2m in front of a 7m background, shifted by 1 pixel per tap acquisition in both the horizontal and vertical directions. Due to motion artifacts, the calculated depth varies between 1.65m and 5.41m along the edge of the object. It should be noted that the measured depth is not only between the foreground and background depths, but may also be outside this depth range. The corresponding pixels can be referred to as flying pixels.

[0090] According to the prior art, there are two main ways to alleviate this problem, both of which use a binning approach. Multiple pixels, such as 4x4 pixels, are considered as binned regions, for which a single depth value is determined. In the first approach, the amplitude samples A of all pixels in the binned region are summed. k Add them together and use Formula 2 (using the sum instead of the individual amplitude samples Ak ) and Formula 5 to calculate a single depth value. This method can be called "tap binning". The result is shown in Figure 7. It is recognized that there are outliers in the depth measured between 0.31m and 6.93m. In other words, there are still depth values ​​outside the depth range of the object 20 and the background, and the flying pixel effect is even increased compared to the high-resolution image of Figure 6. One reason for this increase can be understood from the vector diagram of Figure 8. Increasing the amplitude sample A k This corresponds to the addition of the vectors shown in FIG8 for the first vector representing the depth of the object 20 and the second vector representing the background. Since the phase difference between the two vectors exceeds 180°, the phase of the resulting vector is smaller than the phase of either vector. This results in a depth value outside the depth range.

[0091] According to another approach, a pixel depth value is determined for each individual pixel in the binned area and these pixel depth values ​​are averaged to determine the depth value for the binned area. This approach may be referred to as "pixel binning." The result is shown in FIG9 . Averaging causes the depth values ​​in the range of 1.81 m to 6.9 m to become blurred.

[0092] The above problems are reduced or eliminated by the method of the present invention. Figure 10 is a flow chart illustrating a first embodiment of the method of the present invention.

[0093] After the method starts, a binning region 4 is selected at 100. This can be, for example, a region comprising 4 x 4 pixels 3 (see also Figure 12 ). Next, at 110, a pixel 3 within the binning area 4 is selected. At 120, an amplitude sample A is determined for the sample sequence of the pixel 3. k At 130, based on the amplitude sample A k To calculate the confidence value C. An individual confidence value C is calculated for each sample sequence of the corresponding pixel 3, that is, if there is only one sample sequence, a confidence value C is calculated for each pixel. One possible definition of the confidence value C is as follows:

[0094]

[0095] The amplitude A is calculated according to formula 3, but can be approximated as

[0096]

[0097] According to this definition, the confidence value C is always in the range between 0 and 1, where the highest possible value 1 represents a perfect sine function. At 140, the confidence value C is compared with the first threshold C minThe first threshold value can be calculated, estimated or determined by calibration using a stationary scene. min It can be 0.25. The first threshold C min It can also be called a "motion parameter" because it may be suitable for distinguishing between sample sequences affected by object motion and sample sequences not affected by object motion. min , the corresponding sample sequence is classified as invalid at 190 and is not considered further. On the other hand, if the confidence value C is greater than the first threshold C min , then the corresponding sample sequence is classified as valid at 150. The amplitude value or the first and second difference d 02 ,d 13 are considered as the second and first components of the vector, respectively, and are retained for further processing.

[0098] This process can be viewed as the creation of a binary confidence mask, which is Figure 11 Illustrated graphically in . Figure 11 The upper part of Figure 4 The lower half shows the corresponding confidence mask, wherein the left half is an enlarged view of the portion near the edge of the object 20. Black represents pixels that are considered invalid, while white represents pixels with sample sequences that are considered valid. It is recognized that the area where the tap is blurred is masked by the confidence mask.

[0099] Figure 12 The construction and binning process of the confidence mask for a binning area 4 of 4×4 pixels 3 is further illustrated, where for simplicity a single sample sequence for each pixel 3 is assumed. First, as shown in a), a separate amplitude sample is determined for each pixel (wherein the different shading represents the sampling number or time point, respectively). Then, as shown in b), a confidence value is determined for each pixel (wherein dark shades represent high confidence values). The confidence mask is shown at c), where black represents invalid pixels (or sample sequences, respectively) and white represents valid pixels. Using this confidence mask together with the individual amplitude samples effectively produces binned amplitude samples for the entire binning area 4, as shown in d), where the different shading again represents the sampling number or time point, respectively). If several sequences corresponding to several integration times are considered, a confidence mask can be constructed for each integration time.

[0100] At 160, a first difference d is determined. 02If the sign is positive, the sample sequence and its vector are assigned to the first group at 170, and if the sign is negative, the sample sequence and its vector are assigned to the second group at 180. As indicated by the dotted arrows, steps 160, 170 and 180 can also be skipped in a simplified version of the method.

[0101] The steps mentioned so far are repeated for all pixels 3 in the binned region 4 and, where applicable, for all sample sequences for each pixel 3. When it is determined at 200 that the last pixel 3 has been processed, the method continues at 210 by calculating the first set of vectors r by adding the vectors in the first and second sets, respectively. P =[x P ,y P ] and the second set of vectors r M =[x M ,y M ]. In other words, the pair has a positive first difference d 02 Sum all vectors of , and find the vectors with negative first difference d 02 Therefore, the first and second set of vectors r P ,r M The components of are calculated as follows, where summation over multiple integration times is optional:

[0102]

[0103]

[0104] It's important to note that in both the first and second groups, only vectors with valid sample sequences are summed, while invalid sample sequences are ignored for the binning process. The phase difference between any two vectors in the first group is less than 180°, so adding these vectors will not result in flying pixels. The same applies to vectors in the second group. The fact that the summed vectors have a phase difference of less than 180° ensures that the resulting group vector is unaffected by binning effects.

[0105] At 220, the phase difference ΔΦ between the second set of vectors and the first set of vectors is calculated (assuming both phases are between 0° and 360°) and compared with the second threshold Φ max In particular, the second threshold Φ max It can be equal to 180°. If the phase difference ΔΦ is small, such as Figure 13 As shown in the example, the first and second groups of vectors r P ,r M The binned vector r is simply summed at 230 to calculate b =[x b ,y b ],Right now:

[0106] r b =r P +r M (Formula 10)

[0107] If the phase difference ΔΦ is large, such as Figure 14 As shown in the example of , this can indicate that these groups correspond to pixels 3 of the background and foreground objects, respectively. In either case, the two group vectors r P ,r M For these reasons, a group vector r is selected at 240. P ,r M vector r as bins b , that is, the group vector of the larger group (that is, the group with a larger number of valid sample sequences):

[0108]

[0109] Among them, N P ,N M are the number of valid sample sequences in the first and second groups respectively. Finally, the vector r based on the binning is calculated using Formula 2 and Formula 5. b To determine the depth value D of the bin b .

[0110] If several sample sequences with several integration times It = 1, 2, .. n are recorded, then the binned values, e.g. component x b ,y b Can be standardized as:

[0111]

[0112] Among them, N It is the number of pixels 3 that have a valid sample sequence during a specific integration time, and T It is the length of the integration time. This produces a normalized amplitude that does not show artificial jumps, since it is independent of the number of pixels considered in the bin, and thus allows the application of standard image processing methods such as stray light compensation on the bin taps or amplitude.

[0113] In a simplified version of the method indicated by the dashed line, at 250, all vectors of valid pixels 3 are summed to determine the bin vector r b :

[0114]

[0115] Afterwards, based on the binned vector r b Determine the depth value D of the bin b . Figure 15The results of this simplified version are shown in Figure 7. Compared to Figure 7, a significant improvement can be seen, but there are still anomalous "flying pixels" with depths outside the depth range of binned pixel 3. This problem is reduced if the vectors are assigned to the first and second groups and processed separately, as shown in Figure 7. Figure 16 Although Figure 16 The results for a single integration time are shown, Figure 17 Results for two integration times are shown, but the first integration time is 4 times longer than the second. In this case, the number of flying pixels is reduced to an almost negligible amount.

[0116] There are two possible alternative ways to check the inequality of the phase difference ΔΦ at 220. First, the following relationship can be checked:

[0117] x M y P <x P y M (Formula 13)

[0118] If yes, the method continues at 230, and if no, it continues at 240. This condition is related to the slope of the vector (x, y), which is proportional to y / x. For the critical case of distinguishing whether the angle between two vectors is less than or greater than 180°, one of the vectors must be in quadrant 1 and the other in quadrant 4, or one vector must be in quadrant 2 and the other in quadrant 4. For all other cases, the distinction is trivial. If one vector is in quadrant 1 and the other in quadrant 4, the angle between the two vectors is significantly greater than 180°. If one vector is in quadrant 2 and the other in quadrant 3, the angle between the two vectors is significantly less than 180°.

[0119] Secondly, we can determine x P and x M Are both negative, which means that the object depth is close to half of the blur depth. If so, the method continues at 230, and if not, continues at 240.

[0120] Figure 18 1 is a flow chart illustrating a second embodiment of the method of the present invention. Steps 100, 110, 120, 130, and 140 are identical to those of the first embodiment and are not described again for the sake of brevity. If the sample sequence is classified as valid at 150, then at 155, a pixel depth value D is determined according to Equations 2 and 5. After processing all pixels 3 in the bin region 4, the depth value D of the bin is determined by averaging the pixel depth values ​​D. b :

[0121]

[0122] If multiple sample sequences are acquired for each pixel 3, the pixel depth value is determined separately for each sample sequence and Equation 8 must be modified to average over all valid sample sequences for all pixels 3 or all valid pixels 3 for all integration times. Calculate the depth value D for the bin b The result is a low-resolution depth image, where the depth value D of the bin b Represents the arithmetic mean of the valid pixels in each bin area 4. Figure 19 An example of a depth image calculated in this embodiment is shown. Compared to FIG9 , which shows the result of the averaging process without distinguishing between valid and invalid pixels, the effect of flying pixels is reduced. However, it should be noted that this second embodiment of the method of the present invention requires calculating the pixel depth value D for each valid pixel (and each sample sequence, where applicable) on the full high-resolution image, which may result in increased computational workload and / or increased memory requirements.

[0123] Reference Signs List

[0124] 1. Time of Flight (TOF) camera

[0125] 2 sensor arrays

[0126] 3 pixels

[0127] 4 binning areas

[0128] 5 Memory

[0129] 10 Light

[0130] 11 Reflected Light

[0131] 20 objects

Claims

1. A method for depth measurement using a time-of-flight camera (1) using amplitude-modulated continuous light (10), the method comprising: - acquiring (120) at least one sample sequence for each pixel of a plurality of pixels (3) of a sensor array (2) of the camera (1) at a sampling frequency higher than a modulation frequency of the amplitude-modulated continuous light, the at least one sample sequence comprising at least four amplitude samples (A0, A1, A2, A3); Characterized in that the method further comprises: - determining (130) for each sample sequence for each pixel (3) a confidence value (C), said confidence value indicating how well said amplitude sample (A0, A1, A2, A3) corresponds to the sinusoidal time evolution of the amplitude; - determining a first difference (d) between a first amplitude sample (A0) and a third amplitude sample (A2) of a sequence of samples of a pixel (3) 02 ), and assigning a positive first difference value (d 02 ) and its vector, and assigning a negative first difference value (d 02 )’s sample sequence and its vector; and - determining (260) a depth value (D) of a bin for each of a plurality of binning regions (4) based on amplitude samples (A0, A1, A2, A3) of a sample sequence of pixels (3) from said binning region (4); b ), each of the plurality of binned regions comprises a plurality of pixels (3), wherein the depth value (D b ) depends on its confidence value (C), —wherein the first set of vectors (r) is calculated by adding the vectors in the first set and the second set respectively P ) and the second set of vectors (r M ), and based on the first set of vectors (r P ) and the second set of vectors (r M ) to determine the depth value (D b ), and wherein the first group of vectors is a linear combination of the confidence values ​​(C) of the vectors in the first group based on the corresponding sample sequences, and the second group of vectors is a linear combination of the confidence values ​​(C) of the vectors in the second group based on the corresponding sample sequences.

2. The method according to claim 1, characterized in that The method comprises acquiring (120) four amplitude samples (A0, A1, A2, A3) at a sampling frequency four times higher than a modulation frequency of the amplitude modulated continuous light (10).

3. The method according to claim 1 or 2, characterized in that The method further comprises acquiring a plurality of sample sequences for at least one pixel (3).

4. The method according to claim 1 or 2, characterized in that The confidence value (C) is determined by the relationship between the amplitude samples (A0, A1, A2, A3) of the individual sample sequences.

5. The method according to claim 1 or 2, characterized in that The method further comprises: classifying (150, 190) each sample sequence as valid if the confidence value (C) satisfies a predefined criterion, otherwise classifying (150, 190) each sample sequence as invalid if the confidence value (C) does not satisfy the predefined criterion; and The amplitude samples (A0, A1, A2, A3) of the sample sequence are used to determine the depth value (D) of the bin only if the sample sequence is valid. b ).

6. The method according to claim 5, characterized in that The sample sequence is based on the confidence value (C) and the first threshold (C min ) relationship and are classified.

7. The method according to claim 1 or 2, characterized in that The depth value of the bin (D b ) is determined based on a linear combination of amplitude samples (A0, A1, A2, A3) of sample sequences of pixels (3) from said binned area (4), wherein said contribution of each sample sequence to said linear combination depends on a confidence value (C) of the corresponding sample sequence.

8. The method according to claim 1 or 2, characterized in that The depth value of the bin (D b ) is determined by averaging pixel depth values ​​(D) of a sample sequence of pixels (3) from the binned area (4), wherein a weight of each pixel depth value (D) depends on a confidence value of a corresponding sample sequence of the corresponding pixel (3), and wherein the pixel depth value (D) is determined based on amplitude samples (A0, A1, A2, A3) of the sample sequence of the pixel (3).

9. The method according to claim 1, characterized in that The method further comprises: The vector is defined to have a second difference (d ) between the second amplitude sample (A1) and the fourth amplitude sample (A3). 13 ) as the first component, and having the first difference (d 02 ) as the second component; and Based on the second set of vectors (r M ) and the first set of vectors (r P ) to determine the depth value (D b ).

10. The method according to claim 9, characterized in that The method further comprises: If the phase difference (Δφ) is lower than the second threshold (φ max ), then based on the first set of vectors (r P ) and the second set of vectors (r M ) to determine the depth value (D b );as well as If the phase difference (Δφ) is higher than the second threshold (φ max ), then only based on the first set of vectors (r P ) and the second set of vectors (r M ) to determine the depth value (D b ).

11. The method according to claim 10, characterized in that The second threshold (φ max ) is 180°.

12. The method according to claim 10 or 11, characterized in that If the phase difference (Δφ) is higher than the second threshold (φ ma ), then the depth value of the bin (D b ) is a group vector based on the group with more valid sample sequences (r P ,r M ) confirmed.

13. The method according to claim 1 or 2, characterized in that The method further comprises: If two group vectors (r P ,r M ) in the first component (x P ,x M ) are all negative, then based on the first set of vectors (r P ) and the second set of vectors (r M ) to determine the depth value (D b );as well as If at least one first component (x P ,x M ) is positive, then only based on the first set of vectors (r P ) and the second set of vectors (r M ) to determine the depth value (D b ).