Circuitry and method
The method enhances depth map generation by updating a 3D model based on current sparse depth data and previous data, addressing issues of invalid areas and low resolution in existing ToF imaging techniques.
Patent Information
- Application Number
- PCT/EP2024/082338
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-16
- Filing Date
- 2024-11-14
- Publication Date
- 2025-05-22
AI Technical Summary
Existing depth map generation techniques using time-of-flight (ToF) imaging face challenges such as invalid areas due to undetected light dots and low resolution, which affect the accuracy and detail of the depth maps.
The proposed circuitry and method involve obtaining sparse depth data at a first time frame, estimating the positional relation between the sparse depth data and a 3D model based on previous sparse depth data, updating the 3D model, and generating a depth map based on the updated model.
This approach improves the generation of depth maps by reducing invalid areas and increasing resolution, providing a more accurate and detailed representation of the scene.
Smart Images

Figure EP2024082338_22052025_PF_FP_ABST
Abstract
Description
[0001]Sony Semiconductor Solutions Corporation et al. 1 CIRCUITRY AND METHOD TECHNICAL FIELD The present disclosure generally pertains to circuitry and a method and, in particular, to circuitry and a method for generating a depth map. TECHNICAL BACKGROUND It is generally known to generate a depth map based on time-of-flight (ToF) imaging. ToF imaging includes emitting light into a scene, receiving a reflection of the light from the scene, and determining a respective depth for a plurality of points in the scene based on a respective duration between emitting the light and receiving a reflection of the light from the respective points. Although there exist techniques for generating a depth map, it is generally desirable to provide improved circuitry and an improved method for generating a depth map. SUMMARY According to a first aspect, the disclosure provides circuitry for generating a depth map, wherein the circuitry is configured to: obtain first sparse depth data acquired by a depth sensor at a first time frame, wherein the first sparse depth data indicate three-dimensional points of a scene in the first time frame; estimate a first positional relation between a representation of the scene of the first sparse depth data and a representation of the scene of a three-dimensional model, wherein the three-dimensional model is based on second sparse depth data acquired by the depth sensor at a second time frame prior to the first time frame; update the three-dimensional model based on the first sparse depth data and the first positional relation; and generate a depth map of the scene based on the updated three-dimensional model. According to a second aspect, the disclosure provides a method for generating a depth map, wherein the method includes: obtaining first sparse depth data acquired by a depth sensor at a first time frame, wherein the first sparse depth data indicate three-dimensional points of a scene in the first time frame; estimating a first positional relation between a representation of the scene of the first sparse depth data and a representation of the scene of a three-dimensional model, wherein the three-dimensional model is based on second sparse depth data acquired by the depth sensor at a second time frame prior to the first time frame; updating the three-dimensional model Sony Semiconductor Solutions Corporation et al. 2 based on the first sparse depth data and the first positional relation; and generating a depth map of the scene based on the updated three-dimensional model. Further aspects are set forth in the dependent claims, the drawings and the following description. BRIEF DESCRIPTION OF THE DRAWINGS Embodiments are explained by way of example with respect to the accompanying drawings, in which: Fig.1 illustrates an embodiment of a circuitry; Fig.2 illustrates an embodiment of a method; Fig.3 illustrates an embodiment of an RGB-D fusion pipeline; Fig.4 illustrates an embodiment of a pipeline in an RGB-D odometry block; Fig.5 illustrates a concept of an RGB-D odometry block according to an embodiment; Fig.6 illustrates a first embodiment of an RGB-D odometry block with an inertial measurement unit; Fig.7 illustrates a second embodiment of an RGB-D odometry block with an inertial measurement unit; Fig.8 illustrates an embodiment of an RGB-D odometry block with an external pose input; Fig.9 illustrates embodiments of a temporal buffer behavior; Fig.10 illustrates an embodiment of a timing relationship for generating depth maps based on a 3D model in a sequence of time frames; Fig.11 illustrates an embodiment of generating a depth map stream based on sparse depth data and image data acquired at different frame rates; and Fig.12 illustrates an embodiment of a general-purpose computer. DETAILED DESCRIPTION OF EMBODIMENTS Before a detailed description of the embodiments under reference of Fig.1 is given, general explanations are made. As mentioned in the outset, generating a depth map based on time-of-flight (ToF) includes emitting light into a scene, receiving a reflection of the light from the scene (e.g., detecting the reflection with a ToF sensor), and determining a respective depth for a plurality of points in the Sony Semiconductor Solutions Corporation et al. 3 scene based on a respective duration between emitting the light and receiving a reflection of the light from the respective points in the scene. For example, a depth map stream may be generated by consecutively generating and outputting respective depth maps for consecutive time frames. For example, the emitted light may include a plurality of light dots emitted into the scene, and the reflection may include reflected dots of respective ones of the plurality of light dots. However, in some instances, a depth may not be unambiguously determined for one or more points in the scene and / or a reflection of one or more of the emitted light dots is not detected by a ToF sensor (e.g., because the respective light dots may be reflected into another direction, because an amplitude of the respective reflected light dots may be too low to be detected, and / or because the respective reflected light dots may be distorted or subject to multipath or other device interference), which may cause one or more invalid areas in the generated depth map. Further, in some instances, a resolution of the generated depth map is low (e.g., the depth map may not indicate certain details of the scene, and / or a precision of a depth indicated by the depth map may be low) due to a limited number of dots in the plurality of emitted light dots. It has been recognized that, in a depth map stream, light dots received at one or more previous time frames may be used for avoiding an invalid area in a generated depth map and / or for increasing a resolution of a generated depth map. Consequently, some embodiments of the disclosure pertain to circuitry for generating a depth map, wherein the circuitry is configured to: obtain first sparse depth data acquired by a depth sensor at a first time frame, wherein the first sparse depth data indicate three-dimensional (3D) points of a scene in the first time frame; estimate a first positional relation between a representation of the scene of the first sparse depth data and a representation of the scene of a 3D model, wherein the 3D model is based on second sparse depth data acquired by the depth sensor at a second time frame prior to the first time frame; update the depth model based on the first sparse depth data and the first positional relation; and generate a depth map of the scene based on the updated 3D model. The circuitry may include a processing unit, e.g., programmed microprocessor, a field- programmable gate array (FPGA), an application-specific integrated circuit (ASIC) or the like. The circuitry may be configured to execute instructions (e.g., included in software and / or firmware stored on a non-transitory storage medium) that cause the circuitry to provide a functionality for generating a depth map, as described herein. Sony Semiconductor Solutions Corporation et al. 4 The circuitry may include a storage that may store the instructions and / or data required for providing the functionality described herein (e.g., the first sparse depth data, the second sparse depth data, the 3D model, an indication of the first positional relation, the generated depth map etc.). The storage may include a solid-state drive (SSD), a hard-disk drive (HDD), an electrically-erasable programmable read-only memory (EEPROM), a flash memory, a magnetic memory, a dynamic random-access memory (DRAM), a static random-access memory (SRAM), or the like. The circuitry may include an artificial intelligence (AI) unit that executes a machine learning model, for example, an artificial neural network. For example, the artificial neural network may include a Feed-Forward Network, a Residual Network (ResNet), a Recurrent Neural Network (RNN), a Convolutional Neural Network (CNN), a Generative Adversarial Network (GAN), a Transformer Neural Network and / or any other suitable neural network architecture. The skilled person may find a suitable architecture for the artificial neural network based on their general knowledge. The AI unit may include a graphics processing unit (GPU) and / or a tensor processing unit (TPU) for executing the machine learning model. The circuitry may further include a communication unit for receiving data from another device and / or for transmitting data to another device. For example, the communication unit may include a camera serial interface (CSI), a universal serial bus (USB), a peripheral component interconnect (PCI), a controller area network (CAN), Ethernet, Bluetooth, serial port (RS-232), parallel port (IEEE 1284) or the like. For example, the communication unit may be configured to receive the first and second sparse depth data from the depth sensor, e.g., in a case where the depth sensor is not included in the circuitry. For example, the communication unit may be configured to transmit the generated depth map to the other device, which may perform predetermined processing on the depth map and / or on a depth map stream. For example, the circuitry may include the general-purpose computer 250, as described with respect to Fig.12. The depth sensor may be included in the circuitry or may be provided separately from the circuitry, wherein, in the latter case, the circuitry may receive the first and second depth data from the depth sensor via the communication unit. The depth sensor may include a time-of-flight (ToF) sensor (e.g., indirect ToF (iToF), direct ToF (dToF), “lidar”, or the like). The depth sensor may include an array of photosensitive elements (“pixels”; e.g., photodiodes, single-photon avalanche diodes (SPADs) or the like), which may be Sony Semiconductor Solutions Corporation et al. 5 based on complementary metal-oxide-semiconductor (CMOS) and / or on charge-coupled device (CCD) technology. For acquiring depth data, such as the first and second sparse depth data, a light signal may be emitted onto the scene. The light signal may include one or more light pulses and may include an array of light dots (e.g., laser dots). The light signal may be generated by a light emission unit. The light emission unit may be included in the circuitry or, alternatively, may be provided separately from the circuitry. The light emission unit may be synchronized with the ToF sensor such that a round-trip time of the light signal from the light emission unit to the scene and, as a reflection, back to the ToF sensor may be determined based on the acquired depth data. The light emission unit may include a light-emitting diode (LED), a vertical-cavity surface-emitting laser (VCSEL) (e.g., an array of VCSELs for emitting the array of laser dots) or the like. The light signal may be reflected by the scene (e.g., by an object or a surface in the scene). The ToF sensor may receive the reflected light signal (e.g., via a lens for focusing the reflected light signal on the ToF sensor) and may detect the reflected light signal (e.g., may detect one or more photosensitive elements / pixels of the ToF sensor that receive a reflection of one or more of the emitted laser dots). The ToF sensor may generate the depth data (such as the first and / or second sparse depth data) and transmit the depth data to the circuitry (or to the processing unit of the circuitry). The depth data may include an indication of a respective round-trip time of one or more reflected laser dots, of a respective corresponding depth (e.g., distance to the reflecting object / surface), and / or of a position in the scene (e.g., an azimuth angle and / or an elevation angle as seen from the ToF sensor) where the respective laser dot has been reflected. For example, the first and second sparse depth data may indicate 3D points of the scene, and the 3D points may represent a respective distance / depth and position in the scene. For example, the 3D points may correspond to points in a 3D space of the scene and may have 3D coordinates that indicate the respective points in the 3D space. For example, each 3D point may correspond to a point in the 3D space of the scene from which a reflection of a laser dot is received. The first and second sparse depth data may be sparse in that they indicate a depth for the respective 3D points of the scene instead of the whole scene. For example, each indicated 3D point may correspond to a respective laser dot. A number and / or density of the emitted laser dots (and / or of respective photosensitive elements (pixels) that detect corresponding reflected laser dots) may be smaller than a number and / or density of photosensitive elements (pixels) of the ToF sensor, and the first and second sparse depth data may indicate a depth for pixels that have detected a reflected laser dot, but may not indicate a depth for pixels that have not detected a Sony Semiconductor Solutions Corporation et al. 6 reflected laser dot. For example, the first and second sparse depth data may not indicate a respective depth for a portion of the scene that does not correspond to one of the 3D points indicated by the first and / or second sparse depth data. The depth sensor may acquire depth data (e.g., the first and second sparse depth data) at a plurality of points in time (which may be represented by timestamps), wherein each point in time at which the depth sensor acquires depth data may correspond to a time frame or timestamp. For example, the depth sensor may acquire depth data at a predefined frame rate (e.g., 10 frames per second (fps), 20 fps, 30 fps, 50 fps, 100 fps or the like, without limiting the disclosure to these frame rates). The depth sensor may acquire the first sparse depth data at the first time frame subsequent to the second time frame, at which the depth sensor may acquire the second sparse depth data. The depth sensor may be moved with respect to the scene between the first time frame and the second time frame (e.g., between acquiring the first sparse depth data and acquiring the second sparse depth data), for example, because a user who may be holding or wearing the depth sensor may be moving, or because a vehicle or machine to which the depth sensor may be attached may be moving. Thus, when acquiring the first sparse depth data, the depth sensor may view the scene from a (e.g., slightly) different perspective than when acquiring the second sparse depth data, and respective representations of the scene according to the first and second sparse depth data may indicate the scene from (e.g., slightly) different perspectives. The first positional relation may correspond to a difference between these perspectives. For example, the first positional relation may correspond to a difference in a position (e.g., coordinate), an orientation (e.g., rotation around one or more Euler angles) or the like of the depth sensor and / or of respective representations of the scene by the first and second depth data. For example, the different perspectives may correspond to respective poses of the user who may be holding or wearing the depth sensor, and / or of the vehicle or machine to which the depth sensor may be attached, and the first positional relation may indicate a respective pose change. Accordingly, the estimating of the first positional relation may be based on odometry. The 3D model may represent the scene. The circuitry may generate the 3D model based on sparse depth data from the depth sensor and may update the 3D model for each time frame according to respective sparse depth data acquired at the time frame. The updating of the 3D model may include rigid or non-rigid registration of the 3D model such that a representation of Sony Semiconductor Solutions Corporation et al. 7 the scene of the respective time frame and a representation of the scene of the 3D model may coincide (e.g., within a predefined tolerance). Thus, the circuitry may estimate the first positional relation for determining by what amount the 3D model should be rotated, translated, scaled and / or sheared for updating according to the respective time frame, and may update the 3D model accordingly. The 3D model may include a 3D representation of the scene such that the 3D model can be transformed to correspond to a perspective from which the respective time frame represents the scene. The updating may also include merging depth information indicated by the respective sparse depth data into the 3D model. For example, the circuitry may merge the 3D points of the scene that are indicated by the respective sparse depth data into the 3D model. Thus, the updating of the 3D model may increase a detailedness of the 3D model. The circuitry may then generate a depth map of the scene based on the updated 3D model. For example, the generated depth map may view the scene from a perspective of the depth sensor at the respective time frame. For example, the 3D model may include a 3D representation of the scene. The circuitry may generate a projection of the updated 3D model into a two-dimensional (2D) depth map, e.g., into a 2D array of image pixels, wherein each image pixel indicates a depth value of the scene at a corresponding position in the scene (e.g., in a direction that corresponds to the image pixel). The circuitry may interpolate between 3D points of the scene that are indicated by the first and / or second sparse depth data such that a respective dense depth may be assigned to every image pixel of the generated depth map. For example, the circuitry may perform a depth completion for every time frame. The generated depth map may also include one or more image pixels for which the circuitry cannot determine a depth unambiguously. The circuitry may assign a dummy value to a pixel for which a depth cannot be determined unambiguously. The dummy value may include, e.g., zero, infinity, not-a-number (NaN) or the like, and may indicate that a depth cannot be determined for the respective image pixel. The circuitry may output (e.g., transmit via the communication unit) the generated depth map as a depth map that corresponds to the respective (e.g., first or second) time frame in a depth map stream. The circuitry may then obtain further sparse depth data of subsequent time frame, estimate a positional relation between respective representations of the scene of the 3D model and of the Sony Semiconductor Solutions Corporation et al. 8 further sparse depth data, update the 3D model accordingly, and generate an according depth map for the subsequent time frame in the depth map stream. For example, the circuitry may generate (and output) depth maps for the depth map stream at a predefined frame rate (e.g., 10 frames per second (fps), 20 fps, 30 fps, 50 fps, 100 fps or the like, without limiting the disclosure to these frame rates). The frame rate for generating (and outputting) depth maps for the depth map stream may, for example, correspond to a frame rate at which the depth sensor acquires sparse depth data. However, in some embodiments, the frame rate for generating depth maps for the depth map stream is different (e.g., lower) than the frame rate at which the depth sensor acquires sparse depth data. In some embodiments, the estimating of the first positional relation includes verifying the first positional relation based on data acquired by another sensor, wherein a positional relation between the other sensor and the depth sensor is fixed. For example, the data from the other sensor may provide a more accurate estimate of the first positional relation (e.g., of a pose change between the second time frame and the first time frame). For example, the verifying of the first positional relation may include determining a deviation of the first positional relation and a corresponding value indicated by (or determined based on) the data from the other sensor. The circuitry may decide, based on the deviation, whether the first positional relation is accurate (e.g., whether the deviation is smaller than a predefined threshold). If the first positional relation is accurate (e.g., the deviation is smaller than the predefined threshold), the circuitry may update the 3D model based on the first sparse depth data and the first positional relation. If the first positional relation is not accurate (e.g., the deviation exceeds the predefined threshold), the circuitry may discard (e.g., delete) the 3D model and generate a new 3D model based on the first sparse depth data. The circuitry may determine the deviation separately for certain portions of the 3D model and / or of the first sparse depth data, and may discard only a portion of the 3D model for which the deviation exceeds the predefined threshold. For example, the scene may change between the second time frame and the first time frame (e.g., a person or object in the scene may move), such that the deviation may exceed the predefined threshold for at least a portion of the 3D model, and the circuitry may update and / or regenerate the 3D model such that the 3D model represents the changed scene. The positional relation between the other sensor and the depth sensor may correspond to a relative position of the other sensor with respect to the depth sensor. A connection between the Sony Semiconductor Solutions Corporation et al. 9 depth sensor and the other sensor may be rigid such that a change of a position (coordinate and / or orientation) of the other sensor may cause a corresponding change of a position of the depth sensor and vice versa. Since the positional relation between the other sensor and the depth sensor is fixed, any movement (position change) of the depth sensor may cause a corresponding movement of the other sensor. Thus, data from the other sensor that indicate a movement of the other sensor between the second time frame and the first time frame may allow determining a corresponding movement of the depth sensor. For example, an electronic device that includes the depth sensor and the other sensor may be calibrated, which may include determining the relative position between the depth sensor and the other sensor. In some embodiments, the other sensor includes an image sensor; and the verifying of the first positional relation includes: obtaining first image data acquired by the image sensor at the first time frame; estimating a second positional relation between an image of the scene represented by the first image data and an image of the scene represented by second image data acquired by the image sensor at the second time frame; and comparing the first positional relation with the second positional relation. The image sensor may include a plurality of photosensitive elements (imaging pixels) that are configured to detect an amount of incident light for a predefined wavelength or range of wavelengths (e.g., in a visible range and / or in an infrared range). Image data (e.g., the first image data and the second image data) acquired by the image sensor may represent a respective image of the scene. For example, the image data may represent a red-green-blue (RGB) image, a grayscale image, an infrared image, or the like. The first image data may correspond to the first time frame in that the image sensor may acquire the first image data at the first time frame. Alternatively, if the image sensor acquires image data at a predefined frame rate that differs from a frame rate at which the depth sensor acquires depth data, the image sensor may acquire the first image data close to the first time frame. For example, the first image data may be the last image data acquired by the image sensor before the first time frame, the earliest image data acquired by the image sensor after the first time frame, or an interpolation thereof. Likewise, the image sensor may acquire the second image data at the second time frame, or the second image data may be the last image data acquired by the image sensor before the second time frame, the earliest image data acquired by the image sensor after the second time frame, or an interpolation thereof. Sony Semiconductor Solutions Corporation et al. 10 For example, the first image data and the second image data may represent the scene (and, thus, a pose / position / orientation of the image sensor) at a point in time that is sufficiently close to the first or second time frame, respectively, that an assumption that the first or second image data correspond to the first or second time frame, respectively, may hold within a predefined tolerance. Since the positional relation between the image sensor and the depth sensor may be fixed, the circuitry may determine an accuracy of the first positional relation based on comparing the second positional relation and the first positional relation. In some embodiments, the estimating of the second positional relation includes performing image registration for registering the first sparse depth data and the first image data. The image registration may include identifying corresponding portions of respective representations of the scene in the first sparse depth data and the first image data. For example, the image registration may include determining a transformation between coordinates in the first sparse depth data and coordinates in the first image data. The image registration may allow the comparison between the first positional relation and the second positional relation. In some embodiments, the image registration is based on a machine learning model configured to perform image registration. The machine learning model may include an artificial neural network, as described above. For performing the image registration, the circuitry may input the first sparse depth data and the first image data to the machine learning model, and may receive, as output, an indication (e.g., one or more parameters) of a coordinate transformation between the first sparse depth data and the first image data. In some embodiments, the estimating of the second positional relation is based on a photometric error between the first image data and the second image data. The photometric error may include a difference in an RGB value, grayscale value, hue, saturation, lightness, brightness, or the like, of pixels at corresponding positions in the first and second image data, respectively. Accordingly, the photometric error may be a measure for a quality of a color match between the first image data and the second image data. For example, the determining of the second positional relation may include determining a transform between an image represented by the second image data and an image represented by Sony Semiconductor Solutions Corporation et al. 11 the first image data such that corresponding portions of the scene coincide in the respective images represented by the second image data and the first image data and that the photometric error is minimized. In some embodiments, the other sensor includes an inertial measurement unit (IMU); and the verifying of the first positional relation includes: obtaining inertial measurement data from the IMU, wherein the inertial measurement data indicate a change in position of the depth sensor between the second time frame and the first time frame; and comparing the first positional relation with the change in position indicated by the inertial measurement data. The inertial measurement data may indicate a rotation and / or a translation of the IMU (and, thus, of the depth sensor) between the second time frame and the first time frame based on a detected acceleration. For example, the IMU may include an accelerometer and / or a gyroscope, and the inertial measurement data may indicate a specific force measured by the accelerometer and / or an angular velocity measured by the gyroscope. The translation of the IMU (and, thus, of the depth sensor) may be determined based on the measured acceleration and on a previously determined translation velocity of the IMU (and of the depth sensor). The previously determined translation may be based on further measurements (e.g., registration between image frames and / or depth frames). The comparing of the first positional relation with the change in position and rotation indicated by the inertial measurement data may include determining whether the first positional relation corresponds to the change in position and rotation indicated by the inertial measurement data within a predefined tolerance threshold. If the first positional relation corresponds to the change in position indicated by the inertial measurement data within the predefined tolerance threshold, the circuitry may update the 3D model based on the first sparse depth data and the first positional relation. If the first positional relation does not correspond to the change in position indicated by the inertial measurement data within the predefined tolerance threshold, the circuitry may discard the 3D model and generate a 3D model based on the first sparse depth data. In some embodiments, the estimating of the first positional relation is based on an iterative closest point (ICP) algorithm. The ICP algorithm may determine a coordinate transform between a coordinate system of the first sparse depth data and a coordinate system of the 3D model in which a deviation between a representation of the scene in the first sparse depth data and a representation of the scene in the 3D model is minimized. For example, the ICP algorithm may minimize a sum of surface normals Sony Semiconductor Solutions Corporation et al. 12 from the respective 3D points indicated by the first sparse depth data on a surface of the 3D model. The circuitry may perform several iterations of the ICP algorithm, wherein an iteration may optimize a result of a previous iteration. For example, a first iteration of the ICP algorithm may be initialized with the 3D model and with the 3D points as indicated by the first sparse 3D model. A second iteration may be initialized with a result of the first iteration, and so on. The circuitry may perform as many iterations of the ICP algorithm until the sum of surface normal is smaller than a predefined threshold or until a predefined number of iterations has been performed, for example. The ICP algorithm may be weighted. The circuitry may determine respective weights for the 3D points indicated by the first sparse depth data, and may account for the 3D points in the sum of surface normals according to its respective weight. For example, the ICP algorithm may perform a minimization that corresponds to equation (1). (^^^^, Ω^^^^) ≔ argmin ∑ Ξ,Ω^^^^∈Points (Sparse) Here, ^^ denotes an iteration of the ICP algorithm. Q^^denotes a result of the ^^-th iteration of the ICP algorithm, e.g., a rigid transform (without limiting the disclosure to a rigid transform; some embodiments perform a non-rigid transform) such that the 3D model and the first sparse depth data match (e.g., such that the sum of surface normals through respective 3D points of the first sparse depth data is minimized) after applying the transform. Ω^^^^denotes a mapping (e.g., a “matching”) between the respective 3D points of the first sparse depth data and corresponding (possibly interpolated) points at which the respective surface normals intersect the 3D model after applying the transform Q^^. Ξ denotes a pose guess (e.g., an estimated position of the depth sensor at the first time frame). ^^ is an iteration variable that identifies the respective 3D points ^^^s^ indicated by the first sparse depth data (as indicated by the superscript s). ^^^^,Ω^^(^^)denotes the weights of the respective 3D points ^^^^indicated by the first sparse depth data. denotes the (possibly interpolated, hence the superscript i) points at which the respective surface normals from the respective points ^^^s^ intersect the 3D model. ^^Ωk(^^)represents the respective surface normals from the respective points ^^^s^ to the respective points ^^Ωi^^(^^)of the 3D model according to the mapping Ω^^^^. Sony Semiconductor Solutions Corporation et al. 13 Accordingly, the estimating of the first positional relation may be based on a weighted point-to- plane ICP algorithm. The minimization may be performed based on a least-squares algorithm, a maximum-likelihood algorithm, a gradient descent algorithm or any other regression algorithm (including an artificial neural network).The transform Q^^ may, for example, depend on a rotation ^^(^^, ^^, ^^) around three Euler angles^^, ^^ and ^^ and on the time frame ^^: ^^^^−1→^^ ≔ (^^(^^, ^^, ^^), ^^) (2)The circuitry may transform a sparse representation of the scene, as represented by the first sparse depth data, to a dense 3D model based on a rotation and translation according to Q^^and merging the first sparse depth data with the 3D model from one or more previous time frames (e.g., the second time frame). For example, the circuitry may transform the 3D points indicated by the first sparse depth data according to Q^^to match the 3D model from the previous time frame(s), which may be computationally less expensive than transforming the 3D model to match the 3D points indicated by the first sparse depth data, and may still provide a well-posed 3D model. The circuitry may then generate the depth map based on the 3D model and the transformed 3D points. However, in some embodiments, the 3D model from the previous time frame(s) may be transformed according to Q^^to match the 3D points indicated by the first sparse depth data. As mentioned, the circuitry may interpolate points in the 3D model, e.g., for determining surface normals from the 3D points indicated by the first sparse depth data on the 3D model (and for determining the respective length of the surface normals). The circuitry may also interpolate points in the 3D model for increasing a smoothness of the 3D model and / or for reconstructing invalid portions of the 3D model for which the circuitry has not obtained (unambiguous) depth data. The interpolation of points in the 3D model may be based on the image data (e.g., on the first and / or second image data) acquired by the image sensor. For example, the circuitry may interpolate points in the 3D model according to a color value, grayscale value, brightness value and / or gradient of the image data. Sony Semiconductor Solutions Corporation et al. 14 In some embodiments, the iterative closest point (ICP) algorithm includes weighting the 3D points indicated by the first sparse depth data with respective weights according to respective confidences of the respective 3D points. The weighting may include determining a specific value for ^^^^,Ω^^(^^)of equation (1) for each 3D point ^^ separately. For example, the circuitry may choose a value of in the range from 0 to 1, without limiting the disclosure to this range. The higher the value of is, the stronger the respective point ^^ may be weighted, and the higher the impact of the respective point may be on the result ^^^^. On the other hand, the lower the value of ^^^^,Ω^^(^^)is, the less the respective point ^^ may be weighted, and the lower the impact of the respective point may be on the result ^^^^. For example, the ICP algorithm may ignore a 3D point ^^^s^ if a respective weight ^^^^,Ω^^(^^)= 0. The weighting of the 3D points may allow for handling poor matches, outliers and / or noisy data.For example, the circuitry may choose a low value for ^^^^,Ω^^(^^) (e.g., < 1 or ^^^^,Ω^^(^^) = 0,without limiting the disclosure to these values) for 3D points ^^ that are identified as poor matches, outliers and / or noisy data. For example, the weight ^^s of a point ^^^^ may be configured as a function ^^(^^^^ , ^^^^ , ^^^^ , … ) ofa color ^^^^of the image data acquired by the image sensor at a position the corresponds to the point ^^^s^, of a distance / depth ^^^^indicated by the point ^^^s^, and / or of an amplitude ^^^^of a reflected light dot (e.g., an infrared amplitude of a reflected laser dot) that corresponds to the point ^^^s^, as described below in more detail. Additionally or alternatively, the weight of a point ^^^s^ may depend on a further suitable quantity. However, the disclosure is not limited to weighting the 3D points indicated by the first sparse depth data. Instead, in some embodiments, the 3D points are not weighted. For example, the circuitry may choose a same value for thes for all 3D points ^^^^ = 1 for allpoints ^^^s^, without limiting the disclosure to this value), such that all 3D points are weighted with a same weight. In some embodiments, the weights are based on respective depths that correspond to the respective 3D points. For example, the weights for a respective 3D point ^^^^may be a function of a respective depth ^^^^indicated by the 3D point. For example, the circuitry may determine whether the Sony Semiconductor Solutions Corporation et al. 15 respective point ^^^s^ is a poor match based on a length of a respective surface normal from the point to the 3D model as compared to an average among the respective lengths of respective surface normals from its respective neighboring 3D points to the 3D model. For example, the circuitry may determine whether the respective point ^^^s^ is an outlier based on a deviation of the depth represented by the point ^^^s^ as compared to an average among the respective depths represented by its neighboring 3D points. For example, the circuitry may determine whether the respective point ^^^s^ includes strong noise (e.g., whether the first sparse depth data are noisy) based on a standard deviation or variance of depths indicated by neighboring 3D points ^^^s^. In some embodiments, the depth sensor is configured to acquire the first sparse depth data based on light emitted into the scene; and wherein the weights of the respective 3D points are based on a respective amplitude of reflected light from the scene. As mentioned, the depth sensor may acquire the first sparse depth data based on ToF imaging, e.g., based on emitting an array in infrared laser dots into the scene and detecting reflections thereof. A higher amplitude (e.g., intensity) of the reflected light from the scene may correspond to an object or surface in the scene that has a higher reflectivity and / or a lower distance than for a lower amplitude. For example, the circuitry may weight a 3D point with a higher amplitude stronger than a 3D point with a lower amplitude. In some embodiments, the weights are based on a photometric error between first image data of the scene acquired by an image sensor at the first time frame and second image data of the scene acquired by the image sensor at the second time frame, a positional relation between the image sensor and the depth sensor being fixed. The first and second image data may correspond to the first and second image data described above, and the photometric error may include a difference in an RGB value, grayscale value, hue, saturation, lightness, brightness, or the like, of pixels at corresponding positions in the first and second image data, respectively, as described above. A high photometric error (e.g., a high color difference or a low quality of a color match between the first and the second image data) may indicate a change of the scene between the second time frame and the first time frame. Thus, the circuitry may weight 3D points ^^^s^ that correspond to a Sony Semiconductor Solutions Corporation et al. 16 position with a high photometric error with a lower weight than 3D points ^^^s^ that correspond to a position with a lower photometric error. In some embodiments, the estimating of the first positional relation is based on a machine learning model configured to determine the first positional relation. The machine learning model may include an artificial neural network, as described above. For estimating the first positional relation, the circuitry may input the first sparse depth data and the 3D model from the previous time frame(s) to the machine learning model, and may receive, as output, an indication of the result ^^^^and / or Ω^^^^of the ICP algorithm. In some embodiments, the updating of the 3D model is based on a machine learning model configured to merge the first sparse depth data into the 3D model. The machine learning model may include an artificial neural network, as described above. For updating the 3D model, the circuitry may input the first sparse depth data and the 3D model from the previous time frame(s) (and, possibly, an indication of the result ^^^^and / or Ω^^^^of the ICP algorithm) to the machine learning model, and may receive, as output, an updated 3D model that is based on (e.g., includes) depth information from the first sparse depth data and that represents the scene at the first time frame. In some embodiments, the updating of the 3D model includes deleting, from the 3D model, depth information that is based on sparse depth data which are older than a predefined number of time frames. For example, the circuitry may be configured to “forget” depth information after a predefined number of time frames, e.g., in order to avoid an accumulation of a mistake or error, and / or in order to avoid having to perform a simultaneous localization and mapping (SLAM) algorithm. SLAM may bundle adjustment and pose graph optimization. Avoiding SLAM may allow a less complex and less costly generating of a depth map. For example, the 3D model may include a set of 3D points indicated by (e.g., the first and second) sparse depth data acquired by the depth sensor in subsequent time frames. The circuitry may store an age indication for each 3D point included in the 3D model (e.g., a timestamp of acquisition of the respective 3D point, or a sequential number of a time frame at which the respective 3D point has been acquired). When updating the 3D model, the circuitry may identify and ignore, based on the age indication, 3D points of the 3D model that are older than the Sony Semiconductor Solutions Corporation et al. 17 predefined number of time frames. For example, the circuitry may delete the old 3D points, or may mark the old 3D points as ignored. Thus, the 3D model may be based on sufficiently recent depth information. For example, the circuitry may save memory that might otherwise be occupied by redundant depth information and / or by a representation of a portion of the scene that is not represented by the first sparse depth data (e.g., because a field-of-view of the depth sensor has moved or shifted away from the respective portion of the scene). However, the predefined number of time frames after which depth information is deleted may be chosen sufficiently high such that the circuitry can populate invalid areas in the first sparse depth data with depth information (e.g., 3D points) from the previous frame(s) and / or that the circuitry can cope with an occlusion based on image data from the image sensor and / or based on a respective photometric check. For example, the predefined number of time frames after which depth information may be removed from the 3D model may be chosen to be 4 frames, 8 frames, 10 frames, or the like, without limiting the disclosure to these values. In some embodiments, the 3D model includes a set of 3D points that represents the scene, wherein the set of 3D points is based at least on the second sparse depth data. The 3D points of the set of 3D points may correspond to respective points in the scene at which respective corresponding laser dots have been reflected. The set of 3D points may include 3D points indicated by the second sparse depth data, and may further include 3D points indicated by sparse depth data acquired by the depth sensor at other (e.g., preceding) time frames. Thus, the 3D model may be configured as (or may include) a 3D point model that includes the 3D points of the set of 3D points. Apart from transforming the 3D model according to the first positional relation, the updating the 3D model may include inserting the 3D points indicated by the first sparse depth data in the set of 3D points of the 3D model and, possibly, deleting 3D points that are older than the predefined number of time frames from the set of 3D points of the 3D model. In some embodiments, the 3D model includes a grid that represents the scene, wherein values assigned to respective positions of the grid are based at least on the second sparse depth data. The grid may be configured as a two-dimensional (2D) or three-dimensional (3D) grid in which unit cells may be arranged along two or three axes, respectively. A unit cell of a 2D grid may be referred to as a (model) pixel, and a unit cell of a 3D grid may be referred to as a (model) voxel. Sony Semiconductor Solutions Corporation et al. 18 Each unit cell of the grid may correspond to a respective position in the scene, and may be assigned a value that indicates whether an object / surface is detected at the scene. The circuitry may further assign an age indication (e.g., a timestamp or sequential number of a time frame) to each unit cell of the grid. The age indication may indicate a latest time frame at which an object / surface has been detected at a position in the scene that corresponds to the respective unit cell. Apart from transforming the 3D model according to the first positional relation, the updating of the 3D model may include setting the value of each unit cell that corresponds to a respective 3D point indicated by the first sparse depth data such that the value of the respective unit cell indicates that an object / surface has been detected at the corresponding position. The circuitry may also update the age indication of the respective unit cells to indicate that an object / surface has been detected for the respective unit cell at the first time frame. The updating of the grid model may further include setting the value of each unit cell whose age indication indicates depth information that is older than the predefined number of time frames such that the respective value does not indicate an object / surface in the frame. Configuring the 3D model as a grid model instead of a point model may save computing resources in some embodiments, but may reduce a precision of the generated depth map in some embodiments. Further, a 3D grid model may be able to cope with a rotation of the 3D model better than a 2D grid model. In some embodiments, the depth sensor is configured to acquire the first sparse depth data based on a time-of-flight measurement (ToF). As described, the ToF measurement may include an iToF and / or a dToF measurement and may be based on emitting light (e.g., an array of infrared laser dots) into the scene, receiving a reflection of the emitted light from the scene, and determining a depth, distance and / or round-trip time indicated by the first sparse depth data based on the received reflection of the emitted light. In some embodiments, the circuitry is further configured to generate a depth map stream of subsequent time frames, wherein the depth map stream includes the generated depth map for the first time frame. As described, the depth map stream may include a sequence of depth maps for respective time frames. The circuitry may generate a respective depth map based on the (updated) 3D model for the subsequent time frames. In some embodiments, the estimating of the first positional relation is based on an external pose input that indicates a movement of the depth sensor between a previous time frame to which the Sony Semiconductor Solutions Corporation et al. 19 3D model corresponds and the first time frame, wherein the previous time frame corresponds to a time prior to the first time frame. The external pose input may be generated by a device or algorithmic pipeline external to the circuitry and may be received by the circuitry from the external device. The external device may determine the movement of the depth sensor with respect to the scene between the first time frame at which the depth sensor has acquired the first sparse depth data and the previous time frame (or several previous time frames prior to the first time frame) at which data (e.g., sparse depth data) have been acquired according to which the 3D model has been updated. The movement of the depth data indicated by the external pose input may include a translation, a panning, a rotation or the like relative to the scene and may cause the first sparse depth data acquired by the depth sensor after the movement to represent the scene from a different perspective than sparse depth data acquired by the depth sensor before the movement (e.g., at the one previous time frame or several previous time frames to which the 3D model corresponds). Thus, the external pose input may allow the circuitry to estimate the first positional relation. The external pose input may include pose updates that indicate the movement of the depth sensor for a short temporal window (e.g., up to four time frames, up to eight time frames or the like, without limiting the disclosure to these values) before the first time frame. The short temporal window may be chosen such that a drift (error integral) of a pose (position) of the depth sensor with respect to the scene is expected to remain below a predefined threshold during the short temporal window. This may allow to mitigate the drift of the pose indicated by the externa pose input. The external device may generate the external pose input based on any suitable software that performs SLAM and / or any other kind of odometry. The external device may transform pose data from any suitable source into a pose update included in the external pose input. For example, the external device may generate the external pose input based on a visual-inertial (IMU) odometry system based on any suitable traditional method. For example, the external device may generate the external pose input based on a complex RGB-D SLAM pipeline that may or may not use an IMU. The skilled person may find further suitable sources of pose data. In some embodiments, the external pose input is based on inertial measurement data from an IMU that is fixed to the depth sensor. The external device may generate the external pose input based on the inertial measurement data. Because the IMU is fixed to the depth sensor, thus, the IMU may move together with the depth Sony Semiconductor Solutions Corporation et al. 20 sensor, such that the inertial measurement data may represent the movement of the depth sensor. Accordingly, IMU integration may be used for estimating an accurate pose of the depth sensor in the short temporal window. Thus, the external pose input may be inputted to the circuitry, for example an IMU pose input. Estimating a long-term pose with IMU may be subjected to drift in some instances. Therefore, a pose update may be estimated in a short time window in which no drift is expected to occur. The IMU pose input may be used complementary to RGB as input source. RGB may be used to filter out IMU errors and reset the pose input using photometric checks. Moreover, the IMU pose input may be improved based on advanced filtering, e.g., based on a Kalman filter. For example, an Extended and / or Unscented Kalman filter may be used, which may receive a noisy gyroscope velocity and a specific force, and may return an integrated pose. For example, the Kalman filter may depend on updated internal states, and a machine learning model may estimate and update Kalman filter parameters (e.g., according to the Tight Learned Inertial Odometry (TLIO) approach). In some embodiments, the circuitry is further configured to: obtain sensor data that have been acquired by a further sensor at a third time frame later than the first time frame, wherein a positional relation between the further sensor and the depth sensor is fixed; estimate a third positional relation between a representation of the scene of the 3D model and a representation of the scene of the sensor data; update the 3D model based on the third positional relation; and generate a depth map based on the updated 3D model. The positional relation between the further sensor and the depth sensor may be fixed such that a movement of the further sensor corresponds to a movement of the depth sensor, and the movement of the depth sensor may be inferred from sensor data measured by the further sensor. The further time frame may be a time frame later than the first time frame. For example, the 3D model may have been updated according to the first sparse depth data from the first time frame, and a further output frame that corresponds to the third time frame may be generated for the depth map stream without acquiring sparse depth data at the third time frame. Thus, the 3D model that has been updated according to the first sparse depth data may be transformed (e.g., translated, rotated, scaled, sheared or the like) and / or interpolated to represent the scene from a perspective of the other sensor at the third time frame. A depth map of the further output frame, Sony Semiconductor Solutions Corporation et al. 21 which may correspond to the third time frame, may then be generated based on the 3D model that represents the scene from the perspective of the other sensor at the third time frame. In some embodiments, the circuitry is configured to generate a depth map at an output frame rate that is higher than a depth frame rate at which the depth sensor acquires sparse depth data; and a sensor frame rate at which the further sensor acquires sensor data is at least as high as the output frame rate. Therefore, a power consumption may be reduced by emitting less light into the scene (e.g., by the light emission unit 5 of Fig.1) for a ToF measurement. For time frames at which no sparse depth data are acquired, the 3D model may be transformed based on the sensor data from the further sensor. In some embodiments, the further sensor includes an image sensor, and the sensor data include image data. In some embodiments, the further sensor includes an IMU, and the sensor data include inertial measurement data. Thus, a change of a position of the depth sensor since acquiring the first sparse depth data may be determined based on the image data from the image sensor and / or based on the inertial measurement data from the IMU, such that the further output frame that corresponds to the third time frame may be generated without acquiring sparse depth data and, thus, without emitting light into the scene. Some embodiments pertain to a method for generating a depth map, wherein the method includes: obtaining first sparse depth data acquired by a depth sensor at a first time frame, wherein the first sparse depth data indicate 3D points of a scene in the first time frame; estimating a first positional relation between a representation of the scene of the first sparse depth data and a representation of the scene of a 3D model, wherein the 3D model is based on second sparse depth data acquired by the depth sensor at a second time frame prior to the first time frame; updating the 3D model based on the first sparse depth data and the first positional relation; and generating a depth map of the scene based on the updated 3D model. The method may be configured similar to the circuitry described above, and any feature described with respect to the circuitry may apply accordingly to a respective feature of the method. The features of the circuitry and / or of the method may be combined in any suitable way. Sony Semiconductor Solutions Corporation et al. 22 Accordingly, in some embodiments, the present technology allows obtaining a better depth map from a sparse depth map stream and an image (e.g., RGB) stream that may have an equal or slightly different frame rate. In some embodiments, the present technology may provide a depth map without invalid areas due to using reflected laser dots observed at previous time frames. In some embodiments, the present technology may provide a smoother depth map because of having more dots to average with an RGB-D fusion (depth completion) filter and, thus, a better signal-to-noise ratio (SNR) by averaging. In some embodiments, the circuitry and / or method is capable of coping with occlusions based on using an image (e.g., RGB) and / or photometric check. In some embodiments, the circuitry and / or method leverages at most a few time frames (“short memory”) to avoid accumulating errors and / or requiring a complex and costly SLAM block. For example, the circuitry and / or method may use a smooth motion of a handheld device (e.g., smartphone, tablet, notebook, camera etc.) or of a wearable device (e.g., head-mounted display (HMD), smartglasses, smartwatch etc.) between subsequent time frames. For example, the present technology may be used in augmented reality (AR) applications, virtual reality (VR) applications, indoor / outdoor scanning, computational photography, robotics / industrial / automotive perception tasks, or the like. The methods as described herein are also implemented in some embodiments as a computer program causing a computer and / or a processor to perform the method, when being carried out on the computer and / or processor. In some embodiments, also a non-transitory computer- readable recording medium is provided that stores therein a computer program product, which, when executed by a processor, such as the processor described above, causes the methods described herein to be performed. Returning to Fig.1, Fig.1 illustrates an embodiment of a circuitry 1. The circuitry 1 includes a processing unit 2, an artificial intelligence (AI) unit 2a, a storage unit 3 and a communication unit 4. The processing unit 2 includes a CPU and is configured to control a function of the circuitry 1 and to perform data processing related to generating a depth map. The processing unit 2 provides input data to the AI unit 2a and receives output data from the AI unit 2a. The AI unit 2a includes a GPU and executes a machine learning model that performs an estimation related to generating a depth map according to input from the processing unit 2. The storage unit 3 includes an SSD and a DRAM, and stores data related to generating a depth map, including instructions for the processing unit 2 and the AI unit 2a, and data generated by the Sony Semiconductor Solutions Corporation et al. 23 processing unit 2 or the AI unit 2a, such as a 3D model and a generated depth map. The communication unit 4 is configured to exchange data with an external device and to transmit a depth map generated by the processing unit 2 to the external device. The circuitry 1 further includes a light emission unit 5, a depth sensor 6, an image sensor 7, and an inertial measurement unit (IMU) 8. The light emission unit 5 is configured to emit an array of infrared laser dots 9a into a scene 10 for ToF imaging. The depth sensor 6 detects a reflection 9b of the infrared laser dots 9a from the scene for ToF imaging and, thus, acquires sparse depth data. The image sensor 7 acquires an RGB image (an example of image data) of the scene 10. The IMU 8 measures a movement (translation and rotation) of the circuitry 1 and, thus, acquires inertial measurement data. A positional relation between the depth sensor 6, the image sensor 7 and the IMU 8 is fixed. The storage unit 3 stores the depth data acquired by the depth sensor 6, the image data acquired by the image sensor 7 and the inertial measurement data acquired by the IMU 8. The circuitry 1 is configured to perform the method of Fig.2. It is noted that, in some embodiments, the light emission unit 5 and the depth sensor 6 are provided separately from the circuitry 1, and the circuitry 1 receives the sparse depth data from the depth sensor 6 via the communication unit 4. It is also noted that, in some embodiments, the image sensor 7 is provided separately from the circuitry 1, and the circuitry 1 receives the image data from the image sensor 7 via the communication unit 4. It is further noted that, in some embodiments, the IMU 8 is provided separately from the circuitry 1, and the circuitry 1 receives the inertial measurement data from the IMU 8 via the communication unit 4. In addition, it is noted that, in some embodiments, the image sensor 7 and / or the IMU 8 is omitted, and the circuitry 1 generates a depth map without using the image data from the image sensor 7 and / or without using the inertial measurement data from the IMU 8. Furthermore, in some embodiments, the circuitry 1 does not include the AI unit 2a, and the circuitry 1 may generate the depth map without using an artificial neural network. Fig.2 illustrates an embodiment of a method 20. The method 20 is an example of a method performed by the circuitry 1 of Fig.1. At 21 of the method 20, the processing unit 2 obtains first sparse depth data acquired by the depth sensor 6 at a first time frame. The first sparse depth data indicate 3D points of the scene 10 in the first time frame. The depth sensor 6 acquires the first sparse depth data based on a ToF measurement, which includes emitting light 9a into the scene 10. The first time frame Sony Semiconductor Solutions Corporation et al. 24 corresponds to a current iteration of the method 20 (i.e., of the processing from 21 to 35 of the method 20). At 22, the processing unit 2 of the circuitry 1 estimates a first positional relation between a representation of the scene 10 of the first sparse depth data and a representation of the scene 10 of a 3D model. The 3D model is based on second sparse depth data acquired by the depth sensor 6 at a second time frame prior to the first time frame. The second time frame corresponds to a previous iteration of the method 20 (i.e., of the processing from 21 to 35 of the method 20) preceding to the current iteration. The estimating of the first positional relation at 22 is based on a machine learning model configured to determine the first positional relation, and the processing unit 2 causes the AI unit 2a to execute the machine learning model for estimating the first positional relation. The estimating of the first positional relation at 22 includes performing, at 23, an iterative closest point (ICP) algorithm. The ICP algorithm is based on equation (1) and includes weighting the 3D points ^^s^^ indicated by the first sparse depth data with respective weights according to respective confidences of the respective 3D points, as described with respect toequation (1). The weights are based on respective depths ^^^^ that correspond tothe respective 3D points. The weights of the respective 3D points are alsobased on a respective amplitude ^^^^of the reflected light 9b from the scene 10. The weights ^^^^) are further based on a photometric error ^^^^ between first image data of thescene 10 acquired by the image sensor 7 at the first time frame (and, thus, obtained by the circuitry 1 at 25 of the method 20) and second image data of the scene 10 acquired by the image sensor 6 at the second time frame (i.e., at 25 of the method 20 of an iteration of the method 20 that corresponds to the second time frame). The ICP algorithm is performed at 23 by the processing unit 2 and the AI unit 2a. The estimating of the first positional relation at 22 includes verifying, at 24, the first positional relation based on data acquired by other sensors (i.e., based on image data acquired by the image sensor 7 and based on inertial measurement data acquired by the IMU 8). As mentioned, a positional relation between the other sensors (i.e., the image sensor 7 and the IMU 8) and the depth sensor 6 is fixed. The verifying at 24 is performed by the processing unit 2. The verifying of the first positional relation at 24 includes obtaining, at 25, with the processing unit 2 the first image data acquired by the image sensor 7. As mentioned, the first image data correspond to the first time frame. Sony Semiconductor Solutions Corporation et al. 25 The verifying of the first positional relation at 24 includes estimating, at 26, a second positional relation between an image of the scene 10 represented by the first image data and an image of the scene 10 represented by the second image data. As mentioned, the image sensor 7 has acquired the second image data at 25 of an iteration of the method 20 that corresponds to the second time frame. Accordingly, the second image data correspond to the second time frame. The estimating of the second positional relation at 26 is based on a photometric error between the first image data and the second image data. The estimating of the second positional relation at 26 includes performing, at 27, an image registration for registering the first sparse depth data and the first image data. The image registration is based on a machine learning model that is configured to perform image registration. Thus, the processing unit 2 causes the AI unit 2a to execute the machine learning model for performing the image registration. The verifying of the first positional relation at 24 includes comparing, at 28, the first positional relation (which has been estimated at 22) with the second positional relation (which has been estimated at 26). The comparing at 28 is performed by the processing unit 2. The verifying of the first positional relation at 24 further includes obtaining, at 29, with the processing unit 2, inertial measurement data from the IMU 8. Due to the fixed positional relation between the depth sensor 6 and the IMU 8, the inertial measurement data indicate a change in position of the depth sensor 6 between the second time frame and the first time frame. The verifying of the first positional relation at 24 then includes comparing, at 30, the first positional relation with the change in position indicated by the inertial measurement data. The comparing at 30 is performed by the processing unit 2. At 31, the processing unit 2 updates the 3D model based on the first sparse depth data (which the processing unit 2 has obtained at 21) and the first positional relation (which the processing unit 2 has estimated at 22). The updating at 31 is based on a machine learning model configured to merge the first sparse depth data into the 3D model. Thus, the processing unit 2 causes the AI unit 2a to execute the machine learning model for updating the 3D model. The machine learning model executed by the AI unit 2a merges, at 32, the first sparse depth data (which the processing unit 2 has obtained at 21) and, thus, the 3D points indicated by the first sparse depth data, into the 3D model. The updating of the 3D model at 33 further includes deleting, at 33, from the 3D model, depth information that is based on sparse depth data which are older than a predefined number of time Sony Semiconductor Solutions Corporation et al. 26 frames (i.e., which has been obtained at 21 of a in iteration of the method 20 since which the method 20 has been iterated a predefined number of times). The processing unit 2 determines an age of respective depth information in the 3D model based on a respective age indication stored in the storage unit 3. At 34, the processing unit 2 generates a depth map of the scene 10 based on the updated 3D model, which has been updated at 31. At 35, the processing unit 2 generates a depth map stream of subsequent time frames, wherein the depth map stream includes the generated depth map for the first time frame, which has been generated at 34. The generating of the depth map stream at 35 includes outputting the depth map with an indication that the depth map is included in the depth map stream. Thus, the depth map stream includes depth maps generated at 34 of subsequent iterations of the method 20. The 3D model includes a set of 3D points that represents the scene 10. The set of 3D points is based at least on the second sparse depth data. The merging of the first sparse depth data into the 3D model at 32 includes inserting the 3D points of the scene 10 indicated by the first sparse depth data into the set of 3D points of the 3D model. The deleting of old depth information at 33 includes removing, from the set of 3D points, respective 3D points that have been acquired at least the predefined number of time frames earlier and that thus represent the depth information that is older than the predefined number of time frames. Alternatively, in some embodiments, the 3D model includes a grid that represents the scene 10, wherein values assigned to respective positions of the grid are based at least on the second sparse depth data. In such a case, the merging of the first sparse depth data into the 3D model at 32 includes assigning an occupation value to respective positions of the grid that correspond to the 3D points of the scene 10 which are indicated by the first sparse depth data. Further, in such a case, the deleting of old depth information at 33 includes assigning a vacancy value to respective positions of the grid for which no corresponding 3D points of the scene 10 have been indicated by sparse depth data of the last predefined number of time frames (i.e., of the last predefined number of iterations of the method 20). Here, the occupation value indicates that an object or surface has been detected at a portion of the scene 10 that corresponds to the respective position in the grid, and the vacancy value indicates that no object or surface has been detected at a portion of the scene 10 that corresponds to the respective position in the grid. The method 20 (i.e., the processing at 21 to 35) is performed in several iterations, wherein the 3D model updated at 31 of an iteration is further updated at 31 of a subsequent iteration. Further, Sony Semiconductor Solutions Corporation et al. 27 the (first) sparse depth data obtained at 21 of an iteration corresponds to the second sparse depth data in a subsequent iteration. It is noted that the order in which the processing at 21 to 35 of the method 20 is shown in Fig.2 need not correspond to a chronological order in which the processing at 21 to 35 of the method 20 is performed. For example, when the weights of the ICP algorithm are based on the photometric error ^^^^of the image data, the circuitry 1 performs the obtaining of the image data at 25 before the ICP algorithm at 23. Further, for example, the obtaining of the inertial measurement data at 29 may be performed before the comparing of the first positional relation with the second positional relation at 28, before the estimating at 26, and / or before the obtaining of the image data at 25. Likewise, the comparing of the first positional relation with the inertial measurement data at 30 may be performed before the comparing of the first positional relation with the second positional relation at 28, before the estimating at 26, and / or before the obtaining of the image data at 25, wherein the obtaining of the initial measurement data at 29 may be performed before the comparing of the first positional relation with the inertial measurement data at 30. Further, in some embodiments, the deleting of old depth information at 33 is performed before the merging of the first sparse depth data into the 3D model at 32. The skilled person may find further variations of a chronological order in which the processing at 21 to 35 of the method 20 may be performed. It is also noted that, in some embodiments, some processing described for the method 20 is omitted. For example, the obtaining of image data at 25, the estimating of the second positional relation at 26 and the comparing of the first positional relation with the second positional relation may be omitted, and the verifying of the first positional relation at 24 may be based on the inertial measurement data, but not on the image data. Likewise, the obtaining of the inertial measurement data at 29 and the comparing of the first positional relation with the inertial measurement data at 30 may be omitted, and the verifying of the first positional relation at 24 may be based on the image data, but not on the inertial measurement data. In some embodiments, the whole verifying of the first positional relation at 24 is omitted. Further, in some embodiments, the deletion of old depth information at 33 is omitted, and depth information are accumulated beyond the predefined number of time frames. In some embodiments, the generating of the depth map stream at 35 is omitted, and the circuitry may, for example, generate (and update / refine) one depth map, or may store depth maps generated at 34 of subsequent iterations of the method 20 in the storage unit 3 (e.g., as separate files). In some embodiments, the estimating of the first positional relation at 22 is not based on the ICP algorithm 23, or the Sony Semiconductor Solutions Corporation et al. 28 ICP algorithm at 23 does not include weighting the 3D points according to the weights Fig.3 illustrates an embodiment of an RGB-D fusion pipeline 40. The pipeline 40 is an example of a method performed by the circuitry 1 of Fig.1. According to the pipeline 40, RGB data 41 (an example of image data acquired at 25 of Fig.2 by the image sensor 7 of Fig.1) and sparse depth data 48 (an example of sparse depth data acquired at 21 of Fig.2 by the depth sensor 6 of Fig.1 based on an array of emitted laser dots) are fused for generating a depth map 72. At 42, the circuitry 1 performs an RGB prefiltering. The RGB prefiltering 42 includes reducing a size of the RGB data 41 at 43, performing a downscaling at 44 and applying a memory map converter at 45. The prefiltered RGB data from the RGB prefiltering 42 are then provided to an RGB-D odometry block 46 and to an RGB-D fusion block 47. At 49, the circuitry 1 performs a sparse depth prefiltering of the sparse depth data 48. The sparse depth data 48 are represented in cartesian coordinates. The sparse depth prefiltering 49 includes cropping a region-of-interest (ROI) at 50, removing isolated spots at 51, and applying a sparse bilateral filter at 52. The applying 52 of the sparse bilateral filter includes a conversion from cartesian to radial coordinates, a conversion from DA to IQ, the sparse bilateral filter, a conversion from IQ to DA, and a conversion from radial coordinates to cartesian coordinates. At 53, the circuitry 1 performs an RGB-D registration, which includes a ToF to RGB registration at 54. A result of the RGB-D registration 53 is provided to the RGB-D fusion block 47. The RGB-D odometry block 46 includes estimating, at 46a, a pose of the depth sensor 6, which is an example of the estimating of the first positional relation at 22 of Fig.2. The RGB-D odometry block 46 further includes accumulating points at 46b, which is an example of the updating of the 3D model at 31 of Fig.2 (and, more specifically, of the merging into the 3D model at 32). The accumulating of points at 46b is based on the estimated pose estimated at 46a, and the estimating of the pose at 46a is based on the accumulated points accumulated at 46b. The RGB-D fusion block 47 includes a depth completion at 47a, which interpolates between points of the 3D model for generating a dense 3D model. The RGB-D fusion block 47 outputs an interpolated dense 3D model 55 in radial coordinates and an interpolated dense 3D model 56 in cartesian coordinates. Sony Semiconductor Solutions Corporation et al. 29 The RGB-D fusion block 47 is based on a result of a previous run of the RGB-D odometry block 46, and the RGB-D odometry block 46 is based on a result of a previous run of the RGB-D fusion block 47. At 57, the circuitry 1 performs a postfiltering. The postfiltering 57 includes an amplitude / reflectance filter 58, to which the interpolated dense 3D model 55 in radial coordinates is provided. The postfiltering 57 further includes a conversion 59 from cartesian to radial coordinates, whose result is provided to the amplitude / reflectance filter 58. The postfiltering 57 then includes a small holes recovery 60 and a conversion 61 from radial to cartesian coordinates. An interpolated depth in cartesian coordinates is then provided from the conversion 61 to a 3x3 median filter 62. On a result of the 3x3 median filter 62, a killing of invisible pixels is applied at 63, and the result is provided to a recursive Gaussian 64. Further, on the result of the 3x3 median filter 62, a normal-based killing of flying pixels is applied at 65, and the result is also provided to the recursive Gaussian 64. A result of the recursive Gaussian 64 is provided to the RGB-D odometry block 46 and to an RGB-D small hole filling 66. The RGB-D small hole filling 66 further receives the prefiltered RGB data from the RGB prefiltering 42. After the RGB-D small hole filling 66, a temporal filter 67 and a 3x3 median filter 68 are applied. Then, invalid pixels are updated at 69. The conversion 61 further provides an interpolated amplitude to the updating 69 of invalid pixels. The invalid pixels include pixels for which there is any degree of uncertainty about the depth data. This uncertainty may be related to the fact that interpolated depth map data draw on very distant pixels; or interpolated depth map data lie close to discontinuities in the depth map; or the interpolated depth map data are in a high-variance neighborhood where the invalidated pixel is a background pixel (and therefore not visible). A result of the updating 69 of invalid pixels is provided to an RGB-D large hole filling 70 and to an output projection 71. The RGB-D large hole filling 70 further receives the prefiltered RGB data from the RGB prefiltering 42, and a result of the RGB-D large hole filling 70 is provided to the output projection 71. Sony Semiconductor Solutions Corporation et al. 30 The output projection 71 performs an intrinsics scaling (i.e., camera intrinsics matrix scaling such that, upon changing a pixel resolution, principal coordinates and a focal length are scaled accordingly) and generates an interpolated dense depth map 72 in cartesian coordinates. The interpolated dense depth map 72 includes invalid pixels. Fig.4 illustrates an embodiment of a pipeline in the RGB-D odometry block 46 of Fig.3. The RGB-D odometry block 46 is also referred to as odometry core 46. The odometry core 46 receives data from a current frame RGB-D input 73, which includes the RGB prefiltering 42 of Fig.3. The odometry core 46 also receives data from the depth completion 47a of Fig.3 and provides data to the depth completion 47a.A DA buffer 74 for a previous time frame (^^ − 1) of the odometry core 46 receives a result of therecursive Gaussian 64 of Fig.3. An output of the DA buffer 74 is provided to an amplitude / reflectance filter 75 of the odometry core 46. A result of the amplitude / reflectance filter 75 is provided to an updating 76 of invalid pixels. A result of the updating 76 of invalid pixels is provided to a point buffer 77 for past dense points of the odometry core 46, and to a normals generation 78. A result of the normals generation 78 is also provided to the point buffer 77 for past dense points. The prefiltered RGB data from the RGB prefiltering 42 of the current frame RGB-D input 73 areprovided to a color buffer 79 for a previous time frame (^^ − 1) of the odometry core 46. Anoutput of the color buffer 79 is also provided to the point buffer 77 for past dense points. An output of the point buffer 77 for past dense points is provided to an RGB-D ICP 80. A result of the depth completion 47a is provided to an amplitude / reflectance filter 81 of the odometry core 46, and a result of the amplitude / reflectance filter 81 is provided to a point buffer 82 for current sparse points of the odometry core 46. The point buffer 82 for current sparse points further receives the prefiltered RGB data from the RGB prefiltering 42 of the current frame RGB-D input 73, and provides its output to a frame integration 83 of the odometry core 46, to a merging 84 of DA with checks, and to the RGB-D ICP 80. The RGB-D ICP 80 is an example of an ICP algorithm as performed at 23 of the method 20 of Fig.2. A result of the RGB-D ICP 80 (drawn as dashed arrow) corresponds to a pose update (e.g., of a pose of the depth sensor 6 of Fig.1), and is provided to a transformable buffer 85 of the odometry core 46. Sony Semiconductor Solutions Corporation et al. 31 The transformable buffer 85 corresponds to a copy of past dense points (e.g., a 3D model). The transformable buffer 85 further receives an output of the point buffer 77 for past dense points. An output of the transformable buffer 85 corresponds to a registered color and is provided to photometric checks 86. The photometric checks 86 correspond to the estimating of a pose at 46a of Fig.3 (and, e.g., of the pose of the depth sensor 6 of Fig.1), and are an example of the estimating of the second positional relation at 26 of the method 20 of Fig.2. A result of the photometric checks 86 (an estimated pose, drawn as dashed arrow) is provided to the frame integration 83. The frame integration 83 includes the accumulating of points at 46b of Fig.3, and is an example of the updating of the 3D model at 31 of Fig.2 as well as of the merging of the first sparse depth data into the 3D model at 32 of Fig.2. A result of the frame integration 83 is provided to the merging 84 of the DA (i.e., of the dense 3D model) with the checks. A result of the merging 84 is provided to the depth completion 47a. Fig.5 illustrates a concept of the RGB-D odometry block 46 of Fig.3 according to an embodiment. The RGB-D odometry block 46 receives dense RGB-D data 90 of a previous time frame and sparse RGB-D data 91 of a current time frame. The dense RGB-D data 90 of the previous frame are provided to a reflectance filter 92 of the RGB-D odometry block 46. A result of the reflectance filter 92 is provided to a PC 93 of the RGB-D odometry block 46. An output of the PC 93 is provided to a sparse-to-dense RGB-D ICP 94. Likewise, the sparse RGB-D data 91 of the current time frame are provided to a reflectance filter 95 of the RGB-D odometry block 46. A result of the reflectance filter 95 is provided to a PC 96 of the RGB-D odometry block 46. An output of the PC 96 is also provided to the sparse- to-dense RGB-D ICP 94. An output of the sparse-to-sense RGB-D ICP 94 corresponds to a pose update according to ≔(^^, ^^), and is provided to an RGB-D memory 97 of the RGB-D odometry block 46, toa block 98 of the RGB-D odometry block 46, and to a block 99 of the RGB-D odometry block 46. Likewise, the output of the PC 96 is also provided to the RGB-D memory 97, to the block 98, and to the block 99. Sony Semiconductor Solutions Corporation et al. 32 An output of the RGB-D memory 97 includes a sparse depth frame based on a memory and is provided to the block 98. At the block 98, a rigid transform is applied, photometric checks are performed, and culling is performed. A result of the block 98 includes a sparse depth frame based on a memory and is provided to the block 99 and to the RGB-D memory 97. At the block 99, DA frames are projected, a z-buffer is applied, and edge checks are performed. A result of the block 99 includes a sparse depth frame based on a current frame and on a memory, and corresponds to a temporally integrated sparse DA frame 100. In the RGB-D odometry block 46, all states are reset (i.e., depth information of previous time frames is flushed) if the photometric checks fail or if the ICP fails, such that the temporally integrated sparse DA frame 100 is generated based on sparse depth data of a current time frame without using depth information from a previous time frame. Accordingly, at, worst, the RGB-D odometry block 46 reverts to single-frame processing. Thus, the RGB-D odometry block 46 registers a current RGB frame and a current sparse depth frame to the RGB-D memory 97 (e.g., past RGB and fused depth, which corresponds to atemporal buffer) by guessing pose updates (^^ − 1 → ^^).The RGB-D odometry block 46 then integrates (accumulates) over time the current RGB frame and the current sparse depth frame with the RGB-D memory. The integration may be based on a grid or on finite memory. In some embodiments, a grid-based integration allows an integration over a longer time and / or a more complete integration, whereas a finite-memory based integration allows an integration over a shorter time and raises less occlusion issues. In a case of a finite memory, accumulated colored point cloud data in the temporal buffer may be projected to a current frame). Thus, a density may be increased, and a SNR may be increased by summing dots collected in different frames. Photometric checks may ensure at all stages that past points with RGB color associated to them are not used when visually mismatched with a present frame (because then, a determined matching may be too uncertain). Further, edge checks may verify that when a dot from the temporal buffer falls on a detected color edge, it is discarded (because also then, a determined matching may be too uncertain). In some embodiments, the integration over time may fill holes and increase an SNR by densifying a fusion core input. Sony Semiconductor Solutions Corporation et al. 33 Fig.6 illustrates a first embodiment of an RGB-D odometry block 110 with an IMU. Apart from processing an IMU pose, the RGB-D odometry block 110 is configured similar to the RGB-D odometry block 46 of Fig.3, 4 and 5, and may be provided by the circuitry 1 of Fig.1. The RGB-D odometry block 110 receives a last color 111, a past dense point cloud 112, a past dense amplitude 113, a past color 114, an ICP pose 115 (which may be determined by the ICP algorithm at 23 of the method 20 of Fig.2), and an IMU pose 116 (which may be determined by the obtaining of inertial measurement data at 29 of the method 20 of Fig.2). A transforming 117 of target buffers receives the past dense point cloud 112, the past dense amplitude 113, the past color 114, and the ICP pose 115. A transforming 118 of target buffers receives the past dense point cloud 112, the past dense amplitude 113, the past color 114, and the IMU pose 116. The last color 111 is provided to a binning 119 and a binning 122. A result of the transforming 117 of target buffers is provided to a binning 120 and a binning 121, and a result of the transforming 118 of target buffers is provided to a binning 123 and a binning 124. The result provided from the transforming 117 of target buffers to the binning 120 includes an indication of a dense color, and the result provided from the transforming 117 of target buffers to the binning 121 includes an indication of a dense depth. Likewise, the result provided from the transforming 118 of target buffers to the binning 123 includes an indication of a dense color, and the result provided from the transforming 118 of target buffers to the binning 124 includes an indication of a dense depth. Respective results of the binning 119, the binning 120 and the binning 121 are provided to a calculation 125 of a photometric error. Respective results of the binning 122, the binning 123 and the binning 124 are provided to a calculation 126 of a photometric error. Results of the respective calculations 125 and 126 of a photometric error include a bad percentage (i.e., a percentage of image pixels in the respective image data in which the photometric error exceeds a predefined threshold and, thus, indicates a low correspondence), and are provided to a selecting 127 of a minimum error target buffer. At 127, the minimum error target buffer is selected based on a result of the transforming 117 of target buffers, on a result of the transforming 118 of target buffers, and on the respective results of the calculation 125 and 126 of a photometric error. Sony Semiconductor Solutions Corporation et al. 34 A result 128 of the selecting 127 of the minimum error target buffer includes a registered depth, a registered amplitude, a registered color and a registered point cloud. In some embodiments, adding an IMU-based pose estimation, as shown in Fig.6, improves a pose accuracy. However, in the embodiment of Fig.6, a photometric error is calculated twice (at 125 and at 126), once for the ICP pose 115 and once for the IMU pose 116. A heavy transformation of target buffers is performed twice (at 117 and at 118), as shown by the dashed box in Fig.6, and a corresponding binning is performed twice. The double cost of the double photometric error calculation may be optimized, e.g., as shown in Fig.7. Fig.7 illustrates a second embodiment of an RGB-D odometry block 130 with an IMU. Apart from processing an IMU pose, the RGB-D odometry block 130 is configured similar to the RGB-D odometry block 46 of Fig.3, 4 and 5, and may be provided by the circuitry 1 of Fig.1. In the RGB-D odometry block 130, a cost of a photometric error calculation is optimized with respect to the RGB-D odometry block 110 of Fig.6. Like the RGB-D odometry block 110, the RGB-D odometry block 130 receives the last color 111, the past dense point cloud 112, the past dense amplitude 113, the past color 114, the ICP pose 115 (which may be determined by the ICP algorithm at 23 of the method 20 of Fig.2), and the IMU pose 116 (which may be determined by the obtaining of inertial measurement data at 29 of the method 20 of Fig.2). The last color 111 is provided to a binning 131, the past dense point cloud 112 is provided to a binning 132, and the past color 114 is provided to a binning 133. A transforming 134 of small buffers receives a result of the binning 132, a result of the binning 133, and the ICP pose 115. A result of the transforming 134 of small buffers includes an indication of a small dense color and a small dense depth (wherein the term “small” refers to a buffer size, which is reduced due to the binning 132 and 133), and is provided to a calculation 136 of a photometric error. A transforming 135 of small buffers receives a result of the binning 132, a result of the binning 133, and the IMU pose 116. A result of the transforming 135 of small buffers includes an indication of a small dense color and a small dense depth (wherein the term “small” refers to a Sony Semiconductor Solutions Corporation et al. 35 buffer size, which is reduced due to the binning 132 and 133), and is provided to a calculation 137 of a photometric error. The calculation 136 of a photometric error receives a result of the binning 131 and the result of the transforming 134 of small buffers, and the calculation 137 of a photometric error receives a result of the binning 131 and the result of the transforming 135 of small buffers. Respective results of the calculation 136 and 137 of a photometric error include a bad percentage (i.e., a percentage of image pixels in the respective image data in which the photometric error exceeds a predefined threshold and, thus, indicates a low correspondence), and are provided to a selecting 138 of a minimum error target buffer. At 138, the minimum error target buffer is selected based on the ICP pose 115, on the IMU pose 116, and on the respective results of the calculation 136 and 137 of a photometric error. A result of the selecting 138 of a minimum error target buffer includes a better pose (i.e., a pose among the ICP pose 115 and the IMU pose 116 for which a smaller photometric error is calculated at 136 or 137, respectively) and is provided to a transforming 139 of target buffers. The transforming 139 of target buffers receives the past dense point cloud 112, the past dense amplitude 113, the past color 114 and the result of the selecting 138 of a minimum error target buffer. A result 140 of the transforming 139 of target buffers includes a registered depth, a registered amplitude, a registered color and a registered point cloud. As shown in Fig.7, the binning 131, 132 and 133, respectively, is performed once, in contrast to Fig.6, where a respective binning 119, 120, 121, 122, 123 and 124 is performed twice (once for the ICP pose 115 and once for the IMU pose 116). Further, a transforming of small buffers is performed twice (at 134 and at 135). Since the binning 132 and 133 is performed before the transforming 134 and 135, the transforming 134 and 135 is performed on small buffers instead of the target buffers, such that the transforming twice at 134 and 135 is computationally light, as indicated by the dashed box in Fig.7. Thus, a computationally heavy transforming 139 of target buffers is performed only once in the RGB-D odometry block 130, as indicated by the dashed box around the transforming 139. As described, an IMU input may improve a pose estimation. The IMU input may be used complementary to RGB (or other image data) as input source. Fig.8 illustrates an embodiment of an RGB-D odometry 150 with an external pose input. The RGB-D odometry 150 includes an RGB-D fusion pipeline 151 that receives sparse depth Sony Semiconductor Solutions Corporation et al. 36 data 152 and RGB data 153. The RGB-D fusion pipeline 151 is configured similar to the RGB-D fusion block 47 of Fig.3. The RGB-D fusion pipeline 151 may be performed by the circuitry 1 of Fig.1 and is, in some embodiments, included in the estimating of the first positional relation at 22 of Fig.2. The sparse depth data 152 are an example of the sparse depth data acquired at 21 of Fig.2 by the depth sensor 6 of Fig.1 based on an array of emitted laser dots. The RGB data 153 are an example of image data acquired at 25 of Fig.2 by the image sensor 7 of Fig.1. Additionally, the RGB-D fusion pipeline 151 receives an external pose update 154, which is an example of an external pose input, from a SLAM / odometry software 155. Based on the sparse depth data 152, the RGB data 153 and the external pose update 154, the RGB-D fusion pipeline 151 generates dense RGB-D data 156, in which the sparse depth data 152 and the RGB data 153 are fused according to the external pose update 154. Additionally, the RGB-D fusion pipeline 151 generates dense depth data for the dense RBG-D data 156 by interpolating the sparse depth data 152 according to the RGB data 153. The RGB-D fusion pipeline 151 provides the dense RGB-D data 156 to the SLAM / odometry software 155. The external pose input 154 indicates a movement of a depth sensor (that has acquired the sparse depth data 152, e.g., the depth sensor 6 of Fig.1) between a previous time frame and a latest time frame in which the depth sensor has acquired the sparse depth data 152. The previous time frame corresponds to a time prior to the latest time frame. The previous time frame is a time frame at which a 3D model, which should be updated based on the dense RGB-D data 156, has been updated (i.e., changed). Thus, the 3D model corresponds to the previous time frame. The SLAM / odometry software 155 performs odometry using a SLAM approach based on pose data 157. The SLAM / odometry software 155 generates the external pose update 154 based on the dense RGB-D data 156 and on the pose data 157. The SLAM / odometry software 155 further provides a pose update to an application 158. The pose data 157 include inertial measurement data from an IMU (e.g., from the IMU 8 of Fig.1) that is fixed to the depth sensor (e.g., the depth sensor 6 of Fig.1) which has acquired the sparse depth data 152. Accordingly, the external pose input 154 is based on inertial measurement data from an IMU that is fixed to the depth sensor which has acquired the sparse depth data 152. Fig.9 illustrates embodiments 160 and 170 of a temporal buffer behavior. Circles represent time frames, wherein filled circles correspond to time frames at which sparse depth data are acquired (e.g., by the depth sensor 6 of Fig.1), and empty circles correspond to time frames at which no Sony Semiconductor Solutions Corporation et al. 37 sparse depth data are acquired. Therefore, the depth sensor ignores timestamps that correspond to empty circles and captures timestamps that correspond to filled circles. Boxes around the circles represent a temporal buffer. The temporal buffer is drawn with thinner lines for earlier time frames, and the bolder a box of the temporal buffer is drawn, the later the time frame to which the box corresponds. The temporal buffer includes a 3D representation (e.g., the 3D model updated at 31 of Fig.2) of a scene based on data (sparse depth data and / or other sensor data, such as image data and / or inertial measurement data) that have been acquired at the time frames whose corresponding circles are included in the respective box of the temporal buffer. The 3D model of the temporal buffer is updated for each time frame. Therefore, the temporal buffer is configured as a sliding buffer. An output frame with a depth map is generated (e.g., at 34 of Fig.2) for each time frame based on the 3D model of the temporal buffer at the respective time frame. Thus, a density of circles (and of boxes) on a timeline in Fig.9 corresponds to an output frame rate of a depth map stream (e.g., the depth map stream generated at 35 of Fig.2). In Fig.9, the temporal buffer includes data from up to twelve subsequent time frames, such that 3D points in the temporal buffer for subsequent time frames overlap by eleven time frames. However, the disclosure is not limited to a temporal buffer of twelve time frames. The temporal buffer includes data from two, three, four, six, eight, 16, 20 or any other suitable number of time frames in some embodiments. Moreover, the temporal buffer has a tunable time window that can extend to use one or more frames in the past, depending on a window length. In Fig.9, the temporal buffer is configured as point cloud buffers that store, as the 3D model, 3D points of the scene that are indicated by sparse depth data (e.g., the sparse depth data obtained at 21 of Fig.2). In the embodiment 160, the depth sensor acquires sparse depth data at each one of the subsequent time frames, and the 3D model of the temporal buffer includes 3D points of the scene from all twelve time frames.3D points from a time frame that is earlier than a temporal buffer are ignored in the 3D model of the temporal buffer. Accordingly, the embodiment 160 corresponds to a “no skip” mode in which an acquisition of sparse depth data is not skipped at any one of the twelve time frames of the temporal buffer. Thus, sparse depth data are acquired for each output frame. The embodiment 170 corresponds to a synchronous mode in which an acquisition of sparse depth data is skipped for a predefined number N of time frames between subsequent acquisitions Sony Semiconductor Solutions Corporation et al. 38 of sparse depth data. In the embodiment 170, an acquisition of sparse depth data is skipped for N=7 frames between subsequent acquisitions of sparse depth data. However, the disclosure is not limited to N=7. The number N of skipped time frames may be set to one, two, three, four, eight, twelve, or any other suitable value. Accordingly, a depth frame rate at which the depth sensor acquires depth data is lower than the output frame rate of the depth map stream. Thus, a depth frame rate (e.g., a ToF frame rate) is reduced, such that sparse depth data are acquired every N+1 output frames. At time frames at which sparse depth data are not acquired (empty circles), sensor data are acquired by an image sensor such as the image sensor 7 of Fig.1 and by an IMU such as the IMU 8 of Fig.1, and the 3D model is transformed (e.g., translated, rotated, scaled, sheared etc.) and / or interpolated as necessary for representing the scene according to the acquired sensor data. Therefore, in the embodiment 170, circuitry (e.g., the circuitry 1 of Fig.1) further obtains sensor data acquired by a further sensor at a third time frame. A positional relation between the further sensor and the depth sensor is fixed. After updating the 3D model at 31 of Fig.2 according to sparse depth data that have been acquired at a first time frame that is prior to the third time frame (such that the 3D model corresponds to the first time frame), the circuitry further estimates a third positional relation between a representation of the scene of the 3D model (that corresponds to the first time frame) and a representation of the scene of the sensor data from the third time frame. The circuitry then updates the 3D model based on the third positional relation; and generates a depth map based on the 3D model that has been updated based on the third positional relation. In the embodiment 170, the circuitry is configured to generate a depth map at the output frame rate that is higher than the depth frame rate at which the depth sensor acquires sparse depth data; and a sensor frame rate at which the further sensor acquires sensor data is at least as high as the output frame rate. Thus, the circuitry obtains sensor data for each output frame such that the 3D model, on which a depth map of the output frame is based, can be updated for each output frame even if no sparse depth data are acquired that correspond to the output frame. The further sensor includes an image sensor, and the sensor data include image data. The further sensor also includes an IMU, and the sensor data include inertial measurement data. However, in some embodiments, the further sensor includes the image sensor but not the IMU, and in some embodiments, the further sensor includes then IMU but not the image sensor. Sony Semiconductor Solutions Corporation et al. 39 In the embodiment 170, a minimal amount of frames from a ToF sensor (an example of the depth sensor 6 of Fig.1) is used, while a RGB frame rate from an image sensor (an example of the image sensor 7 of Fig.1) is kept high. Therefore, the embodiment 170 corresponds to “low frame rate ToF”. A key performance index of the embodiment 170 is an RGB frame rate divided by a ToF frame rate. The higher this ratio, the less power is spent on ToF capture (i.e., on acquiring sparse depth data). The embodiment 170 is based on the following principles. Sparse depth data from a past frame is used in a current frame, by using a pose estimate provided by one of the remaining data sources, which include visual, inertial, and / or visual-inertial odometry. The resulting accumulated sparse depth frame (temporal buffer) is densified by running a main RGB-D fusion block (e.g., the RGB-D fusion block 47 of Fig.3 or the RGB-D fusion pipeline 151 of Fig.8). A depth completion (e.g., the depth completion 47a of Fig.3) includes an interpolation. The embodiment 170 may be combined with using an IMU pose input, such as the IMU pose input 116 of Fig.6 or 7, and / or the pose data 147 of Fig.8. Further, the depth estimate may be produced by more than one RGB frame (an example of image data form more than one time frames) and sparse depth data from one time frame (multi-view RGB with sparse depth). Fig.10 illustrates an embodiment of a timing relationship for generating depth maps based on a 3D model in a sequence of time frames. A 3D model 181 of time frame ^^ indicates 3D points of a scene according to sparse depth data acquired by a depth sensor (e.g., by the depth sensor 6 of Fig.1) at time frame ^^. The 3D points from time frame ^^ are illustrated as circles in Fig.10.3D points (coordinates) of the scene for which the sparse depth data of time frame ^^ do not indicate a depth (e.g., because no reflected illumination light was received from the coordinates, or because an uncertainty of a measured depth exceeds a predefined threshold) are not represented in the sparse depth data of time frame ^^. A depth completion 182 is performed on the 3D model 181 for determining dense depth data according to a predefined resolution. The depth completion 182 includes interpolating between points of the 3D model 181 and extrapolating in a neighborhood (i.e., in a predefined periphery) of the points of the 3D model 181. Sony Semiconductor Solutions Corporation et al. 40 Based on the 3D model 181 and on the depth completion 182, a depth map 183 is generated (e.g., by the processing unit 2 of Fig.1). The depth map 183 includes a 2D array of pixels. A hatched portion of the depth map 183 illustrates a portion of the 2D array of pixels for which a depth value is determined based on the 3D model 181 and / or based on the depth completion 182. A white portion of the depth map 183 illustrates a portion of the 2D array of pixels for which a depth value is not determined.At time frame ^^ + 1, the depth sensor acquires sparse depth data (e.g., at 21 of Fig. 2), and a poseis estimated at 184 for registering the 3D model 181 of time frame ^^ and the sparse depth data oftime frame ^^ + 1 (an example of the estimating of the first positional relation at 22 of Fig. 2).Points are then accumulated according to the estimated pose in an updated 3D model 185 of timeframe ^^ + 1 such that the 3D model 185 indicates both the 3D points of the scene that areindicated by the sparse depth data of time frame ^^ and 3D points of the scene that are indicatedby the sparse depth data of time frame ^^ + 1 (an example of the updating of the 3D model at 31of Fig. 2). The 3D points from time frame ^^ + 1 are illustrated as squares in Fig. 10.On the updated 3D model 185, a depth completion 186 is performed and a depth map 187 isgenerated (e.g., at 34 of Fig. 2) based on the updated 3D model 185 of time frame ^^ + 1 and onthe depth completion 186. Due to the higher number and density of points in the 3D model 185of time frame ^^ + 1 than in the 3D model 181 of time frame ^^ + 1, a depth value is determinedfor a larger portion of the 2D array of pixels of the depth map 187 of time frame ^^ + 1 than forthe depth map 183 of time frame ^^. In the case of Fig.10, a depth value is determined for each pixel of the depth map 187, as indicated by the hatching. Fig.11 illustrates an embodiment of generating a depth map stream based on sparse depth data and image data acquired at different frame rates. At time frame ^^, sparse depth data are acquired in a ToF measurement by a depth sensor (e.g., by the depth sensor 6 of Fig.1). A processing unit (e.g., by the processing unit 2 of Fig.1) obtains the sparse depth data and generates a 3D model based on the sparse depth data. Further, an RGB image (an example of image data) is acquired by an image sensor (e.g. by the image sensor 7 of Fig.1). A depth map is generated (e.g., by the processing unit 2) based on the 3D model (and, thus on the sparse depth data of time frame ^^) and on the RGB image of time frame ^^.At time frame ^^ + 1, the image sensor acquires a further RGB image. The processing unitobtains the RGB image and generates a depth map based on the 3D model (and, thus on thesparse depth data of time frame ^^) and on the RGB image of time frame ^^ + 1. Sony Semiconductor Solutions Corporation et al. 41Likewise, at time frame ^^ + 2, the image sensor acquires a further RGB image. The processingunit obtains the RGB image and generates a depth map based on the 3D model (and, thus on thesparse depth data of time frame ^^) and on the RGB image of time frame ^^ + 2.At time frame ^^ + 3, the depth sensor acquires further sparse depth data in a ToF measurementand the image sensor acquires a further RGB image. The processing unit acquires the further sparse depth data (e.g., at 21 of Fig.2) and the further RGB image (e.g., at 25 of Fig.2), estimates (e.g., at 22 of Fig.2) a positional relation between a representation of a scene of the further sparse depth data and a representation of the scene of the 3D model from time frame ^^, and updates (e.g., at 31 of Fig.2) the 3D model by merging the sparse depth data of time frame^^ + 3 into the 3D model such that the 3D model is based on the sparse depth data of time frame ^^and the sparse depth data of time frame ^^ + 3. The processing unit then generates (e.g., at 34 ofFig.2) a depth map based on the updated 3D model (and, thus on the sparse depth data of timeframes ^^ and ^^ + 3) and on the RGB image of time frame ^^.Thus, sparse depth data are acquired at a lower frame rate than image data, and a depth map stream is generated by generating a depth map for each time frame, wherein a depth map for a time frame in which no sparse depth data are acquired is based on a 3D model from a last time frame at which sparse depth data have been acquired. Thus, Fig.11 shows additional channel’s dynamics of a low frame rate depth. Accordingly, the present disclosure provides using sparse iToF / dToF fusion with odometry in short-time. A camera may be moved, e.g., by panning the camera and / or by moving the camera through a scene to scan it by accumulating points. A main tracking source may be a sparse depth at a present frame and may be fused with a dense depth of a past frame, wherein RGB (or other image data) may be used as data source to validate whether the pose is acceptable (e.g., is consistent within a predefined tolerance with the RGB or other image data). A photometric check may be used to verify that the past- and present-frame RGB images can be overlaid up to projection with the estimated depth. An IMU may be used as a completely different source of pose, and in some embodiments, an IMU source of pose can be integrated in short-time (e.g., 3 to 4 frames, without limiting the disclosure thereto) so that the relative pose may be estimated with higher accuracy and / or lower uncertainty. According to the present disclosure, an RGB-D fusion pipeline with temporal accumulation of sparse ToF points uses a temporal buffer with checks to handle an uncertainty of points. Sony Semiconductor Solutions Corporation et al. 42 An external pose input may be used to obtain a smooth dense depth. The external pose input may be obtained from visual-inertial odometry, visual odometry, and / or classic SLAM. Further, inertial (IMU-only) odometry may allow IMU integration as a pose source for short temporal windows. Additionally, the external pose input may be based on a Kalman filter algorithm or on a machine learning-based extension of a Kalman filter algorithm on IMU data. A smooth depth mode may be obtained as byproduct of a temporal accumulation of sparse ToF points. A dense map may be a result of a densification of very sparse point clouds by using a temporal axis. A ToF frame rate may be reduced by temporal accumulation of sparse ToF points. Thus, a number of exposures / captures of ToF per second may be reduced while keeping an RGB frame rate high, as long as a pose of a depth sensor is well estimated. Moving objects may be handled by the temporal buffer with checks. Fig.12 illustrates an embodiment of a general-purpose computer 250. The general-purpose computer 250 is an example of an information processing apparatus that includes circuitry that is configured to perform the method according to the present technology (e.g., the method 20 of Fig.2). The computer has components 251 to 261, which can form a circuitry, such as any one of the circuitry 1, the processing unit 2, or the like, as described herein. Embodiments which use software, firmware, programs or the like for performing the methods as described herein can be installed on computer 250, which is then configured to be suitable for the concrete embodiment. The computer 250 has a CPU 251 (Central Processing Unit), which can execute various types of procedures and methods as described herein, for example, in accordance with programs stored in a read-only memory (ROM) 252, stored in a storage 257 and loaded into a random-access memory (RAM) 253, stored on a medium 260 which can be inserted in a respective drive 259, etc. Furthermore, the computer 250 includes an artificial intelligence (AI) processor 251a. The AI processor 251a may include a graphics processing unit (GPU) and / or a tensor processing unit (TPU). The AI processor 251a may be configured to execute an AI model (e.g., an artificial neural network). For example, the AI processing 251a may provide the AI unit 2a of Fig.1. The CPU 251, the ROM 252 and the RAM 253 are connected with a bus 261, which in turn is connected to an input / output interface 254. The number of CPUs, memories and storages is only Sony Semiconductor Solutions Corporation et al. 43 exemplary, and the skilled person will appreciate that the computer 250 can be adapted and configured accordingly for meeting specific requirements which arise when it functions as an information processing apparatus according to the present technology. At the input / output interface 254, several components are connected: an input 255, an output 256, the storage 257, a communication interface 258 and the drive 259, into which a medium 260 (compact disc (CD), digital video disc (DVD), universal serial bus (USB) flash drive, secure digital (SD) card, CompactFlash (CF) memory, or the like) can be inserted. The input 255 can be a pointer device (mouse, graphic table, or the like), a keyboard, a microphone, a camera, a touchscreen, an eye-tracking unit etc. The output 256 can have a display (liquid crystal display (LCD), cathode ray tube (CRT) display, light-emitting diode (LED) display, electronic paper, etc.; e.g., included in a touchscreen), loudspeakers, etc. The storage 257 can have a hard disk drive (HDD), a solid-state drive (SSD), a flash drive and the like. The communication interface 258 can be adapted to communicate, for example, via universal serial bus (USB), a serial port (RS-232), parallel port (IEEE 1284), a local area network (LAN; e.g., ethernet), wireless local area network (WLAN; e.g., Wi-Fi, IEEE 802.11), mobile telecommunications system (GSM, UMTS, LTE, NR etc.), Bluetooth, near-field communication (NFC), ZigBee, infrared, etc. It should be noted that the description above only pertains to an example configuration of computer 250. Alternative configurations may be implemented with additional or other sensors, storage devices, interfaces or the like. For example, the communication interface 258 may support other radio access technologies than the mentioned UMTS, LTE and NR. It should be recognized that the embodiments describe methods with an exemplary ordering of method steps. The specific ordering of method steps is however given for illustrative purposes only and should not be construed as binding. Changes of the ordering of method steps may be apparent to the skilled person. Please note that the division of the circuitry 1 into units 2 to 8 is only made for illustration purposes and that the present disclosure is not limited to any specific division of functions in specific units. For instance, the circuitry 1 could be implemented by a respective programmed processor, field programmable gate array (FPGA) and the like. It is further noted that, although Sony Semiconductor Solutions Corporation et al. 44 the circuitry 1 is shown as one apparatus, the circuitry according to the disclosure may as well be implemented by a plurality of apparatuses (e.g., of processors) that are connected (e.g., via a communication network and / or via a bus) to provide the present technology. The method 20 of Fig.2 can also be implemented as a computer program causing a computer and / or a processor, such as the circuitry 1 and / or the processing unit 2 discussed above, to perform the method, when being carried out on the computer and / or processor. In some embodiments, also a non-transitory computer-readable recording medium is provided that stores therein a computer program product, which, when executed by a processor, such as the processor described above, causes the method described to be performed. All units and entities described in this specification and claimed in the appended claims can, if not stated otherwise, be implemented as integrated circuit logic, for example on a chip, and functionality provided by such units and entities can, if not stated otherwise, be implemented by software. In so far as the embodiments of the disclosure described above are implemented, at least in part, using software-controlled data processing apparatus, it will be appreciated that a computer program providing such software control and a transmission, storage or other medium by which such a computer program is provided are envisaged as aspects of the present disclosure. Note that the present technology can also be configured as described below. (1) Circuitry for generating a depth map, the circuitry being configured to: obtain first sparse depth data acquired by a depth sensor at a first time frame, wherein the first sparse depth data indicate three-dimensional points of a scene in the first time frame; estimate a first positional relation between a representation of the scene of the first sparse depth data and a representation of the scene of a three-dimensional model, wherein the three- dimensional model is based on second sparse depth data acquired by the depth sensor at a second time frame prior to the first time frame; update the three-dimensional model based on the first sparse depth data and the first positional relation; and generate a depth map of the scene based on the updated three-dimensional model. (2) The circuitry of (1), wherein the estimating of the first positional relation includes verifying the first positional relation based on data acquired by another sensor, a positional relation between the other sensor and the depth sensor being fixed. Sony Semiconductor Solutions Corporation et al. 45 (3) The circuitry of (2), wherein the other sensor includes an image sensor; and wherein the verifying of the first positional relation includes: obtaining first image data acquired by the image sensor, wherein the first image data correspond to the first time frame; estimating a second positional relation between an image of the scene represented by the first image data and an image of the scene represented by second image data acquired by the image sensor, wherein the second image data correspond to the second time frame; and comparing the first positional relation with the second positional relation. (4) The circuitry of (3), wherein the estimating of the second positional relation includes performing image registration for registering the first sparse depth data and the first image data. (5) The circuitry of (4), wherein the image registration is based on a machine learning model configured to perform image registration. (6) The circuitry of any one of (3) to (5), wherein the estimating of the second positional relation is based on a photometric error between the first image data and the second image data. (7) The circuitry of any one of (2) to (6), wherein the other sensor includes an inertial measurement unit; and wherein the verifying of the first positional relation includes: obtaining inertial measurement data from the inertial measurement unit, wherein the inertial measurement data indicate a change in position of the depth sensor between the second time frame and the first time frame; and comparing the first positional relation with the change in position indicated by the inertial measurement data. (8) The circuitry of any one of (1) to (7), wherein the estimating of the first positional relation is based on an iterative closest point algorithm. (9) The circuitry of (8), wherein the iterative closest point algorithm includes weighting the three-dimensional Sony Semiconductor Solutions Corporation et al. 46 points indicated by the first sparse depth data with respective weights according to respective confidences of the respective three-dimensional points. (10) The circuitry of (9), wherein the weights are based on respective depths that correspond to the respective three-dimensional points. (11) The circuitry of (9) or (10), wherein the depth sensor is configured to acquire the first sparse depth data based on light emitted into the scene; and wherein the weights of the respective three-dimensional points are based on a respective amplitude of reflected light from the scene. (12) The circuitry of any one of (9) to (11), wherein the weights are based on a photometric error between first image data of the scene acquired by an image sensor at the first time frame and second image data of the scene acquired by the image sensor at the second time frame, a positional relation between the image sensor and the depth sensor being fixed. (13) The circuitry of any one of (1) to (12), wherein the estimating of the first positional relation is based on a machine learning model configured to determine the first positional relation. (14) The circuitry of any one of (1) to (13), wherein the updating of the three-dimensional model is based on a machine learning model configured to merge the first sparse depth data into the three-dimensional model. (15) The circuitry of any one of (1) to (14), wherein the updating of the three-dimensional model includes deleting, from the three- dimensional model, depth information that is based on sparse depth data which are older than a predefined number of time frames. (16) The circuitry of any one of (1) to (15), wherein the three-dimensional model includes a set of three-dimensional points that represents the scene, wherein the set of three-dimensional points is based at least on the second sparse depth data. (17) The circuitry of any one of (1) to (16), wherein the three-dimensional model includes a grid that represents the scene, wherein Sony Semiconductor Solutions Corporation et al. 47 values assigned to respective positions of the grid are based at least on the second sparse depth data. (18) The circuitry of any one of (1) to (17), wherein the depth sensor is configured to acquire the first sparse depth data based on a time-of-flight measurement. (19) The circuitry of any one of (1) to (18), further configured to generate a depth map stream of subsequent time frames, wherein the depth map stream includes the generated depth map for the first time frame. (20) The circuitry of any one of (1) to (19), wherein the estimating of the first positional relation is based on an external pose input that indicates a movement of the depth sensor between a previous time frame to which the three- dimensional model corresponds and the first time frame, wherein the previous time frame corresponds to a time prior to the first time frame. (21) The circuitry of (20), wherein the external pose input is based on inertial measurement data from an inertial measurement unit that is fixed to the depth sensor. (22) The circuitry of any one of (1) to (21), wherein the circuitry is further configured to: obtain sensor data acquired by a further sensor at a third time frame later than the first time frame, a positional relation between the further sensor and the depth sensor being fixed; estimate a third positional relation between a representation of the scene of the three- dimensional model and a representation of the scene of the sensor data; update the three-dimensional model based on the third positional relation; and generate a depth map based on the updated three-dimensional model. (23) The circuitry of (22), wherein the circuitry is configured to generate a depth map at an output frame rate that is higher than a depth frame rate at which the depth sensor acquires sparse depth data; and wherein a sensor frame rate at which the further sensor acquires sensor data is at least as high as the output frame rate. Sony Semiconductor Solutions Corporation et al. 48 (24) The circuitry of (22) or (23), wherein the further sensor includes an image sensor, and the sensor data include image data. (25) The circuitry of any one of (22) to (24), wherein the further sensor includes an inertial measurement unit, and the sensor data include inertial measurement data. (26) A method for generating a depth map, the method comprising: obtaining first sparse depth data acquired by a depth sensor at a first time frame, wherein the first sparse depth data indicate three-dimensional points of a scene in the first time frame; estimating a first positional relation between a representation of the scene of the first sparse depth data and a representation of the scene of a three-dimensional model, wherein the three-dimensional model is based on second sparse depth data acquired by the depth sensor at a second time frame prior to the first time frame; updating the three-dimensional model based on the first sparse depth data and the first positional relation; and generating a depth map of the scene based on the updated three-dimensional model. (27) The method of (26), wherein the estimating of the first positional relation includes verifying the first positional relation based on data acquired by another sensor, a positional relation between the other sensor and the depth sensor being fixed. (28) The method of (27), wherein the other sensor includes an image sensor; and wherein the verifying of the first positional relation includes: obtaining first image data acquired by the image sensor, wherein the first image data correspond to the first time frame; estimating a second positional relation between an image of the scene represented by the first image data and an image of the scene represented by second image data acquired by the image sensor, wherein the second image data correspond to the second time frame; and comparing the first positional relation with the second positional relation. (29) The method of (28), wherein the estimating of the second positional relation includes performing image registration for registering the first sparse depth data and the first image data. Sony Semiconductor Solutions Corporation et al. 49 (30) The method of (29), wherein the image registration is based on a machine learning model configured to perform image registration. (31) The method of any one of (28) to (30), wherein the estimating of the second positional relation is based on a photometric error between the first image data and the second image data. (32) The method of any one of (27) to (31), wherein the other sensor includes an inertial measurement unit; and wherein the verifying of the first positional relation includes: obtaining inertial measurement data from the inertial measurement unit, wherein the inertial measurement data indicate a change in position of the depth sensor between the second time frame and the first time frame; and comparing the first positional relation with the change in position indicated by the inertial measurement data. (33) The method of any one of (26) to (32), wherein the estimating of the first positional relation is based on an iterative closest point algorithm. (34) The method of (33), wherein the iterative closest point algorithm includes weighting the three-dimensional points indicated by the first sparse depth data with respective weights according to respective confidences of the respective three-dimensional points. (35) The method of (34), wherein the weights are based on respective depths that correspond to the respective three-dimensional points. (36) The method of (34) or (35), wherein the depth sensor acquires the first sparse depth data based on light emitted into the scene; and wherein the weights of the respective three-dimensional points are based on a respective amplitude of reflected light from the scene. (37) The method of any one of (34) to (36), wherein the weights are based on a photometric error between first image data of the Sony Semiconductor Solutions Corporation et al. 50 scene acquired by an image sensor at the first time frame and second image data of the scene acquired by the image sensor at the second time frame, a positional relation between the image sensor and the depth sensor being fixed. (38) The method of any one of (26) to (37), wherein the estimating of the first positional relation is based on a machine learning model configured to determine the first positional relation. (39) The method of any one of (26) to (38), wherein the updating of the three-dimensional model is based on a machine learning model configured to merge the first sparse depth data into the three-dimensional model. (40) The method of any one of (26) to (39), wherein the updating of the three-dimensional model includes deleting, from the three- dimensional model, depth information that is based on sparse depth data which are older than a predefined number of time frames. (41) The method of any one of (26) to (40), wherein the three-dimensional model includes a set of three-dimensional points that represents the scene, wherein the set of three-dimensional points is based at least on the second sparse depth data. (42) The method of any one of (26) to (41), wherein the three-dimensional model includes a grid that represents the scene, wherein values assigned to respective positions of the grid are based at least on the second sparse depth data. (43) The method of any one of (26) to (42), wherein the depth sensor acquires the first sparse depth data based on a time-of-flight measurement. (44) The method of any one of (26) to (43), further comprising generating a depth map stream of subsequent time frames, wherein the depth map stream includes the generated depth map for the first time frame. (45) The method of any one of (26) to (44), wherein the estimating of the first positional relation is based on an external pose input that indicates a movement of the depth sensor between a previous time frame to which the three- Sony Semiconductor Solutions Corporation et al. 51 dimensional model corresponds and the first time frame, wherein the previous time frame corresponds to a time prior to the first time frame. (46) The method of (45), wherein the external pose input is based on inertial measurement data from an inertial measurement unit that is fixed to the depth sensor. (47) The method of any one of (26) to (46), wherein the method further comprises: obtaining sensor data acquired by a further sensor at a third time frame later than the first time frame, a positional relation between the further sensor and the depth sensor being fixed; estimating a third positional relation between a representation of the scene of the three- dimensional model and a representation of the scene of the sensor data; updating the three-dimensional model based on the third positional relation; and generating a depth map based on the updated three-dimensional model. (48) The method of (47), wherein the method comprises generating a depth map at an output frame rate that is higher than a depth frame rate at which the depth sensor acquires sparse depth data; and wherein a sensor frame rate at which the further sensor acquires sensor data is at least as high as the output frame rate. (49) The method of (47) or (48), wherein the further sensor includes an image sensor, and the sensor data include image data. (50) The method of any one of (47) to (49), wherein the further sensor includes an inertial measurement unit, and the sensor data include inertial measurement data. (51) A computer program comprising program code causing a computer to perform the method according to anyone of (26) to (50), when being carried out on a computer. (52) A non-transitory computer-readable recording medium that stores therein a computer program product, which, when executed by a processor, causes the method according to anyone of (26) to (50) to be performed.
Claims
Sony Semiconductor Solutions Corporation et al. 1 CLAIMS 1. Circuitry for generating a depth map, the circuitry being configured to: obtain first sparse depth data acquired by a depth sensor at a first time frame, wherein the first sparse depth data indicate three-dimensional points of a scene in the first time frame; estimate a first positional relation between a representation of the scene of the first sparse depth data and a representation of the scene of a three-dimensional model, wherein the three- dimensional model is based on second sparse depth data acquired by the depth sensor at a second time frame prior to the first time frame; update the three-dimensional model based on the first sparse depth data and the first positional relation; and generate a depth map of the scene based on the updated three-dimensional model.
2. The circuitry of claim 1, wherein the estimating of the first positional relation includes verifying the first positional relation based on data acquired by another sensor, a positional relation between the other sensor and the depth sensor being fixed.
3. The circuitry of claim 2, wherein the other sensor includes an image sensor; and wherein the verifying of the first positional relation includes: obtaining first image data acquired by the image sensor, wherein the first image data correspond to the first time frame; estimating a second positional relation between an image of the scene represented by the first image data and an image of the scene represented by second image data acquired by the image sensor, wherein the second image data correspond to the second time frame; and comparing the first positional relation with the second positional relation.
4. The circuitry of claim 3, wherein the estimating of the second positional relation is based on a photometric error between the first image data and the second image data.
5. The circuitry of claim 2, wherein the other sensor includes an inertial measurement unit; and wherein the verifying of the first positional relation includes: obtaining inertial measurement data from the inertial measurement unit, wherein the inertial measurement data indicate a change in position of the depth sensor between the second time frame and the first time frame; andSony Semiconductor Solutions Corporation et al. 2 comparing of the first positional relation with the change in position indicated by the inertial measurement data.
6. The circuitry of claim 1, wherein the estimating of the first positional relation is based on an iterative closest point algorithm; in particular wherein the iterative closest point algorithm includes weighting the three- dimensional points indicated by the first sparse depth data with respective weights according to respective confidences of the respective three-dimensional points.
7. The circuitry of claim 1, wherein the updating of the three-dimensional model is based on a machine learning model configured to merge the first sparse depth data into the three-dimensional model.
8. The circuitry of claim 1, wherein the updating of the three-dimensional model includes deleting, from the three- dimensional model, depth information that is based on sparse depth data which are older than a predefined number of time frames.
9. The circuitry of claim 1, wherein the estimating of the first positional relation is based on an external pose input that indicates a movement of the depth sensor between a previous time frame to which the three- dimensional model corresponds and the first time frame, wherein the previous time frame corresponds to a time prior to the first time frame.
10. The circuitry of claim 1, wherein the circuitry is further configured to: obtain sensor data acquired by a further sensor at a third time frame later than the first time frame, a positional relation between the further sensor and the depth sensor being fixed; estimate a third positional relation between a representation of the scene of the three- dimensional model and a representation of the scene of the sensor data; update the three-dimensional model based on the third positional relation; and generate a depth map based on the updated three-dimensional model.
11. A method for generating a depth map, the method comprising: obtaining first sparse depth data acquired by a depth sensor at a first time frame, wherein the first sparse depth data indicate three-dimensional points of a scene in the first time frame; estimating a first positional relation between a representation of the scene of the firstSony Semiconductor Solutions Corporation et al. 3 sparse depth data and a representation of the scene of a three-dimensional model, wherein the three-dimensional model is based on second sparse depth data acquired by the depth sensor at a second time frame prior to the first time frame; updating the three-dimensional model based on the first sparse depth data and the first positional relation; and generating a depth map of the scene based on the updated three-dimensional model.
12. The method of claim 11, wherein the estimating of the first positional relation includes verifying the first positional relation based on data acquired by another sensor, a positional relation between the other sensor and the depth sensor being fixed.
13. The method of claim 12, wherein the other sensor includes an image sensor; and wherein the verifying of the first positional relation includes: obtaining first image data acquired by the image sensor, wherein the first image data correspond to the first time frame; estimating a second positional relation between an image of the scene represented by the first image data and an image of the scene represented by second image data acquired by the image sensor, wherein the second image data correspond to the second time frame; and comparing the first positional relation with the second positional relation.
14. The method of claim 13, wherein the estimating of the second positional relation is based on a photometric error between the first image data and the second image data.
15. The method of claim 12, wherein the other sensor includes an inertial measurement unit; and wherein the verifying of the first positional relation includes: obtaining inertial measurement data from the inertial measurement unit, wherein the inertial measurement data indicate a change in position of the depth sensor between the second time frame and the first time frame; and comparing the first positional relation with the change in position indicated by the inertial measurement data.
16. The method of claim 11, wherein the estimating of the first positional relation is based on an iterative closest point algorithm;Sony Semiconductor Solutions Corporation et al. 4 in particular wherein the iterative closest point algorithm includes weighting the three- dimensional points indicated by the first sparse depth data with respective weights according to respective confidences of the respective three-dimensional points.
17. The method of claim 11, wherein the updating of the three-dimensional model is based on a machine learning model configured to merge the first sparse depth data into the three-dimensional model.
18. The method of claim 11, wherein the updating of the three-dimensional model includes deleting, from the three- dimensional model, depth information that is based on sparse depth data which are older than a predefined number of time frames.
19. The method of claim 11, wherein the estimating of the first positional relation is based on an external pose input that indicates a movement of the depth sensor between a previous time frame to which the three- dimensional model corresponds and the first time frame, wherein the previous time frame corresponds to a time prior to the first time frame.
20. The method of claim 11, wherein the method further comprises: obtaining sensor data acquired by a further sensor at a third time frame later than the first time frame, a positional relation between the further sensor and the depth sensor being fixed; estimating a third positional relation between a representation of the scene of the three- dimensional model and a representation of the scene of the sensor data; updating the three-dimensional model based on the third positional relation; and generating a depth map based on the updated three-dimensional model.
Citation Information
Patent Citations
Simultaneous localization and mapping with reinforcement learning
US20180174038A1
Electronic device and method for adaptive time-of-flight sensing based on a 3D model reconstruction
WO2023072707A1