Sensor device and method for operating a sensor device for generating RGBD images
The sensor device addresses high latency and power consumption in RGBD image generation by using an inertial measurement unit to control depth data capture based on sensor movement, ensuring efficient and low-latency RGBD image production.
Patent Information
- Application Number
- PCT/EP2025/057534
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-20
- Filing Date
- 2025-03-19
- Publication Date
- 2025-09-25
AI Technical Summary
Existing RGBD image generation techniques suffer from high latency and power consumption, particularly in portable devices like smartphones and head-mounted displays, due to insufficient data processing during movements of the field of view.
A sensor device comprising an inertial measurement unit, image capturing unit, and depth measuring unit, controlled by a control unit that activates and deactivates the depth measuring unit based on predetermined positional and orientational changes, ensuring low latency and reduced power consumption by only capturing depth data when necessary.
The solution achieves low latency and reduced power consumption in generating RGBD images by controlling depth data acquisition based on sensor movement, optimizing power usage and maintaining image quality.
Smart Images

Figure EP2025057534_25092025_PF_FP_ABST
Abstract
Description
[0001] SENSOR DEVICE AND METHOD FOR OPERATING A SENSOR DEVICE FOR GENERATING
[0002] RGBD IMAGES
[0003] FIELD OF THE INVENTION
[0004] The present disclosure relates to a sensor device for generating RGBD images and a method for operating such a sensor device as well as to a head mounted display. In particular, the present disclosure is related to capturing RGBD images of an object.
[0005] BACKGROUD
[0006] In recent years devices for generating RGBD images have been provided, i.e. for generating images that contain per pixel in addition to color information of a captured object such as red (R), green (G), and blue (B) also depth information (D), i.e. the distance between the image capturing device and the object. Although the name RGBD image suggests red, green, and blue color information, the RGBD image can be based on any representation of the color space such as YUV, YCbCr, HSV or the like. Further, also grayscale images that are supplemented with depth information are termed RGBD images. Such RGBD images are for example used for image classification, since the combination of depth and color information can ease segmentation of different objects. Further, RGBD images are used as input to simultaneous localization and mapping, SLAM, algorithms, since both depth and color information can help to localize the device in space that captures the RGBD images.
[0007] SUMMARY OF THE INVENTION
[0008] However, present techniques for generating RGBD images suffer from insufficient latency between capturing and processing of data, in particular during or after movements of the field of view of image capturing. Moreover, the usefulness of RGBD images for SLAM algorithms is highest in portable devices such as smart phones, tablets, head mounted displays, and the like, which have limited power storage capacities. It is therefore desirable to reduce the latency of RGBD image generation during movements of the field of view as well as the power consumption of RGBD image generation.
[0009] To this end, a sensor device for generating RGBD images of a scene is provided, which sensor device comprises an inertial measurement unit that is configured to detect movement data indicating accelerations and rotations of the sensor device, an image capturing unit that is configured to capture for a field of view image data for a visible light image of the scene, a depth measuring unit that is configured to detect for the same field of view depth data for a depth map of the scene, and a control unit that is configured to determine a position and / or orientation of the sensor device in space based on the movement data of the inertial measurement unit, and to control activation of the image capturing unit and / or the depth measuring unit. Here, the latency of data exchange between the inertial measurement unit and the control unit as well as the latency of data exchange between the control unit and each of the image capturing unit and the depth measuring unit is smaller than a predetermined latency, the predetermined latency being preferably 2 ms, more preferably 1 ms, and more preferably 0.5 ms. The control unit is configured to activate the depth measuring unit if the position and / or orientation of the sensor device that is determined by the control unit has changed by more than a first predetermined amount and is configured to deactivate the depth measuring unit if the position and / or orientation of the sensor device that is determined by the control unit has changed by less than a second predetermined amount.
[0010] Further, a method for operating an according sensor device to generate RGBD images of a scene is provided, the method comprising: by the inertial measurement unit, detecting movement data indicating accelerations and rotations of the sensor device; by the image capturing unit, capturing for a field of view image data for a visible light image of the scene; by the depth measuring unit, detecting for the same field of view depth data for a depth map of the scene; by the control unit, determining a position and / or orientation of the sensor device in space based on the movement data of the inertial measurement unit, and controlling activation of the image capturing unit and / or the depth measuring unit; by the control unit, activating the depth measuring unit if the position and / or orientation of the sensor device that is determined by the control unit has changed by more than a first predetermined amount and deactivating the depth measuring unit if the position and / or orientation of the sensor device that is determined by the control unit has changed by less than a second predetermined amount. Here, the latency of data exchange between the inertial measurement unit and the control unit as well as the latency of data exchange between the control unit and each of the image capturing unit and the depth measuring unit is smaller than a predetermined latency, the predetermined latency being preferably 2 ms, more preferably 1 ms, more preferably 0.5 ms.
[0011] The sensor device provides besides the data necessary for the generation of RGBD images, i.e. besides the image data and the depth data also movement data that indicate accelerations and rotations of the sensor device and allow deduction of relative motions of the sensor device in space. Capturing of the RGBD image data is controlled by a control unit that communicates with low latency with the inertial measurement unit, the image capturing unit, and the depth measuring unit. Based on the relative motions of the sensor device, which affect the field of view of RGBD image data acquisition, the control unit decides whether to activate or deactivate the depth measuring unit. Due to the low latency connection between the inertial measurement unit, the control unit and the depth measuring unit, this control can be carried out with low latency. Moreover, by only using the depth measuring unit for sufficiently large changes of the sensor device’s position / orientation, i.e. for sufficiently large changes of the field of view, the power consumption necessary for depth measurements and for integrating the image data and the depth data can be reduced. In this manner it is guaranteed that RGBD image capturing can be controlled with low latency and reduced power consumption.
[0012] BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Fig. 1 is a simplified block diagram of a sensor device for generating RGBD images;
[0014] Figs. 2A and 2B show in a simplified manner movements of a sensor device relative to a captured scene; Fig. 3 shows in a simplified manner a determination of operation modes of a sensor device;
[0015] Fig. 4 shows a simplified block diagram of another sensor device;
[0016] Fig. 5 shows in a simplified manner a selection of a part of a field of view;
[0017] Fig. 6 shows in a schematic manner an indication for unreliable depth estimates for a region of a captured scene;
[0018] Fig. 7 shows in a schematic manner a two-dimensional confidence map;
[0019] Fig. 8 shows in a simplified manner a determination of operation modes of a sensor device;
[0020] Fig. 9 shows in a simplified manner a sensor device including a time of flight sensor;
[0021] Fig. 10 shows in a simplified manner a head mounted display;
[0022] Fig. 11 schematically shows a process flow for operating a sensor device for generating
[0023] RGBD images; and
[0024] Figs. 12A and 12B show schematically different exemplary applications of a sensor device.
[0025] Fig. 1 is a schematic illustration of a sensor device 100 for generating RGBD images of a scene 200 containing various objects.
[0026] Here, RGBD images are combinations of conventional images, i.e. visible color images containing visible color information as e.g. in red (R), green (G), blue (B) channels or in any other color space. These conventional images are supplemented with depth information for each image point, e.g. for each pixel of the RGBD image. This means, besides the intensities of visible light of different colors each pixel indicates the distance of the device that captured the image to the point of an object of the observed scene from which the received intensities did come. In general, RGBD images are known, and a detailed description can be omitted here.
[0027] The sensor device 100 comprises an inertial measurement unit 110, an image capturing unit 120, a depth measuring unit 130, and a control unit 140. These components might be arranged in a single device, i.e. in a housing of the device, or may be parts of different devices. In the latter case the sensor device 100 is constituted by a system.
[0028] The inertial measurement unit 110 is configured to detect movement data indicating accelerations and rotations of the sensor device 100. To this end, the inertial measurement unit 110 comprises preferably an accelerometer 112 and a gyroscope 114. The accelerometer 112 is configured to detect accelerations along three independent, preferably orthogonal axes in space. The gyroscope 114 is configured to detect rotations around three independent, preferably orthogonal axes in space (which might but need not be identical to the measurement axes of the accelerometer 112). Thus, the inertial measurement unit 110 is configured to detect positional changes of the sensor device 100 along all six degrees of freedom in the form of longitudinal or transversal accelerations. Based on such movement data, it is possible to determine changes in position and / or orientation of the sensor device 100 by integrating the respective equations of motion over time. The accelerometer 112 and the gyroscope 114 may for example by microelectromechanical systems, MEMS, that allow generation of the movement data with sufficient precision. However, also any other sensor units can be used in the inertial measurement unit 110 that allow detecting movement data from which relative position and orientation changes can be determined in an iterative manner. Since such inertial measurement units 110 are in principle known, a more detailed description of them can be omitted here.
[0029] The image capturing unit 120 is configured to capture for a field of view F image data for a visible light image of the scene 200. That is, the image capturing unit 120 receives visible light that is emitted / reflected from the scene 200 within the field of view F of the imaging unit 120. The field of view F is here for example defined by imaging optics 121 of the image capturing unit 120. The light received by the imaging optics 121 is projected onto a pixel array 125 consisting of a, preferably two- dimensionally arranged, plurality of pixels 126, and converted into electrical signals representing the received intensity, which electrical signals constitute the image data. The image capturing unit 120 may detect the absolute intensities of the received light in a synchronized, frame-based manner. The image capturing unit 120 may additionally also detect changes of the intensities and output event data indicating that a change of intensity above an event threshold has occurred at a specific time at a specific pixel 126. Thus, the image capturing unit 120 may operate as a conventional image sensor and additionally also as an event-based vision sensor. In principle, the structure of the image capturing unit 120 is arbitrary, as long as it generates image data that allow reconstructing a color image of the scene 200 within the field of view F. Since such image capturing devices 120 are in principle known, a more detailed description thereof can be omitted here.
[0030] The depth measuring unit 130 is configured to detect for the same field of view F depth data for a depth map of the scene 200. That is, the depth measuring unit 130 observes the same field of view F as the image capturing unit 120. The field of view F may be identical, i.e. the image capturing unit 120 and the depth measuring unit 130, or at least their optical paths, may be aligned such that both units have the same point of view and the same viewing range. However, in principle it is sufficient that it is possible to assign to each pixel captured by the image capturing unit 120 a depth value measured by the depth measuring unit 130. To this end, it will be sufficient if the image capturing unit 120 and the depth measuring unit 130 observe the same region of the scene 200 from different points of view, but with a known or deducible positional relationship to each other. This allows then to match the fields of view F computationally such that each color pixel of the resulting color image can be provided with a corresponding depth value. Accordingly, the same field of view F is understood in the present context to also include the possibility to perform such a computational matching. The depth measuring unit 130 may in principle by constituted arbitrarily, as long as the depth data captured by it allow a deduction of the distance between the observed objects in the scene 200 and the depth measuring unit 130. For example, the depth measuring unit 130 may be a stereo camera. One camera of this stereo camera may then be constituted by the image capturing unit 120. The stereo camera may, however, also be provided additionally to the image capturing unit 120. The depth measuring unit 130 may also be a structured light system that projects a light pattern such as one or a plurality of lines, a checkerboard pattern or the like onto the scene 200 and deduces the distances to the scene from the distortion of the emitted light patterns and by use of triangulation. Further, the depth measuring unit 130 may be constituted as a direct time of flight sensor that emits light pulses into the field of view F and determines the distance to the scene 200 from the time it takes until the light pulses are received again. For structured light systems or time of flight sensors, the light receiving parts may be aligned or interleaved with the pixel array 125 of the image capturing unit 120. For example, the light receiving part of the depth measuring unit 130 may operate with infrared light and may be provided along an optical path behind the pixel array 125 and behind a filter that allows only transmission of infrared light (see also the description of Fig. 9 below). Then, the field of view F of the image capturing unit 120 and the depth measuring unit 130 can be made identical and there is no necessity to align different fields of view by computation.
[0031] The control unit 140 is configured to determine a position and / or orientation of the sensor device 100 in space based on the movement data of the inertial measurement unit 110, and to control activation of the image capturing unit 120 and / or the depth measuring unit 130 based on the determined position / orientation. Here, the control unit 140 may be a processor, like e.g. a microprocessor, a CPU, a GPU, a FPGA or the like. The control unit 140 may be any kind of hardware, software or combination of both that allows carrying out the functions of the control unit 140 described herein.
[0032] The control unit 140 receives the movement data from the inertial measurement unit 110, i.e. data regarding accelerations and rotations of the sensor device 100, and determines from these iteratively changes in the position and the orientation of the sensor device 100 in a in principle known manner. Based on this determination the control unit 140 controls the image capturing unit 120 and in particular the depth measuring unit 130.
[0033] Specifically, the control unit 140 is configured to activate the depth measuring unit 130 if the position and / or orientation of the sensor device 100 that is determined by the control unit 140 has changed by more than a first predetermined amount and is configured to deactivate the depth measuring unit 130 if the position and / or orientation of the sensor device 100 that is determined by the control unit 140 has changed by less than a second predetermined amount. Accordingly, whenever the field of view F of the image capturing unit 120 and the depth measuring unit 130 changes considerably due to a large change of the position and / or orientation of the sensor device 100 in space, the depth measuring unit 130 is turned on and generates depth data. During depth acquisition both the image capturing device 120 and the depth measuring unit 130 may operate at high frame rates that are e.g. equal to the rate at which the changes in position / orientation are determined. When the amount of position / orientation change per unit time is reduced, i.e. when the field of view F becomes more and more static, the depth measuring unit 130 is turned off again.
[0034] In this manner, depth data are only acquired, when the field of view F is directed to a new part of the observed scene 200, for which it is necessary to supplement the captured color image data with depth data. When the field of view does no longer move (or when the movement is sufficiently slow and / or small), newly acquired depth data would be redundant to the depth data that have already been acquired. Thus, depth data acquisition can be omitted in this case without deteriorating the eventually generated RGBD images. Accordingly, switching the depth measuring unit 130 on, if the position / orientation changes by more than a first predetermined amount, and switching the depth measuring unit 130 off again, if the changes become smaller than a second predetermined amount (that may, but is not necessarily equal to the first predetermined amount), reduces the power necessary for operating the depth measuring unit 130 and the power necessary for sampling the depth data with the image data. In addition, the control unit 140 may also switch on and off the image capturing unit 120 concurrently with the depth measuring unit 130 in order to further increase reduction of power consumption.
[0035] Figs. 2A and 2B show schematically different examples of a change of position and orientation of the sensor device 100 that might trigger depth data capturing. Fig. 2A shows a change in position and orientation that keeps the sensor device 100 focused on the same part of the scene 200, however, from a different viewing angle. This means that although the same object is captured before and after the change of position / orientation, different parts of the object are visible. Thus, although the observed object stays the same, the field of view has changed considerably. Accordingly, the depth measuring unit 130 will be switched on in order to provide at least depth data for the parts of the object that were not visible before the change of position / orientation.
[0036] Fig. 2B shows a change of position without a change of orientation. In such a case the provision of depth data is necessary, since a totally new part of the scene is observed. Here, it should be noted that also a shift of the position along the optical axis should trigger capturing of depth data, since by such a movement, parts of the scene 200 at the periphery of the field of view F could become visible. Further, although not illustrated, it is clear that also mere rotations of the sensor device 100, i.e. changes of the orientation without changes of the position will trigger depth data acquisition.
[0037] The control unit 140 may compare consecutively determined positions and / or orientations of the sensor device 100 in order to determine how much the position and / or orientation of the sensor device 100 has been changed, and the changes between consecutively determined positions and / or orientations are compared to the first predetermined amount and / or the second predetermined amount. This is schematically illustrated in Fig. 3.
[0038] Fig. 3 shows the amount of rotation / shift versus time. The amount of rotation and the amount of shift are provided in an iterative manner in accordance with the operation principle of integrating equations of movements based on the movement data provided from the inertial measurement unit 130. The control unit 140 determines a new position / orientation of the sensor device 100 with a specific update cycle. By comparing these positions / orientations, the amounts of shift / rotations can be obtained with the same update cycle. The amount of shift is for example quantified by the absolute value of the vector connecting consecutively determined positions of the sensor device 100 in space. The amount of rotation is for example quantified by the angle between consecutively determined pointing directions of the optical axis of the sensor device 100. Thresholds may be defined for the amount of shift and the amount of rotation separately. Both quantities might also be mathematically combined and compared to a single threshold.
[0039] Fig. 3 shows the first predetermined amount, i.e. a first threshold Tl, as a dash-dotted line, and the second predetermined amount, i.e. a second threshold T2, as a dashed line. As explained above, thresholds may be set for shift and rotation separately. Then, Fig. 3 has to be understood to refer to either the shift amount or the rotation amount. The size of the thresholds will depend on the specific usage of the sensor device 100. For usages where the scene 200 is located close to the sensor device 100 the threshold values will be smaller than for usages where the scene 200 is located farther away, since for a smaller distance between sensor device 100 and scene 200 smaller amounts of shifts and / or rotations will cause changes of the observed parts of the scene 200 that make capturing of depth data necessary.
[0040] As illustrated in Fig. 3 by region A, once the change in position / orientation becomes larger than the first predetermined amount / the first threshold Tl, depth data acquisition starts. Depth acquisition may be carried out every time the control unit 140 determines that the rotation / shift amount is larger than the first threshold Tl. Then, the frame rate of depth acquisition is equal to the update cycle of the position / orientation of the sensor device 100. However, the frame rate may in principle be set arbitrary during depth acquisition. It may for example also depend on the size of the determined rotations / shifts amounts and become higher the larger these amounts are.
[0041] Depth data acquisition is stopped again, once the changes drop below the second predetermined amount / the second threshold T2. Here, as illustrated in Fig. 3, it might be preferable that the second threshold T2 is smaller than the first threshold Tl in order to avoid a too early switch off of depth data (and image data) acquisition. However, the two thresholds may also be equal.
[0042] By using the above described technique, it is possible to restrict the operation time of the depth measuring unit 130 to times where the acquisition of depth data is necessary to generate RGBD images. During times where depth data acquisition becomes redundant, it is switched off. In this manner power consumption can be reduced.
[0043] The latency of data exchange between the inertial measurement unit 110 and the control unit 140 as well as the latency of data exchange between the control unit 140 and each of the image capturing unit 120 and the depth measuring unit 130 is smaller than a predetermined latency, the predetermined latency being preferably 2 ms, more preferably 1 ms, and even more preferably 0.5 ms. This means that the control unit 140 received the movement data very fast from the inertial measurement unit 110 and provides its control signals also very fast to the image capturing unit 120 and the depth measuring unit 130. In particular, the entire procedure from detection of movement data to activating / deactivating the depth measuring unit 130 may be effected within below 5 ms, below 2 ms, or preferably below 1 ms. This is preferably achieved by providing the inertial measurement unit 110, the image capturing unit 120, the depth measuring unit 130, and the control unit 140 on the same chip 150, preferably in the same integrate circuit, as schematically illustrated in Fig. 1. This ensures that communication paths can be specifically designed to realize the above discussed latency values. However, an integration of the inertial measurement unit 110, the image capturing unit 120, the depth measuring unit 130 and the control unit 140 is not necessary, if the low latency condition can be provided otherwise.
[0044] As illustrated in Fig. 4, the sensor device 100 may additionally comprise a processing unit 160 that is different from the control unit 140, and which is configured to receive the image data and the depth data in order to generate visible light images, depth maps, and RGBD images that combine the information of the visible light images and the depth maps. That is, RGBD images may be generated by the processing unit 160 from the image data and the depth data. In addition, the processing unit 160 may also separately generate the visible light image from the image data and the dept map from the depth data. The processing unit 160 may also first generate the visible light image and the depth map and generate the RGBD image by combining / synthesizing / fusing the visible light image and the depth map. The image data and the depth data may be converted into a format specific to the processing unit 160, if necessary.
[0045] Just as the control unit 140 the processing unit 160 may be a processor, like e.g. a CPU, a GPU, a FPGA or the like. The processing unit 160 may be any kind of hardware, software or combination of both that can carry out the functions of the processing unit 160 described herein. However, preferably, the processing unit 160 is an application processor of the sensor device 100.
[0046] In contrast to the control unit 140 that is provided specifically for fast control of the image capturing unit 120 and the depth measuring unit 130, the processing unit 160 is provided for more general and complex tasks, like e.g. the generation of the RGBD images. Further, the processing unit 160 may also receive the movement data and carry out a SUAM algorithms based on the movement data, the image data and the depth data. While generally provided with more computing power than the control unit 140, the processing unit 160 might in effect need longer to carry out the tasks executed by the control unit 140, since these tasks would need to be scheduled in order to allow the processing unit 160 to carry out all its tasks. The provision of the processing unit 160 in addition to the control unit 140 allows therefore low latency control of the inertial measurement unit 110, the image capturing unit 120, and the depth measuring unit 130 as well as execution of more complex tasks like RGBD image generation, executing SUAM algorithms, or the like.
[0047] Here, a latency of data exchange of communications between the processing unit 160 and the control unit 140 is larger than the predetermined latency, i.e. it takes more time to direct commands from the processing unit 160 to the control unit 140 than it takes the control unit 140 to control and / or communicate with the inertial measurement unit 110, the image capturing unit 120, and the depth measuring unit 130. Also, the provision of data from the inertial measurement unit 110, the image capturing unit 120 and the depth measuring unit 130 to the processing unit 160 may have a latency that is larger than the predetermined latency. This is illustrated in Fig. 4 by the thick arrows, which indicate high latency compared to the thin arrows representing low latency.
[0048] In particular, as also indicated schematically, the processing unit 160 is not part of the chip 150 on which the inertial measurement unit 110, the image capturing unit 120, the depth measuring unit 130, and the control unit 140 are located. This naturally increases the latency between the processing unit 160 and any component on the chip 150. However, it allows for more flexibility in designing the processing unit 160.
[0049] In the following it will be described how the control unit 140 and the processing unit 160 can interact with each other and the other components of the sensor device such as to further improve the latency and power consumption during the generation of RGBD images.
[0050] The depth measuring unit 130 may be configured to select a part P of the field of view F and to detect depth data only for the selected part P of the field of view F. This is schematically illustrated in Fig. 5, where the side margins at the right and at the bottom of the field of view F are selected as the part P. Such a selection is for example advantageous after a movement of the field of view F to the lower right, i.e. in a situation where the part P came new into the field of view F of the image capturing unit 120 and the depth measuring unit 130. Then, depth data is only acquired in the newly visible area of the field of view F, while for the already captured parts of the field of view F no redundant data are generated. Thus, depth data are only acquired if they provide additional information. This reduces power consumption further.
[0051] Of course, the selected part P may in principle have any location, shape, and size. Which part P of the field of view F to select will also depend on the circumstances. For example, besides providing depth data for parts of the scene 200 that newly came into the field of view F also already captured parts may be of further interest. For example, regions of interest may be defined within the field of view F, e.g. by the control unit 140, that indicate parts P to be selected for further, if necessary more thorough examination. For example, if the sensor device 100 is used for a face or gesture recognition process, parts of the field of view F that show faces or hands may be considered regions of interest and depth detection may not switched off in these parts at all, or may be switched on and off according to different conditions than for the rest of the field of view F.
[0052] The selection is performed according to the specific type of depth measuring unit 130 that is used. If e.g. a stereo camera is used, triangulation processing can be restricted to the selected part P, i.e. only for the selected part P the depth is determined from the stereo images. Then, it will also be sufficient to output only stereo image data pertaining to this selected part P. If a structured light system or a time of flight sensor is used as depth measuring unit 130, the emission of light can be restricted to the selected part P of the field of view F. Thus, light will only be reflected from objects within the selected part P of the field of view F. Accordingly, depth data will only be available for the selected part P.
[0053] The selection of the part P is preferably controlled by the control unit 140. This means the control unit 140 provides commands to the depth measuring unit 130 which indicate the part P of the field of view F for which depth data are needed. This guarantees due to the low latency of the communication between control unit 140 and the depth measuring unit 130 that the selection is performed quickly. In particular, when the selection becomes necessary due to changes of the position / orientation of the sensor device 100 that cause a shift of the field of view F, it is advantageous to quickly set a part P of the field of view F from which depth data are to be acquired. It the time between the change of position / orientation and the selection of the part P is too long, the selected part P does probably not fit to the field of view that is observed when the selection is applied. This may lead to a situation where the depths of regions of interest where not detected and / or where redundant depth data are gathered.
[0054] The control unit 140 may determine which part P to select by itself, e.g. based on the movement data from the inertial measurement unit 110. However, since this might bind computing resources of the control unit 140 that are better used for low latency control of the depth measuring unit 130, it is preferably that the processing unit 160 generates control information and stores them in a memory unit 141 of the control unit 140. The control unit 140 is then configured to control the selection of the part P of the field of view F based on the control information. For example, the processing unit 160 may provide a look up table to the control unit 140, in which it is indicated which part P is to be selected for which kind of movement data.
[0055] For example, if the changes of position / orientation that are deducible from the movement data indicate a shift of the field of view F into a particular direction, the control information may indicate to select a margin of the new field of view F located in this direction. Thus, in the example of Fig. 5, a shift of the field of view F towards the lower right leads to the selection of the part P lying at the right and the bottom of the field of view F, since new information can be obtained in this area. Similarly, a shift upwards leads to selection of the upper part, and so on. The width of the margin can here be determined by the speed of the shift, i.e. the amount of change per unit time.
[0056] Further, on a more advanced level, the control information indicates for a region S of the scene 200 that is larger than the field of view F the reliability of the depth map, and the control unit 140 is configured to control selection of the part P of the field of view F such that depth data of areas of the scene 200 are generated for which the reliability of the depth map is below a reliability threshold. The control information indicates therefore not in a general manner where in the field of view F depth data are to be acquired, but provides reliability values that indicate how reliable a depth map is that has already been generated, e.g. by the processing unit 160. Based on these reliability values the control unit 140 controls the selection of the parts P of the field of view F in which depth data are to be acquired.
[0057] Moreover, the reliability is not only measured in the field of view F that is currently observed by the sensor device 100, but for a region S of the scene 200 that is larger than the field of view F. This ensures that the control unit 140 can operate based on the control information, i.e. the reliability values, without the necessity to request new control information from the processing unit 160, even if the field of view F is shifted. Accordingly, due to the low latency connection between control unit 140 and depth measuring unit 130, the control unit 140 can quickly control the selection of the parts P for which depth data are to be acquired even if the field of view F changes. Only if considerably changes of the field of view F occur that lead the field of view out of the region S for which reliability values are provided, it is necessary to request new control information. Additionally or alternatively, the processing unit 160 may send new control information from time to time, e.g. periodically, to the control unit 140 as an update without necessitating an interruption of the control by the control unit 140.
[0058] In this manner it is possible to decouple the computational complex generation and evaluation of a depth map from the low latency control of the data acquisition without deteriorating the performance of the system.
[0059] A schematic illustration of the control information is given in Fig. 6. Here, two consecutive fields of views F and F’ are shown that each cover a part of the scene 200. In addition, the region S is shown that is larger than the fields of view F, F’. Due to the size of the region S, it is possible to change the field of view F within the region S, i.e. the same control information is provided for control with respect to the field of view F as well as with respect to the new field of view F’. The reliability of the depth map that has already been obtained is indicated in Fig. 6 as differently thick lines. Thin lines indicate high reliability, thicker lines indicate less reliability. Accordingly, based on a comparison with the reliability threshold, the control unit 140 can easily set those areas of the region S that are indicated to have reduced reliability, i.e. for which the probability that the depth map is not sufficiently correct or precise is high, as the part P of the new field of view F’ for which depth data are to be acquired.
[0060] Here, the control information is preferably a two-dimensional confidence map that indicates a confidence level for correctness of the depth map. An example of such a confidence map is highly schematically shown in Fig. 7, where different shades of grey indicate different confidence levels of the depth map, i.e. the probability that the respective parts of the depth map are correct. The confidence level can for example be generated by counting the number of points per unit volume in a point cloud constituting a three-dimensional depth map or model of the scene 200. The more points per unit volume are present, the higher is the confidence level. The confidence map is generated by projecting the three-dimensional depth map or model of the scene 200 along the line of sight of the sensor device 100 onto a two-dimensional plane such that the confidence map resembles what will be seen in the field of view of the image capturing unit 120 and the depth measuring unit 130. Since the generation of confidence maps is in principle well-known, a more detailed description can be omitted here. Based on the confidence map the control unit 140 can easily and quickly set the regions with low confidence as the regions where additional depth data are to be acquired.
[0061] The region S for which the reliability / confidence values are provided to the control unit 140 may be between 50% and 150%, preferably between 75% ad 125% larger than the field of view F. For example, when the field of view has a horizontal size of 80° and a vertical size of 60°, the region S may have a horizontal size of 140° and a vertical size of 120°. This ensures that smaller shifts of the field of view can be handled immediately by the control unit 140 and only shifts that lead out of the region S require new control information.
[0062] Of course, the above described look-up table based control, the region of interest based control, and the reliability value based control can be combined with each other. Further, also other manners of controlling the selection of the part P for which depth data are to be acquired could be implemented, as long as it is guaranteed that the control is carried out in a low latency manner by the control unit 140, while more complex control information is prepared outside the control unit 140 and provided to the control unit 140 upon request or from time to time without interrupting the control by the control unit 140.
[0063] The control unit 140 may further be configured to activate the depth measuring unit 130 and to control selection of the part P of field of view F if the position and / or orientation of the sensor device 100 has been changed by more than a third predetermined amount and by less than a fourth predetermined amount. This means, if the field of view of the sensor device 100 has changed sufficiently much, but not too much, depth acquisition is carried out, but only for parts P of the field of view F.
[0064] This is shown in Fig. 8 that shows rotation / shift amounts versus time in the same manner as Fig. 3. Once the changes in position / orientation become larger than the third predetermined amount, represented by a third threshold T3, the control unit 140 switches the depth measuring unit 130 on and operates it in the partial depth sensing mode (region B in Fig. 8). Here, the frame rate of partial depth measuring may be equal to the update cycle of the position / orientation determination or may be set based on the rotation / shift amount. In Fig. 8, it is indicated that the third threshold may be equal to the first threshold of Fig. 3. However, the two quantities may also be unrelated.
[0065] Once the position and / or orientation of the sensor device 100 changes by more than the fourth predetermined amount, the control unit 140 is configured to request new control information from the processing unit 160 and to control the depth measuring unit 130 to detect depth data for the entire field of view F. That is, once the change of the field of view becomes so rapid that the control information provided by the processing unit 160 must be considered insufficient for a control of partial depth acquisition in the new field of view F’, the control unit 140 requests new control information from the processing unit 160. Also, the control unit 140 controls the depth measuring unit 130 to capture depth data for the full field of view F’, since it has to be expected that no depth data for this new field of view F’ is present. The frame rate of the full depth measuring may be equal to the update cycle of the position / orientation determination or may be set based on the rotation / shift amount.
[0066] This is also shown in Fig. 8, where after the transgression of a fourth threshold T4 representing the fourth predetermined amount of change the full depth acquisition mode is entered (region A). Here, it should be noted that also the fourth threshold T4 might be equal to the first threshold T1 of Fig. 3. Then, just as discussed with respect to Fig. 3, full depth acquisition is executed for shifts / rotations by more than the first threshold T1 (= the fourth threshold T4). For smaller shifts / rotations that are, however, larger than the third threshold T3, the partial depth sensing mode is performed.
[0067] When the change rate drops again below the fourth predetermined amount / the fourth threshold T4, the partial depth sensing mode is entered again (region C). Depth acquisition is ended, once the change rate drops below the second predetermined amount / the second threshold T2 that is indicated to be lower than the other thresholds in Fig. 8. However, the second threshold T2 might also be equal to the third threshold T3. Then ,the same threshold controls initiation and ending of the partial depth sensing mode.
[0068] In this manner it is possible to select a mode for acquiring depth data that is appropriate for the movements detected by the inertial measurement unit 110. If only little movements are detected, energy and processing power is saved by switching the depth measuring unit 130 off, i.e. by not carrying out depth measurements. In an intermediate range of movement only parts of the observed field of view are provided with depth data, thereby saving energy and processing power compared to a full operation. Only if the movements become too rapid, the entire field of view is sampled for depth data in order to maintain the possibility to provide depth maps / RGBD images with high reliability.
[0069] As discussed above, the depth measuring device 130 may be a direct time of flight sensor that is configured to determine for each of a plurality of light rays emitted into the field of view F the time until a reflection of the respective light ray from the scene 200 is received. A schematic configuration of an according sensor device 100 is shown in Fig. 9. The direct time of flight sensor comprises a light emitting unit 131, e.g. LED or laser diode array, or a VCSEL array that is configured to emit light pulses into the field of view F. The direct time of flight sensor comprises further a single photon avalanche diode, SPAD, array 132, in which a plurality of SPAD pixels are arranged to detect the reflection of the light emitted by the light emitting unit 131. In the example of Fig. 9 the pixel array 125 and the SPAD array 132 are stacked onto each other and receive light through the imaging optics 121. Preferably, the light emitting unit 131 emits infrared light and a filter 133 that passes only infrared light is provided between the pixel array 125 and the SPAD array 132. In this manner it is ensured that fields of view F of the direct time of flight sensor and the image capturing unit 120 are identical. The control of the selection of the parts P for which depth data are to be acquired is carried out by controlling the light emitting unit 131 to emit light only into the respective parts of the field of view. This provides a sensor device 100 that is capable to carry out the above described functions.
[0070] As illustrated in Fig. 10, the sensor device 100 as described above may be part of a head mounted display 300 that is worn on the head of a user 400 and that comprises a display 310 in front of the eyes of the user 400 that is configured to display virtual, augmented and / or mixed reality contents based on the movement data, the image data, and / or the depth data. In particular, these data may be used by the head mounted display 300 to carry out a SLAM algorithm based on which functions of the application / software executed by the head mounted display 300 are controlled. By using the sensor device 100 as described above, the latency between movements of the user 400 and execution of the controls to be triggered by said movements can be reduced. Moreover, the power consumption can be reduced by switching off / deactivating the depth measuring unit 130 as described above. In this manner a portable head mounted display 300 can be provided whose operation time without recharging is extended and which has a low latency between user movements and an execution of the intended controls.
[0071] The above described method for operating the sensor device 100 (included in the head mounted display 300) is summarized again in Fig. 11. At SI 10 movement data indicating accelerations and rotations of the sensor device 100 are detected by the inertial measurement unit 110. At S120 image data for a visible light image of the scene 200 are captured by the image capturing unit 120 for a field of view F. At S130 depth data for a depth map of the scene 200 are captured by the image capturing unit for the same field of view F.
[0072] At S140 the control unit 140 determines a position and / or orientation of the sensor device 100 in space based on the movement data of the inertial measurement unit 110 to control activation of the image capturing unit 120 and / or the depth measuring unit (130). At SI 50 the control unit 140 activates the depth measuring unit 130 if the position and / or orientation of the sensor device 100 that is determined by the control unit 140 has changed by more than a first predetermined amount. At SI 60 the control unit 140 deactivates the depth measuring unit 130 if the position and / or orientation of the sensor device 100 that is determined by the control unit 140 has changed by less than a second predetermined amount. In all these processes the latency of data exchange between the inertial measurement unit 110 and the control unit 140 as well as the latency of data exchange between the control unit 140 and each of the image capturing unit 120 and the depth measuring unit 130 is smaller than a predetermined latency, the predetermined latency being preferably 2 ms, more preferably 1 ms, and more preferably 0.5 ms.
[0073] In this manner it is possible to produce the data necessary for generation of an RGBD image with low latency and with reduced power consumption.
[0074] The sensor device 100 and the method for operating it can be used in any technical area that is in need of generation of RGBD images. Some examples for uses are indicated in Figs. 12A and 12B.
[0075] For example, as shown in Fig. 12A sensor device 100 may be integrated into smart glasses to allow generation of RGBD images during an inspection process, e.g. during a quality check, during a material detection process, during a security check at an airport, or the like. By using the RGBD images image classification may be eased, since the additional provision of depth eases segmentation of images.
[0076] Another possible application is the integration of the sensor device 100 into a mobile terminal as shown in Fig. 12B to enhance image classification at the mobile terminal or to improve SLAM algorithms caried out thereon. For example, the mobile terminal might be used in computer games or augmented / virtual reality applications.
[0077] In general, the sensor device 100 may be applied to any field of real -world image classification or SLAM processing as implemented in various technical areas.
[0078] The present technology can also be configured as described below:
[0079] [1] A sensor device (100) for generating RGBD images of a scene (200), the sensor device (100) comprising: an inertial measurement unit (110) that is configured to detect movement data indicating accelerations and rotations of the sensor device (100); an image capturing unit (120) that is configured to capture for a field of view (F) image data for a visible light image of the scene (200); a depth measuring unit (130) that is configured to detect for the same field of view (F) depth data for a depth map of the scene (200); and a control unit (140) that is configured to determine a position and / or orientation of the sensor device (100) in space based on the movement data of the inertial measurement unit (110), and to control activation of the image capturing unit (120) and / or the depth measuring unit (130); wherein the latency of data exchange between the inertial measurement unit (110) and the control unit (140) as well as the latency of data exchange between the control unit (140) and each of the image capturing unit (120) and the depth measuring unit (130) is smaller than a predetermined latency, the predetermined latency being preferably 2 ms, more preferably 1 ms, more preferably 0.5 ms; and the control unit (140) is configured to activate the depth measuring unit (130) if the position and / or orientation of the sensor device (100) that is determined by the control unit (140) has changed by more than a first predetermined amount and is configured to deactivate the depth measuring unit (130) if the position and / or orientation of the sensor device (100) that is determined by the control unit (140) has changed by less than a second predetermined amount.
[0080] [2] The sensor device (100) according to [1], wherein the control unit (140) is configured to compare consecutively determined positions and / or orientations of the sensor device (100) in order to determine how much the position and / or orientation of the sensor device (100) has been changed; and the changes between consecutively determined positions and / or orientations are compared to the first predetermined amount and / or the second predetermined amount.
[0081] [3] The sensor device (100) according to [1] or [2], wherein the inertial measurement unit (110), the image capturing unit (120), the depth measuring unit (130), and the control unit (140) are located on the same chip (150), preferably in the same integrate circuit, in order to provide the latency below the predetermined latency.
[0082] [4] The sensor device (100) according to any one of [1] to [3], further comprising a processing unit (160) that is different from the control unit (140); wherein the processing unit (160) is configured to receive the image data and the depth data in order to generate visible light images, depth maps, and RGBD images that combine the information of the visible light images and the depth maps.
[0083] [5] The sensor device (100) according to [4], wherein the processing unit (160) is configured to receive the movement data, the image data and the depth data in order to execute a simultaneous localization and mapping, SLAM, algorithms based on them.
[0084] [6] The sensor device (100) according to [4] or [5], wherein the processing unit (160) is configured to communicate with the control unit (140); and the latency of data exchange between the processing unit (160) and the control unit (140) is larger than the predetermined latency.
[0085] [7] The sensor device (100) according to any one of [4] to [6], wherein the depth measuring unit (130) is configured to select a part (P) of the field of view (F) and to detect depth data only for the selected part (P) of the field of view (F).
[0086] [8] The sensor device (100) according to [7], wherein the control unit (140) is configured to control the selection of the part (P) of the field of view (F) by the depth measuring unit (130).
[0087] [9] The sensor device (100) according to [8], wherein the control unit (140) comprises a memory unit (141); the processing unit (160) is configured to store control information in the memory unit (141) of the control unit; and the control unit (140) is configured to control the selection of the part (P) of the field of view (F) based on the control information.
[0088]
[0010] The sensor device (100) according to [9], wherein the control information indicates for a region (S) of the scene (200) that is larger than the field of view (F) the reliability of the depth map; and the control unit (140) is configured to control selection of the part (P) of the field of view (F) such that depth data of areas of the scene (200) are generated for which the reliability of the depth map is below a reliability threshold.
[0089]
[0011] The sensor device (100) according to [9] or
[0010] , wherein the control information is a two-dimensional confidence map that indicates a confidence level for correctness of the depth map.
[0090]
[0012] The sensor device (100) according to any one of [9] to
[0011] , wherein the control unit (140) is configured to activate the depth measuring unit (130) and to control selection of the part (P) of field of view (F) if the position and / or orientation of the sensor device (100) has been changed by more than a third predetermined amount and by less than a fourth predetermined amount; and the control unit (140) is configured to request new control information from the processing unit (160) and to control the depth measuring unit (130) to detect depth data for the entire field of view (F) if the position and / or orientation of the sensor device (100) has been changed by more than the fourth predetermined amount.
[0091]
[0013] The sensor device (100) according to any one of [1] to
[0012] , wherein the depth measuring device (130) is a direct time of flight sensor that is configured to determine for each of a plurality of light rays emitted into the field of view (F) the time until a reflection of the respective light ray from the scene (200) is received.
[0092]
[0014] A head mounted display (300) comprising: the sensor device (100) according to any one of [1] to
[0013] ; and a display (310) configured to display virtual, augmented and / or mixed reality contents based on the movement data, the image data, and the depth data.
[0093]
[0015] A method for operating the sensor device (100) according to any one of [1] to
[0013] or the head mounted display (300) according to
[0014] to generate RGBD images of a scene (200), the method comprising: by the inertial measurement unit (110), detecting movement data indicating accelerations and rotations of the sensor device (100); by the image capturing unit (120), capturing for a field of view (F) image data for a visible light image of the scene (200); by the depth measuring unit (130), detecting for the same field of view (F) depth data for a depth map of the scene (200); by the control unit (140), determining a position and / or orientation of the sensor device (100) in space based on the movement data of the inertial measurement unit (110), and controlling activation of the image capturing unit (120) and / or the depth measuring unit (130); by the control unit (140), activating the depth measuring unit (130) if the position and / or orientation of the sensor device (100) that is determined by the control unit (140) has changed by more than a first predetermined amount and deactivating the depth measuring unit (130) if the position and / or orientation of the sensor device (100) that is determined by the control unit (140) has changed by less than a second predetermined amount; wherein the latency of data exchange between the inertial measurement unit (110) and the control unit (140) as well as the latency of data exchange between the control unit (140) and each of the image capturing unit (120) and the depth measuring unit (130) is smaller than a predetermined latency, the predetermined latency being preferably 2 ms, more preferably 1 ms, and more preferably 0.5 ms.
Claims
Claims1. A sensor device for generating RGBD images of a scene, the sensor device comprising: an inertial measurement unit that is configured to detect movement data indicating accelerations and rotations of the sensor device; an image capturing unit that is configured to capture for a field of view image data for a visible light image of the scene; a depth measuring unit that is configured to detect for the same field of view depth data for a depth map of the scene; and a control unit that is configured to determine a position and / or orientation of the sensor device in space based on the movement data of the inertial measurement unit, and to control activation of the image capturing unit and / or the depth measuring unit; wherein the latency of data exchange between the inertial measurement unit and the control unit as well as the latency of data exchange between the control unit and each of the image capturing unit and the depth measuring unit is smaller than a predetermined latency, the predetermined latency being preferably 2 ms, more preferably 1 ms, and more preferably 0.5 ms; and the control unit is configured to activate the depth measuring unit if the position and / or orientation of the sensor device that is determined by the control unit has changed by more than a first predetermined amount and is configured to deactivate the depth measuring unit if the position and / or orientation of the sensor device that is determined by the control unit has changed by less than a second predetermined amount.
2. The sensor device according to claim 1, wherein the control unit is configured to compare consecutively determined positions and / or orientations of the sensor device in order to determine how much the position and / or orientation of the sensor device has been changed; and the changes between consecutively determined positions and / or orientations are compared to the first predetermined amount and / or the second predetermined amount.
3. The sensor device according to claim 1, wherein the inertial measurement unit, the image capturing unit, the depth measuring unit, and the control unit are located on the same chip, preferably in the same integrate circuit, in order to provide the latency below the predetermined latency.
4. The sensor device according to claim 1, further comprising a processing unit that is different from the control unit; wherein the processing unit is configured to receive the image data and the depth data in order to generate visible light images, depth maps, and RGBD images that combine the information of the visible light images and the depth maps.
5. The sensor device according to claim 4, whereinthe processing unit is configured to receive the movement data, the image data and the depth data in order to execute a simultaneous localization and mapping, SLAM, algorithms based on them.
6. The sensor device according to claim 4, wherein the processing unit is configured to communicate with the control unit; and the latency of data exchange between the processing unit and the control unit is larger than the predetermined latency.
7. The sensor device according to claim 6, wherein the depth measuring unit is configured to select a part of the field of view and to detect depth data only for the selected part of the field of view.
8. The sensor device according to claim 7, wherein the control unit is configured to control the selection of the part of the field of view by the depth measuring unit.
9. The sensor device according to claim 8, wherein the control unit comprises a memory unit; the processing unit is configured to store control information in the memory unit of the control unit; and the control unit is configured to control the selection of the part of the field of view based on the control information.
10. The sensor device according to claim 9, wherein the control information indicates for a region of the scene that is larger than the field of view the reliability of the depth map; and the control unit is configured to control selection of the part of the field of view such that depth data of areas of the scene are generated for which the reliability of the depth map is below a reliability threshold.
11. The sensor device according to claim 9, wherein the control information is a two-dimensional confidence map that indicates a confidence level for correctness of the depth map.
12. The sensor device according to claim 9, wherein the control unit is configured to activate the depth measuring unit and to control selection of the part of field of view if the position and / or orientation of the sensor device has been changed by more than a third predetermined amount and by less than a fourth predetermined amount; and the control unit is configured to request new control information from the processing unit and to control the depth measuring unit to detect depth data for the entire field of view if the position and / or orientation of the sensor device has been changed by more than the fourth predetermined amount.
13. The sensor device according to claim 1, wherein the depth measuring device is a direct time of flight sensor that is configured to determine for each of a plurality of light rays emitted into the field of view the time until a reflection of the respective light ray from the scene is received.
14. A head mounted display comprising: the sensor device according to claim 1 ; and a display configured to display virtual, augmented and / or mixed reality contents based on the movement data, the image data, and the depth data.
15. A method for operating the sensor device according to claim 1 to generate RGBD images of a scene, the method comprising: by the inertial measurement unit, detecting movement data indicating accelerations and rotations of the sensor device; by the image capturing unit, capturing for a field of view image data for a visible light image of the scene; by the depth measuring unit, detecting for the same field of view depth data for a depth map of the scene; by the control unit, determining a position and / or orientation of the sensor device in space based on the movement data of the inertial measurement unit, and controlling activation of the image capturing unit and / or the depth measuring unit; by the control unit, activating the depth measuring unit if the position and / or orientation of the sensor device that is determined by the control unit has changed by more than a first predetermined amount and deactivating the depth measuring unit if the position and / or orientation of the sensor device that is determined by the control unit has changed by less than a second predetermined amount; wherein the latency of data exchange between the inertial measurement unit and the control unit as well as the latency of data exchange between the control unit and each of the image capturing unit and the depth measuring unit is smaller than a predetermined latency, the predetermined latency being preferably 2 ms, and more preferably 1 ms, more preferably 0.5 ms.
Citation Information
Patent Citations
Environmental Model Maintenance Using Event-Based Vision Sensors
US20220060622A1
Methods and systems for providing depth maps with confidence estimates
US20220111839A1