Electronic device, method and computer program
Patent Information
- Application Number
- EP2024794834
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-30
- Filing Date
- 2024-10-28
- Publication Date
- 2026-09-09
AI Technical Summary
Existing electronic devices face challenges in autofocusing due to issues like object occlusion, rotation, and changes in lighting conditions, which complicate the detection and tracking of objects in 2D image data.
An electronic device is configured to obtain depth image data from a Time-of-Flight (ToF) sensor and visual image data from a visual image sensor, detect objects in the depth data, select the object based on user input or autonomous criteria, track the object's spatial position, and perform autofocusing on the selected object.
This solution enables accurate and efficient autofocusing, even when objects are occluded, rotating, or experiencing changes in lighting, by utilizing depth data to track the object's spatial position and adjust the lens system accordingly.
Smart Images

Figure EP2024080365_08052025_PF_FP_ABST
Abstract
Description
[0001] ELECTRONIC DEVICE, METHOD AND COMPUTER PROGRAM
[0002] TECHNICAL FIELD
[0003] The present disclosure generally pertains to an electronic device, a method and a computer program.
[0004] TECHNICAL BACKGROUND
[0005] Known techniques for autofocusing of electronic devices are based on detection of objects in 2D image data. Once objects are detected, their movements are tracked frame by frame. Once an object is detected, the algorithm tries to predict where it could be in the next frame, usually based on its current measured velocity, and then tries to detect the same object in the next frame somewhere in the general area of the predicted new location. This poses a lot of challenges. The object may get occluded, it may rotate, it may go from sun to shadow which will change its color. Deciding whether the object we see in a new frame is the same one from the last one or any preceding one is a complex problem.
[0006] SUMMARY
[0007] According to a first aspect the disclosure provides an electronic device in accordance with independent claim 1. According to a second aspect the disclosure provides a method in accordance with independent claim 28. According to a third aspect the disclosure provides a computer program in accordance with independent claim 29.
[0008] Further aspects are set forth in the dependent claims, the drawings and the following description.
[0009] BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Embodiments are explained by way of example with respect to the accompanying drawings, in which:
[0011] Fig. 1 schematically illustrates an autofocusing method; and
[0012] Fig. 2 schematically illustrates a method for obtaining an object selection based on user input; and
[0013] Fig. 3 schematically illustrates a method for obtaining an object selection based on predetermined selection criteria; and
[0014] Fig. 4 schematically illustrates a method for presenting visual image data to a user in order to obtain user input; and Fig. 5 schematically illustrates a spatial position and spatial movement; and
[0015] Fig. 6 schematically illustrates visual image data that may be presented to the user to obtain user input; and
[0016] Fig. 7 (A) schematically illustrates a first example of a first embodiment of visual image data and input fields that may be presented to the user; and
[0017] Fig. 7 (B) schematically illustrates a second example of a first embodiment of visual image data and input fields that may be presented to the user; and
[0018] Fig. 7 (C) schematically illustrates a third example of a first embodiment of visual image data and input fields that may be presented to the user; and
[0019] Fig. 8 (A) schematically illustrates a first example of a second embodiment of visual image data and input fields that may be presented to the user; and
[0020] Fig. 8 (B) schematically illustrates a second example of a second embodiment of visual image data and input fields that may be presented to the user; and
[0021] Fig. 8 (C) schematically illustrates a third example of a second embodiment of visual image data and input fields that may be presented to the user; and
[0022] Fig. 9 schematically illustrates fields of view of a visual sensor and a ToF sensor; and
[0023] Fig. 10 schematically illustrates fields of view of a combined RGB-ToF sensor; and
[0024] Fig. 11 schematically illustrates movement of a depth detected object towards the field of view of the visual image sensor; and
[0025] Fig. 12 schematically illustrates movement of a depth detected object away from the field of view of the visual image sensor; and
[0026] Fig. 13 (A) illustrates an example of a visual image that illustrates a scene on which autofocus may be performed; and
[0027] Fig. 13 (B) illustrates an example of a point cloud and a point cloud segment; and
[0028] Fig. 13 (C) illustrates an example of a spatial position being calculated based on a point cloud segment; and
[0029] Fig. 14 is a schematic illustration of the electronic device configured to perform the disclosed methods. DETAILED DESCRIPTION OF EMBODIMENTS
[0030] Before a detailed description of the embodiments under reference of Fig. 1 is given, general explanations are made.
[0031] The present disclosure provides an electronic device comprising circuitry configured to obtain depth image data from a ToF-sensor and visual image data from an visual image sensor; detect an object in the depth image data; select the depth detected object based on user input or autonomously based on a predetermined selection criterion; track a spatial position of the selected depth detected object, based on the depth image data; and perform autofocusing of a lens system on the spatial position of the selected depth detected object.
[0032] Circuitry may be a mobile terminal a computer and / or a processor or another electronic data processing apparatus. In some embodiments, also a non-transitory computer-readable recording medium is provided that stores therein a computer program product, which, when executed by a processor, such as the processor described above, causes the method described to be performed.
[0033] The electronic device may be an imaging system. The electronic device may be a camera, or an imaging system included in a camera. The electronic device may be a camera included in an electronic apparatus. The electronic apparatus may be a mobile phone, a smartphone, a laptop PC, a desktop PC, a video camera, a head mounted display, a virtual reality device, an augmented reality device or an imaging device included in another mechanism, such as a vehicle. The electronic device may also be seen as the electronic apparatus.
[0034] The visual image data is image data that describes an image in a wavelength range corresponding to the average human visual range. The visual image sensor is an image sensor sensitive to the wavelength range corresponding to the average human visual range. The visual image data may describe an image in greyscale, or in color. Color image data may be RGB image data. The visual image may be called a 2D image. Hereinbelow, presentation of visual image data should be understood to mean presentation of the image described by the visual image data.
[0035] The RGB image data is the data describing an RGB image. The RGB image being a composite image of spectral components corresponding to red (R), green (G) and blue (B), as can be provided by a Bayer filter arrangement.
[0036] The lens system may be autofocused by adjusting the lens system, in a case where the lens system comprises a single lens, to position the lens at a suitable lens position. The suitable lens position can be calculated based on the focus distance according to known visual relations. The suitable lens position can, for example, be calculated according to known principles of optics, for example using the known equation (1)
[0037] Wherein L is the focal length of the lens, u is the focus distance and v is the distance of the lens 50 to the photosensitive surface of the visual image sensor. The lens position can, for example, be based on the quantity v. The parameter L is a known quantity related to the structure and type of the lens. The parameter u can be determined based on the focus distance. The equation (1) is, in this example, solved for the parameter v.
[0038] Depth information, which may be called depth image data, may be provided using known Time- of-Flight techniques. Such known techniques include the direct time-of-flight (dToF) and the indirect time-of-flight (dToF) techniques. Either of dToF or iToF may be used in the device according to the present disclosure. In the following, whenever the phrase “ToF” is used, the possible use of either of dToF or iToF is implicit.
[0039] In sensors working according to the iToF -principle, for example, depth information is obtained by determining a phase angle c|) of incident light in pixels of the ToF sensor as compared to a modulation signal. The modulation signal is provided by an emitter, for example a laser, that emits a signal to be reflected from a reflection surface, such as an IR laser beam. Each pixel of a ToF sensor may typically include two distinct taps (e.g. tap A and tap B), each tap providing gain information (e.g. GA for tap A) in a 7t / 2 phase offset as well as phase-dependent intensity information (e.g. S(0) for the a phase of zero and S(K / 2) for a phase of 7t / 2). Using the gain information and the phase-dependent intensity, variables I and Q can be calculated using the relations
[0040] I = 2(GA- GA)S(0) (2) and
[0041] Q = 2(GA- GA)S(TT / 2) . (3)
[0042] The phase angle c|) for each pixel comprising the ToF sensor can then be calculated using
[0043] 4> = —atan2 Q, ). (4)
[0044] Executing this calculation for each pixel comprising the ToF sensor yields a phase image associating each pixel with phase information. Depth information is calculated from the phase information by correlating the phase information with phase information of emitted laser beams, such as a vertical cavity surface-emitting laser (VCSEL) array or an infrared laser (IR), generating a plurality of points. The correlation of the phase information detected by the pixel with phase information of emitted IR laser beam yields a phase offset A0 between the emitted laser beams and the incident radiation sensed by the pixel. A distance D of the pixel to the surface reflecting the emitted laser beams can then be calculated using where c is the speed of light in an atmosphere and f is a frequency of the emitted light.
[0045] For example, if the emitted light has a frequency of =20 MHz and the calculated phase offset A0 is 22°, the distance to the surface is about 0.45m. The distance is depth information. If a plurality of points is generated, depth information is generated for each point. Correspondingly, the distance to each point is known.
[0046] Thus, by correlating the phase angle of the incident radiation with the phase angle of the point pattern, the distance of the pixel to the point can be calculated.
[0047] Furthermore, by determining the distribution of pixels detecting points, a spatial distribution of reflection locations of the points can be calculated.
[0048] In sensors working according to the dToF -principle, a photon counting (PC) technique pulse method is used.
[0049] Photon counting ToF systems like the dToF system record a photon histogram. DToF systems may, for example, use single photon avalanche diodes (SPADs) as detectors. In dToF, depth information is obtained, for example, by measuring the time a signal pulse of a defined duration, emitted by the device using dToF, takes to reach an object, be reflected on a reflection surface, and return to the device to be sensed by the dToF sensor. The signals are then correlated to obtain the distance.
[0050] A pixel of a dToF-type sensor may comprise two distinct switches SA and SB and two memory elements MA and MB. The switches SA and SB route an incident signal, which is the reflection of the emitted signal, received by the pixel to the corresponding memory element MA and MB. The switches are controlled by a control signal cadenced to coincide with the emitted signal pulse. During exposure, the incident signal is first routed to the first memory element, for example MA. After a time of the length of the emitted signal pulse has elapsed, the incident signal is routed to the second first memory element, for example MB. The ratio of a number of photons of the incident signal routed to either MA or MB is related to the length of time it took the signal to travel to the reflection surface and back. The number of photons define an intensity, with, for example, IA or IB corresponding to the signals routed to either MA or MB.
[0051] The distance of the reflection surface can then be calculated where c is the speed of light in an atmosphere and to is the duration of the emitted pulse.
[0052] For example, if the emitted pulse has a duration of 25 ns and the ratio of intensities is IA / IB = 2 / 3, then the distance to the surface is about 3.75 m. The distance is depth information.
[0053] Calculating a distance to an object requires a choice of which points detected by the ToF sensor to use to calculate the depth information. A description of the choice of which points detected by the ToF sensor to use to calculate the depth information is presented hereinbelow.
[0054] Hereinbelow, ToF image data is to be understood to comprise depth image data. Whenever, based on ToF image data, distance information is to be obtained, the depth image data included in the ToF image data may be used.
[0055] The ToF image data describes a spatial position of an object. The ToF image data may therefore be called 3D image data.
[0056] Is should be noted that acquisition of depth image data from a ToF sensor using a ToF method is only one possible option. Other sources may be used according to the present disclosure. For example, a depth estimation using Dual-Pixels may be used. If alternative methods are used, the term “depth image data” and “3D image data” refers to the data provided by the alternative method.
[0057] Detecting a depth detected object comprises detecting, in the depth image data, depth image data that is associated with an object. For example, detecting an object may comprise detecting IR , infrared, light reflected from an object and determining the depth image data associated with this reflected light. The object that the light is reflected by is the depth detected object. A plurality of depth detected objects may be detected at the same time.
[0058] Furthermore, detecting the depth detected object may comprise determining a position of the depth detected object with respect to a coordinate system. The coordinate system may by a coordinate system relative to the device. The coordinate system may, alternatively or in addition, be a coordinate system relative to a space that the electronic device is located in. The coordinate system may further be a coordinate system relative to another object in the scene, a specific marker, such as one or more fiducial markers, in the scene with a known specific location or coordinates relative to the marker.
[0059] Selecting the depth detected object comprises determining the depth detected object in such a way that it is distinguished from other depth detected objects. Selecting the depth detected object does not require the depth detected object to be identified.
[0060] Information may be provided that designates, upon selection, the selected depth detected object as the selected depth detected object. The designation may, for example, be accomplished by determining a position of the selected depth detected object.
[0061] Tracking the position of the selected depth detected object requires that, at least in a plurality of frames of ToF image data captured by the ToF sensor, the position of the selected depth detected object determined. The determination may, when the selected depth detected object is tracked, be carried over to later frames without the depth detected object being required to be selected again. A later frame is a frame captured at a later time. A later frame is not required to be a consecutive frame. There may be an interval of time or an interval of frames between a current frame and the later frame.
[0062] Performing autofocusing on the tracked depth detected object may be called continuously performing autofocusing on the tracked depth detected object.
[0063] The tracking, autofocusing and obtaining visual image data and ToF image data may be provided in a frame-by-frame manner, which can also be called continuous.
[0064] The tracking, autofocusing and obtaining visual image data and ToF image data may, alternatively, be provided for some, but not all, of the consecutive frames. For example, tracking, autofocusing may be provided for every tenth frame. Furthermore, quality criteria may be applied to select individual frames for which autofocusing is provided. Autofocusing and tracking may also be called continuous if provided only for some, but not all, consecutive frames in this manner.
[0065] Continuous autofocusing may be accomplished by refreshing the autofocus with each newly captured image by either the visual image sensor or the ToF sensor or both the visual image sensor and the ToF sensor.
[0066] The autofocusing may be applied as long as the selected depth detected object remains in a field of view (FOV) of either the visual image sensor or the ToF sensor or both the visual image sensor and the ToF sensor. A picture mode, which is a mode to capture individual pictures, and a video mode, which is a mode to capture video, may be provided. The picture mode or the video mode may be used with the disclosure as desired.
[0067] The lens system may comprise one lens or plurality of lenses. The lens system is provided to focus incident visible light on a photosensitive surface of the visual image sensor. The visual image sensor generates, based on the incident light, visual image data. The visual image data may be an image. By focusing incident light on the photosensitive surface of the visual image sensor, objects that are in focus of the lens will produce a sharp image on the visual image sensor, such that a high quality image of the objects that are in focus of the lens can be produced.
[0068] The lens system may be called an visual assembly. The lens system may, instead of lenses, comprise reflective surfaces, such as mirrors and / or curved mirrors that also fulfill the task of focusing incident visible light on a photosensitive surface of the visual image sensor.
[0069] The lens system is provided to be adjustable. Autofocusing is performed by adjusting the lens system. The lens system may comprise actuators which, upon receiving an input signal, may perform adjusting of the lens system.
[0070] Adjusting the lens system may comprise moving any or all of the lenses included in the lens system with respect to the image sensor. Adjusting the lens system may also comprise moving the image sensor with respect to any or all of the lenses included in the lens system.
[0071] By performing autofocusing on the tracked depth detected object, the tracked depth detected object is kept in focus of the visual image sensor through the lens system.
[0072] Some embodiments further comprise the ToF sensor, the visual image sensor and the lens system, the visual image sensor receiving light through the lens system, wherein an FOV of the ToF-sensor overlaps an FOV of the visual image sensor.
[0073] The ToF sensor and the visual image sensor may be provided in the electronic device. The ToF sensor and the visual image sensor are disposed such that the FOV of the ToF sensor overlaps the FOV of the visual image sensor.
[0074] The FOV is a volume such that any object located in the volume and that has a clear line-of-sight to an image sensor, such as the visual image sensor or the ToF sensor, can be detected by the image sensor. The FOV of the ToF sensor overlaps the FOV of the visual image sensor such that at least some objects that have a clear line-of-sight to the visual image sensor can be detected by the ToF sensor. Such a disposition may be provided, for example, by the ToF sensor and the visual image sensor being provided in a codirectional disposition.
[0075] According to the present disclosure, a depth to lens position model is applied continuously in a live view and in conjunction with live object tracking. If an object for focusing is selected, and that object moves with respect to the electronic device, or the electronic device is moved, as long as the selected object is successfully tracked, it will stay in focus, until tracking is disabled or another object is selected.
[0076] Tracking may be disabled by a user. Another object may be selected by the user.
[0077] Alternatively, tracking may be disabled autonomously, or another object may be selected autonomously.
[0078] In this way a user experience may be improved.
[0079] In some embodiments the FOV of the ToF-sensor exceeds the FOV of the visual image sensor. An FOV subtends a solid angle that delimits the volume of the FOV. The FOV of the ToF sensor exceeds the FOV of the visual image sensor of the solid angle subtended by the FOV of the ToF sensor is larger than the solid angle subtended by the FOV of the visual image sensor. This way, information provided by the ToF sensor that is not available to the visual image sensor can be used in performing autofocusing.
[0080] In some embodiments, detecting the object comprises segmenting a point cloud obtained from the depth image data.
[0081] A point cloud is a plurality of points detected in the depth image data. The depth detected object corresponds to an isolated structure of points in the point cloud. The isolated structure of points in the point cloud may be called a point cloud segment. Thus, the depth detected object corresponds to a point cloud segment.
[0082] The segmented point cloud is segmented in order to produce a point cloud segment. The point cloud segment is determined by determining that a plurality of points in the depth image data is associated with each other. Determining that the plurality of points is associated with each other can be accomplished using a suitable selection algorithm or a clustering algorithm. Clustering algorithms can detect clusters based on simple properties like density or matching shapes. For example, known algorithms that may be used are RANSAC, which is able to detect e.g. planes (wall, floor table,..) and DBSCAN, which is able to detect clusters of points with similar density. (Basically all other non-planar objects.) In some embodiments, an identification of objects may be provided. In these cases, semantic ML based algorithms may be used. Semantic ML based algorithms are neural networks that try understand what the points are based on training processes using vast amounts of known objects. This is similar to an ML based approach in 2D, but instead of images, it works on 3D point clouds. A known algorithm is VoxNet.
[0083] In some embodiments, detecting the depth detected object comprises determining a spatial position of the depth detected object.
[0084] The spatial position of the depth detected object may comprise a range of the depth detected object to the electronic device. The range may be a distance. The spatial position of the depth detected object may instead comprise a coordinate that allows the range or the distance to be calculated. The spatial position further comprises coordinates that describe the relative position of the depth detected object to the electronic device or that allow the relative position of the depth detected object to the electronic device to be calculated.
[0085] In some embodiments, the spatial position of the depth detected object is a geometric center of the points of a segment of the point cloud.
[0086] The segment of the point cloud is a point cloud segment.
[0087] The geometric center of the point cloud segment can be calculated by calculating an average position of all points of a point cloud segment. Alternatively, a center of mass of the point cloud segment may be determined. Alternatively again, the position of the point of the point cloud segment closest to the lens may be determined to be the position of the point cloud segment. The position of the point cloud segment is a position in three dimensions.
[0088] In some embodiments, detecting the depth detected object comprises determining a spatial movement of the depth detected object.
[0089] Determining the spatial movement of the depth detected object is based on the position of the depth detected object. For example, the spatial position of the depth detected objects at two different times may be recorded. The spatial movement of the depth detected object is then calculated as a velocity vector of the depth detected object.
[0090] For example, the well-known equation (7) may be used:
[0091] In equation the spatial position at a first time t , is the spatial position at a later, second time t2and v is the spatial movement. The position of the depth detected object at times later than the second time may be calculated based on the spatial movement. More sophisticated methods, such as a physical trajectory or multi-position interpolation and / or extrapolation may be used.
[0092] By calculating the movement of the depth detected object, a more efficient and accurate autofocusing may be achieved.
[0093] In some embodiments, the electronic device is further configured to perform autofocusing based on the spatial movement of the depth detected object.
[0094] The movement of the depth detected object may be used in selection of the depth detected object.
[0095] In some embodiments, the electronic device is further configured to predict a position of the selected depth detected object based on the spatial movement of the selected depth detected object and perform autofocusing based on the predicted spatial position of the selected depth detected object.
[0096] The predicted spatial position of the selected depth detected object at times later than the second time may be calculated based on the spatial movement of the selected depth detected object. By performing autofocusing based on the predicted spatial position of the selected depth detected object, the selected depth detected object may be kept in focus even if the relative position of the selected depth detected object changes with respect to the lens.
[0097] In some embodiments, the ToF-sensor is a ToF-component of a combined ToF-RGB sensor and the visual image sensor is an RGB -component of a combined ToF-RGB sensor.
[0098] All-in-one RGB / ToF sensor module
[0099] The combined ToF-RGB sensor may, for example, be an RGB+SPAD sensor. However, this choice is only one possibility and should be understood as non-limiting.
[0100] Use of the ToF-RGB sensor will provide for the FOV of the ToF-sensor and the visual image sensor to be substantially coaxial.
[0101] However, any ToF-component of the coaxial RGB-ToF sensor capable of generating depth image data can be considered according to the present disclosure.
[0102] For example, a device as described in patent document WO2021117642, specifically as shown in Fig. 12a therein, may be provided. Alternatively, for example, a single-chip RGB-ToF camera, as described in non-patent literature SHINMURA et al. Estimation of Human Orientation using Coaxial RGB-Depth Images. In: Proceedings of the 10th International Conference on Computer Vision Theory and Applications (VISAPP-2015), pages 113-120, ISBN: 978-989-758-090-1, especially section 2, may be provided.
[0103] By providing a combined RGB-ToF sensor, a more compact electronic device may be provided.
[0104] In some embodiments, the electronic device further comprises a point pattern emission assembly, the point pattern emission assembly being configured to generate a point pattern in the FOV of the ToF sensor.
[0105] The point pattern emission assembly may be an IR laser or any other emitter capable of emitting electromagnetic radiation that can be detected by the ToF sensor. Furthermore, the emitter should be capable of providing a point pattern and the modulated signal required to produce depth image data by the ToF sensor.
[0106] In some embodiments, the point pattern emission assembly is a vertical cavity surface-emitting laser (VCSEL) array.
[0107] A VCSEL array is an array of VCSEL emitters. Known VCSEL arrays can be used to provide the spot pattern. The VCSEL array can provide one or more laser beams and the spot pattern comprising one or more spots. The VCSEL array can provide the laser beam in the IR wavelengths.
[0108] In some embodiments, the electronic device is further configured to obtain user input from a user, wherein selecting the depth detected object is based on the user input.
[0109] The user input may be any manipulation of the electronic device, using an electronic assembly provided for this purpose, wherein the electronic assembly generates an electronic signal that can be interpreted by the electronic device such that a depth detected object is designated.
[0110] In some embodiments, the electronic device further comprises a user input assembly and an image output assembly.
[0111] The user input assembly is an electronic assembly provided for the purpose of receiving user input. The user input assembly may be a user interface. Image output assembly may be a display.
[0112] In some embodiments, the user input assembly is a physical button.
[0113] The physical button may be provided in a location such that the user may manipulate the button while using the electronic device. The physical button provides an electronic signal that can be interpreted by circuitry.
[0114] In some embodiments, the user input assembly and the image output assembly are comprised by a touch- screen display. The touch-screen display provides an easy means for the user to manipulate the device. The touch screen may provide information visually to the user.
[0115] In some embodiments, obtaining user input comprises presenting the visual image data to the user; presenting an input option associated with the depth detected object to the user; and acquiring the user selection of the input option as the user input
[0116] The visual image data is the visual image data acquired by the visual image sensor. The visual image may be called a live output. The live output may be an output of the live visual image based on the visual image data captured by the visual image sensor.
[0117] At least some of the visual image may be included in a user interface (UI).
[0118] In other words, the UI is in a live view: not offline processed, i.e., not editing images that are already captured and saved in the device.
[0119] The visual image data describes an image. Thus, presenting the visual image data to the user means presenting an image to the user that is based on the visual image data.
[0120] A UI setting may be selectable according to a desired focus setting for the FOV.
[0121] A focus setting may be provided such that the focus setting changes according to depth image data as the depth detected object corresponding to the selection moves in the visual image sensor’s FOV.
[0122] In some embodiments, obtaining user input further comprises presenting, associated with the input option, a distance or a range of the selected depth detected object to the user or a distance or a range of the selected depth detected object to sensor and / or the lens. The distance or range may be an estimated distance or range.
[0123] The presented visual image may be updated whenever a new frame is captured by the visual image sensor. The presented visual image may also be updated according to a predetermined update frequency independent of frame capture. A user input option may be an input field that visually represents a shape of a button to the user on the touch screen display. Furthermore, the input field may present to the user the distance to depth detected objects. The user may then select one of the input options. Once the user selects the input option, the device performs autofocusing on the position of the depth detected object. By presenting the distances to depth detected objects to the user, a scene structure becomes visible to the user. Specifically, the user can understand at which distances objects included in the scene, as captured by the visual image sensor, are located, such that the user obtains greater control over the image they want to cause the visual image sensor to capture. In this way, manipulating the electronic device in order to autofocus the object is rendered less complicated for the user.
[0124] Thus, in some embodiments, an visual image sensor with a lens system, which may be a lens arrangement, and a depth sensor is provided.
[0125] The device, in some embodiments, uses the depth sensor to detect objects in the FOV and to track them. The device further presents to the user an visual image of the FOV on a device and uses the depth detected objects, fused with corresponding objects in the visual image to allow selection of an object, sets the visual image sensor lens arrangement focus to the depth detected by the depth sensor for the selected depth detected object and to modify the focus as the tracked depth changes for the selected object. A fused object may be an associated object. The associated object may be identified in metadata.
[0126] In some embodiments, obtaining user input further comprises presenting a background input option associated with a scene background to the user; wherein selection of the background input option causes the electronic device to perform autofocusing on the scene background as the selected depth detected object.
[0127] By indicating that, upon choosing input option associated with the scene background, the background will be focused, the user can be made further aware of the scene structure. The background may, for example, be a wall at a far side of a room that the user is located in, or a cityscape in the distance, or the horizon, or the sky.
[0128] If the background is ambiguous, an input option associated with an infinity focus may be provided in addition to the input option associated with the scene background. The input option associated with the infinity focus causes the device to set a focus distance of the lens assembly to infinity.
[0129] For example, in a scene composed of a wall with a window aperture, the background input option may cause the device to focus on the wall, whereas the infinity focus option may cause the device to set the focus distance to infinity, such that objects visible through the window aperture are in focus of the lens assembly. In another example in a scene comprising a wire fence, the background input option may cause the device to focus on the wire fence, the wire fence recognized by object recognition, whereas the infinity focus option may cause the device to set the focus distance to infinity, such that objects visible through the openings of the wire fence are in focus of the lens assembly. There are embodiments wherein alternatives to providing an input option that focusses on a background and an input option that focusses infinity are provided. For example, an input option that causes the lens assembly to be focused on an object in the foreground is provided. Such an object may be an extended object, such as the fence or the pane of glass, that is a visual foreground to the background, the background being, for example, the sky or the like.
[0130] An input option labelled “foreground”, that is associated with the foreground as described, may be provided such that the lens assembly is focused on the fence and an input option associated with the scene background is provided which causes the lens assembly to be focused on infinity.
[0131] Note that all labels, such as “foreground”, “background” etc. are to be understood as nonlimiting examples and other labels descriptive of the scene may be provide according to the present disclosure.
[0132] Note that, in all embodiments, a contextual re-selection of the depth detected object may be provided. For example, when recording a video, the video may start with the fence in focus and then smoothly transition, such that the background is in focus.
[0133] In some embodiments, the association of each provided input option may be configured to change based on contextual information of the scene. The contextual information may, for example, comprise information on whether there is an object that would be labeled “background”, but that would not require the lens assembly to be set to focus on infinity, such as a wall on the far side of a room, or if the object is not present and the “background” is, instead, the sky, which would require the lens assembly to be set to infinity. If, in the present example, the sky is present, selecting the option associated with the scene background causes the lens assembly to be focused on infinity. If, in the present example, the wall on the far side of the room is present, selecting the option associated with the scene background causes the lens assembly to be focused on the wall.
[0134] Changing of the association of input options may also be provided between frames, such that the association of an input option changes is a suitable object appears in the FOV of the ToF sensor or the FOV of the visual image sensor. In some embodiments, obtaining user input further comprises presenting a distance or range scale to the user. The range scale may be shown along a predetermined edge of the image output assembly. The image output assembly may be rectangular. Alternatively, the image output assembly may be circular, oval or irregular. The image output assembly may be curved in three dimensions.
[0135] The distance or range scale may comprise a slider that the user may manipulate in order to change the focus distance of the lens assembly. By showing the range scale to the user, the user is further made aware of the scene structure. Thus, manipulation of the electronic device is made more efficient.
[0136] In some embodiments, obtaining user input further comprises presenting an input slider associated with the range scale to the user, the slider indicating a focus distance; wherein manipulation of the slider by the user causes the electronic device to set the focus distance based on the position of the slider recalculate the autofocus based on the focus distance.
[0137] By allowing the user to set a focus distance using a slider, manipulation of the electronic device is rendered less complicated for the user.
[0138] In some embodiments, obtaining user input further comprises presenting the input option in visual association with the distance or range scale.
[0139] The range scale may indicate distances in units of meters, centimeters, feet, or the like. The range scale may be shown in association with input fields, such that each input field is located, in relation to the range scale, at the range of the associated depth detected object. By showing the range scale to the user in this way, the user is further made aware of the scene structure. Thus, manipulation of the electronic device is eased.
[0140] The input fields may be provided as virtual buttons.
[0141] In some embodiments, obtaining user input further comprises detecting the depth detected object in the visual image data and visually indicating the depth detected object to the user. Detecting the depth detected object may be based on analyzing the visual image data using object detection algorithms. To detect objects, this approach may use one or more algorithms from a family of methods that include traditional computer vision algorithms and machine learning algorithms like convolutional networks, single shot detectors and transformers, that can be trained to recognize objects in the 2D visual image. This approach may be based on trained machine learning. This results in a bounding box of where the object are in the image. This bounding box may be presented to the user together with the visual image.
[0142] In some embodiments, obtaining user input further comprises recognizing the depth detected object in the visual image data and identifying the depth detected object to the user.
[0143] Objects may be detected and recognized in the visual image in order to understand what object it is and to offer tracking to the user, then the 3D detection and tracking may be used to keep the object in focus as it moves. Alternatively, both methods may be used in parallel to fill blind spots of each method with the other one depending on conditions. Also, in beneficial conditions, such as good lighting with clearly detected objects, 2D may be used. In detrimental conditions, if light conditions deteriorate or tracking in 2D is not suitable, 3D tracking may be used. In order to recognize objects, object recognition algorithms, as described hereinabove, may be used.
[0144] In some embodiments, wherein tracking may be achieved by projecting boundaries of the point cloud segment, as detected in the depth image data in the 2D image and the area in the 2D image corresponding to the boundaries of the point cloud segment are tracked in the 2D image. This way, use object recognition may be avoided.
[0145] In some embodiments, certain objects may be specified such that, if detected as the depth detected objects and identified by object recognition, no input option is presented to the user or a specially indicated input option is presented to the user. Said certain objects may be objects that are not usually a desired target of autofocusing. Said certain object may be a pane of glass or a fence.
[0146] In some embodiments, the electronic device is further configured to acquire and store visual image data while performing autofocusing.
[0147] In this way, a high quality image of the depth detected object can be obtained and stored for further use. The stored visual image data may be a still image or a video.
[0148] In some embodiments, the predetermined selection criteria comprise at least one of: select depth detected object if the depth detected object is located inside the FOV of the ToF-sensor, but outside the FOV of the visual image sensor; select depth detected object if the depth detected object is located inside the FOV of the ToF-sensor, but outside the FOV of the visual image sensor, but moving toward the FOV of the visual image sensor; select depth detected object with fastest movement in FOV of the ToF sensor; select depth detected object located in FOV of the ToF sensor based on user preparation input.
[0149] The present disclosure may provide for the depth detected object to be selected autonomously by the device. For example, if the user is having trouble selecting the detected object using the user input assembly, or the user is unable to keep the object in the FOV of the visual image sensor by turning the electronic device, the electronic device may still focus on it such that, if the object appears in the FOV of the visual image sensor, it is in focus. This may be useful in detrimental lighting conditions. This may also be useful if the object is moving fast. By selecting a depth detected object if it is moving towards the FOV, objects that the user may want to image will be efficiently selected. In this way, acquiring high quality image data of objects that are moving or difficult to visually see is made easier. The selection criteria may comprise additional criteria. The selection may also be based on user preference information. In other words, in some embodiments, the depth sensor has a wider FOV than the visual image sensor FOV, which allows to use the depth sensor to detect objects in the depth sensor FOV, calculates the spatial position, which may, in some embodiments, only comprise depth values, for an object in the depth sensor FOV but outside of the visual image sensor FOV and track the position according to the movement of the depth detected object. The device then detects the object tracked by the depth sensor when it enters the FOV of the visual image sensor and applies a focus setting for the lens arrangement based on the tracked position.
[0150] Furthermore, in some embodiments, a fast moving object is detected in depth sensor FOV. The device then tracks the fast moving object as a priority and stops tracking a previously selected depth detected object. The device may also be configure to track a fastest moving object, whose movement is calculated based on ToF image data, in a scene. This allows to achieve capturing focused images even if the scene is unexpected.
[0151] The selection may also be based on user preparation input which is presented to the user as a consequence of autonomous selection. If an object is detected in the FOV of the TOF sensor but outside the FOV of the visual image sensor, the object is selected if input is provided by the user indicating that the object should be selected.
[0152] In some embodiments the user preparation input is obtained by presenting a preparation input option associated with the non-moving depth detected object to a user such that, if the user selects the preparation input option, the non-moving depth detected object is selected.
[0153] The preparation input option may be presented as a user input option as described hereinabove. Specifically, an input option may be presented that is labeled “prepare” or the like as the preparation input option. If the user selects the input option labeled “prepare”, the lens assembly is focused such that, if the selected object enters the FOV of the visual image sensor, the lens assembly is already focused on it.
[0154] This way, capturing artistic photographs is eased.
[0155] There are embodiments wherein the preparation input option is presented in the live view of the scene at a point where, in relation to the live view, the depth detected object is estimated to enter the FOV of the visual image sensor.
[0156] The point where the depth detected object is estimated to enter the FOV of the visual image sensor may be a point on the edge of the FOV of the visual image sensor that is closest to the depth detected object. Alternatively the point where the depth detected object is estimated to enter the FOV of the visual image sensor may be a point on the edge of the FOV of the visual image sensor where the movement of the depth detected object intersects the edge of the FOV of the visual image sensor.
[0157] Furthermore, in some embodiments, if a selected depth detected objects leaves and re-enters the FOV of the visual image sensor, autofocus may optionally autonomously enter a default mode, but reapply autofocusing to the selected depth detected object if the selected depth detected object re-enters the FOV of the visual image sensor. This may comprise designating the object, such that the device remembers the object as long as it is located inside the FOV of the ToF sensor, but outside the FOV of the visual image sensor. This may be called object memory.
[0158] As an evolution of object memory, in some embodiments, if the tracked object becomes occluded and moves behind the occluding object to a different depth while still being occluded, if the selected depth detected object reappears on one side of the occluding object, check if object the same and, if it is, tracking is reapplied.
[0159] If an asymmetrical object rotates, in some embodiments , the likelihood that it is the same object is determined by comparing shape data with a database of object shapes or machine learning algorithm trained on shapes and retain the tracking and autofocus if there is a positive comparison - for example a bird flying towards a camera may have “V” shaped wing profile, but as it turns in the air for example by 90 degrees, the “V" shaped wing profile becomes less distinct until one of the wings obscures another. It is possible to algorithmically determine that it is the same bird at a substantially same distance from the lens. Similarly, an object with metamorphic properties, such as an explosion or rapid combustion or a balloon being inflated or bursting, may be determined to be a same object with respect the database of object shames or machine learning algorithm.
[0160] Additionally, or alternatively, a rotating object may be recognized by comparing the position of the isolated point cloud segment of the object with the predicted position of the object as predicted based on the tracking speed. Furthermore, a rotating object may be recognized based on the size or shape of the object. This way, use of object recognition may be avoided.
[0161] Additionally, a rotating object may be recognized in the visual image data by recognizing that an object detected in the depth image data exhibits a color variation.
[0162] Furthermore, in some embodiments, if a selected depth detected object becomes occluded by another object, refocusing to the occluding object may be delayed for a predetermined refocusing delay time. If the selected depth detected object does not reappear after the predetermined refocusing delay time, autofocusing is moved to the occluding object. This procedure may comprise object recognition.
[0163] In some embodiments, the electronic device is further configured cease tracking the depth detected object based on predetermined cessation criteria, the predetermined cessation criteria comprising: cease tracking based on a predetermined user input; cease tracking based on a lack of a predetermined user input; cease tracking if a clear movement of the selected depth detected object is not measured.
[0164] Additionally, if the autofocus is lost intentionally or unintentionally, a default focus setting may be adopted. Indicating data representing of the loss of autofocus may be presented to the user on the UI. The default setting may perform autofocusing on a center of the visual image using, for example, contrast detection autofocus using the visual image data. Alternatively or in addition, the default setting may perform autofocusing on face detection.
[0165] By ceasing tracking under certain criteria, computational resources may be conserved. Furthermore, energy may be conserved and manipulation of the electronic device is made easier.
[0166] The present disclosure provides a method comprising: obtain depth image data from a ToF- sensor and visual image data from an visual image sensor; detect a depth detected object in the depth image data; select the depth detected object based on user input or autonomously based on a predetermined selection criterion; track a spatial position of the selected depth detected object, based on the depth image data; and perform autofocusing of a lens system on the spatial position of the selected depth detected object.
[0167] The present disclosure provides a computer program that, if executed by a processor, causes the processor to execute the following method: obtain depth image data from a ToF-sensor and visual image data from an visual image sensor; detect a depth detected object in the depth image data; select the depth detected object based on user input or autonomously based on a predetermined selection criterion; track a spatial position of the selected depth detected object, based on the depth image data; and perform autofocusing of a lens system on the spatial position of the selected depth detected object.
[0168] The methods as described herein are also implemented in some embodiments as a computer program causing a computer and / or a processor to perform the method, when being carried out on the computer and / or processor. In some embodiments, also a non-transitory computer- readable recording medium is provided that stores therein a computer program product, which, when executed by a processor, such as the processor described above, causes the methods described herein to be performed. The computer, and / or processor may also be called circuitry. Fig. 1 shows a procedural block diagram of a method to perform autofocusing according to the present disclosure. This method is to be executed by the device in order to provide the autofocusing functionality according the present disclosure. Based on depth image data and visual image data, an object is selected and the lens moved such that the object is in focus of the imaging device. The autofocusing method comprises an image capturing step SI 1, a first object selection step S12, a determination step S13, a second object selection step S15, a tracking step S14 and a lens movement step S16.
[0169] In the image capturing step SI 1, depth image data and RGB-image data is obtained from the ToF-sensor and the RGB-image sensor or the combined ToF-RGB image sensor. In the first object selection step, a user object selection is obtained based on the depth image data and the visual image data. The user object selection may be based on user input and may comprise a user selection method as described hereinbelow in Fig. 2. In the determination step S13, it is determined whether an object was selected in the first object selection step S12. Whether or not an object was selected may, for example, be determined based on the output of the first object selection step S12. It should be noted that the selection of an object by the user in the first object selection step S12 may be persistent across multiple executions of the first object selection step. This means that the user may only select an object once, but the next time the first object selection step is executed, the same object is still recognized as being selected. The output of the first object selection step may, for example, be a computational flag indicating that an object has been selected.
[0170] If an object has been selected (indicated by “yes”), the determination step S13 is followed by the tracking step S14. If an object has not been selected (indicated by “no”), the determination step S13 is followed by the second object selection step.
[0171] In the second object selection step, an object selection is determined based on an understanding of scene structure. The second object selection step may comprise a selection heuristic method, such as described hereinbelow in Fig. 3. The second object selection step is followed by the tracking step SI 4.
[0172] In the tracking step SI 4, the object selected either in the first object selection step, which leads to the determination “yes” in the determination step SI 3, the selected object is tracked using object tracking. Furthermore, a focus position is calculated such that the selected object is in focus of the visual image sensor or the combined ToF-RGB sensor.
[0173] The tracking step S14 is followed by the lens movement step S16. In the lens movement step, the lens is moved to the focus position calculated in the tracking step S14. The lens movement step is followed, again, by the image capturing step SI 1, such that the method is executed repeatedly, unless interrupted by an outside signal. In this way, the lens is always kept in a position such that the selected object is in focus, even if a relative position of the object in relation to the lens changes between consecutive executions of the method.
[0174] The tracking step S14 may comprise calculating, based on the position of the tracked object in a current frame and one or more previous frames, a spatial movement of the tracked object. The spatial movement may be expressed in terms of a relative position of the tracked object with respect to the lens. Alternatively, the spatial movement may be expressed in terms of a relative position of the tracked object with respect to a coordinate system independent of the lens, such as a coordinate system of the space that the imaging device is located in. A coordinate transformation may be provided to express the position and / or movement of the tracked object in different coordinate systems as required. The spatial movement may be calculated by linear extrapolation of the position of the tracked object or be based on a physical modelling of the movement. The movement may be a trajectory. Calculation of the movement may be accomplished as described hereinbelow in Fig. 5.
[0175] The tracking step may provide for the focus position to be set such that the tracked object is in focus if, due to the calculated movement, it moves out of focus between consecutive executions of the method of Fig. 1.
[0176] Fig. 2 shows a method for obtaining an object selection based on user input comprising presenting visual image data and detected objects to the user. This method may be used to select an object in the object selection step S12 in Fig. 1 hereinabove. This method is to be executed by the device in order to provide the autofocusing functionality according the present disclosure.
[0177] Prior to the beginning of the method (i.e. at the start), visual image data has been acquired by the visual image sensor and ToF image data, comprising depth image data, has been obtained by the ToF sensor. Both the visual image data and the ToF image data are to be used in the method according to Fig. 2.
[0178] The method comprises obtaining in a first step S21, from the depth image data, the scene structure. The scene structure is described by depth image data of the plurality of infrared spots illuminating the scene as detected by the ToF-sensor.
[0179] In a second step S22, objects in the scene structure are detected. The detection may be based in a suitable object detection algorithm. Alternatively, machine learning or a deep neural network may be used. The obtaining of the scene structure and detection of objects in the scene structure in steps S21 and S22 of Fig. 2 is described in greater detail hereinbelow in Fig. 13.
[0180] In a step S23, the detected objects and the visual image data are presented to the user. Presenting the detected objects and an visual image of the visual image data to the user may comprise presenting, to the user, a user input interface. The user input interface may be presented using a display assembly. The user input interface may be configured as described with respect to Figs. 6, 7 and 8 hereinbelow.
[0181] In a step S24, user input is obtained. Obtaining the user input comprises prompting the user to manipulate the user input assembly. Alternatively, obtaining user input may comprise presenting the detected objects and the visual image of the visual image data in such a way to the user that manipulating the user input assembly is suggested. User input may also be a lack of manipulation of the user input assembly by the user.
[0182] In step S25, the object selection is obtained from the user input. For example, the user input may select one of the detected objects. Thus, the selected detected object may be called a selected object. Otherwise, the user input may indicate a background of the scene or no object. Thus, the object selection may also indicate the background or no object.
[0183] If no object is selected, then the determination in Step S13 in Fig. 1 will lead to Step S15 being executed, wherein, based on the structural understanding of the scene is used to select an object. A method to select an object based on the structural understanding of the scene is described hereinabove with reference to Fig. 3.
[0184] Finishing the method of Fig. 2 yields an object selection, which may indicate an object or no object, and which, as described hereinabove, may be used at the appropriate step of the method of Fig. 1.
[0185] Fig. 3 shows a method to select an object based on the structural understanding of the scene. The method of Fig. 3 may be used in Step S15 in Fig. 1 if no object is selected in Step S12 in Fig. 1. This method is to be executed by the device in order to provide the autofocusing functionality according the present disclosure.
[0186] At the start of the method of Fig. 3, depth image data has been obtained and, in the scene structure, object have been detected. The method of Fig. 3 comprises a series of decision steps, S31, S32 and S33, that lead to different detected objects being selected without user input. In order for the method to be executed, information on the size of the FOV of the visual image sensor with respect to the lens and / or the FOV of the ToF sensor may be provided (see Fig. 9 for reference).
[0187] In step S31, it is determined whether there is a detected object that is both in view of the visual image sensor, i.e. located inside the FOV of the visual image sensor, and in motion. If a detected object is both in view of the visual image sensor and in motion (indicated by “yes”), the object is selected in step S311. If a plurality of objects are both in view of the visual image sensor and in motion, the method may comprise selecting the object with the fastest motion. There may be embodiments wherein the selection in step S311 is further based on user preference information. Following step S311, the method is concluded and an object selection is provided to the parent method, such as the method described in Fig. 1.
[0188] It is to be noted that, as described hereinabove, the FOV of the ToF sensor may be larger than the FOV of the ToF sensor. Furthermore, objects are detected in the ToF image data. It is therefore possible that a detected object is located inside the FOV of the ToF sensor and outside the FOV of the visual image sensor. It is, of course, also possible for a detected object to be located both inside the FOV of the ToF sensor and inside the FOV of the visual image sensor (see Fig. 9 for reference).
[0189] If no object is determined to meet the criteria of step S31 (indicated by “no”), a second determination step, S32, is executed. In step S32, it is determined whether there is a detected object that is not in view of the visual image sensor, i.e. located outside the FOV of the visual image sensor but moving towards the FOV of the visual image sensor. Movement may be determined as described with reference to Fig. 5 hereinbelow. Movement towards or away from the FOV of the visual image sensor may be determined as described with reference to Figs. 12 and 13 hereinbelow. Movement of the detected object may be a movement of the detected object relative to the lens, which may arise from the lens being moved or rotated with respect to the surrounding space with the detected object remaining stationary with respect to the surrounding space, or a movement of the detected object with respect to the surrounding space with the lens remaining stationary with respect to the surrounding space. Movement of the detected object relative to the lens may also arise from a superposition of movement of the lens with respect to the surrounding space and movement of the detected object with respect to the surrounding space.
[0190] If a detected object is both not in view of the visual image sensor and moving towards the FOV of the visual image sensor (indicated by “yes”), the object is selected in step S321. If a plurality of objects are both not in view of the visual image sensor and moving towards the FOV of the visual image sensor, the method may comprise, based on a predetermined choice or selection criteria, any of selecting the object with the fastest motion, the object closest to the FOV of the visual image sensor or the object requiring the least time to enter the FOV of the visual image sensor. There may be embodiments wherein the selection in step S321 is further based on user preference information. The motion may also be called a movement.
[0191] If multiple objects meet the criteria of steps S31 or S32, additional selection criteria may be used to select a single object. For example, of the objects meeting the criteria of steps S31 or S32, an object closest to the FOV of the visual image sensor may be selected. Alternatively, of the objects meeting the criteria of steps S31 or S32, an object closest to the center of the FOV of the visual image sensor may be selected. Alternatively, of the objects meeting the criteria of steps S31 or S32, an object to be determined using object recognition to be a predetermined object may be selected. The predetermined object may be a face, a person, an animal or any type of inanimate object that may be of interest. The predetermined object may be predetermined based on user preference information.
[0192] Following step S321, the method is concluded and an object selection is provided to the parent method, such as the method described in Fig. 1.
[0193] No object being determined to meet the criteria of step S32 (indicated by “no”) is equivalent to a determination of no motion being detected, as illustrated by step S33, which leads to step S331. In step S331, conventional methods are used to select an object, or determine a focus. For example, focusing may be based on focusing on a region of interest in the visual image, for example using a contrast detection autofocus procedure.
[0194] It should be noted that the method shown in Fig. 3 may comprise additional methods for object selection. Specifically, conventional methods for selecting an object inside the FOV of the visual image sensor, i.e. in the visual image, may be used. For example, face detection and / or face recognition may be used in order to detect a in the visual image, the face may be selected instead of the moving object. A priority of which type of object is selected according to which criteria may be determined based on user preference information. According to some embodiments, the selection according to steps S31 and / or S32 may be assigned a lower priority than conventional selection methods. The selection according to steps S31 and / or S32 may be assigned a lower priority than conventional selection methods based on user preference information.
[0195] Following step S331, the method is concluded and an object selection is provided to the parent method, such as the method described in Fig. 1. It should be noted that, if desired, steps S31 and S32, and associated steps S311 and S321 may be reversed, leading to a change in priority of object selection options.
[0196] Fig. 4 shows a method for the user to select different focusing methods and objects in accordance with the present disclosure. This method is to be executed by the device in order to provide the object selection functionality and obtain an object selection as required by step S12 in Fig. 1 and / or step S24 in Fig. 2 according the present disclosure. The method to be presented herein may be executed in concert and / or in parallel with the methods of Figs. 1, 2 and 3 to produce the desired effect.
[0197] The method may start with the device entering or caused to enter the focusing method according to the present disclosure and image data is acquired, for example in Step SI 1 of Fig. 1. In a first step S41, objects in the image are detected. This object detection may be identical to the object detection of step S12 in Fig. 1 and / or step S22 in Fig. 2. Alternatively, a separate object detection may be provided. In a second step S42, distances of detected objects and the background are shown in a camera interface. The camera is an embodiment of the imaging device according to the present disclosure. The camera interface is an embodiment of the display apparatus and / or the user input apparatus. Step S42 may be an embodiment of step S12 in Fig. 1 and / or step S23 in Fig. 2. Showing said content may be provided in a persistent manner, such that additional user input is obtained by the user providing additional input in the same interface.
[0198] In a step S43, the user selects a detected object using an associated button or by manipulating the camera interface. The camera interface may be a touch screen display.
[0199] In step S44, the device changes the focus to the selected object based on the depth image data. This may be accomplished according to steps S14 and S16 in Fig. 1.
[0200] As the presenting of information to the user is persistent, the user is able to provide additional input. This input is processed in step S45. If the user selects no other object (i.e. “no”), step S44 is repeated, such that the selected object remains in focus.
[0201] If, instead, another object is selected, step S43 is repeated, such that the new object is in focus.
[0202] If the user selects another option, such as “cancel”, which is provided for this reason, the focus mode changes to a conventional focus mode, such that the focusing method according to the present disclosure is exited.
[0203] Fig. 5 illustrates calculation of the position and / or movement of a detected object with respect to a lens 111. Fig. 5 comprises two dimensions for illustrative purposes, but all calculations and positions are to be understood to be obtained in spatial coordinates. As such, a coordinate axis Al may indicate a lateral position and coordinate axis A2 may indicate a longitudinal position. Both the position along the axis Al and A2 may be provided by the ToF sensor. However, the position along the axis A2 is provided by depth information, while the position along the axis Al is provided by the position of the image of the detected object on the sensor surface, i.e. which pixels detect the image. Note that the positions may be given in terms of positions relative to the lens 111 or in positions relative to a space that the imaging device is located in. The lens 111 may be assigned a position within this coordinate system.
[0204] A detected object may be located at a first time (tA) at the position Pl. The position Pl may indicate the position as described hereinbelow with respect to Fig. 13. At a second time (tn) the same detected object may be located at a position P2. The timespan between tA and tn may, for example, be a time between two subsequent captures of depth image data by the ToF sensor.
[0205] Between tA and tn, the detected object moves from position Pl to position P2, by the movement m. M may be calculated as a velocity vector or a trajectory using established physical methods. For example, the movement m may be obtained by extrapolating the positions Pl and P2 for times later than tn. Furthermore, the movement may also be based on additional positions (not shown) at additional times.
[0206] Fig. 9 is a schematic illustration of the device 100 according to the present disclosure comprising one visual image sensor 110 and one ToF sensor 120 as separate units. The device comprises an visual image sensor 110, a lens 111 and a ToF sensor 120. The visual image sensor 110 is disposed such that incident light first passes the lens 111 and is then detected by the visual image sensor, which then provides, based on the incident light, visual image data. The visual image sensor 110 has a field of view FOV1, such that objects located inside the field of view FOV1 of the visual image sensor 120 can be detected by the visual image sensor 110. A ToF sensor 120 is also provided. The ToF sensor 120 is disposed such that its field of view FOV2 intersects the field of view of the visual image sensor 110. In some embodiments, the field of view FOV2 of the ToF sensor 120 overlaps and exceeds the field of view FOV1 of the visual image sensor 110. In any case, the ToF sensor 120 is provided such that objects located in the field of view FOV1 of the visual image sensor 110 are detectable by the ToF sensor 120.
[0207] Fig. 10 is a schematic illustration of the device 100 according to the present disclosure using a combined RGB-ToF sensor 130 in a single unit. Here, the field of view FOV2 of the ToF component of the combined RGB-ToF sensor 130 overlaps the field of view FOV1 of the RGB component of the combined RGB-ToF sensor. In an embodiment according to Fig. 10, both the visual image data and the ToF image data are provided by the combined RGB-ToF sensor 130.
[0208] Fig. 11 illustrates movement of detected objects towards the field of view FOV1 of the visual image sensor 110. Movement towards the field of view FOV1 of the visual image sensor 110 may be defined as any movement of a detected object that, based on its movement, will intersect the field of view FOV1 of the visual image sensor 110. In Fig. 11, this type of movement is illustrated by a movement Ml of the detected object with a position Pl, a movement M2 of the detected object with a position P2 and a movement M3 of the detected object with a position P3. Note that all detected objects are located inside the field of view FOV2 of the ToF sensor 120. Both the position and the movement of the detected objects are determined based on the ToF image data.
[0209] Fig. 12 illustrates movement of detected objects away from the field of view FOV1 of the visual image sensor 110. Movement away from the field of view FOV1 of the visual image sensor 110 may be defined as any movement of a detected object that, based on its movement, will not intersect the field of view FOV1 of the visual image sensor 110. In Fig. 12, this type of movement is illustrated by a movement Ml of the detected object with a position Pl, a movement M2 of the detected object with a position P2 and a movement M3 of the detected object with a position P3. Note that all detected objects are located inside the field of view FOV2 of the ToF sensor 120. Both the position and the movement of the detected objects are determined based on the ToF image data.
[0210] Fig. 13 illustrates determination of a position of an object based on the ToF image data.
[0211] Fig. 13 (A) illustrates a visual scene 400 as may be detected conventionally by the visual image sensor 110, showing an object 01, which is, here embodied by a toy train 01 running on tracks T, in front of a background B. Note that Fig. 13 (A) is only provided for convenience in order for the reader to understand the scene to be analyzed in a visual sense. The technology according to the present disclosure is described with reference to other figures. Fig. 13 (B) illustrates the same scene, from a different angle, as a scene structure 401, without the object 01. The scene structure is comprised of the depth image data, as provided by the ToF sensor 120, of individual points. Each point is the reflection of an infrared signal emitted by a VCSEL array, which is provided as part of the device to generate the point pattern for the ToF sensor 120. A plurality of point 301 is an example of a point cloud segment, which, in Fig. 13 (B) indicates the train track T in Fig. 13 (A). In Fig. 13 (C) the point cloud segment 300 (indicated by a dashed circle, which is not part of the point cloud) is associated with the object 01 in Fig. 13 (A). As in Fig. 13 (B), the point cloud segment 301 is associated with the train track T. Note that the train track T may also be an object. An algorithm is provided in order for the device to detect points belonging to a point cloud segment. For example, a statistical method or a deep neural network may be provided in order to determine that certain point are associated with a point cloud segment 300. By detecting a point cloud segment 300, objects are detected in the scene structure. Detecting objects in this way may be used in Step S12 in Fig. 1, step S22 in Fig. 2 or step S41 in Fig. 4. The position of the object 01 is determined by determining the position of the associated point cloud segment 300. The flat structure of points is the point cloud segment Bl is, in this example, associated with the background B. Object recognition may be provided to detect objects. For example, object recognition may be provided such that object 01 is recognized as a toy train or that the background B is a scene background. The object recognition may use either ToF image data or visual image data.
[0212] Determination of the position of the point cloud segment 310, illustrated in Fig. 13 (C), is accomplished by calculating, based on the ToF image data of the points associated with the point cloud segment 300, a geometric center of the point cloud segment. The position of the point cloud segment 310 may, for example, be represented by a position of the object Pl in Fig. 5.
[0213] The position 311 is, for example, a position of the point cloud segment associated with the background B. A movement of the object 01, as described with reference to Fig. 5, is determined by calculating the movement of the point cloud segment 300 associated with the object 01.
[0214] Fig. 6 shows a scene 400 that may be presented to the user in order to obtain the user input. Based on the ToF image data, the device 100 presents, to the user, an visual image of the visual image data obtained at a given moment as well as a distance scale 500, located, for example, at the right edge of the scene. The distance scale may indicate, by giving the distance values, the distances of objects in the scene as calculated from the scene structure as described in Fig. 5 and Fig. 13.
[0215] Fig. 7 shows other possible presentations of different scenes 400 that may be presented to the user in order to obtain user input.
[0216] For example, the device may present to the user the visual image of the visual image data. In Fig. 7 (A), the scene 400 includes a detected object 01 and a detected object 02 and the background B. The visual image is an image of the FOV of the visual image sensor. The visual image may be called a scene. The scene may be presented on a camera display, which may be provided as a touchscreen display, which is the user input assembly. The device may present to the user input fields Fl, associated with detected object 01, F2, associated with detected object 02 and FB, associated with the background B. The user may manipulate the touchscreen display by touching the input field Fl. This causes the detected object 01 to be the selected object and the device to track and subsequently focus the object 01. Alternatively or subsequently, the user may manipulate the touchscreen display by touching the input field F2. This causes the detected object 02 to be the selected object and the device to track and subsequently focus the object 02. Alternatively or subsequently, the user may manipulate the touchscreen display by touching the input field FB. This causes the background to be the selected object and the device to focus the background. The input fields may indicate to the user the distance of the associated detected object. For example, the field Fl may display the character string “50 cm” or the like to indicate that object DI is located at a distance of 50 cm, the field F2 may display the character string “1.1 m” to indicate that object DI is located at a distance of 1.1 m. Furthermore, the field FB may, based on object recognition, display the character string “background” or the like to indicate that object B is the background.
[0217] Alternatively to displaying a character string, the fields may display miniature images of the associated depth detected object. The miniature image may be a cropped and resized segment of the visual image data.
[0218] Alternatively again, the fields may display small icons indicating the associated depth detected object.
[0219] There are embodiments wherein virtual buttons presented to the user are located visually separate from the fields, such that, if the user manipulates the virtual button, the depth detected object associated with the input field associated with the virtual button is selected.
[0220] Alternatively, the background option may cause the focus distance to be set to infinity.
[0221] Furthermore, the fields may be colored in different colors in order to more clearly indicate which object they are associated with.
[0222] In particular, the field may be colored in an average or predominant color of the object they are associated with. If, for example, the average color of the object 01 is green, then field Fl may be colored green etc. Calculating the average color may comprise calculating the average color detected by the pixels in the visual image sensor detecting the associated object. The pixels may be identified using object recognition based on the visual image data. It should be recognized that, instead of manipulating a touch screen display, the user may instead manipulate physical buttons, with selections indicated by a cursor or the like. Other selection mechanisms such as voice control are possible.
[0223] The fields Fl, F2, FB etc. may be ordered in order of distance. For example, the field associated with the nearest object may be located near the bottom of the scene as shown on the camera interface, and the furthest object, or background, may be located near the top of the scene as shown on the camera interface.
[0224] The fields are shown located along the right edge of the scene for illustrative purposes, but locations along other edges may be provided. The choice of which edge the fields are located along may change depending on the scene 400 or the scene structure.
[0225] Fields may be presented in their distance order from the background or from the lens. They may be presented adjacent each other or with a separation indicating their distance from one another. In Fig. 6 they are presented in a column, but they might be presented in a row. They may be presented in an array of two or more objects are a same distance away. For example, an object at a same distance as 01, 01 ' may have a horizontally adjacent Field Fl' to Fl .
[0226] However, the field should not be provided in visual association, i.e. next to, their associated object, unless that object is accidentally located next to the intended position of the associated field.
[0227] The scenes shown in Fig. 7 (B) and 8 (C) are to illustrate other scenes and the associated input fields. For example, in Fig. 7 (B), field F3 may be associated with detected object 02 and display the character string “20 cm” to indicate that the detected object 03 is located at a distance of 20 cm. Further, in Fig. 7 (C), field F4 may be associated with detected object 04 and display the character string “90 cm” to indicate that the detected object 03 is located at a distance of 90 cm.
[0228] In the illustrations in Fig. 7, the input fields are shown located next to each other, but other configurations, such as shown in Fig. 8 hereinbelow, may be provided.
[0229] Fig. 8 shows an alternative configuration for the scene to be presented to the user in order to obtain user input. The difference to the embodiment in Fig. 7 is that the input fields F1-F4 and FB are shown in relation to a distance scale, with each field located at the position on the distance scale corresponding to the distance of the associated object. Each field, as in Fig. 7, may display a character string indicating the distance of the associated object and / or a designation obtained using object recognition. The user may press a button or manipulate an input field to focus directly on the associated detected object. The selected object will be tracked for focus until the button is pressed again or another object is selected.
[0230] Using this method, autofocus can be maintained even if there are multiple objects moving, the camera is moving or where there are close objects where it is not always clear which one the user wants to focus on.
[0231] Fig. 14 schematically illustrates an embodiment of a device that implements the methods described in the embodiments. The electronic device 1200 may implement all processes related to the embodiments. The electronic device 1200 may be identical with the electronic device 100. The electronic device 1200 comprises a CPU 1201 as processor. The electronic device 1200 further comprises an image sensor assembly 1210 connected to the processor 1201. The processor 1201 may implement any processes related to code-based runtime estimation as described above. Image sensor assembly 1210 may comprise any or all of a combined RGB-ToF sensor (130 in Fig. 10 above), a RGB-IR sensor (110 in Fig. 9 above) and a ToF sensor (120 in Fig. 9 above) and may generate intensity image data, depth image data, visual image data or grey scale image data. The data generated by the image sensor assembly 1210 may be processed by the processor. The electronic device further comprises an lens system adjustment unit 1211, which, upon receiving instructions from the processor, may mechanically configure the lens system and / or move a lens (111 in Fig. 5 above) from a first position to a second position. The electronic device further comprises a laser array 1213, which, upon receiving instructions from the processor, emits laser beams to generate a point pattern. The laser array 1213 may be the VCSEL array.
[0232] The electronic device 1200 may further comprise a user interface 1212 that is connected to the processor 1201. This user interface 1212 acts as a man-machine interface and enables a dialogue between a user and the electronic device. The user interface may in particular be a display or a touch screen display. For example, the user may make configurations to the system using this user interface 1212. The electronic device 1200 may further comprise a Bluetooth interface 1204 and a WLAN interface 1205. These units 1204, 1205 may act as I / O interfaces for data communication with external devices. For example, other devices with WLAN or Bluetooth connection may be coupled to the processor 1201 via these interfaces 1204 and 1205. The electronic device 1200 further comprises a data storage 1202, and a data memory 1203 (here a RAM). The data storage 1202 is arranged as a long-term storage, e.g. for storing image data, like depth image data, intensity image data, visual image data, or user preference information or the like. The data memory 1203 is arranged to temporarily store or cache data or computer instructions for processing by the processor 1201.
[0233] It should be recognized that the embodiments describe methods with an exemplary ordering of method steps. The specific ordering of method steps is however given for illustrative purposes only and should not be construed as binding. For example the ordering of S31 and S32 in the embodiment of Fig. 3 may be exchanged. Other changes of the ordering of method steps may be apparent to the skilled person.
[0234] All units and entities described in this specification and claimed in the appended claims can, if not stated otherwise, be implemented as integrated circuit logic, which may be called circuitry, for example on a chip, and functionality provided by such units and entities can, if not stated otherwise, be implemented by software.
[0235] In so far as the embodiments of the disclosure described above are implemented, at least in part, using software-controlled data processing apparatus, it will be appreciated that a computer program providing such software control and a transmission, storage or other medium by which such a computer program is provided are envisaged as aspects of the present disclosure.
[0236] Note that the present technology can also be configured as described below.
[0237] (1) An electronic device comprising circuitry configured to obtain depth image data from a ToF- sensor and visual image data from an visual image sensor; detect an object in the depth image data; select the depth detected object based on user input or autonomously based on a predetermined selection criterion; track a spatial position of the selected depth detected object, based on the depth image data; and perform autofocusing of a lens system on the spatial position of the selected depth detected object.
[0238] (2) The electronic device according to (1), further comprising the ToF sensor, the visual image sensor and the lens system, the visual image sensor receiving light through the lens system, wherein an FOV of the ToF-sensor overlaps an FOV of the visual image sensor.
[0239] (3) The electronic device according to any of (1) or (2), wherein the FOV of the ToF-sensor exceeds the FOV of the visual image sensor.
[0240] (4) The electronic device according to any of (1) to (3), wherein detecting the object comprises segmenting a point cloud obtained from the depth image data. (5) The electronic device according to any of (1) to (4), wherein the spatial position of the depth detected object is a geometric center of the points of a segment of the point cloud.
[0241] (6) The electronic device according to any of (1) to (5), wherein detecting the depth detected object comprises determining a spatial movement of the depth detected object.
[0242] (7) The electronic device according to any of (1) to (6), further configured to perform autofocusing based on the spatial movement of the depth detected object.
[0243] (8) The electronic device according to any of (1) to (7), further configured to predict a position of the selected depth detected object based on the spatial movement of the selected depth detected object and perform autofocusing based on the predicted spatial position of the selected depth detected object.
[0244] (9) The electronic device according to any of (1) to (8), wherein the ToF-sensor is a ToF- component of a combined ToF-RGB sensor and the visual image sensor is an RGB-component of a combined ToF-RGB sensor.
[0245] (10) The electronic device according to any of (1) to (9), further comprising a point pattern emission assembly, the point pattern emission assembly being configured to generate a point pattern in the FOV of the ToF sensor.
[0246] (11) The electronic device according to any of (1) to (10), wherein the point pattern emission assembly is a vertical cavity surface-emitting laser (VCSEL) array.
[0247] (12) The electronic device according to any of (1) to (11) further configured to obtain user input from a user, wherein selecting the depth detected object is based on the user input.
[0248] (13) The electronic device according to any of (1) to (12) further comprising a user input assembly and an image output assembly.
[0249] (14) The electronic device according to any of (1) to (13), wherein the user input assembly is a physical button.
[0250] (15) The electronic device according to any of (1) to (14), wherein the user input assembly and the image output assembly are comprised by a touch-screen display.
[0251] (16) The electronic device according to any of (1) to (15), wherein obtaining user input comprises presenting the visual image data to the user; presenting an input option associated with the depth detected object to the user; and acquiring the user selection of the input option as the user input (17) The electronic device according to any of (1) to (16), wherein obtaining user input further comprises presenting, associated with the input option, a distance or a range of the selected depth detected object to the user.
[0252] (18) The electronic device according to any of (1) to (17), wherein obtaining user input further comprises presenting a background input option associated with a scene background to the user; wherein selection of the background input option causes the electronic device to perform autofocusing on the scene background as the selected depth detected object.
[0253] (19) The electronic device according to any of (1) to (18), wherein obtaining user input further comprises presenting a distance or range scale to the user.
[0254] (20) The electronic device according to any of (1) to (19), wherein obtaining user input further comprises presenting an input slider associated with the range scale to the user, the slider indicating a focus distance; wherein manipulation of the slider by the user causes the electronic device to set the focus distance based on the position of the slider recalculate the autofocus based on the focus distance.
[0255] (21) The electronic device according to any of (1) to (20), wherein obtaining user input further comprises presenting the input option in visual association with the distance or range scale.
[0256] (22) The electronic device according to any of (1) to (21), wherein obtaining user input further comprises detecting the depth detected object in the visual image data and visually indicating the depth detected object to the user.
[0257] (23) The electronic device according to any of (1) to (22), wherein obtaining user input further comprises recognizing the depth detected object in the visual image data and identifying the depth detected object to the user.
[0258] (24) The electronic device according to any of (1) to (23), further configured to acquire and store visual image data while performing autofocusing.
[0259] (25) The electronic device according to any of (1) to (24), the predetermined selection criteria comprising at least one of: select depth detected object if the depth detected object is located inside the FOV of the ToF-sensor, but outside the FOV of the visual image sensor; select depth detected object if the depth detected object is located inside the FOV of the ToF-sensor, but outside the FOV of the visual image sensor, but moving toward the FOV of the visual image sensor; select depth detected object with fastest movement in FOV of the ToF sensor; select depth detected object located in FOV of the ToF sensor based on user preparation input. (26) The electronic device according to any of (1) to (25), wherein the user preparation input is obtained by presenting a preparation input option associated with the non-moving depth detected object to a user, such that, if the user selects the preparation input option, the non-moving depth detected object is selected.
[0260] (27) The electronic device according to any of (1) to (26), further configured cease tracking the depth detected object based on predetermined cessation criteria, the predetermined cessation criteria comprising at least one of: cease tracking based on a predetermined user input; cease tracking based on a lack of a predetermined user input; cease tracking if a clear movement of the selected depth detected object is not measured.
[0261] (28) A method comprising: obtain depth image data from a ToF-sensor and visual image data from an visual image sensor; detect a depth detected object in the depth image data; select the depth detected object based on user input or autonomously based on a predetermined selection criterion; track a spatial position of the selected depth detected object, based on the depth image data; and perform autofocusing of a lens system on the spatial position of the selected depth detected object.
[0262] (29) A computer program that, if executed by a processor, causes the processor to execute the following method: obtain depth image data from a ToF-sensor and visual image data from an visual image sensor; detect a depth detected object in the depth image data; select the depth detected object based on user input or autonomously based on a predetermined selection criterion; track a spatial position of the selected depth detected object, based on the depth image data; and perform autofocusing of a lens system on the spatial position of the selected depth detected object.
[0263] (30) A non-transitory computer-readable recording medium that stores therein a computer program product, which, when executed by a processor, causes the method according to (28) to be performed. LIST OF REFERENCE SIGNS
[0264] 110 visual image sensor
[0265] 111 Lens
[0266] 120 ToF sensor
[0267] 130 Combined RGB-ToF sensor
[0268] 300 Point cloud segment
[0269] 301 Point cloud segment
[0270] 310 Position of point cloud segment
[0271] 311 Position of point cloud segment
[0272] 400 Scene
[0273] 1200 Electronic device
[0274] 1201 CPU
[0275] 1202 Storage
[0276] 1203 RAM
[0277] 1204 Bluetooth interface
[0278] 1205 WLAN interface
[0279] 1210 Image sensor assembly
[0280] 1211 Optical assembly adjustment unit
[0281] 1212 User interface
[0282] 1213 Laser array
[0283] 01-4 Object
[0284] B Background
[0285] T Train track
[0286] Fl -4 Input field
[0287] FB Background input field
[0288] FOV1-2 Field of view
[0289] M Movement Ml -3 Movement
[0290] Pl -3 Position
Claims
CLAIMS1. An electronic device comprising circuitry configured to obtain depth image data from a ToF-sensor and visual image data from an visual image sensor; detect an object in the depth image data; select the depth detected object based on user input or autonomously based on a predetermined selection criterion; track a spatial position of the selected depth detected object, based on the depth image data; and perform autofocusing of a lens system on the spatial position of the selected depth detected object.
2. The electronic device according to claim 1, further comprising the ToF sensor, the visual image sensor and the lens system, the visual image sensor receiving light through the lens system, wherein an FOV of the ToF-sensor overlaps an FOV of the visual image sensor.
3. The electronic device according to claim 2, wherein the FOV of the ToF-sensor exceeds the FOV of the visual image sensor.
4. The electronic device according to claim 1, wherein detecting the object comprises segmenting a point cloud obtained from the depth image data.
5. The electronic device according to claim 4, wherein the spatial position of the depth detected object is a geometric center of the points of a segment of the point cloud.
6. The electronic device according to claim 1, wherein detecting the depth detected object comprises determining a spatial movement of the depth detected object.
7. The electronic device according to claim 6, further configured to perform autofocusing based on the spatial movement of the depth detected object.
8. The electronic device according to claim 6, further configured to predict a position of the selected depth detected object based on the spatial movement of the selected depth detected object and perform autofocusing based on the predicted spatial position of the selected depth detected object.
9. The electronic device according to claim 1, wherein the ToF-sensor is a ToF-component of a combined ToF-RGB sensor and the visual image sensor is an RGB-component of a combined ToF-RGB sensor.
10. The electronic device according to claim 1, further comprising a point pattern emission assembly, the point pattern emission assembly being configured to generate a point pattern in the FOV of the ToF sensor.
11. The electronic device according to claim 10, wherein the point pattern emission assembly is a vertical cavity surface-emitting laser (VCSEL) array.
12. The electronic device according to claim 1, further configured to obtain user input from a user, wherein selecting the depth detected object is based on the user input.
13. The electronic device according to claim 12, further comprising a user input assembly and an image output assembly.
14. The electronic device according to claim 13, wherein the user input assembly is a physical button.
15. The electronic device according to claim 13, wherein the user input assembly and the image output assembly are comprised by a touch-screen display.
16. The electronic device according to claim 12, wherein obtaining user input comprises presenting the visual image data to the user; presenting an input option associated with the depth detected object to the user; and acquiring the user selection of the input option as the user input17. The electronic device according to claim 16, wherein obtaining user input further comprises presenting, associated with the input option, a distance or a range of the selected depth detected object to the user.
18. The electronic device according to claim 17, wherein obtaining user input further comprises presenting a background input option associated with a scene background to the user; wherein selection of the background input option causes the electronic device to perform autofocusing on the scene background as the selected depth detected object.
19. The electronic device according to claim 17, wherein obtaining user input further comprises presenting a distance or range scale to the user.
20. The electronic device according to claim 19, wherein obtaining user input further comprises presenting an input slider associated with the range scale to the user, the slider indicating a focus distance; wherein manipulation of the slider by the user causes the electronic device to set thefocus distance based on the position of the slider recalculate the autofocus based on the focus distance.
21. The electronic device according to claim 19, wherein obtaining user input further comprises presenting the input option in visual association with the distance or range scale.
22. The electronic device according to claim 16, wherein obtaining user input further comprises detecting the depth detected object in the visual image data and visually indicating the depth detected object to the user.
23. The electronic device according to claim 22, wherein obtaining user input further comprises recognizing the depth detected object in the visual image data and identifying the depth detected object to the user.
24. The electronic device according to claim 1, further configured to acquire and store visual image data while performing autofocusing.
25. The electronic device according to claim 1, the predetermined selection criteria comprising at least one of: select depth detected object if the depth detected object is located inside the FOV of the ToF-sensor, but outside the FOV of the visual image sensor; select depth detected object if the depth detected object is located inside the FOV of the ToF-sensor, but outside the FOV of the visual image sensor, but moving toward the FOV of the visual image sensor; select depth detected object with fastest movement in FOV of the ToF sensor; select depth detected object located in FOV of the ToF sensor based on user preparation input.
26. The electronic device according to claim 25, wherein the user preparation input is obtained by presenting a preparation input option associated with the non-moving depth detected object to a user, such that, if the user selects the preparation input option, the non-moving depth detected object is selected.
27. The electronic device according to claim 1, further configured cease tracking the depth detected object based on predetermined cessation criteria, the predetermined cessation criteria comprising at least one of:cease tracking based on a predetermined user input; cease tracking based on a lack of a predetermined user input; cease tracking if a clear movement of the selected depth detected object is not measured.
28. A method comprising: obtain depth image data from a ToF-sensor and visual image data from an visual image sensor; detect a depth detected object in the depth image data; select the depth detected object based on user input or autonomously based on a predetermined selection criterion; track a spatial position of the selected depth detected object, based on the depth image data; and perform autofocusing of a lens system on the spatial position of the selected depth detected object.
29. A computer program that, if executed by a processor, causes the processor to execute the following method: obtain depth image data from a ToF-sensor and visual image data from an visual image sensor; detect a depth detected object in the depth image data; select the depth detected object based on user input or autonomously based on a predetermined selection criterion; track a spatial position of the selected depth detected object, based on the depth image data; and perform autofocusing of a lens system on the spatial position of the selected depth detected object.