Apparatus and method for camera positioning guidance
A 2D visualization method using static and movable targets on a screen guides users to align cameras for high-quality 3D reconstruction, addressing inefficiencies in existing methods by simplifying positional and rotational control.
Patent Information
- Application Number
- PCT/EP2025/063232
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-14
- Filing Date
- 2025-05-14
- Publication Date
- 2025-12-26
AI Technical Summary
Existing user guidance methods for image-based 3D reconstruction are inefficient, particularly when real-time 3D reconstruction is not possible or the quality is insufficient, lacking precise control over camera positions and orientations, and often requiring additional devices or complex 2D-to-3D projections.
A 2D visualization approach that uses static and movable targets on a screen to guide users to align the camera with target locations and orientations, employing pattern matching and gamification principles to simplify the alignment process.
Provides intuitive and precise guidance for users to capture high-quality images for 3D reconstruction, reducing the need for real-time processing and additional devices, and enhancing user experience.
Smart Images

Figure EP2025063232_26122025_PF_FP_ABST
Abstract
Description
[0001] APPARATUS AND METHOD FOR CAMERA POSITIONING GUIDANCE
[0002] Technical Field
[0003] The present application relates to a concept for a positioning guidance for a camera user such as for guiding a user to take images suitable for image-based 3D reconstruction.
[0004] 1 Background
[0005] Multi-view stereo [1], also called photogrammetry, neural radiance fields [2], Gaussian splatting [3], and signed distance fields [4] are all technologies that convert multiple photos of an object or a scene into a 3D model of the latter. While the underlying algorithms of those methods may differ significantly, they all have in common, that for high output quality they need to be provided with “proper” input photos. On the one hand, this means that each photo must have high quality. Over- or underexposed image regions, focus blur and motion blur resulting from camera movement while capturing the photo is to be avoided. On the other hand, sufficient photos need to be taken from the “right” locations. In the following, we call the camera position from which a photo has been taken also as capture location.
[0006] Finding the right capture location requires a lot of expertise and experience. While this is not a problem for expert users, people using such technology more occasionally struggle to obtain a high quality 3D reconstruction due to inappropriate input photos. The following sections describe a visualization approach that helps guiding a user to capture the photos from the locations required for a good 3D reconstruction.
[0007] 2 State of the art
[0008] Guiding the user is mostly performed using real-time preview of the current geometry reconstruction. Trnio [5], Display.land [6], and RealityScan [7] display the user a point cloud giving a coarse indication on where the geometry could be incomplete (see Fig. 1), thus requiring additional captures.
[0009] As point clouds only provide coarse geometry and hence an incomplete picture, other apps such as itSeez3D [8] or Canvas [9] return 3D surfaces reconstructed in real-time (see Fig. 2). Similar approaches are described in
[0010] , Alternatively, Scaniverse
[0011] and the work in
[0010]
[0012] show how to highlight regions for which only insufficient geometry information is available. However, due to real-time requirements, the quality of the available mesh data is often limited.
[0011] All those approaches provide a very direct way to inform a user where to scan more. However, they all require real-time 3D reconstruction. Either this causes high workload on the capture device resulting in corresponding battery drain, or good data connection is required for cloud processing. Moreover, real-time 3D reconstruction is difficult for challenging objects such as homogeneous surfaces or periodic structures. It possibly even requires specific depth sensing devices. Finally, for neural radiance fields (NeRFs), the reconstructed geometry is not fully indicative, because NeRFs also require that a surface is captured from various different directions to properly reproduce view-dependent appearance (shading). Moreover, none of these approaches allows precise control on capture locations and camera orientations. This also holds for the approach described in
[0012]
[0013] , It proposes to synthesize the view of the next capture location. The user is then asked to move the camera in such a way that the captured image resembles the synthesized view as closely as possible. While this is again a very indirect way of guiding the user, it suffers on top from the problem that the synthesized view will most likely show poor quality, as otherwise the new capture location would not provide any added value.
[0013] Qlone
[0014] avoids 3D real-time reconstruction. Instead, it overlays the object to scan with a segmented surface. Surface segments that are sufficiently covered by captured photos are removed from the visualization. While this approach is very intuitive, it also does not allow to define precise camera positions or orientations for user guidance. This also prevents proper user guidance in case one segment requires more viewing angles than another segment, because the user cannot be guided to add those additional viewing angles.
[0014] Instead of visualizing hull segments of the object to scan, Luma Al
[0015] finally indicates necessary capture locations in 3D space by a corresponding AR overlay (see Fig. 14). However, with such an approach it is also not possible to guide the user regarding camera orientation. Moreover, finding the precise camera locations is also difficult as the 3D target camera positions need to be projected on a 2D screen. This makes it more difficult for a user to see where exactly he or she needs to move with his or her camera. In
[0016] , a system is proposed to precisely control the capture location and orientation of a smartphone camera using an additional AR headset. While being able to achieve similar goals as our proposal, it differs significantly from our approach: First of all, it requires two devices (capture device and AR headset) to guide the user. Secondly, while also using rectangles for visualization, they are used in a completely different manner. They represent the target camera and the current camera both as rectangles. The user then must align both rectangles in 3D space. This is completely different to our approach, where we project a rectangle generated for the desired field of view of the target camera into the current camera, such that the user can align the target and the current camera field of views relying on a simple 2D screen. Finally, they do not provide a fine-grained rotation user guidance while our approach is able to do so.
[0015] Reference
[0017] presents an approach that is closest to ours (see Fig. 6). They allow for rotation control in the same manner as we do. To this end, they ask the user to rotate the camera in such a way that a green point matches a red point. However, for positional control, they use a different approach. The target position is represented by two rings in the center of the image. The user then needs to move the camera such that a sphere visible on the screen is located insight of the rings. Moreover, they introduce a transparent plane that moves with the camera, and that needs to be placed between the two rings to control the distance between the camera and the object of interest. Our approach differs by replacing the rotational symmetric and opaque sphere of
[0017] by a rectangle frame that has a hole in its center. This allows to represent both rotational and positional control by one symbol on the screen, including the distance between camera and object of interest (see Section 5.6). Consequently, the screen is less cluttered, and the system is easier to use for the user.
[0016] In summary, there is still need for an approach for user guidance which is particularly beneficial in cases where a real-time 3D reconstruction is either not possible, or the provided quality is insufficient. Advantageously, it should be possible to guide the user precisely with respect to both camera locations and orientations. Preferably, the approach should by implementable so that same relies on pure 2D visualization only.
[0017] 3 Summary of the invention
[0018] It is an object of the present invention to provide a concept for guiding a user to move an image capturing device, such as a camera, to a target location and target orientation, which operates by presenting guidance information on a screen, or display, of the image capturing device, and renders the guidance more efficient by being intuitive in user handling.
[0019] This object is achieved by the subject matter of the independent claims.
[0020] The idea underlying the present invention is that a more efficient and intuitive guidance of a user to move an image capturing device to a target location and target orientation by presenting guidance information on a screen, or display, of the image capturing device (e.g. for capturing an object (e.g. in a foreground); e.g. for capturing a (background) scene), may be achieved by visualizing a static target and a movable target on the screen, which form, or contribute to, the guidance information, e.g. by being indicative of a deviation of the current position from the target location or the target location and the target orientation so that, by moving the image capturing device to the intended target location and orientation, the user is able to aim at bringing the visualized targets to be collocated or overlay each other, and by positioning the static and movable targets within the screen using a virtual target, so that the static target is positioned within the screen on a first screen position which corresponds to, via a predetermined mapping between the screen and an image plane of the image capturing device, a first image plane position within the image plane, onto which, when the image capturing device is positioned at the target location and oriented at the target orientation, the virtual target is projected, while the moveable target is positioned within the screen on a second screen position which corresponds to, via the predetermined mapping, a second image plane position within the image plane, onto which, according to a current position of the image capturing device, the virtual target is projected.
[0021] This pattern matching approach has been shown to engage users by triggering a gamification effect. As a result, it simplifies the process for users to align the image capturing device with the desired target position — both in terms of orientation and location.
[0022] In accordance with an embodiment, the virtual target is planar and non-circular, or the virtual target is non-rotational, such as a rectangle in a plane which might, optionally, be arranged perpendicular to the target orientation. Thereby, the static and movable targets may provide hints guiding the user towards the target orientation.
[0023] In accordance with an embodiment, the virtual target is planar and a polygon such as a rectangle and the static and movable targets are visualized as geometric similar polygons such as rectangles even if the movable target would actually appear distorted or deformed relative to the static target. For instance, the second screen position could be computed such that same approximates an exact screen position corresponding to an exact image plane position onto which the virtual target is projected in case of the current orientation of the image capturing device experiencing a deviation from the target orientation. For instance, the second screen position could be computed by projecting the virtual target onto the image plane using a current location of the image capturing device and the target orientation. The guiding hints provided by the two targets would, thus, concentrate on the locational deviation from the target location and a roll deviation, for instance, while the guidance information could, for a guidance in terms of pitch and yaw, comprise a further pair of movable and static target, namely a static rotational target being positioned within the screen on a third screen position which corresponds to, via the predetermined mapping, a third image plane position within the image plane, onto which, when the image capturing device is positioned at the target location and oriented at the target orientation, all points in front of the image capturing device on a target axis leading through the image capturing device are projected, and a moveable rotational target positioned within the screen on a fourth screen position which corresponds to, via the predetermined mapping, a fourth image plane position within the image plane, onto which, according to the current position of the image capturing device, all points on the target axis are projected, or a line or wedge connecting the third and fourth position.
[0024] In accordance with an embodiment, the virtual target is planar and a polygon such as a rectangle and the static and movable targets are polygons, too. They are, for instance, visualized at their exact projected positions. Thereby, the static and movable targets may provide hints guiding the user towards the target orientation in terms of location and orientation.
[0025] In accordance with an embodiment, the virtual target is planar and located in a reference plane which is perpendicular to the target orientation and, for the visualizing the moveable target, a distance between the reference plane and the target location is adapted depending on a length of a projection of the actual distance between the current location and the target location onto an axis perpendicular to the reference plane and leading through the target location so that the distance between the reference plane and the target location is the larger the larger the length is. By this measure, the virtual target may be chosen in a manner so that the static target is format filling in the screen such as extending over more than 40% of the portion of the screen on which the whole field of view of the image capturing device is shown with the movable target nevertheless not being out of the screen when the actual location of the image capturing device is still far away from the target location in terms of distance to the object. The distance between the virtual target and the target location may also be adapted depending on a length of the actual distance between the current location and target location itself.
[0026] In accordance with an embodiment, a distance between the virtual target and the target location may be adapted depending on a current location and current orientation of the image capturing device so that the movable target remains, according to predetermined mapping, within a portion of the screen onto which a field of view of the image capturing device is mapped. For instance, when the current orientation already equals the target orientation, but the current location of the image capturing device gets transversally too offset from the axis perpendicular to the reference plane and leading through the target location, such as more offset than a predetermined limit, and / or gets too close to the object, such as closer than a predetermined limit, then the distance between the reference plane and the target location should be made larger so that the movable target stays within the screen portion.
[0027] In accordance with an embodiment, it may be continuously determined whether the static target and the moveable target are closer to a state of exact mutual overlay than a predetermined threshold and the user is provided with a target-reached signal depending on whether the moveable target and the moveable target are closer to the state of exact mutual overlay than the predetermined threshold. Any suitable distance measure measuring the distance between the targets may be used and be compared to the predetermined threshold to this end.
[0028] In accordance with an embodiment, the image capturing device and the screen may be attached to each other so that the screen is faced towards an opposite side relative to the image capturing device. For example, the screen and the image capturing device (e.g. camera) may be housed in a housing, or in a mobile phone. This allows for an intuitive moving of the image capturing to the target position, i.e. target location and target orientation, device by the user.
[0029] In accordance with an embodiment, for sake of visualizing the movable target, the virtual target may be displaced along an axis leading through the virtual target and the target location to a displaced virtual target location so that a distance between the displaced virtual target location (e.g. position) and the current location becomes equal to a distance between a non-displaced virtual target location of the virtual target and the target location. By this measure, the guidance is made in a manner leaving freedom with respect to the distance of the location to which the user moves the image capturing device along the target location, but merely guides the user laterally. In accordance with a further embodiment, the guidance restricts this distance in a manner to fall into a certain interval, restricted to a certain maximum distance from the target location - both towards deviations away from the object and towards the object. For instance, for visualizing the movable target, a) the virtual target may be displaced along an axis leading through the virtual target and the target location to a displaced virtual target location so that a distance between the displaced virtual target location (e.g. position) and the current location - measured, for instance, along the axis - becomes equal to a maximum of 1) a distance (AJ02) between a non-displaced virtual target location (AG06) of the virtual target and the target location (AA02) and 2) a distance (AJ04) between the non-displaced virtual target location (AG06) of the virtual target and the current location (AJ05) minus a predetermined maximum distance (e.g. AP04) if the displaced virtual target location (e.g. position) gets farer away from the object, and / or b) the virtual target may be displaced along the axis leading through the virtual target and the target location to a displaced virtual target location so that the distance between the displaced virtual target location (e.g. position) and the current location - measured, for instance, along the axis - becomes equal to a minimum of a distance between 1) a non-displaced virtual target location (AG06) of the virtual target and the target location (AA02) and 2) a sum of the distance between the non-displaced virtual target location of the virtual target and the current location on the one hand and a predetermined maximum distance on the other hand if the displaced virtual target location (e.g. position) gets closer to the object.
[0030] In accordance with an embodiment, geometric observation information (e.g. observation volume) may be provided by the apparatus as a component in determining the target location and target orientation, e.g. capturing location or capturing locations, for capturing an image of a scene suitable for generating a scene model based on the image which allows for a generation, or synthesis, of view of the scene from observation points other than the target location and target orientation, wherein a user is interacting with the image capturing device by moving said device to predetermined positions, from which and the apparatus defines as the geometric observation information an observation volume surface of an observation volume within which viewpoints for the views are located, the synthesis of which the scene model shall allow, or the observation volume itself. Using the observation volume surface, a straight forward approach to generate geometric observation information has been found. Geometric observation information can be crucial for applications determining optimal capture points, e.g. future target locations and eventually associated future target orientation, to generate a high quality scene model. Using an apparatus according to this embodiment can yield a higher quality scene model in the end due to the optimized representation of the observation volume using the observation volume surface. The same arguments hold true for an apparatus for providing geometric observation information. Features of an apparatus for presenting guidance information may be combined with any features described in relation to an apparatus for providing geometric observation information.
[0031] In accordance with embodiments of the present invention, the concept is used to guide a user where to capture a photo for a good 3D model. Any reconstruction algorithm may be used for generating the 3D model based on the photo(s) captured.
[0032] Brief description of the drawings
[0033] Embodiments of the present invention are described in the following on the basis of the accompanying drawings, in which:
[0034] Fig. 1 shows user guidance by reconstructed point clouds in Display. land [6] (left), Trnio [5] (middle) and RealityScan [7] (right):
[0035] Fig. 2 shows a realtime preview of mesh reconstructed in real-time for capture guidance, on the left side: itSqueez3D [8], and on the right side: Canvas [9];
[0036] Fig. 3 shows highlighting of image regions with insufficient geometry information in Scaniverse
[0011] ;
[0037] Fig. 4 shows a visualization of surface coverage in Qlone
[0014] ;
[0038] Fig. 5 shows user guidance in Luma Al
[0015] ;
[0039] Fig. 6 shows user guidance according to reference
[0017] ;
[0040] Fig. 7 shows possible degrees of freedom;
[0041] Fig. 8 shows an example object and target capture location and orientation, wherein camera has wrong target location, but correct orientation; Fig. 9 shows a capture device;
[0042] Fig. 10 shows a mapping between capture device sensor and visualization area;
[0043] Fig. 11 shows guiding the user towards the desired camera orientation;
[0044] Fig. 12 shows a computation of location of the target orientation symbol and of the current orientation symbol on the 2D screen;
[0045] Fig. 13 shows guidance of the camera roll angle;
[0046] Fig. 14 shows guiding the user towards the desired capture location;
[0047] Fig. 15 shows a visualization for current camera location being equal to the target camera location;
[0048] Fig. 16 shows a current position frame with dashed contour;
[0049] Fig. 17 shows a computation of the target position frame in the visualization area of the capture device;
[0050] Fig. 18 shows a visualization of the current position frame as a target object;
[0051] Fig. 19 shows textual user instructions;
[0052] Fig. 20 shows a current position frame for a wrong camera rotation;
[0053] Fig. 21 shows distance agnostic user guidance;
[0054] Fig. 22 shows allowing a range for the distance between the camera and the object;
[0055] Fig. 23 shows Guidance cameras;
[0056] Fig. 24 shows a process of 3D reconstruction; Fig. 25 shows an observation volume for an object scan;
[0057] Fig. 26 shows an observation volume defined by a complex observation volume surface shape, for an object of interest;
[0058] Fig. 27 shows two different but essentially equivalent observation volume surfaces and for an object of interest;
[0059] Fig. 28 shows observation volume for an inside-out scan;
[0060] Fig. 29 shows two different and not equivalent observation volume surfaces and for an object of interest;
[0061] Fig. 30 shows an observation volume for a scene consisting of several objects (AH01 , AH02) and background objects, wherein the observation volume is defined by the multiple observation volume surfaces;
[0062] Fig. 31 shows a complex observation volume;
[0063] Fig. 32 shows a definition of cylindrical observation volume;
[0064] Fig. 33 shows a selection of relevant cylinder depending on the scanning mode;
[0065] Fig. 34 shows an irregular observation volume with upright walls;
[0066] Fig. 35 shows a required ray volume (shaded in blue color) to capture for a capture location and an observation volume surface defining the observation volume;
[0067] Fig. 36 shows a capture location located on a surface of a single planar primitive;
[0068] Fig. 37 shows a capture location located on an edge of two planar primitives;
[0069] Fig. 38 shows a spherical projection of the planar primitives onto a face of the spherical coordinate surface cube;
[0070] Fig. 39 shows a benefit of ROI specification for an object to capture; Fig. 40 shows a region of interest definition for a background scan;
[0071] Fig. 41 shows masks for a scanning process where the tree, parts of the city wall and parts of the floor shall be reconstructed;
[0072] Fig. 42 shows exclusion mask for the masks given in Figure 41 ;
[0073] Fig. 43 shows system diagram;
[0074] Fig. 44 shows guided capture workflow, wherein blocks without shading are performed by a computer system control by the expert, and blocks with dashed shading are performed by the user and supervised by the export, the other blocks are performed by the user; and
[0075] Fig. 45 shows a method for presenting guidance information.
[0076] 4 Detailed description of embodiments
[0077] Before describing certain embodiments of the present application, some general thoughts which lead to the underlying idea of the subsequent embodiments are presented.
[0078] For a proper user guidance, embodiments of the invention may need to control at least five out of six degrees of freedom (see Fig. 7). The three translational degrees of freedom (left / right, up / down, backward / forward) as well as the pitch and yaw angles may directly impact which part of the object will be captured. On the other hand, the roll angle may only define whether the image is captured in portrait or landscape mode (or an angle in between these two extremes). This parameter may only indirectly define which part of the object is captured, as the horizontal and vertical field-of-view may differ between landscape and portrait mode. Hence, while the user’s choice needs to be taken into account to compute the necessary capture locations, a guidance may not be absolutely necessary. However, it may be beneficial to guide the user whether to use portrait or landscape mode (or any roll angle in-between), as for some object forms portrait mode may be more beneficial than landscape mode or vise-versa.
[0079] In principle, the necessary capture locations can be easily visualized in a 3D model as done by some related approaches (see Section 2). However, for visualization, this 3D model needs to be projected on a two-dimensional screen of the capture device. This is a big difference to our daily interaction with our 3D surroundings, where thanks to stereo vision by our two eyes we get an understanding of where each object is located in 3D space. Lacking this possibility on a 2D screen, understanding where exactly a desired capture position is located is still hard or at least not natural. And the two rotational degrees of freedom (yaw, pitch) are hard to control by these means as well.
[0080] Consequently, a goal of embodiments of the invention is to guide the user by simple 2D visualizations in such a way that when he or she follows the 2D guidance on a 2D screen, he or she will end up in exactly the desired capture location.
[0081] To guide a user towards the best capture locations and orientations, according to embodiments of the invention there may be two steps. First, an algorithm defines the best possible capture locations. Then, using an embodiment of this invention the user is steered in such a way that he or she can perform the capture from the desired locations.
[0082] Unless stated otherwise, it is assumed, that the algorithm to compute the capture positions is given and a focus is on the question, how to guide the user to these computed camera locations and orientations according to embodiments of the invention.
[0083] 5 2D visualization for user guidance according to embodiments of the invention
[0084] 5.1 Disentanglement of rotational and translational guidance
[0085] Both the translational degrees of freedom as well as the pitch and yaw angle can define which parts of an object are captured by the camera. When, for instance, the camera is moved to the right, the object of interest may move out of the camera cone which represents the volume in 3D space which is projected on the camera sensor (field of view). This can be partially compensated by a rotation of the camera to the left.
[0086] This ambiguity between position and rotation can make a precise guidance of the user more difficult. For these reasons, embodiments of the invention may be configured to opt to separate the guidance for the rotational and translational degrees of freedom. By these means, a user can easily understand whether he or she must change his or her position or the camera angle to reach the desired capture location. Fig. 8 depicts a corresponding example of a target object AA01 and a desired target capture location and orientation (AA02) (e.g. a target location and target orientation of an image capturing device). The target capture location (AA02) can be defined by three coordinates in space, while the capture orientation (AA02) can be defined by the yaw and pitch angle. The guidance according to embodiments of the invention can be performed in such a way that even when the user still has the wrong position (e.g. current position AA03) in 3D space (see AA03 in Fig. 8), he or she can already adjust the camera orientation to match with the desired target values (e.g. adjust the image capturing device’s orientation to match the target orientation).
[0087] Fig. 9 shows a conceptual view of an image capturing device AB01 , having a screen AB02. The screen AB02 may be split into a sub regions AB03, AB04, for example, a visualization area AB03 and a utility area AB04, wherein the visualization area AB03 may depict an image plane or a part thereof and the utility area AB04 may depict utility buttons for interacting with the image capturing device AB01 .
[0088] According to embodiments, an image capture device AB01 can be equipped with a display AB02. Such a capture device AB01 may be a mobile phone, which may have a camera on its back side and a display AB02 on its front side. The display AB02 may be partitioned into different sub-regions AB03, AB04. One of these (sub) regions AB03 may be used for visualizing the captured images or videos and for guiding the user (see Fig. 9). In the following, this region may be called visualization area AB03.
[0089] Fig. 10 shows a conceptual view of a possible predetermined mapping between a screen AB02 that is connected to an image capturing device AB01 and an image plane AE02. For visualization purposes, a fixed mapping AE04, AE05, AE06 and AE07 between the visualization area AB03 and a region AE03 on the sensor (e.g. area / region; e.g. image plane AE03 / AE02; e.g. field of view AE03) may be established.
[0090] In a preferred embodiment, the sensor area AE02 (e.g. image plane), the mapped sensor region AE03 (e.g. field of view) and / or the visualization area AB03 (e.g. of the screen AB02) all have the same aspect ratio. Otherwise, either some sensor content cannot be visualized, or the display of the sensor image may be distorted.
[0091] For brevity and increased readability, going forward the visualization area AB03 may be understood as a synonym for the screen AB02 and the region AE03 may be understood as the image plane AE02. The omitted part of the image plane AE02 and the omitted part of the screen AB02, the region AB04, may not be present on an image capturing device AB01 at all.
[0092] 5.2. Rotational guidance for camera orientation control (yaw and pitch)
[0093] Fig. 11 depicts a visualization of an embodiment of the invention for allowing the user to adjust the camera orientation in form of the pitch and yaw angle displayed on a visualization area AB03 of an image capturing device AB01 .
[0094] To this end, a target orientation symbol (e.g. static rotational target AC01) is visualized on the visualization area AB03 of the (e.g. image) capture device AB01. Moreover, also the current camera orientation by means of a current orientation symbol (e.g. movable rotational target AC02) is visualized. While the target orientation symbol (e.g. static rotational target AC01) remains fixed on the screen, the current orientation symbol (e.g. movable rotational target AC02) changes its position whenever the user changes the yaw or pitch angle of the capture device AB01. By these means, the task of the user becomes very easy. He or she simply changes the camera orientation until the current orientation symbol (e.g. movable rotational target AC02) has the same position on the 2D screen AB03 as the target orientation symbol (e.g. static rotational target AC01).
[0095] Fig. 12 depicts a conceptual 3D view of a computation of guidance according to an embodiment of the invention. There are three possible positions of an image capturing device (e.g. AB01) depicted, a current position AD06, a temporary position AA03 and a target position AA02. For sake of readability and brevity, the word position may be used to summarize location and orientation of an image capturing device.
[0096] Fig. 12 depicts how to compute the position of the target orientation symbol (e.g. static rotational target AC01) and the current orientation symbol (e.g. movable rotational target AC02) on the screen AB03 of the capture device AB01 , for example, to capture an object AA01.
[0097] Camera AA02 (e.g. image capturing device AB01 at target position and target location) represents the desired target position (e.g. target location) and orientation (e.g. target orientation), for example, for the next photo to take. To visualize the target orientation symbol (e.g. static rotational target AC01), a target orientation reference ray AD04a can be selected. For example, this ray AD04a can define the central pixel of the mapped sensor region AE03 of the target camera AA02) By these means, the target orientation reference ray AD04a defines the target orientation of the camera (e.g. image capturing device AB01). On this ray AD04a, a target orientation symbol AD05a (e.g. static rotational target AC01) in 3D space may be placed. The position of the target orientation symbol (e.g. static rotational target AC01) on the 2D screen AB03 can then be computed by projecting the target orientation symbol AD05a (e.g. static rotational target AC01) in 3D space onto the camera sensor (e.g. AE03) using well known camera projection equations
[0018] , During projection, it can be assumed that the camera has the target orientation (AA02, AA03) (e.g. as the image capturing devices AA02, AA03 in Fig. 12), even if an actual orientation AD06 (e.g. actual or current orientation of an image capturing device AB06) is different. The position / distance of the target orientation symbol on the target orientation reference ray AD04a doesn’t matter and can hence be selected arbitrarily. Please note that these steps only need to be performed once, because the target orientation symbol (AC01) (e.g. AC02) has constant position on the visualization area AB03 during capture. Please note that alternatively, the target orientation symbol AC02 can be drawn directly on the screen and then perform inverse projection to place it on the orientation reference ray AD04a.
[0098] To disentangle the orientation guidance from the actual camera location, the position of the target orientation symbol AD05a (e.g. static rotational target AC01) is always relative to the current camera location. In other words, if the camera (e.g. image capturing device AB01) moves from position AA02 to position AA03, the target orientation symbol moves by the same amount from AD05a (e.g. static rotational target AC01) to AD05b.
[0099] To compute the position of the current orientation symbol (e.g. movable rotational target AC02) on the 2D screen AB03, the target orientation symbol AD05c in 3D space (e.g. static rotational target AC01 on the screen AB03) is projected on the camera sensor AE03, AE02 using the actual camera orientation AD06.
[0100] Please note that alternative formulations may be possible to compute the same orientation symbols.
[0101] 5.3 Guidance of the camera roll angle
[0102] According to embodiments, the roll angle can be controlled by the visualization of an electronic spirit level as exemplary depicted in Fig. 13. This can, for instance be achieved by changing the roll angle of the camera (e.g. image capturing device AB01) until an actual rotation symbol AO01 matches with a desired rotation symbol AO02. Alternatively, or additionally, arrows AO03 may indicate a necessary orientation.
[0103] When the user should switch from landscape to portrait mode (or vice versa), the actual rotation symbol AO01 may be rotated by 90 degrees. Guiding the user to a certain roll angle may be beneficial to obtain a better coverage of the object on the sensor AE02, AE03.
[0104] 5.4. Positional guidance
[0105] Fig. 14 depicts a visualisation of guidance presented on a screen AB03 of an image capturing device AB01 according to an embodiment of the invention. The guidance information may be used to guide a user to move the image capturing device AB01 to a target location and target orientation for capturing an object.
[0106] To guide the user towards the desired capture location, a target position frame AF01(e.g. static target) may be drawn on the visualization area AB03 of the capture device AB01 (see Fig. 14). The position of the target position frame AF01 (e.g. static target) in the visualization area AB03 remains constant throughout the capture process (e.g. on a first screen position, which corresponds to, via a predetermined mapping between the screen AB02 and an image plane AE03).
[0107] In addition, a current position frame AF02 (e.g. moveable target) is depicted on the visualization area AE03. Its position and size depend on the deviation between the target camera (e.g. image capturing device AB01) location (e.g. target location) and the current camera (e.g. image capturing device AB01) location. The user can be guided by asking him or her to move the camera (e.g. image capturing device AB01) in such a way that the current position frame AF02 (e.g. movable target) fills the complete target position frame AF01 (e.g. static target).
[0108] For example, to this end, the user can imagine the target position frame AF01 (e.g. static target) to represent the viewfinder of the capture device (e.g. image capturing device AB01). The current position frame AF02 (e.g. movable target) on the other hand may represent a virtual object in 3D space. The task of the user may, for example, hence be described as to “take a photo” of the virtual object and to bring the virtual object as close as possible such that it fills the complete viewfinder. With this explanation in mind, moving the capture device (e.g. image capturing device AB01) becomes obvious to the user. In the example shown in Fig. 14, the user needs to move the capture device AB01 to the right to the shift current position frame AF02 (e.g. movable target) to the left. Similarly, the user needs to shift the capture device AB01 to the bottom to move the current position frame AF02 (e.g. movable target to the top. Finally, the user needs to move the capture device AB01 closer to the object (along the optical axis of the camera (e.g. image capturing device AB01)) to make the current position frame AF02 (e.g. movable target fit the complete target position frame AF01 (e.g. static target (for example as depicted in Fig. 15).
[0109] For a clearer indication that current position frame AL02 (e.g. AF02) (e.g. movable target) is on top of the target position frame AF01 (e.g. static target), the current position frame (e.g. AF02) (e.g. movable target) may use a dashed contour AN02 as depicted in Fig. 16. Alternatively, the target position frame AF01 (e.g. static target) can be drawn on top of the current position frame (e.g. AF02 movable target). In this case, the target position frame (e.g. static target AF01) may use a dashed contour. Or the frames (e.g. static target AF01 and movable target AF02) could change their colour as soon as they have sufficient alignment.
[0110] Fig. 17 depicts a conceptual 3D view of a possible computation of a movable target AF02 in the visualisation area AE03 of an image capturing device AB01 at a current position AG05 for a target position AA02 to capture an object AA01 with an optical axis AG01 . The image capturing device AB01 has a possible current viewing angle AG04 (that may be translated to the image plane of the image capturing device AB01) at the current position AG05 and a possible target viewing angle AG02 (that may be translated to the image plane AE03 of the image capturing device AB01) at the target position AA02. A target position frame AG06 in the 3D view is depicted, whose position can be translated to the position of a movable target AF02 on a screen AB02 of the image capturing device AB01.
[0111] To compute the position of the current position frame AF02 (e.g. movable target) in the visualization area AE03, the target position frame (e.g. AF01) (e.g. target position of the movable target) may first be mapped onto the sensor (e.g. image plane AE03), for example, by using the mapping relation introduced in Fig. 10. Next, the target position frame may be projected into the 3D space (e.g. depicted as AG06) onto a plane (AG03), for example, using the camera projection equations
[0018] (see Fig. 17). In a preferred embodiment, this plane AG03 is orthogonal to the optical axis AG01 of the camera of the capture device (e.g. image capturing device). For the projection itself, the target camera position AA02 may be used of the next image to capture. The projected target position frame AG06 in 3D space is then projected onto the sensor (e.g. image plane AE03) of the current camera position AG05. For example, by means of the mapping relation defined in Fig. 10. The target position frame AG06can then be visualized in the visualization area AB03 of the capture device AB01 as current camera position AF02 (e.g. movable target).
[0112] The distance between the projection plane (AG03) and the target camera position AA02 is a degree of freedom that can be chosen by embodiments of the invention. For example, the larger the distance, the smaller the precision of the location guidance. The smaller the distance, the higher the precision of the location guidance. However, for small distances between the target camera position AA02 and the projection plane AG03, small deviations of the current camera position AG05 from the target camera position AA02 may be sufficient to cause the target position frame AF02 (e.g. movable target) being located outside of the visualization area AB03. This may negatively impact the user experience. Consequently, it may be advantageous to choose a good trade-off between precision and usability. It is also possible to dynamically adapt the distance of the projection plane AG03 from the target camera position AA02 depending on the difference between the target camera position AA02 and current camera position AG05. The larger the difference, the bigger should be the distance of the projection plane AG03 from the target camera position AA02.
[0113] 5.5 Visualization strategies for positional guidance
[0114] 5.5.1 Motivation
[0115] For intuitive usage, it may be advantageous that a user interprets the target position frame AF01 (e.g. static target) as a viewfinder and the current position frame AF02 (e.g. movable target) as a virtual object to capture. This can be achieved by different approaches explained in the following.
[0116] 5.5.2 Use of gamification principles
[0117] By displaying the current position frame (e.g. movable target) as an object that the user intuitively wants to capture, bringing the current position frame (e.g. movable target) into alignment with the target position frame (e.g. static target) may be more self-explanatory. This could be for instance a target mark AH02 or similar objects or symbols as exemplary depicted in Fig. 18. Please note that in a preferred embodiment the current position frame (e.g. movable target) has a rectangular silhouette (or is a polygon) to simplify the mapping with the target position frame (e.g. static target). The static target and the movable target may have the same shape or may be visualized by embodiments of the invention as mutually geometrically similar polygons.
[0118] 5.5.3 Textual annotation
[0119] Alternatively, it is also possible for embodiments of the invention to directly add textual annotations (e.g. to the movable target AF02 or the movable rotational target AC02) explaining the user what to do as depicted in Fig. 19.
[0120] 5.6 Combined positional and rotational guidance
[0121] Although the target AF01 (e.g. static target) and current AF02 (e.g. movable target) position frames may be intended to guide the user towards the desired capture location (e.g. target location), they can also be used to control the rotation (e.g. to the target orientation) of the capture device AB01. The reason is that when the current device rotation does not correspond to the target device rotation (e.g. target orientation), the current position frame (e.g. movable target) can be distorted AI02 (e.g. movable target) as depicted in Fig. 20. Consequently, by trying to match the current position frame AI02 (e.g. movable target) with the target position frame AF01 (e.g. static target), the user automatically also needs to rotate the capture device AB01 appropriately.
[0122] This means that the target orientation symbol AC01 (e.g. static rotational target) and the current orientation symbol AC02 (e.g. movable rotational target) are optional. However, using them can result in higher precision for the user guidance.
[0123] In case the target position frame AF01 (e.g. movable target) is not quadradic, it can also be used for guidance of the roll angle (e.g. to achieve the target orientation). If for instance the next capture should be done in landscape mode and the target position frame is assumed to be landscape as well, but the current orientation (e.g. of the image capturing device) is in portrait mode, then the current position frame would be portrait as well, and the user could not make the two frames fitting without rotating the capture device (e.g. image capturing device).
[0124] 5.7 Optional extension by transparency guidance Similar to
[0017] , the user could be further guided by means of transparency. This transparency indication could either happen for the target position frame AF01 (e.g. static target) or for the current position frame AF02 (e.g. movable target). To do so, the respective frames may be duplicated in a similar manner as
[0017] used two rings. These frames are then projected in 3D space before being rendered, such that it is possible to overlay a transparent frame before rendering.
[0125] 5.8 Method for presenting guidance information
[0126] Embodiments of the invention may be done independently from an apparatus and may be performed as a method MV45 as depicted in Fig. 45. The method MV45 may start with visualizing MV01 the static target and continue with visualizing MV02 the movable target. An inverse order is also possible. Also, a continues update of the moving target in dependence on the location and orientation of the image capturing device in relation to the target location and target orientation, or rather the virtual target may be performed.
[0127] 6 Capture guidance
[0128] 6.1 Guidance for a complete scene or object digitization
[0129] The previous description explained the principles of an apparatus according to embodiments of the invention on how to guide a user to move and rotate the capture device (e.g. image capturing device) such that it fits with a target location and orientation. To guide the user to capture a complete object (e.g. AA01) or scene, all possible target camera positions may have to be computed. Next, they may be ordered in such a way that consecutive camera positions are located closely to each other. Otherwise, the user may spend a lot of time in navigating between the different target camera positions.
[0130] Once the ordered list of target camera positions is available, it can be used to guide the complete capture process. The guidance system according to embodiments may be initialized with a first target camera position. Then, the user navigates to this position. Once the photo has been successfully captured, the target camera position is set to the next desired position in the list, and the procedure repeats accordingly. The guidance system may be configured to automatically determine if the appropriate capture position has been reached 6.2 the of freedom
[0131] In some applications, it may be beneficial to not constrain the distance between the object or scene to capture and the capture location (e.g. of the image capturing device). This can, for example, be achieved by an extension of the algorithm described previously, especially referring to Fig 17, as visualized in Fig. 21. For example, to this end, the projected target position frame AG06(e.g. virtual target) is shifted to a new position AJ06 (e.g. a displaced virtual target location AJ06) in 3D space along an optical axis AG01 of the target camera (e.g. position AA02), such that a distance AJ03 between the shifted projected target position frame AJ06 (e.g. displaced virtual target location AJ06) and the current camera (e.g. position AJ05) equals a distance AJ02 between the original projected target position frame AG06 (e.g. (original) virtual target AG06) and the target camera (e.g. position AA02).
[0132] 6.3 Defining a distance range for capture
[0133] By the approach of the previous section, the user can take an arbitrary distance relative to the object AA01 to scan. If the distance should not be fully arbitrary, but be contained in a certain interval, this can also be achieved by embodiments of the invention. Such an embodiment is depicted in Fig. 22 as a possible extension of Fig 21 . To this end, a distance AJ07 between the shifted target position frame AJ06 (e.g. displaced virtual target location AJ06) and the target camera position AA02 (e.g. the target location and target orientation of the image capturing device) is limited to a certain interval. The target position frame (e.g. virtual target AG06) may not be shifted any further when this interval is exceeded. The shifted target position frame AJ06 (e.g. displaced virtual target location AJ06) is assumed to be contained within a range (AP01 , AP02) (e.g. a range between a maximum AP01 and a minimum AP02). The distance AJ03 between the minimum AP01 and camera position AJ05 may equal the distance AJ02 between the projected target position frame AG06 (e.g. virtual target) and the target camera position AA02 as described in reference to Fig. 21. When the user moves the camera to a position AP03 even further away, then the projected target position frame AJ06 (e.g. displaced virtual target location AJ06) may not further be shifted, but remain at the position AP01 (e.g. remain at the maximum AP01) (see Fig. 22). By these means, the user gets informed that now the admissible distance between object AA01 to scan and the current target position frame (e.g. virtual target AG06) is exceeded.
[0134] 6.4 Handling large deviations between target camera and current camera User guidance as described in Section 5.4 may only work when the deviation between the target camera orientation and the current camera orientation is not too huge. Otherwise, the current orientation symbol AC02 (e.g. movable target) may not be visible in the visualization area AB03. Similarly, when the difference between the target camera position (e.g. AA02) and the current camera position (e.g. AG05) is too large, the current position frame AF02 (e.g. movable target) may not be visible in the visualization area AB03. In both cases, user guidance may not be possible.
[0135] This can be solved by introducing additional guidance camera positions as visualized in Fig. 23. Let’s suppose that we need to move from camera position AA02 (e.g. current camera position AA02) to camera position AK01 (e.g. target camera position AG05). As both cameras (e.g. image capturing device positions) observe the object AA01 from different sides, the guidance procedure described in section 5 may not be directly applied. Instead, some auxiliary camera positions AK02, AK04 may be introduced by embodiments of the invention. By these means, the system guides the user from position AA02 (e.g. the current position) to position AK04, followed by position AK02 until finally reaching the target position AK01. In other words, an apparatus according to embodiments can be configured to utilize auxiliary position (e.g. auxiliary target locations and auxiliary target orientations) to guide the user to the target location and target orientation.
[0136] Alternative approaches may consist in adding additional arrows (e.g. on the screen AB02) for guidance or indications when the projected target position frame is observed from the wrong side.
[0137] Please note that in practice this problem may rarely occur. The reason is that the different camera views usually show significant overlap for a good 3D reconstruction. Combined with a good order of the target camera positions, the user does not move too much from one camera position to the next one.
[0138] 6.5 Extension with other guidance modes
[0139] Although the guidance approach described previously provides all necessary information to precisely guide the user towards the desired capture locations and orientations, some users may desire to have a 3D visualization in addition. To this end, it is important to state that an apparatus according to embodiments may be combined with the visualizations shown in Fig. 1 , Fig. 2, Fig. 3, Fig. 4 and Fig. 5. This may be achieved by overlaying the target position frame AF01 (e.g. static target AF01), the current position frame AF02 (e.g. movable target AF02), the target orientation symbol AC01 (e.g. the static rotational target AC01) and / or the current orientation symbol AC02 (e.g. the movable rotational target AC02) over the corresponding 3D visualization instead of the 2D image captured from the sensor. An apparatus according to an embodiment of the invention may therefore depict an improvement of existing methods to guide the user to a certain target location and target position.
[0140] 7 Use cases for apparatuses according to embodiments offering precise user guidance
[0141] Apparatuses according to embodiments may be utilized, inter alia, in the following use cases:
[0142] • Capture with different distances to the object for being able to reconstruct fine details
[0143] • Guide the user to capture some object parts with more photos than others
[0144] • Create overview images for better camera calibration
[0145] • Better control of texture quality
[0146] • Proper control of the observation volume for background and scene scans
[0147] • Ensuring that the object to scan covers the majority of the sensor
[0148] 8 Further advantages
[0149] • By using a non rotational-symmetric (e.g. symmetric) object for the target and current position frames (e.g. of the static target and the movable target), the projection can also give an indication about the rotation. This is not possible in
[0017] ,
[0150] • If a current camera position gets too close (e.g. current camera position AA02 is too close to an object AA01), the complete current position frame (e.g. movable target AF02) will disappear, which is a clearer indication to move backward than provided in other concepts such as the one of
[0017] ,
[0151] • If a current camera position gets too far away (e.g. current camera position AA02 is too close to an object AA01), the current position frame (e.g. movable target AF02) will get very small, which in contrast to the concept of
[0017] , for instance, is a very clear indication to move forward.
[0152] • The use of a non-opaque target position frame (e.g. static target AF01) allows a user guidance by the size of the current position frame (e.g. movable target AF02) on the screen (e.g. AB02). • Using non-filled line-objects such as rectangles (e.g. as static target AF01 ; e.g. as movable target AF02) occludes only very little information of the visualized image. This is in contrast to
[0017] , where a fully filled sphere occludes a lot of the image.
[0153] • Using non-filled line-objects both for the target (e.g. e.g. static target AF01) and the current position frames (e.g. movable target AF02) delivers better information whether to move forward or backward along the optical axis. This is because it is very easy to see whether the target and current position frame are overlapping or not. This is in contrast to
[0017] , where the sphere occludes the ring. So it is not easy to identify whether the sphere actually appears too big and occludes parts of the ring, which means that the user must change his position along the optical axis. This causes that with the proposed approach the target camera position can be reached with higher precision compared to
[0017] ,
[0154] 9 Summarizing embodiments
[0155] Note that the apparatus for presenting guidance information on a screen AB02 of an image capturing device AB01 presented herein may comprise the image capturing device AB01 such as by being housed in a housing along with the image capturing device AB01. Thus, it would comprise the housing, the image capturing device AB01 and a processor, for instance, for performing the tasks described herein. However, it may also be that the apparatus is separate from the image capturing device AB01 such as communicationally connectable to the image capturing device AB01. Similar statements are true with respect to the screen AB02 or display: The apparatus may comprise same and, for instance, be housed in one housing along with the screen AB02, or may be connectable with the latter. The image capturing device AB01 and the screen AB02 may be attached to each other such as by being commonly housed in a housing. For instance, the apparatus may be a mobile phone that comprises the image capturing device AB01 and the screen AB02. The image capturing device AB01 and the screen AB02 may be attached to each other (or arranged relative to each other in a housing) so that the screen AB02 is faced towards an opposite side relative to the image capturing device AB01. That is, they may be commonly housed in a housing such as a mobile phone in this manner. As described above, the visualization may be done so that the static target AF01 and the moveable target AF02 are visualized on the screen AB02 in a manner overlaid onto a currently captured image, currently captured by the image capturing device AB01 , and in form of a circumference or outline of the static target AF01 and the moveable target AF02. As described above, the apparatus continuously updates the screen positions of the movable target AF02 on the basis of updates of the current position (e.g. AA02) of the image capturing device AB01. To this end, the apparatus may use one or more sensors such as one or more inertia sensors and / or a triangulation sensor and / or a phase or time difference position sensor and / or travel time position sensor of the image capturing device AB01. The latter one or more sensors may, thus, be part of the apparatus in case of the image capturing device AB01 being integral part of the apparatus as described above. Additionally or alternatively, an evaluation of an image capturing device signal of the image capturing device AB01 , i.e. the live video stream output thereby, is used to determine the current position (e.g. AA02), i.e. current location and current orientation.
[0156] 10 Determining positions for capturing images for scene view synthesis or for providing a determination component
[0157] A further aspect of the application is concerned with a concept for determining positions (e.g. target location and target orientation) for capturing images which are suitable for scene view synthesis such as for 3D reconstruction, or a concept for proving a determination component. Features described below can be understood as features of a system for determining capturing positions for capturing images of a scene suitable for generating a scene model based on the images which (e.g. such that the scene model) allows for a synthesis of views of the scene from observation points other than the capturing positions or may also be understood as features of an apparatus for presenting guidance information on a screen of an image capturing device, the guidance information guiding a user to move the image capturing device to a target location and target orientation for capturing an object that is configured to receive the positions for determining the images. As previously, unless stated otherwise, position may be understood as the combination of location and orientation of the image capturing device or of a portable device.
[0158] The further also be described as capturing position determination shared between remote device and portable device.
[0159] 11 State of the art of determination for scene or
[0160] Building upon the discussions presented in the state of the art chapter the following discusses state of the art in respect to the determination of the position. In order to guide a user to capture the required photos for a good 3D reconstruction, there exist multiple smartphone-based solutions
[0023] ,
[0024] ,
[0025] ,
[0026] ,
[0027] ,
[0028] ,
[0029] ,
[0030] that precompute the required capture locations and then try to guide a user towards these locations for capture. However, in general the guidance is not very precise. Moreover, to compute capture locations, the apps typically make some generalizing assumptions on the scene, which in practice may not be true. This causes that either the proposed capture locations are not sufficient, or they simply cannot be attained due to blocking objects. Moreover, even if the capture locations were perfect, there are still a multitude of other factors that need to be controlled such as proper illumination of the object, avoidance of movement, avoidance of disturbing objects, proper camera settings etc. In other words, even when strictly following the proposed capture locations, a good 3D reconstruction is not guaranteed.
[0161] To overcome this problem, an iterative capture and reconstruction process can be followed. After having performed all captures assumed to be sufficient and perfect, the 3D reconstruction can be done. When the results are not satisfactory, new captures can be made. Unfortunately, in practice it is not obvious to understand the reasons for the quality limitations. It could be a lack of input captures, but also inconsistent illumination, moving objects, blurry images, wrong camera parameters and many more. Consequently, such an iterative approach is not of great help for unexperienced users. Moreover, it is very timeconsuming, as 3D reconstruction is a slow process requiring several hours of computation. In practice, it is hence not possible to continue the capture after such a long waiting period, because for instance the illumination has changed. Consequently, 3D digitization needs to be started from scratch again.
[0162] Methods trying to compute the reliability of the reconstruction
[0031] ,
[0032] to find regions with insufficient coverage are able to better identify regions with insufficient coverage. Still, they cannot solve the problem of long waiting times and the difficulties to analyse why the available images are not sufficient to generate a good 3D reconstruction.
[0163] Experienced users, on the other hand, collected enough knowledge to predict during acquisition, which capture locations are required. Moreover, by inspecting the images, they can also judge in how far the images have sufficient quality. Based on this knowledge, companies offer customers to send them their objects to digitize such that they take the photographs and 3D reconstruction. Unfortunately, the physical transfer of an object is cumbersome, time-consuming, or sometimes even impossible.
[0164] Alternatively, users are asked to capture the photos by themselves. Due to limited experience, the captured photos are either not sufficient in number or show quality imperfections. This makes automatic 3D model creation far more difficult. As a result, lots of manual cleaning is necessary, or the model even needs to be created from scratch.
[0165] 12 Fundamentals of scanning for 3D reconstruction
[0166] Fig. 24 shows possible (e.g. may be necessary) steps for creating a 3D representation for an object (e.g. AA01) or a scene. Those steps are detailed in the following sections.
[0167] 12.1 Scene setup
[0168] Creation of a 3D representation from a physical object (e.g. AA01) or scene starts by setting up the objects (e.g. AA01) and all required equipment. For an object (e.g. AA01) scan, the object (e.g. AA01) needs to be placed at an appropriate location and the illumination needs to be configured as desired. Also, the background visible in the photos should be carefully arranged.
[0169] For a background scan, the objects of interest typically cannot move. Still, choosing appropriate illumination and weather conditions can be extremely helpful.
[0170] 12.2 Definition of the observation volume and the object volume
[0171] The observation volume defines the region in 3D space from which a scene or an object (e.g. AA01) should be observed in the application where the 3D object (e.g. AA01) is used. The space not being part of the observation volume may contain one or several objects. Consequently, we call this volume object volume in the following. The definition of the observation volume can have a huge impact on the required capture efforts. Interestingly, although having a crucial impact on the 3D digitization, it is practically never considered in related work.
[0172] 12.3.1 Observation volume for an object scan (Outside-in scan)
[0173] When creating a 3D model of an object AA01 , the observation volume can be defined by a surface AC02 that separates the object AA01 from the observation volume AC03 (see Fig.
[0174] 25). In the following, we call this surface observation volume surface. While the observation volume surface AC02 may have complex shape (observation volume surface AD02 in Fig.
[0175] 26), typically more easy forms as those in Fig. 25 are used. Fig. 26 shows an observation volume AC03 defined by a complex observation volume surface AD02 shape, wherein AD01 is the object of interest AA01.
[0176] From the light-field theory
[0035] ,
[0036] ,
[0037] it is known that when the light-field on a surface is known, then the images of camera locations “seeing” the surface can also be computed. Consequently, placing the capture cameras on the observation volume surface is a possible strategy for digitizing an object. In other words, embodiments of the invention can be configured to determine the target location and target orientation of the image capturing device in proximity or on the observation volume surface AC02.
[0177] Fig. 27 shows two different but essentially equivalent observation volume surfaces AC02 and AE01 for an object of interest AA01. For convex observation volume surfaces AC02 the size and form of the surface essentially may define the texture resolution (in combination with the focal length). For instance, in Fig. 27, two possible observation volume surfaces AE01 and AC02 are depicted. As the object of interest AA01 does not intersect with any of the observation volume surfaces AE01 and AC02, observation volume defined by AC02 is admissible when observation volume defined by AE01 is admissible. Please note that this is not true anymore when the observation volume surface is not convex anymore. In other words, when capturing the object AA01 from the observation volume surface AE01 , the user is still allowed to move to the observation volume surface AC02, even if no capture camera has been placed on this surface AC02. This is why we call those two observation volumes equivalent in the following. The only drawback he or she must accept is that the cameras (e.g. image capturing device AB01) located on the observation volume surface AE01 deliver a lower texture resolution than those that would have been placed on the observation volume surface AC02 (assuming same focal length and sensor properties of the cameras (e.g. image capturing device AB01)). This is because the cameras (e.g. image capturing device AB01) are closer to the object AA01. On the other hand, as the field-of-view may be limited, moving closer to the object AA01 may require more camera images (e.g. of the image capturing device AB01) to capture the complete light-field information, which again makes the determination of camera parameters more challenging.
[0178] For all these reasons, defining the observation volume surface AC02, AD02 or AE01 for an object AA10 scan is recommended. Although not strictly necessary, it allows for better quality and effort control. Please note furthermore that for non-convex observation volumes, the above analysis does not hold and a precise tracking of the observation volume gets even more important.
[0179] 12.3.2 Observation volume for a background scan (Inside-out scan)
[0180] For inside-out scans (e.g. for capturing a scene) as conceptually depicted in Fig. 28, not a single object but a complete background AF03 should be scanned. The observation volume AF02 is again defined by an observation volume surface AF01.
[0181] In contrast to an object AA01 scan as conceptually depicted in Fig. 29, a larger observation volume surface AG01 is not equivalent to a smaller observation volume surface AF01. In other words, if a user places the capture cameras (e.g. image capturing device AB01) on the observation surface volume AF01 , the background AF03 cannot be observed from the observation volume surface AG01 without risking for missing information due to occlusions. This may happen, because there exist rays, that intersect both the background object AF03 and the observation volume surface AG01 without intersecting the observation volume surface AF01. Hence, proper definition of the observation volume AF02 gets even more important, so that all relevant capture locations are taken into account (see Figs. 5 and 6).
[0182] 12.3.3 Observation volume for space scan
[0183] For virtual tours (e.g. for capturing a scene), an observer wants to virtually move through a digitized space, such as a building, a room or a landscape. This requires the most complex observation volume definition, as an observation volume should never intersect with a real object. Consequently, multiple observation volume surfaces may be necessary to define the observation volume (see Fig. 7).
[0184] Fig. 30 shows an observation volume AH07 for a scene consisting of several objects AH01 and AH09 and background objects AH08 and AH03, wherein the observation volume AH07 is defined by the multiple observation volume surfaces AH04, AH05 and AH06.
[0185] Alternatively, a single observation volume AH09 with a very complex form may be sufficient (see Fig. 31). Fig. 31 shows a complex observation volume AH09.
[0186] 12.4 Determination of observation volumes with predefined forms For a user, it is easiest to describe the observation volume by defining exemplary positions from which an object e.g. AA01 , a background or a scene shall be observed. Unfortunately, for general forms this can become quite laborious, as many sample points are necessary to describe the form of the observation volume. This can, for example, be simplified by predefining the general form of the observation volume. The following subsections define a few examples.
[0187] 12.4.1 Cylindrical observation volume surface
[0188] Let’s assume an orthonormal coordinate system, where the x- and y-axes define the floor, and the z-axis defines the upright direction. Assuming further that the axis of the cylinder DH04 shall be parallel to the z-axis. Then, a cylinder DH04 can be defined by two points in space in combination with the assumed viewing directions as shown in Fig. 32. The two cameras (e.g. image capturing devices AH01 , AH02) define both the radius, the height and the position of the cylinder DH04. To this end, one capture location (e.g. position of image capturing device DH01) is assumed to define the top-plane of the cylinder DH04. The other capture location (e.g. position of image capturing device DH02) is assumed to define the bottom plane of the cylinder DH04.
[0189] Please note that a priori two points in 3D space result in two possible cylinders DH04 and AH03 (see Fig. 33). This ambiguity can be resolved by the scanning type. For instance, for an object AA01 scan, the cameras DH01 , DH02 are assumed to point towards the object AA01 that should be located within the cylinder DH04. Consequently, cylinder DH04 in Fig. 33 is the one to choose. For a background scan, on the other hand, the cameras DH01 , DH02 are assumed to be oriented towards the background. Consequently, cylinder DH03 in Fig. 33 is to be chosen.
[0190] In practice, defining a cylinder DH04, DH03 from two points may result in instable results. Consequently, the user may be advised to define a couple of points. The radius of the cylinder DH04, DH03 can then be computed by solving a least square error problem. The top-plane of the cylinder DH03, DH04 may be given by the upmost camera position, while the bottom plane may be given by the lowest camera position.
[0191] 12.4.2 Spherical observation volume A sphere (e.g. the cylinder DH03, DH04 in Fig. 33 may also be viewed as a sphere) in 3D space has four degrees of freedom:
[0192] • The center of the sphere (3 degrees of freedom)
[0193] • The radius of the sphere (1 degree of freedom)
[0194] Consequently, a sphere in 3D space can be defined in a similar way as described in Section 12.4.1 by two points in 3D space (6 constraints) (e.g. DH01 , DH02 of Fig. 33),. Again, more measurements are advised to avoid instable results.
[0195] 12.4.3 Observation volumes in the form of a cuboid
[0196] Cuboids in 3D space have 8 (or 9) degrees of freedom:
[0197] • Center of rectangle (3 degrees of freedom)
[0198] • Height, width, and length of the rectangle (3 degrees of freedom)
[0199] • Rotation of the rectangle (1 degree of freedom, if we assume that one side of the rectangle shall be parallel to the floor, otherwise 2 degrees of freedom)
[0200] Rectangular observation volumes can hence be defined by at least 3 points in space in a similar way as described in Section 12.4.1. In case of axis alignment, two points are sufficient. Similar approaches can be applied for elliptical surfaces.
[0201] 12.4.4 Irregular surfaces parallel to the z-axis
[0202] Irregular observation volumes AH09 like those depicted in Fig. 34 can also be easily described by a user when we assume that all walls shall be perpendicular to the floor. Then, the user can simply move the camera (e.g. image capturing device AB01) along the desired observation volume. The camera positions are continuously tracked. All points are then projected onto the xy-plane to define the form of the observation volume AH09. Similar to Section 12.4.1 , the bottom and top planes (not shown in Fig. 34) can be defined by the highest and lowest camera position that the user moved to.
[0203] 12.5 Computation of the capture locations
[0204] Once the observation volumes AC03, AF02, AH09 are defined, the capture locations for sampling the light-fields defined by the observation volume surfaces can be computed. From a theoretical point of view, those cameras are best located on the observation volume surfaces AC02, AH04 AH06, AH05, AF01 , AG01 , AF01 , AD02. The precise selection of the capture locations on this surface AC02, AH04 AH06, AH05, AF01 , AG01 , AF01 , AD02 may depend on the selected strategy such as the one presented in
[0038] ,
[0205] 12.6 Volume to scan
[0206] 12.6.1 Relevant rays to scan
[0207] Having defined the observation volume (e.g. AC03, AF02, AH09, AH07), the light-field theory defines the necessary capture locations (e.g. the target locations and target orientations of an image capturing device, wherein the user is provided with guidance by an apparatus according to an embodiment). A priori, the observation volume surface should be densely sampled by corresponding camera positions. In other words, the user must move the camera (e.g. the image capturing device) in such a way that the nodal point of the camera hits each and every point of the observation volume surface (e.g. AC02, AH04 AH06, AH05, AG01 , AD02).
[0208] In practice, this may be intractable, as this would require an infinite number of photos to take. Consequently, only a sparse sampling of the observation volume surface is done. More details about light-field sampling are described for instance in
[0039] ,
[0209] Fig. 35 shows a required ray volume AI03 (shaded in blue color) to capture for a capture location AI02 and an observation volume surface AI01 defining the observation volume AI04.
[0210] For each capture position AI02, all rays AI03 that originate in the nodal point of the corresponding camera AI02 and which intersect the object volume (see Section 12.3) when leaving the camera (see Fig. 35) may need to be captured. (Rays that first intersect the observation volume and then intersecting the object volume may not need to be considered).
[0211] All rays to capture can be described by their spherical coordinates relative to the currently considered capture location AI02. Those spherical coordinates can then be visualized by a point on a convex and closed volume encompassing the current capture location. Typically, either a sphere AI05 or a cube is chosen for this purpose. In the following, we call this surface simply spherical coordinate surface. For reasons of simplicity, let’s assume that the observation volume surface is either described or approximated by a set of planar primitives, such as triangles or quads. Then, the three different scenarios to compute the rays to capture can be differentiated:
[0212] • Capture location is located on a surface of a single planar primitive.
[0213] • Capture location is located on an edge of two planar primitives.
[0214] • Capture location is located on a vertex of three or more planar primitives.
[0215] Those cases are described in the following subsections.
[0216] 12.6.2 Capture location located on a surface of a single planar primitive
[0217] Fig. 36 shows a capture location AM03 located on a surface AM01 of a single planar primitive AM01. In this case, the rays to capture form a hemisphere AM03 aligned with the planar primitive AM01 (see Fig. 36).
[0218] 12.6.3 location located on an of two
[0219] Fig. 37 shows a capture location AN03a, AN03b located on an edge of two planar primitives AN01a, AN02a, AN01 b, AN02b. Similar to Section 12.6.2, each planar primitive AN01 , AN02a and AN01 b, AN02b, defines a hemisphere aligned with the primitive. In case the angle between the two planes is smaller than 180° for the object volume (larger than 180° for the observation volume), the rays to capture belong to the intersection of these two spheres. Otherwise, the rays to capture belong to the union of these two spheres (see Fig. 37).
[0220] 12.6.4 location located on a vertex of three or more
[0221] This is the most complex case and is easiest solved when we assume that the spherical coordinate surface is represented by a cube. Each planar primitive can then be projected onto this cube using a spherical projection (the angular coordinates of the projected vertex remain constant). As the planar primitive intersects the capture location, on each face of the cube each projection AO04, AO05, AO06 is represented by a line with a certain length. Each line separates the points belonging to the observation volume AO03 and those of the object volume AO02. Moreover, for each line we know on which side is which volume (observation or object volume). This allows to compute all rays being part of the object volume using for instance a region filling algorithm. Fig. 38 shows a corresponding example. Fig 38 shows a spherical projection of the planar primitives AO04, AO05, AO06 onto a face AO07 of the spherical coordinate surface cube. It shows one face (out of six) (e.g. AO07) of the spherical coordinate surface cube. The camera then needs to be rotated in such a way, that it records all points in the area AO02. By doing so, all relevant rays are captured in the camera.
[0222] While in the most general case this is a rather complex algorithm, it must be noted that in practice it can be significantly simplified, because for predefined forms like those described in Section 12.4, the number of planar primitives intersecting a vertex can be bounded by a low number such as three. This makes the whole algorithm easier to implement and faster to execute.
[0223] 12.7 Determination of region of interest
[0224] The rays to capture may exceed the field of view of the camera (e.g. of the image capturing device). In such a case, the camera needs to be rotated to make several images from the same capture location with different rotations. Overall, there may hence be many captured images necessary to record the full light-field. To reduce the number of photos, specification of a region of interest may be beneficial.
[0225] 12.7.1 Object scan
[0226] Fig. 39 depicts the benefit of defining a region of interest for an object AT01 scan (e.g. the object AT01 may be the object AA01). It shows an object AT01 together with an assumed spherical observation volume surface AT02. According to Section 12.6, for a given capture location AT03, the ray cone to capture encompasses a complete hemisphere AT07 to sample all possible rays. If, however, only the object AT01 is of interest, the cone AT05 may be sufficient. As the object of interest AT01 may have a complex form, it is in practice not possible to only capture those rays that intersect with it. Nevertheless, the user can define a volume AT04 (e.g. the region of interest (Rol; ROI) AT04) with a simple form that contains the object of interest AT01. By these means, the rays to capture can be easily limited to cone AT06, being much more efficient than capturing a complete hemisphere AT07).
[0227] 12.7.2 Background scan In case of a background scan, the main purpose of the region of interest (ROI) definition may consist in excluding unwanted areas of the background. For instance, in Fig. 40, only a small subpart AR01 of the complete background AF07 may be of relevance. This allows to exclude some capture positions completely. For others like AR03, the necessary camera cone AR02 (e.g. field of view of image capturing device) can be restricted.
[0228] 12.7.3 Defining the region of interest in 3D
[0229] In most scanning applications, the region of interest is defined by placing a virtual volume around the object of interest (e.g. AA01 ; AT01). This volume can be visualized using augmented reality technigues (e.g. on the screen of the image capturing device). In more detail, depending on the user’s position, the camera image may be overlaid with the assumed volume and visualized on a 2D screen of the capture device.
[0230] Due to the 2D visualization, this may only work appropriately when the user can walk around the object of interest (e.g. AA01 ; AT01). For background scans, however, this is very difficult. Conseguently, the following subsection describes a method according to an embodiment how this can be achieved in 2D. This not only simplifies the application for a user, but it is also directly applicable for background scans.
[0231] 3.7.4 Defining the region of interest in 2D space
[0232] The region of interest can be defined by marking it in one or multiple images using one or multiple inclusion masks. This is exemplified in Fig. 41. A user captures an image AL01 and marks all regions with possibly overlapping inclusion masks AL02, AL03, AL04 that contain objects he or she wants to reconstruct. It is important that object or object parts that shall be reconstructed must be part of at least one inclusion mask.
[0233] A point p in 3D space then belongs to the region of interest, if and only if the following holds: For all cameras where p belongs to the camera frustum and where the user has drawn at least one mask, p is projected onto a masked pixel.
[0234] This is eguivalently expressed by the following pseudo code, where P(p) is a function that computes the 2D pixel coordinates imaging the point p in 3D space: Def isRoi (p) : / / function returns True if p belongs to ROI
[0235] For all cameras c :
[0236] If camera c does not have any associated mask continue if p is visible in image of camera c : for all masks m in camera c : if P(p) e m : return True / / p belongs to region of interest . return False / / p does not belong to region of interest , return True / / p belongs to region of interest .
[0237] 12.7.5 Computing a 3D volume from 2D ROI masks
[0238] Fig. 42 shows exclusion mask AL05 for the masks given in Figure 41 (AL02, AL03, AL04). Given the ROI masks of an image, an exclusion mask AL05 encompassing all those pixels not being part of any ROI mask (e.g. AL02, AL03, AL04) can be computed (see Fig. 42). The projection of the exclusion mask AL05 into 3D space using the pin-hole camera matrix
[0040] defines all those points in 3D space that are projected onto the exclusion mask AL05 and hence do not belong to the region of interest.
[0239] Starting from the complete 3D space, we can then iteratively compute the ROI volume. To this end, the complete 3D space can be started from as ROI. Then, for any image where the user has drawn a mask, the corresponding projection volume of the exclusion mask can be subtracted from the current ROI volume.
[0240] The resulting ROI volume may contain 3D points in space that are not part of any userdrawn mask. This can be avoided by a second processing step. To this end, each ROI mask drawn by the user may be projected into 3D space using the pin-hole camera matrix
[0040] , Then, the current ROI volume can be intersected with all projected ROI mask volumes and only those parts of the ROI volume are kept that intersect at least with one ROI mask volume.
[0241] 12.8 Refining the volume to scan
[0242] Given a capture location, Section 12.6 computed the rays to capture. Consideration of the region of interest as defined in Section 12.7 allows to further reduce their number. Only those rays intersecting with the region of interest need to be considered. 14 Additional embodiments and
[0243] In the following, additional embodiments and aspects of the further aspect of the invention will be described which can be used individually or in combination with any of the features and functionalities and details described herein.
[0244] According to a first aspect there is provided a system for determining capturing positions for capturing images of a scene suitable for generating a scene model based on the images which (e.g. scene model) allows for a synthesis of views of the scene from observation points other than the capturing positions, comprising a portable device (e.g. mobile phone with a corresponding application thereon), wherein the portable device is portable by a user and comprises an image capturing device for capturing images of the scene, a user input interface (e.g. button and / or touch screen and / or microphone), a user output interface (e.g. screen and / or loudspeaker), and a first communication interface (e.g. wireless connection module), and a remote device comprising a second communication interface for a communication of the remote device with the portable device via the first communication interface, wherein the remote device (e.g. control system in description) and the portable device are configured to cooperate via the communication to determine the capturing positions (e.g. AK03) depending on user input of the user obtained via the user input interface.
[0245] (e.g. audio and / or visual user-expert communication allowing cooperation between user and expert)
[0246] According to a second aspect when referring back to the first aspect the remote device is user-controlled by an expert user and the portable device and the remote device are configured to enable a user communication (e.g. audio and / or video communication) between the user and expert user.
[0247] (user and expert inputs both contribute to the final result)
[0248] According to a third aspect when referring back to the first or second aspect, the remote device comprises a further user input interface for further user input of an expert user and the remote device (e.g. control system in description) and the portable device are configured to cooperate via the communication to determine the capturing positions (e.g. AK03) depending on the user input of the user obtained via the user input interface and the further user input of the expert user obtained via the further user input interface (e.g. because the user coarsely defines the Rol and the observation volume / surface such by means of moving the portable device to certain (e.g. predetermined) positions and selecting same using the input user interface, and defining, via these positions, one or more predefined volume primitive forms in order to define Rol and / or observation volume / surface, respectively, wherein the expert then adjusts and renders more precise the Rol and / or observation volume / surface, whereupon, based on the Rol and observation volume / surface, the capturing positions are determined at the remote device or the portable device).
[0249] (e.g. capturing position determination involves Rol which may be defined by user and / or expert input)
[0250] According to a fourth aspect when referring back to any one of the first to third aspects the remote device comprises a further user input interface ( e.g. key board, mouse, one or more buttons or any other computer-user-input interface) for further user input of the expert user and the cooperation via the communication to determine the capturing positions (e.g. AK03) comprises defining, via the user input interface and / or the further user input interface, geometric region of interest information indicating a region of interest (e.g. AT04) in the scene, with respect to which the scene model shall allow for the synthesis of the views, and using the geometric region of interest information for the determination of the capturing positions (e.g. AK03).
[0251] (e.g. capturing position determination involves observation volume / surface which may be defined by user and / or expert input)
[0252] According to a fifth aspect when referring back to any one of the first or third aspects, the remote device comprises a further user input interface (e.g. key board, mouse, one or more buttons or any other computer-user-input interface) for further user input of the expert user and the cooperation via the communication to determine the capturing positions (e.g. AK03) comprises defining, via the user input interface and / or the further user input interface, geometric observation information indicating an observation volume surface (e.g. AC02) of an observation volume within which viewpoints for the views are locatable, the synthesis of which the scene model shall allow, or the observation volume (e.g. AH07) itself, and using the geometric observation information for the determination of the capturing positions (e.g. AK03). (e.g. 3D scene model is generated at portable and / or remote device)
[0253] According to a sixth aspect when referring back to any one of the first to fourth aspects, the cooperation via the communication to determine the capturing positions (e.g. AK03) comprises generating, at the portable device and / or the remote device, a 3D scene model based on an image signal of the image capturing device.
[0254] (e.g. positional live tracking is performed)
[0255] According to a seventh aspect when referring back to any one of the first to sixth aspects, the portable device and / or the remote device is configured to, in real time, track a position of the image capturing device using one or more sensors (e.g. one or more inertia sensor and / or a triangulation sensors and / or a phase or time difference position sensor and / or travel time position sensor) of the portable device, and / or an evaluation of an image capturing device signal of the image capturing device.
[0256] (e.g. predetermined positions defined using user input may be used to define Rol using expert input)
[0257] According to an eight aspect when referring to any one of the first to seventh aspects, the remote device comprises a graphical 3D processing tool e.g. an application running on a computer of the remote device, the application allowing 3D modeling) and a further user input interface (e.g. key board, mouse, one or more buttons or any other computer-user- input interface) for further user input of an expert user and the cooperation via the communication to determine the capture positions (e.g. AK03) comprises generating, at the remote device, a 3D scene model based on an image signal of the image capturing device received by the remote device via the communication, receiving at the remote device information on predetermined positions (e.g. AH01 , AH02) obtained using the user input interface (e.g. pressing a button when the image capturing apparatus is at a position intended by the user, and / or using a touch screen such as a soft button, and / or using a microphone so as to, for instance, inform the remote expert when the image capturing apparatus is at a position intended by the user with the remote device having access to the live position of the image capturing apparatus either by performing a positional tracking pf the image capturing apparatus itself or by receiving the live position from the portable device via the communication, wherein the expert then actually determines the predetermined position in the 3D scene model via the further user input interface) , and defining (e.g. see sections 12.3 and 12.4 noting that the volume definition described there may likewise be used to define the Rol - the remote device uses the predetermined positions and the expert improves the region definition), using the further user input interface and by feeding the graphical 3D processing tool with the 3D scene model and the predetermined positions, geometric region of interest information indicating a region of interest (e.g. AT04) in the scene, with respect to which the scene model shall allow for the synthesis of the views, and using the geometric region of interest information for the determination (e.g. see sections 12.7 - the expert and / or a tool of the remote device automatically distributing the capture positions) of the capture positions (e.g. AK03).
[0258] (e.g. predetermined positions defined using user input may be used to define observation volume / surface using expert input)
[0259] According to a ninth aspect when referring back to any one of the first to eight aspects, the remote device comprises a graphical 3D processing tool e.g. an application running on a computer of the remote device, the application allowing 3D modeling) and a further user input interface (e.g. key board, mouse, one or more buttons or any other computer-user- input interface) for further user input of an expert user and the cooperation via the communication to determine the capture positions (e.g. AK03) comprises generating, at the remote device, a 3D scene model based on an image signal of the image capturing device received by the remote device via the communication, receiving at the remote device information on predetermined positions (e.g. AH01 , AH02) obtained using the user input interface (e.g. pressing a button when the image capturing apparatus is at a position intended by the user, and / or using a touch screen such as a soft button, and / or using a microphone so as to, for instance, inform the remote expert when the image capturing apparatus is at a position intended by the user with the remote device having access to the live position of the image capturing apparatus either by performing a positional tracking pf the image capturing apparatus itself or by receiving the live position from the portable device via the communication, wherein the expert then actually determines the predetermined position in the 3D scene model via the further user input interface) , and defining (e.g. see sections 12.3 and 12.4 - the remote device uses the predetermined positions and the expert improves the region definition), using the further user input interface and by feeding the graphical 3D processing tool with the 3D scene model and the predetermined positions, geometric observation information indicating an observation volume surface (e.g. AC02) of an observation volume within which viewpoints for the views are locatable, the synthesis of which the scene model shall allow, or the observation volume (e.g. AH07) itself, and using the geometric observation information for the determination (e.g. the expert and / or a tool of the remote device automatically distributing the capture positions) of the capture positions (e.g. AK03).
[0260] (e.g. 2D masks in images captured by the image capturing device from different viewpoints may be used to define Rol and Rol is used to determine capturing positions; here the masks are defined by the user)
[0261] According to a tenth aspect when referring back to any one of the first to ninth aspects, the user input interface of the portable device comprises a touch screen for user-based definition of 2D masks in images of the scene, captured by the image capturing device from different viewpoints, and the cooperation via the communication to determine the capture positions (e.g. AK03) comprises defining (e.g. see sections 12.7.4 and 12.7.5 - at the remote device or the portable device), using the 2D masks, geometric region of interest information indicating a region of interest (e.g. AT04) in the scene, with respect to which the scene model shall allow for the synthesis of the views, and using the geometric region of interest information for the determination (e.g. see sections 12.7 - the expert and / or a tool of the remote device automatically distributing the capture positions) of the capture positions (e.g. AK03).
[0262] (e.g. 2D masks in images captured by the image capturing device from different viewpoints may be used to define Rol and Rol is used to determine capturing positions; here the masks are defined by the expert)
[0263] According to an eleventh aspect when referring back to any one of the first to ninth aspects, the remote device comprises a further user input interface for further user input of an expert user and the cooperation via the communication to determine the capture positions (e.g. AK03) comprises defining 2D masks in images of the scene, captured by the image capturing device from different viewpoints, by means of the further user input device (e.g. see sections 13.5 - the expert draws the 2D mask outline on a screen portion showing the images) of the capture positions (e.g. AK03), defining (e.g. see sections 12.7.4 and 12.7.5 - at the remote device or the portable device), using the 2D masks, geometric region of interest information indicating a region of interest (e.g. AT04) in the scene, with respect to which the scene model shall allow for the synthesis of the views, and using the geometric region of interest information for the determination (e.g. see sections 12.7 - the expert and / or a tool of the remote device automatically distributing the capture positions) of the capture positions (e.g. AK03).
[0264] (e.g. predetermined positions assumed by the image capturing device (or portable device) may be used to define observation volume / surface and the latter is used to determine capturing positions; here the user determines the predetermined positions such as be button pressing; maybe guided by the expert using audio communication)
[0265] According to a twelfth aspect when referring back to any one of the first to eleventh aspects the cooperation via the communication to determine the capture positions (e.g. AK03) comprises defining predetermined positions by the user moving the image capturing device to the predetermined positions and using the user input interface, and using the predetermined positions, defining geometric observation information indicating an observation volume surface (e.g. AC02) of an observation volume within which viewpoints for the views are locatable, the synthesis of which the scene model shall allow, or the observation volume (e.g. AH07) itself, and using the geometric observation information for the determination (e.g. see sections 12.5 and 12.6 - the expert and / or a tool of the remote device automatically distributing the capture positions) of the capture positions (e.g. AK03).
[0266] (e.g. predetermined positions assumed by the image capturing device (or portable device) may be used to define observation volume / surface and the latter is used to determine capturing positions; here the user only moves the portable device (e.g. and the image capturing device, respectively) and the expert determines the predetermined positions such as be button pressing; maybe the movement by the user is guided by the expert using audio communication)
[0267] According to a thirteenth aspect when referring back to any one of the first to eleventh aspects, the remote device comprises a further user input interface for further user input of an expert user and the cooperation via the communication to determine the capture positions (e.g. AK03) comprises defining predetermined positions by the user moving the image capturing device to the predetermined positions and using the further user input interface, and using the predetermined positions, defining geometric observation information indicating an observation volume surface (e.g. AC02) of an observation volume within which viewpoints for the views are locatable, the synthesis of which the scene model shall allow, or the observation volume (e.g. AH07) itself, and using the geometric observation information for the determination (e.g. see sections 12.5 and 12.6 - the expert and / or a tool of the remote device automatically distributing the capture positions) of the capture positions (e.g. AK03).
[0268] (e.g. the user is guided to the determined capturing positions)
[0269] According to a fourteenth aspect when referring back to any one of the first to thirteenth aspects, the portable device is configured to present to the user hints on a screen (e.g. display) of the portable device which guide the user to move the portable device so that the image capturing device sequentially assumes the capture positions (e.g. AK03; e.g. the capture positions are sequentially processed to prompt the user with hints leading the user to move the image capturing device to the current position whereupon the process continues with a subsequent capture position (e.g. by presenting hints until the user moves the image capturing device to that subsequent position and so forth)).
[0270] (e.g. defining an observation volume / surface as a component of capturing position determination; the apparatus may be solely implemented in the portable device or parts thereof may be implemented on a remote device - no expert is needed)
[0271] According to fifteenth aspect there is provided an apparatus for providing geometric observation information as a component in determining capturing positions for capturing images of a scene suitable for generating a scene model based on the images which (e.g. scene model) allows for a synthesis of views of the scene from observation points other than the capturing positions, the apparatus comprising an image capturing device for capturing images of the scene; and a user input interface (e.g. button and / or touch screen and / or microphone), and configured to define predetermined positions by the user moving the image capturing device to the predetermined positions and using the user input interface, and using the predetermined positions, define as the geometric observation information an observation volume surface (e.g. AC02) of an observation volume within which viewpoints for the views are locatable, the synthesis of which the scene model shall allow, or the observation volume (e.g. AH07) itself.
[0272] (e.g. the own use the observation volume / surface for capturing position determination or sending same to a remote device for performing the capturing position determination)
[0273] According to a sixteenth aspect when referring back to the fifteenth aspect, the apparatus is configured to use the geometric observation information, determine (e.g. see sections 12.5 and 3.6 - the expert and / or a tool of the remote device automatically distributing the capture positions) the capture positions (e.g. AK03), or send the geometric observation information to a remote device for a determination (e.g. see sections 12.5 and 12.6 - the expert and / or a tool of the remote device automatically distributing the capture positions) of the capture positions (e.g. AK03) using the geometric observation information.
[0274] (e.g. defining an Rol as a component of capturing position determination)
[0275] According to a seventeenth aspect when referring back to any one of the fifteenth or sixteenth aspects, the user input interface comprises a touch screen for user-based definition of 2D masks in images of the scene, captured by the image capturing device from different viewpoints, and the apparatus is configured to define (e.g. see sections 12.7.4 and
[0276] 3.7.5 - at the remote device or the portable device), using the 2D masks, geometric region of interest information indicating a region of interest (e.g. AT04) in the scene, with respect to which the scene model shall allow for the synthesis of the views.
[0277] (The own use of the Rol for capturing position determination or sending same to a remote device for performing the capturing position determination)
[0278] According to an eighteenth aspect when referring back to the seventeenth aspect, the apparatus is configured to use the geometric region of interest information, determine (e.g. see sections 12.5 and 12.6 - the expert and / or a tool of the remote device automatically distributing the capture positions) the capture positions (e.g. AK03), or send the geometric region of interest information to a remote device for a determination (e.g. see sections 12.5 and 12.6 - the expert and / or a tool of the remote device automatically distributing the capture positions) of the capture positions (e.g. AK03) using the geometric region of interest information.
[0279] (e.g. the apparatus may itself perform the guidance to the capturing positions)
[0280] According to a nineteenth aspect when referring back to any one of the fifteenth to eighteenth aspects,, the apparatus is configured to present to the user hints on a screen (e.g. display) of the image capturing device which guide the user to move the image capturing device so that the image capturing device sequentially assumes the capturing positions (e.g. AK03; e.g. the capture positions are sequentially processed to prompt the user with hints leading the user to move the image capturing device to the current position whereupon the process continues with a subsequent capture position (e.g. by presenting hints until the user moves the image capturing device to that subsequent position and so forth)).
[0281] According to a twentieth aspect, there is provided a method for determining capturing positions for capturing images of a scene suitable for generating a scene model based on the images which (e.g. scene model) allows for a synthesis of views of the scene from observation points other than the capturing positions, the method being performed on a portable device (e.g. mobile phone with a corresponding application thereon), wherein the portable device is portable by a user and comprises an image capturing device for capturing images of the scene, a user input interface (e.g. button and / or touch screen and / or microphone), a user output interface (e.g. screen and / or loudspeaker), and a first communication interface (e.g. wireless connection module), and a remote device comprising a second communication interface for a communication of the remote device with the portable device via the first communication interface, wherein method comprises the remote device (e.g. control system in description) and the portable device cooperating via the communication to determine the capturing positions (e.g. AK03) depending on user input of the user obtained via the user input interface.
[0282] According to a twenty-first aspect, there is provided a method for providing geometric observation information as a component in determining capturing positions for capturing images of a scene suitable for generating a scene model based on the images which (e.g. scene model) allows for a synthesis of views of the scene from observation points other than the capturing positions, using an apparatus comprising an image capturing device for capturing images of the scene; and a user input interface (e.g. button and / or touch screen and / or microphone), the method comprising: defining predetermined positions by the user moving the image capturing device to the predetermined positions and using the user input interface, and using the predetermined positions, define as the geometric observation information an observation volume surface (e.g. AC02) of an observation volume within which viewpoints for the views are locatable, the synthesis of which the scene model shall allow, or the observation volume (e.g. AH07) itself. 15 alternatives:
[0283] Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a feature of a method step. Analogously, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps may be executed by (or using) a hardware apparatus, like for example, a microprocessor, a programmable computer or an electronic circuit. In some embodiments, one or more of the most important method steps may be executed by such an apparatus.
[0284] Depending on certain implementation requirements, embodiments of the invention can be implemented in hardware or in software. The implementation can be performed using a digital storage medium, for example a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium may be computer readable.
[0285] Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.
[0286] Generally, embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer. The program code may for example be stored on a machine readable carrier.
[0287] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.
[0288] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer. A further embodiment of the inventive methods is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein. The data carrier, the digital storage medium or the recorded medium are typically tangible and / or non- transitionary.
[0289] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals may for example be configured to be transferred via a data communication connection, for example via the Internet.
[0290] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.
[0291] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.
[0292] A further embodiment according to the invention comprises an apparatus or a system configured to transfer (for example, electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may, for example, be a computer, a mobile device, a memory device or the like. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.
[0293] In some embodiments, a programmable logic device (for example a field programmable gate array) may be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor in order to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware apparatus.
[0294] The apparatus described herein may be implemented using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
[0295] The apparatus described herein, or any components of the apparatus described herein, may be implemented at least partially in hardware and / or in software. The methods described herein may be performed using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
[0296] The methods described herein, or any components of the apparatus described herein, may be performed at least partially by hardware and / or by software.
[0297] The above described embodiments are merely illustrative for the principles of the present invention. It is understood that modifications and variations of the arrangements and the details described herein will be apparent to others skilled in the art. It is the intent, therefore, to be limited only by the scope of the impending patent claims and not by the specific details presented by way of description and explanation of the embodiments herein.
[0298] 16 References
[0299] [1] Y. Furukawa and C. Hernandez, Multi-view stereo: a tutorial, in Foundation and trends in computer graphics and vision, no. 9,1 / 2. Boston Delft: Now, 2015.
[0300] [2] B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, ‘NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis (arXiv)’. Mar. 19, 2020. doi: https: / / doi.org / 10.48550 / arXiv.2003.08934.
[0301] [3] B. Kerbl, G. Kopanas, T. Leimkuehler, and G. Drettakis, ‘3D Gaussian Splatting for Real-Time Radiance Field Rendering’, ACM Trans. Graph., vol. 42, no. 4, pp. 1-14, Aug. 2023, doi: 10.1145 / 3592433.
[0302] [4] P. Wang, L. Liu, Y. Liu, C. Theobalt, T. Komura, and W. Wang, ‘NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view Reconstruction’, arXiv, Jun. 2021 , Accessed: Jun. 23, 2021. [Online], Available: http: / / arxiv.org / abs / 2106.10689
[0303] [5] Trnio 3D Scanning Tutorial, (Aug. 17, 2020). Accessed: Mar. 21 , 2024. [Online Video], Available: https: / / youtu.be / 3C8WKxNTxbQ?feature=shared&t=45
[0304] [6] How to Scan in 3D with Display. land, (Mar. 19, 2020). Accessed: Mar. 21 , 2024.
[0305] [Online Video], Available: https: / / youtu.be / wep9AGwzPWY?feature=shared&t=13
[0306] [7] How To Use RealityScan, (Dec. 01 , 2022). Accessed: Mar. 21 , 2024. [Online Video], Available: https: / / youtu.be / HVkvHZCmVjU?feature=shared&t=86
[0307] [8] itSeez3d tutorial on face scanning with Structure Sensor or iSense, (Jun. 25, 2014).
[0308] Accessed: Mar. 21 , 2024. [Online Video], Available: https: / / youtu.be / aH J UgfWPhhQ?feature=shared&t=59 [9] Canvas - 3D Capture Spaces Within Minutes and Convert Them Into 3D Files, (Apr.
[0309] 26, 2022). Accessed: Mar. 21 , 2024. [Online Video], Available: https: / / youtu.be / jWACqA0CrM4?feature=shared&t=144
[0310]
[0010] C. Hoppe et al., ‘Online Feedback for Structure- from-Motion Image Acquisition’, in Procedings of the British Machine Vision Conference 2012, Surrey: British Machine Vision Association, 2012, p. 70.1-70.12. doi: 10 / ghjkj3.
[0311]
[0011] How to Capture & Share in 3D with Scaniversel, (Jan. 28, 2022). Accessed: Mar. 21 , 2024. [Online Video], Available: https: / / youtu.be / ELXLFMjLOx8?feature=shared&t=28
[0312]
[0012] D. Zhang et al., ‘Virtual Reality Aided High-Quality 3D Reconstruction by Remote Drones’, ACM Trans. Internet Technol., vol. 22, no. 1 , pp. 1-20, Feb. 2022, doi: 10.1145 / 3458930.
[0313]
[0013] F. Langguth and M. Goesele, ‘Guided Capturing of Multi-view Stereo Datasets’,
[0314] Eurographics 2013 - Short Papers, p. 4 pages, 2013, doi:
[0315] 10.2312 / CON F / EG2013 / SHORT / 093-096.
[0316]
[0014] How to use Qlone?, (Jun. 15, 2017). Accessed: Mar. 21 , 2024. [Online Video], Available: https: / / youtu.be / XkTaCOQ_Ojl?feature=shared&t=66
[0317]
[0015] Getting Started with Luma Al: Guided Capture Mode, (Nov. 04, 2022). Accessed: Mar.
[0318] 21 , 2024. [Online Video], Available: https: / / www.youtube.com / watch?v=JXb_a3ZIGnl
[0319]
[0016] D. Andersen, P. Villano, and V. Popescu, ‘AR HMD Guidance for Controlled Hand- Held 3D Acquisition’, IEEE Transactions on Visualization and Computer Graphics, vol. 25, no. 11 , pp. 3073-3082, Nov. 2019, doi: 10 / gpjddn.
[0320]
[0017] C. Birklbauer and O. Bimber, ‘Active guidance for light-field photography on smartphones’, Computers & Graphics, vol. 53, pp. 127-135, Dec. 2015, doi: 10 / f3s2ff.
[0321]
[0018] ‘Camera matrix’, Wikipedia. Jun. 28, 2023. Accessed: Mar. 20, 2024. [Online], Available: https: / / en. wikipedia. org / w / index.php?title=Camera_matrix&oldid=1162278810
[0322]
[0019] Y. Furukawa and C. Hernandez, Multi-view stereo: a tutorial, in Foundation and trends in computer graphics and vision, no. 9,1 / 2. Boston Delft: Now, 2015.
[0323]
[0020] B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, ‘NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis (arXiv)’. Mar. 19, 2020. doi: https: / / doi.org / 10.48550 / arXiv.2003.08934.
[0324]
[0021] B. Kerbl, G. Kopanas, T. Leimkuehler, and G. Drettakis, ‘3D Gaussian Splatting for Real-Time Radiance Field Rendering’, ACM Trans. Graph., vol. 42, no. 4, pp. 1-14, Aug. 2023, doi: 10.1145 / 3592433.
[0022] P. Wang, L. Liu, Y. Liu, C. Theobalt, T. Komura, and W. Wang, ‘NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view Reconstruction’, arXiv, Jun. 2021 , Accessed: Jun. 23, 2021. [Online], Available: http: / / arxiv.org / abs / 2106.10689
[0325]
[0023] Trnio 3D Scanning Tutorial, (Aug. 17, 2020). Accessed: Mar. 21 , 2024. [Online Video], Available: https: / / youtu.be / 3C8WKxNTxbQ?feature=shared&t=45
[0326]
[0024] ‘Learn. Display. Land’. Accessed: Jul. 20, 2020. [Online], Available: https: / / learn.display.land /
[0327]
[0025] How To Use RealityScan, (Dec. 01 , 2022). Accessed: Mar. 21 , 2024. [Online Video], Available: https: / / youtu.be / HVkvHZCmVjU?feature=shared&t=86
[0328]
[0026] itSeez3d tutorial on face scanning with Structure Sensor or iSense, (Jun. 25, 2014).
[0329] Accessed: Mar. 21 , 2024. [Online Video], Available: https: / / youtu.be / aH J UgfWPhhQ?feature=shared&t=59
[0330]
[0027] Canvas - 3D Capture Spaces Within Minutes and Convert Them Into 3D Files, (Apr.
[0331] 26, 2022). Accessed: Mar. 21 , 2024. [Online Video], Available: https: / / youtu.be / jWACqA0CrM4?feature=shared&t=144
[0332]
[0028] How to Capture & Share in 3D with Scaniversel, (Jan. 28, 2022). Accessed: Mar. 21 , 2024. [Online Video], Available: https: / / youtu.be / ELXLFMjLOx8?feature=shared&t=28
[0333]
[0029] How to use Qlone?, (Jun. 15, 2017). Accessed: Mar. 21 , 2024. [Online Video], Available: https: / / youtu.be / XkTaCOQ_Ojl?feature=shared&t=66
[0334]
[0030] Getting Started with Luma Al: Guided Capture Mode, (Nov. 04, 2022). Accessed: Mar.
[0335] 21 , 2024. [Online Video], Available: https: / / www.youtube.com / watch?v=JXb_a3ZIGnl
[0336]
[0031] X. Pan, Z. Lai, S. Song, and G. Huang, ‘ActiveNeRF: Learning where to See with Uncertainty Estimation’. arXiv, Sep. 18, 2022. Accessed: Sep. 26, 2022. [Online], Available: http: / / arxiv.org / abs / 2209.08546
[0337]
[0032] N. Sunderhauf, J. Abou-Chakra, and D. Miller, ‘Density-aware NeRF Ensembles: Quantifying Predictive Uncertainty in Neural Radiance Fields’. arXiv, Sep. 18, 2022. Accessed: Oct. 04, 2022. [Online], Available: http: / / arxiv.org / abs / 2209.08718
[0338]
[0033] A. Schafer, G. Reis, and D. Stricker, ‘A Survey on Synchronous Augmented, Virtual, andMixed Reality Remote Collaboration Systems’, ACM Comput. Surv., vol. 55, no. 6, pp. 1-27, Jul. 2023, doi: 10.1145 / 3533376.
[0339]
[0034] A. Aschauer, I. Reisner-Kollmann, and J. Wolfartsberger, ‘Creating an Open-Source
[0340] Augmented Reality Remote Support Tool for Industry: Challenges and Learnings’, Procedia Computer Science, vol. 180, pp. 269-279, 2021 , doi:
[0341] 10.1016 / j. procs.2021.01.164.
[0035] M. Levoy and P. Hanrahan, ‘Light field rendering’, in Proceedings of the 23rd annual conference on Computer graphics and interactive techniques - SIGGRAPH ’96, Not Known: ACM Press, 1996, pp. 31-42. doi: 10.1145 / 237170.237199.
[0342]
[0036] T. Herfet, K. Chelli, and M. Le Pendu, ‘Chapter 7 - Light field representation: The dimensions in light fields’, in Immersive Video Technologies, G. Valenzise, M. Alain, E. Zerman, and C. Ozcinar, Eds., Academic Press, 2023, pp. 173-199. doi: 10.1016 / B978-0-32-391755-1 .00013-4.
[0343]
[0037] A. Davis, M. Levoy, and F. Durand, ‘Unstructured Light Fields’, Computer Graphics Forum, vol. 31 , no. 2pt1 , pp. 305-314, 2012, doi: 10.1111 / j.1467-8659.2012.03009.x.
[0344]
[0038] C. Birklbauer and O. Bimber, ‘Active guidance for light-field photography on smartphones’, Computers & Graphics, vol. 53, pp. 127-135, Dec. 2015, doi: 10 / f3s2ff.
[0345]
[0039] ‘Plenoptic Sampling’, presented at the SIGGRAPH, New Orleans, LA, USA, 2000.
[0346]
[0040] ‘Camera matrix’, Wikipedia. Jun. 28, 2023. Accessed: Mar. 20, 2024. [Online], Available: https: / / en. wikipedia. org / w / index.php?title=Camera_matrix&oldid=1162278810
[0347]
[0041] A. Macario Barros, M. Michel, Y. Moline, G. Corre, and F. Carrel, ‘A Comprehensive Survey of Visual SLAM Algorithms’, Robotics, vol. 11 , no. 1 , p. 24, Feb. 2022, doi: 10 / gpjddp.
[0348]
[0042] J. P. M. Covolan, A. C. Sementille, and S. R. R. Sanches, ‘A mapping of visual SLAM algorithms and their applications in augmented reality’, in 2020 22nd Symposium on Virtual and Augmented Reality (SVR), Nov. 2020, pp. 20-29. doi: 10 / gppvnj.
[0349]
[0043] Z. Teed and J. Deng, ‘DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D Cameras’. arXiv, Feb. 02, 2022. doi: 10.48550 / arXiv.2108.10869.
[0350]
[0044] G. Younes, D. Asmar, E. Shammas, and J. Zelek, ‘Keyframe-based monocular SLAM: design, survey, and future directions’, Robotics and Autonomous Systems, vol. 98, pp. 67-88, Dec. 2017, doi: 10 / gck36z.
[0351]
[0045] J. Engel, T. Schdps, and D. Cremers, ‘LSD-SLAM: Large-scale direct monocular SLAM’, in European conference on computer vision, 2014, pp. 834-849.
[0352]
[0046] Z. Yang and D. Shi, ‘Mapping Technology in Visual SLAM: A Review’, in Proceedings of the 2018 2nd International Conference on Computer Science and Artificial Intelligence, in CSAI ’18. New York, NY, USA: Association for Computing Machinery, Dezember 2018, pp. 291-295. doi: 10 / gpd2cj.
[0353]
[0047] R. Li, ‘Ongoing Evolution of Visual SLAM from Geometry to Deep Learning: Challenges and Opportunities’, Cogn Comput, p. 15, 2018, doi: 10 / ghjpcg.
[0048] R. Mur-Artal, J. M. M. Montiel, and J. D. Tardos, ‘ORB-SLAM: A Versatile and Accurate Monocular SLAM System’, IEEE Transactions on Robotics, vol. 31 , no. 5, pp. 1147— 1163, Oct. 2015, doi: 10 / f7xpqn.
[0354]
[0049] R. Mur-Artal and J. D. Tardos, ‘ORB-SLAM2: an Open-Source SLAM System for Monocular, Stereo and RGB-D Cameras (arXiv)’, IEEE Trans. Robot., vol. 33, no. 5, pp. 1255-1262, Oct. 2017, doi: 10 / gdz797.
[0355]
[0050] C. Campos, R. Elvira, J. J. G. Rodriguez, J. M. M. Montiel, and J. D. Tardos, ‘ORBSLAMS: An Accurate Open-Source Library for Visual, Visual-Inertial, and Multimap SLAM’, IEEE Transactions on Robotics, vol. 37, no. 6, pp. 1874-1890, Dec. 2021 , doi: 10 / gkzqvv.
[0356]
[0051] ‘Global Positioning System’, Wikipedia. Mar. 13, 2024. Accessed: Mar. 25, 2024.
[0357] [Online], Available: https: / / en.wikipedia.Org / w / index. php?title=Global_Positioning_System&oldid=121356 7746
[0358]
[0052] ‘Inertial measurement unit’, Wikipedia. Mar. 10, 2024. Accessed: Mar. 25, 2024.
[0359] [Online], Available: https: / / en.wikipedia.org / w / index.php?title=lnertial_measurement_unit&oldid=1213033 229
Claims
Claims1. Apparatus for presenting guidance information on a screen (AB02) of an image capturing device (AB01), the guidance information guiding a user to move the image capturing device (AB01) to a target location and target orientation for capturing an object (AA01), the apparatus configured to visualize a static target (AF01) on the screen (AB02), the static target (AF01) being positioned within the screen (AB02) on a first screen (AB02) position which corresponds to, via a predetermined mapping between the screen (AB02) and an image plane (AE02) of the image capturing device (AB01), a first image plane position within the image plane (AE02), onto which, when the image capturing device (AB01) is positioned at the target location and oriented at the target orientation, a virtual target (AG06) is projected, and visualize a moveable target (AF02) on the screen (AB02), the moveable target (AF02) being positioned within the screen (AB02) on a second screen position which corresponds to, via the predetermined mapping between the screen (AB02) and the image plane (AE02) of the image capturing device (AB01), a second image plane position within the image plane (AE02), onto which, according to a current position of the image capturing device (AB01), the virtual target (AG06) is projected, the static target (AF01) and the moveable target (AF02) forming the guidance information.
2. Apparatus of claim 1 , configured to visualize the moveable target (AF01) and the moveable target (AF02) as mutually geometrically similar polygons.
3. Apparatus of claim 2, the mutually geometrically similar polygons are rectangles.
4. Apparatus of any of claims 1 to 3, configured to compute the second screen (AB02) position by projecting the virtual target (AG06) onto the image plane (AE02) using a current location and current orientation of the image capturing device (AB01).
5. Apparatus of any of claims 1 to 3, configured to visualize the static target (AF01) and the moveable target (AF02) as mutually geometrically similar polygons, and compute the second screen (AB02) position such that same approximates an exact screen (AB02)position corresponding to an exact image plane position onto which the virtual target (AG06) is projected in case of the current orientation of the image capturing device (AB01) experiencing a deviation from the target orientation.
6. Apparatus of any of claims 1 to 3, configured to compute the second screen (AB02) position by projecting the virtual target (AG06) onto the image plane (AE02) using a current location of the image capturing device (AB01) and the target orientation.
7. Apparatus of any of claims 1 to 6, wherein the virtual target (AG06) is located between the object (AA01) and the target location or, relative to the target location, behind the object (AA01).
8. Apparatus of any of claims 1 to 7, wherein the virtual target (AG06) is planar and located in a reference plane which is perpendicular to the target orientation.
9. Apparatus of any of claims 1 to 8, configured to, for the visualizing the moveable target (AF02), adapt a distance between the virtual target (AG06) and the target location depending on a current location and current orientation of the image capturing device (AB01) so that the movable target remains, according to predetermined mapping, within a portion (AB03) of the screen (AB02) onto which a field of view of the image capturing device (AB01) is mapped.
10. Apparatus of any of claims 1 to 9, configured to, for the visualizing the moveable target (AF02), adapt a distance between the virtual target (AG06) and the target location depending on a length of an actual distance between the current location and target location.
11. Apparatus of any of claims 1 to 10, configured to continuously determine whether the static target (AF01) and the moveable target (AF02) are closer to a state of exact mutual overlay than a predetermined threshold and to provide the user with a target-reached signal depending on whether the static target (AF01) and the moveable target (AF02) are closer to the state of exact mutual overlay than the predetermined threshold.
12. Apparatus of any of claims 1 to 11 , configured to continuously determine whether the static target (AF01) and the moveable target (AF02) are closer to a state of exact mutual overlay than a predetermined threshold and to automatically initiate an image capturingdepending on whether the static target (AF01) and the moveable target (AF02) are closer to a state of exact mutual overlay than a predetermined threshold.
13. Apparatus of any of claims 1 to 11 , configured to continue with presenting further guidance information on the screen (AB02) of the image capturing device (AB01) responsive to the user having moved the image capturing device (AB01) to be located at the target location and oriented at the target orientation, the further guidance information relating to a next pair of target location and target orientation.
14. Apparatus of any of claims 1 to 13, configured to track the current positions in terms of current location and current orientation by means of one or more sensors of the image capturing device (AB01), and / or an evaluation of an image capturing device signal of the image capturing device (AB01).
15. Apparatus of any of claims 1 to 14, wherein the apparatus is housed in a housing along with the image capturing device (AB01), or the apparatus is communicationally connectable to the image capturing device (AB01).
16. Apparatus of any of claims 1 to 15, wherein the image capturing device (AB01) and the screen (AB02) are attached to each other.
17. Apparatus of any of claims 1 to 16, wherein the image capturing device (AB01) and the screen (AB02) are attached to each other so that the screen (AB02) is faced towards an opposite side relative to the image capturing device (AB01).
18. Apparatus of any of claims 1 to 17, wherein the apparatus is configured to visualize the static target (AF01) and the moveable target (AF02) on the screen (AB02) in a manner overplayed onto a currently captured image, currently captured by the imagecapturing device (AB01), and in form of a circumference of the static target (AF01) and the moveable target (AF02).
19. Apparatus of any of claims 1 to 18, wherein the apparatus is further configured to, visualize, as additional component of the guidance information, a static rotational target (AC01) on the screen (AB02), the static rotational target (AC01) being positioned within the screen (AB02) on a third screen (AB02) position which corresponds to, via the predetermined mapping between the screen (AB02) and an image plane (AE02) of the image capturing device (AB01), a third image plane position within the image plane (AE02), onto which, when the image capturing device (AB01) is oriented at the target orientation, all points in front of the image capturing device (AB01) on a target axis leading through the image capturing device (AB01) are projected, and visualize a moveable rotational target (AC02) on the screen (AB02), the moveable rotational target (AC02) being positioned within the screen (AB02) on a fourth screen (AB02) position which corresponds to, via the predetermined mapping between the screen (AB02) and the image plane (AE02) of the image capturing device (AB01), a fourth image plane position within the image plane (AE02), onto which, according to the current orientation of the image capturing device (AB01), all points on the target axis are projected, or a line or wedge connecting the third and fourth position.
20. Apparatus of any of claims 1 to 19, wherein the apparatus is configured to, in visualizing the movable target, displace the virtual target (AG06) along an axis leading through the virtual target (AG06) and the target location to a displaced virtual target (AJ06) location so that a distance between the displaced virtual target (AJ06) location and the current location becomes equal to a distance between a non-displaced virtual target location of the virtual target (AG06) and the target location.
21. Apparatus of any of claims 1 to 20, wherein the apparatus is configured to, in visualizing the movable target,displace the virtual target (AG06) along an axis leading through the virtual target (AG06) and the target location to a displaced virtual target (AJ06) location so that a distance between the displaced virtual target (AJ06) location and the current location becomes equal to a maximum of a distance between a non-displaced virtual target location of the virtual target (AG06) and the target location and the distance between the non-displaced virtual target location of the virtual target (AG06) and the current location minus a predetermined maximum distance on if the displaced virtual target (AJ06) location gets farer away from the object (AA01).
22. Apparatus of any of claims 1 to 21 , wherein the apparatus is configured to, in visualizing the movable target, displace the virtual target (AG06) along an axis leading through the virtual target (AG06) and the target location to a displaced virtual target (AJ06) location so that a distance between the displaced virtual target (AJ06) location and the current location becomes equal to a minimum of a distance between a non-displaced virtual target location of the virtual target (AG06) and the target location and a sum of the distance between the non-displaced virtual target location of the virtual target (AG06) and the current location on the one hand and a predetermined maximum distance on the other hand if the displaced virtual target (AJ06) location gets closer to the object (AA01).
23. Apparatus of any of the claims 1 to 22, wherein the apparatus is configured to provide geometric observation information as a component in determining capturing positions (AI02, AM02, AN03a, AN03b, DH01 , DH02, AR03) for capturing an image of a scene suitable for generating a scene model based on the image which allows for a synthesis of views of the scene from observation points other than the capturing positions (AI02, AM02, AN03a, AN03b, DH01 , DH02, AR03), wherein the image capturing device comprises a user input interface, wherein the apparatus is configured to define predetermined positions by the user moving the image capturing device to the predetermined positions and using the user input interface, andusing the predetermined positions, define as the geometric observation information an observation volume surface (AC02, AH04 AH06, AH05, AG01 , AD02) of an observation volume (AC03, AF02, AH09, AH07) within which viewpoints for the views are locatable, the synthesis of which the scene model shall allow, or the observation volume (AC03, AF02, AH09, AH07) itself.
24. Apparatus of claim 23, configured to use the geometric observation information, determine the capture positions (AI02, AM02, AN03a, AN03b, DH01 , DH02, AR03), or send the geometric observation information to a remote device for a determination of the capture positions (AI02, AM02, AN03a, AN03b, DH01 , DH02, AR03) using the geometric observation information.
25. Apparatus of claim 23 or 24, wherein the user input interface comprises a touch screen for user-based definition of 2D masks in images of the scene, captured by the image capturing device from different viewpoints, and the apparatus is configured to define, using the 2D masks, geometric region of interest information indicating a region of interest (AT04) in the scene, with respect to which the scene model shall allow for the synthesis of the views.
26. Apparatus of claim 25, configured to use the geometric region of interest information, determine the capture positions (AI02, AM02, AN03a, AN03b, DH01 , DH02, AR03), or send the geometric region of interest information to a remote device for a determination of the capture positions (AI02, AM02, AN03a, AN03b, DH01 , DH02, AR03) using the geometric region of interest information.
27. Apparatus of any of claims 23 to 26, configured to present to the user guidance on the screen of the image capturing device which guide the user to move the image capturingdevice so that the image capturing device sequentially assumes the capturing positions (AI02, AM02, AN03a, AN03b, DH01 , DH02, AR03).
28. Method (MV45) for presenting guidance information on a screen (AB02) of an image capturing device (AB01), the guidance information guiding a user to move the image capturing device (AB01) to a target location and target orientation for capturing an object (AA01), the method comprising: visualizing (MV01) a static target (AF01) on the screen (AB02), the static target (AF01) being positioned within the screen (AB02) on a first screen (AB02) position which corresponds to, via a predetermined mapping between the screen (AB02) and an image plane (AE02) of the image capturing device (AB01), a first image plane position within the image plane (AE02), onto which, when the image capturing device (AB01) is positioned at the target location and oriented at the target orientation, a virtual target (AG06) is projected, and visualizing (MV02) a moveable target (AF02) on the screen (AB02), the moveable target (AF02) being positioned within the screen (AB02) on a second screen (AB02) position which corresponds to, via the predetermined mapping between the screen (AB02) and the image plane (AE02) of the image capturing device (AB01), a second image plane position within the image plane (AE02), onto which, according to a current position of the image capturing device (AB01), the target is projected, the static target (AF01) and the moveable target (AF02) forming the guidance information.
29. A computer program for performing the method according to claim 28 when the computer program runs on a computer.