Surface heading estimation for points in images
The method estimates surface headings in 2D images by projecting 3D models onto 2D planes, addressing human error and resource inefficiencies, enabling accurate and efficient remote site inspection and planning.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
- Filing Date
- 2025-01-27
- Publication Date
- 2026-07-30
AI Technical Summary
Existing methods for determining the geospatial heading of objects in 2D images are prone to human error, rely on incorrect assumptions about site plans, and require substantial computational resources, making them inefficient and inaccurate.
A method and device for estimating the surface heading of a user-selected point in a 2D image by projecting a 3D model onto the 2D image plane, calculating a surface heading vector in local coordinates, and converting it to geospatial coordinates, without relying on human judgment or full 3D BIM models.
Provides accurate and efficient estimation of geospatial headings, reducing computational and bandwidth requirements, and enabling lightweight implementations suitable for user devices, facilitating remote site inspection and planning.
Smart Images

Figure EP2025051905_30072026_PF_FP_ABST
Abstract
Description
[0001] SURFACE HEADING ESTIMATION FOR POINTS IN IMAGES
[0002] TECHNICAL FIELD
[0003] Embodiments presented herein relate to a method, an image processing device, a computer program, and a computer program product for estimating surface heading of a user-selected point within a two-dimensional image.
[0004] BACKGROUND
[0005] As an introductory example, digital twin (DT) portals can be used to collect raw scans (in terms of two-dimensional (2D) images), of real-world sites and generate different representations of them. Many scans are obtained by guiding scanner device to capture a plurality of images, defining an image set, along one or more orbits around a scene of interest. For example, in the telecommunication industry, cell site installations, ranging from large radio towers to telecommunication equipment mounted on building rooftops and / or walls, can be scanned and the resulting images can be uploaded to a DT portal, so that a three-dimensional (3D) pointcloud, spin and / or SkyBox views, business information modelling (BIM) models of the scene of interest, etc. can be generated. Other examples of scenes of interest are power equipment sites, industry sites, cultural heritage sites, just to mention a few. These scans may complement non-scan documentation, such as architectural plans, Radio Frequency Datasheets (RFDS), etc. in creating a complete knowledge base of the real-world site.
[0006] Scanner devices can in this way be used for scanning a variety of types of sites.
[0007] Unmanned Aerial Vehicles (UAVs), also known as drones, as well as camera-equipped smartphones are commonly used as scanning devices. These types of scanning devices may embed geolocation - commonly represented by global positioning system (GPS) coordinates - as image metadata. From a set of images, representing a scan of the site, a 3D structure of the site can be reconstructed via techniques such as Structure-from-Motion or neural volumetric representations, such as Neural Radiance Fields (NeRFs). From the metadata and the 3D reconstruction, an overall heading (Azimuth) of each image can also be estimated.
[0008] Often, the users of DT portals need high-resolution, high-fidelity reference images to facilitate remote site inspection, acceptance, and / or planning, to reference whenanalyzing positioning algorithm behavior in real-world deployments, to facilitate creation of BIM models from 3D pointclouds, etc. For example, for some infrastructure sites, objects (such as equipment, furniture), but also infrastructure parts (such as walls, ceilings, floors, or even whole buildings) must be correctly geolocated in both position (e.g., latitude, longitude, altitude) and geospatial heading (e.g., azimuth, elevation).
[0009] Image-based viewers showing photos of a site (e.g., sequential viewers or SkyBox rendered environments) allow the user to inspect sites and perform point-to-point measurements based on the underlying 3D model. Such viewers also can indicate an overall geospatial heading associated with a particular image (e.g. image 1 facing due North, image 2 facing 18 degrees North-West, etc.).
[0010] However, for some site deployments and / or for planning, or validating, some site deployments, it is relevant to retrieve the geospatial heading of a particular piece of equipment (e.g., an antenna) or infrastructure piece (e.g., a building wall) seen in the image. In the context of DT portals, this implies that the user would need to (i) manually find an image that seems to be facing in the same direction as the object (or surface) of interest, or (ii) reference some site plan, or (iii) load and interpret a full geospatially-referenced 3D BIM model of the site. All these alternatives come with disadvantages. Alternative (i) is based on human approximation, and hence on human judgement, and thus prone to error. Alternative (ii) is based not only on the assumption that the site plan is correct, but also on the presumption that the real-world site has not been altered with respect to the site plan, which is not always the case. The full geospatially-referenced 3D BIM model in alternative (iii) may be complicated to view for non-expert users. Generating the full geospatially-referenced 3D BIM model also requires substantial bandwidth, time, and computational power.
[0011] SUMMARY
[0012] An object of embodiments described herein is to address the above issues in terms of retrieving the geospatial heading of objects in 2D images.
[0013] According to a first aspect there is presented a method for estimating surface heading of a user-selected point within a 2D image. The method is performed by an image processing device. The method comprises obtaining the user-selected point in the 2Dimage. The 2D image is depicting a scanned scene, is associated with a 2D image plane, and is provided in a geospatial coordinate space. The method comprises obtaining a 3D model of the scanned scene. The 3D model is composed of 3D points and is provided in a local coordinate space. The method comprises projecting the 3D model onto the 2D image plane. The method comprises estimating a surface heading vector in the local coordinate space for the user-selected point as a normal to a local surface given by depth values of the user-selected point and 3D points neighbouring the user-selected point. The method comprises estimating the surface heading by converting the surface heading vector from the local coordinate space to the geospatial coordinate space.
[0014] According to a second aspect there is presented an image processing device for estimating surface heading of a user-selected point within a 2D image. The image processing device comprises processing circuitry. The processing circuitry is configured to cause the image processing device to obtain the user-selected point in the 2D image. The 2D image is depicting a scanned scene, is associated with a 2D image plane, and is provided in a geospatial coordinate space. The processing circuitry is configured to cause the image processing device to obtain a 3D model of the scanned scene. The 3D model is composed of 3D points and is provided in a local coordinate space. The processing circuitry is configured to cause the image processing device to project the 3D model onto the 2D image plane. The processing circuitry is configured to cause the image processing device to estimate a surface heading vector in the local coordinate space for the user-selected point as a normal to a local surface given by depth values of the user-selected point and 3D points neighbouring the user-selected point. The processing circuitry is configured to cause the image processing device to estimate the surface heading by converting the surface heading vector from the local coordinate space to the geospatial coordinate space.
[0015] According to a third aspect there is presented a computer program for estimating surface heading of a user-selected point within a 2D image. The computer program comprises computer code which, when run on processing circuitry of an image processing device, causes the image processing device to perform actions. One action comprises the image processing device to obtain the user-selected point in the 2D image. The 2D image is depicting a scanned scene, is associated with a 2D image plane, and is provided in a geospatial coordinate space. One action comprises theimage processing device to obtain a 3D model of the scanned scene. The 3D model is composed of 3D points and is provided in a local coordinate space. One action comprises the image processing device to project the 3D model onto the 2D image plane. One action comprises the image processing device to estimate a surface heading vector in the local coordinate space for the user-selected point as a normal to a local surface given by depth values of the user-selected point and 3D points neighbouring the user-selected point. One action comprises the image processing device to estimate the surface heading by converting the surface heading vector from the local coordinate space to the geospatial coordinate space.
[0016] According to a fourth aspect there is presented a computer program product comprising a computer program according to the third aspect and a computer readable storage medium on which the computer program is stored. The computer readable storage medium could be a non-transitory computer readable storage medium.
[0017] Advantageously, these aspects do not suffer from the above issues.
[0018] Advantageously, these aspects are not based on human approximation, and thus are not prone to error.
[0019] Advantageously, these aspects are neither based on the assumption of correct site plans nor on the presumption that the real-world site has not been altered with respect to the site plan.
[0020] Advantageously, these aspects are not based on displaying, or even generating, any full geospatially-referenced 3D BIM model, thus reducing bandwidth, time, and computational power.
[0021] Advantageously, these aspects enable a lightweight implementation suitable for user devices, such as laptops, etc. (or at least reduce the bandwidth for streaming data needed for the computation of the surface heading).
[0022] Advantageously, these aspects enable verification of site deployment, in terms of estimating surface headings of different types of objects in a scanned scene, in turn enabling accurate remote site inspection, site and radio signal modelling, site acceptance, and equipment deployment planning.Advantageously, these aspects can be used as a complement, or reference, to site plans, RFDSs, etc.
[0023] Advantageously, by providing the surface heading vector in the geospatial coordinate space to an image processing application, these aspects enable the surface heading to be provided to the user in an easily understandable way from real-world site scans.
[0024] Other objectives, features and advantages of the enclosed embodiments will be apparent from the following detailed disclosure, from the attached dependent claims as well as from the drawings.
[0025] Generally, all terms used in the claims are to be interpreted according to their ordinary meaning in the technical field, unless explicitly defined otherwise herein. All references to "a / an / the element, apparatus, component, means, module, step, etc." are to be interpreted openly as referring to at least one instance of the element, apparatus, component, means, module, step, etc., unless explicitly stated otherwise. The steps of any method disclosed herein do not have to be performed in the exact order disclosed, unless explicitly stated.
[0026] BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The inventive concept is now described, by way of example, with reference to the accompanying drawings, in which:
[0028] Fig. 1 is a block diagram of an image processing device according to embodiments;
[0029] Fig. 2 schematically illustrates user interface views according to an embodiment;
[0030] Fig. 3 is a flowchart of methods according to embodiments;
[0031] Figs. 4, 5, and 6 are block diagrams of an image processing device according to embodiments;
[0032] Fig. 7 is a schematic illustration of the relation between a 3D model and a 2D image according to an embodiment;
[0033] Fig. 8 is a schematic diagram showing structural units of an image processing device according to an embodiment; andFig. 9 shows one example of a computer program product comprising computer readable storage medium according to an embodiment.
[0034] DETAILED DESCRIPTION
[0035] The inventive concept will now be described more fully hereinafter with reference to the accompanying drawings, in which certain embodiments of the inventive concept are shown. This inventive concept may, however, be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided by way of example so that this disclosure will be thorough and complete, and will fully convey the scope of the inventive concept to those skilled in the art. Like numbers refer to like elements throughout the description. Any step or feature illustrated by dashed lines should be regarded as optional.
[0036] Fig. i is a block diagram of an image processing device too configured for estimating surface heading of a user-selected point within a 2D image according to an embodiment. Dashed-border areas indicate groupings of steps, or actions, that can be repeated in a loop, depending on user interaction. User-interaction is shown in rounded solid-line rectangles, computations are shown in sharp-corner solid-line rectangles. Arrows indicate process flow. The image processing device too is configured to receive as input an image 1112 depicting a scanned site or a part thereof, as shown to the user of an image-based viewer, and as indicated via user selection in block 110. The image can be a color image, a grayscale image, a thermal image or any other type of 2D image. The image processing device 100 is further configured to receive as input a 3D model, or geometry, £1 (e.g., a pointcloud), depicting the same scanned site as shown, partly of fully, in the image 1. The image processing device 100 is further configured to receive as input the position t, and rotation / ?, of the image 1 in the coordinate system of the 3D model £1. The image processing device 100 is further configured to receive as input an intrinsic matrix K, depicting the camera properties of image 1. As an example, the intrinsic matrix / can be given as:
[0037]
[0038] which is often visualized as the camera frustum. Here, fx, fyare the horizontal and vertical focal lengths, respectively, s is the skew factor, and cx, cyare the horizontal and vertical principal points, respectively. The parameters K, R, t may be obtained in the same process (e.g., SfM) as used to reconstruct fl from a set of images, such as set of images including the image 1. However, the actual generation of the 3D model, or the initial alignment between the 3D model and the 2D images are outside the scope of this disclosure. The image processing device 100 is further configured to receive as input a geolocated position gpsz= [Lat;, Lon;, Alt;] and a geolocated heading ; = [Azimuth , Elevation ] of the image I. These can be obtained from image metadata and / or be calculated using other means. Specific methods for obtaining a geolocated heading and position for the image 1 are, however, outside the scope of this disclosure. In block 114 the 3D geometry is projected to the image plane of the selected image 1, creating a depthmap D, 116. The image processing device 100 is further configured to receive as input a user-designated point pu, via block 118, that lies in the image I. Depths for all points in the image near, or surrounding, the user-designated point puare retrieved from £>; in block 120. In block 122 a surface heading vector is estimated in the local coordinate space for the user-selected point pu. In block 124 the surface heading vector is converted to the geospatial coordinate space. The image processing device 100 is configured to provide as output a geolocated heading hu= [Azimuth,,, Elevation,,] that, for example, can be displayed to the user, via block 126, in relation to pu.
[0039] In Fig. 2 is provided a high-level illustration of user-interaction. At (a), an image 200a with some object 210 (e.g., communications tower) is displayed to a user in an image-based viewer. At (b), the user selects a point 220 on the image 200a (e.g., main antenna face). At (c), a geolocated heading 230 of the object’s surface at the point of user selection is rendered on top of the image at the point of user selection, and an alphanumeric description of the heading is provided in a text box 240 in the image 200c.
[0040] Fig. 3 is a flowchart illustrating embodiments of methods for estimating surface heading of a user-selected point within a 2D image. The methods are performed by the image processing device 100 described with reference to the block diagram in Fig.
[0041] 1. Some alternative implementations of the image processing device for implementingthe methods will be disclosed below. The methods are advantageously provided as computer programs 920.
[0042] S102: The image processing device 100 obtains the user-selected point puin the 2D image. The 2D image is depicting a scanned scene, is associated with a 2D image plane, and is provided in a geospatial coordinate space.
[0043] S104: The image processing device 100 obtains a 3D model, or geometry, of the scanned scene. The 3D model is composed of 3D points and is provided in a local coordinate space.
[0044] S106: The image processing device 100 projects the 3D model onto the 2D image plane.
[0045] S108: The image processing device 100 estimates a surface heading vector in the local coordinate space for the user-selected point puas a normal to a local surface given by depth values of the user-selected point and 3D points neighbouring the user-selected point. This local surface may be a planar surface, or even a spheroid surface, or a saddle-shaped surface. In any case, and as will be disclosed in further detail below, a fitted plane can here be used. That is, the plane being the closest -fit approximation of the local surface given by the depth values.
[0046] S110: The image processing device 100 estimates the surface heading by converting the surface heading vector from the local coordinate space to the geospatial coordinate space.
[0047] Embodiments relating to further details of estimating surface heading of a user-selected point within a 2D image as performed by the image processing device 100 will now be disclosed with continued reference to Fig. 3.
[0048] As in Fig. 1 and Fig. 2, the heading hucan be indicated to the user as alphanumeric supplement or as a graphical overlay on the image 1 in the image-based viewer.
[0049] Hence, in some embodiments, the image processing device 100 is configured to perform (optional) step S112.
[0050] S112: The image processing device 100 provides the surface heading vector huin the geospatial coordinate space to an image processing application.The image processing application may be any of: a digital twin application, a 3D modelling application, a Business Information Modelling (BIM) application, a computer aided design (CAD) application.
[0051] Further aspects of how the 3D model can be projected onto the 2D image plane will be disclosed next.
[0052] In some embodiments, the 3D model is projected onto the 2D image plane as a depthmap D, of same width and height as the 2D image. In this respect, an empty “image” of same width and height as 1 can be generated as part of creating a 2D depthmap
[0053]
[0054] As the user selects new points puin the image 1, the depthmap D, can be reused and does not need to be re-estimated. When the user selects a new image 1, a new depthmap D, needs to be generated.
[0055] Further, the 3D model fl may be projected onto the 2D image plane by a 3D-to-2D projection of each 3D point P[%, Y, Z] of the 3D model to a 2D pixel coordinate p = [x, y] . That is, the 3D model fl can be projected onto the image plane of D,, by a 3D-to-2D projection of each 3D point P = [A, Y,Z] of fl to a 2D pixel coordinate [x,y]:
[0056]
[0057] where 1 is an arbitrary scale factor used to recover the homogeneous vector [x, y, 1].
[0058] At the position [x, y], the depth d is stored (depth = Euclidean distance between P and image position t: d = ||P - 11|). That is, in some embodiments, the depth value d = ||P — 11| is stored for the position P with 2D pixel coordinate p = [x, y], where the depth value represents Euclidean distance between the position P and image position t in the local coordinate space. In further detail, the image position t in the local coordinate space is given from the process of 3D model generation from the set of images (as in step S104). If the 3D model is generated not from images (e.g., from a LIDAR scanner), then the 2D-3D alignment between the 2D images and the 3D model may be implemented as part of step S104 or at least before step S106.
[0059] In cases where the 3D model fl is a mesh of polygons (e.g., triangles), each polygon can be projected onto the image plane by applying the above projection equation tothe corners (vertices) of the polygon, and interpolating from the vertex coordinates the projection coordinates for each point within the polygon.
[0060] Further aspects of how the local surface given by depth values of the user-selected point and 3D points neighbouring the user-selected point can be estimated will be disclosed next.
[0061] Generally, the user-selected point puhas pixel coordinates [xu, yu], and an associated depth duin the depthmap
[0062]
[0063] Depths for all points in the image near pucan be retrieved from D, in a window ranging in width from xu— m to xu+ m , and ranging in height from yu— n to yu+ n, where m, n > 0. Here, each of m, n typically takes a low value, such as 1 or 2. This selection window can be a full rectangle (where all pixels in the width and height range are selected), or a cross pattern (where only the pixels directly above / below / leftwards / rightwards of puare selected). The collection (or set) of all the points from the selection window is represented by the parameter Y2Dof size k (i.e., Y2D= {{p1(d , {p2ld2, {p3ld3, ... , {pu, du}, ... , {pkldk}}.
[0064] The 3D points neighbouring the user-selected point can be estimated by a reverse 2D-to-3D projection being applied to 2D points in the 2D image that are neighbouring the user-selected point. As above, assuming that the user-selected point puhas pixel coordinates [xu, yu] and an associated depth duin the depthmap D,, the depth values of the 3D points neighbouring the user-selected point can be retrieved from the depthmap D, in a window ranging in width from xu— m to xu+ m, and ranging in height from yu— n to yu+ n, where m, n > 0 are integers. As an example, from the collection (or set) of points Y2D, for each point pa= [xa,ya] the corresponding 3D position Pa= [Xa, Ya, Za] can be estimated via reverse 2D-to-3D projection, e.g.:
[0065]
[0066] which gives a set of 3D points, Y3D= {P1, P2, P3, ... , Pu, ... , Pk}. The 3D points in Y3D(i.e., the 3D points neighbouring the user-selected point) approximate the surface of an object placed at the user-selected point pu.
[0067] Further aspects of how the surface heading vector for the user-selected point pucan be estimated in the local coordinate space will be disclosed next.In some embodiments, the surface heading vector nuin the local coordinate space is estimated as a normal to a plane fitting the surface.
[0068] For example, from the set of 3D points Y3Dcorresponding to pu, the surface heading vector nu, or surface normal, can be estimated by first estimating a local surface that best fits Y3D, and then computing the normal of that local surface.
[0069] For example, a covariance matrix C can be computed to estimate the local surface and compute the normal:
[0070]
[0071] is the y-th eigenvalue of C, and Vj is the y-th eigenvector of C.
[0072] Then, e is selected from vx, vy, vzas the vector v with the greatest norm, and is then normalized:
[0073] e = v / ||v||.
[0074] Next, each of vx, vy, vzis normalized to e by u = v — (v •
[0075]
[0076] giving respective vectors ux, uy, uz. Then, e2is selected from ux, uy, uzas the vector u with the greatest norm, and is then normalized. Finally, the normal vector can be calculated:
[0077] nu= ei x e2,
[0078] where nuis the surface heading vector in the local coordinate space.
[0079] Further aspects of how the surface heading vector for the user-selected point pucan be estimated in the geospatial coordinate space will be disclosed next.
[0080] The surface heading vector nuis in the local coordinate space of the 3D model £1, because it is calculated from Y3D. If £1 is already aligned to the same geospatial coordinate space that gives gps, and hj, then the geospatial heading is hu= nu.
[0081] However, if £1 is not geospatially aligned, then the geospatial heading huneeds to bealigned from the local space to the geospatial space. In this respect, in some embodiments, converting the surface heading vector from the local coordinate space to the geospatial coordinate space comprises (i) obtaining a reverted surface heading vector by reverting local image rotation and position of the surface heading vector from the local coordinate space, and (ii) applying a global image rotation and position to the reverted surface heading vector. For example, the geospatial heading humay be aligned from the local space to the geospatial space by:
[0082]
[0083] where 7?, is a 3-by-3 rotation matrix form of
[0084]
[0085] Some alternative implementations of the image processing device for performing the herein disclosed embodiments will be disclosed next with reference to the block diagrams in Fig. 4, Fig. 5, and Fig. 6. For example, the process of creating a 3D geometry of a site from 2D images, the process of geolocating the 3D geometry, the process of projecting 3D geometry onto the image planes of all images, and the process of estimating local headings of all pixels of all images, can be performed a-priori, prior to any user input. Conversely, each of these processes could also be performed on-demand for a specific image and a specific selection, depending on computational ability, storage availability, etc. of the image processing device.
[0086] The block diagram of Fig.4 represents an example of an image processing device 400 configured for executing the computational steps of the herein disclosed methods during user interaction. The image processing device 400 can be implemented in a digital portal. This example represents an embodiment where the 3D model is projected onto the 2D image plane upon the user-selected point puhaving been obtained. The image processing device 400 is configured to receive as input a set of images 1410 depicting a site or a part thereof, as shown to the user of an image-based viewer. Each image can be a color image, a grayscale image, a thermal image or any other type of 2D image. The image processing device 400 is further configured to generate a 3D geometry 424 related at least one of the images, as selected via user input 428 via an image-based viewer block 438. The 3D geometry 424 can be generated in different ways, depending on whether a 3D scan of the site exists or not, as checked in box 412. If a 3D scan 414 exists, the 3D geometry 424 canbe generated based on 2D-3D alignment of the 2D scan 414 and at least the selected image, possibly in combination with more images from the set of images 410. Else, the 3D geometry 424 can be obtained from a 3D reconstruction 416 of at least the selected image, possibly in combination with more images from the set of images 410. The 3D geometry 424 is stored in a storage 426 together with geolocated poses 422 as extracted in a geolocation extraction block 418 from the selected image. In block 432 a surface heading vector is estimated in the local coordinate space for the user-selected point pu434 based on the 3D geometry 430 and image poses. In block 436 the surface heading vector is converted to the geospatial coordinate space and then provided to the image-based viewer block 438.
[0087] The block diagram of Fig.5 represents an example of an image processing device 500 where as many computational steps as possible are executed prior to user interaction. The image processing device 500 can be implemented in a digital portal. This example represents an embodiment where a respective surface heading vector in the geospatial coordinate space is estimated for a plurality of points in the 2D images in the set of 2D images ahead of the user-selected point pubeing obtained. Estimation of the surface heading of the user-selected point can then be based on the estimated surface heading vector for at least one of the plurality of points in the 2D images in the set of 2D images. The image processing device 500 is configured to generate a 3D geometry 524 related to a set of images 510 depicting a site or a part thereof. The images can be color images, grayscale images, thermal images or any other type of 2D images. This can be achieved in different ways, depending on whether a 3D scan of the site exists or not, as checked in box 512. If a 3D scan 514 exists, the 3D geometry 524 can be generated based on 2D-3D alignment of the 3D scan 514 and the images 510. Else, the 3D geometry 524 can be obtained from a 3D reconstruction 516 of the images 510. In block 532 surface heading vectors are estimated in the local coordinate space for a set of geolocated poses 522 as extracted in a geolocation extraction block 618 from the set of images 510 and based on the 3D geometry 530. Surface heading vectors 530 as converted to the geospatial coordinate space are then provided to a storage 528. The image processing device 500 is further configured to receive as input selection of one image 538 in the set of images 510 and a user-selected point pu534. In block 532 the geospatial heading 536 for the image 538 and the user-selected point pu534 is retrieved from the storage 528 and provided to the image-basedviewer block 540. In this example, as many steps as possible are performed ahead-of-time for a whole set, or some subset, of all possible points 622 in all possible images 610. Then, during user interaction the image processing device 600 only needs to then retrieve the precalculated corresponding heading 536 for the user-requested point 534. If only a subset of all possible points have pre-calculated headings, then during user-interaction, any missing heading for a point 534 can be quickly approximated from the pre-calculated headings 530 of two or more neighbouring points. For example, a missing heading huat point pucan be approximated as a weighted average between the precalculated headings
[0088]
[0089] hu+1) of neighbouring points (pu-npu+1) hu= shu-1+ (1 - s)hu+1, with some weight s e range(O.l) (e.g. E = 0.5).
[0090] The block diagram of Fig. 6 represents an example of an image processing device 600 where some of the computational steps are executed prior to user interaction, and some of the computational steps are executed during user interaction. For example, estimating the depthmaps can be performed ahead-of-time, thus splitting the estimation process in ahead-of-time and real-time processing. The image processing device 600 can be implemented in a digital portal. This example represents an embodiment where the 2D image is part of a set of 2D images depicting the scanned scene, where each 2D image in the set of 2D images is associated with a respective 2D image plane, and the 3D model is projected onto each of the 2D image planes ahead of the user-selected point pubeing obtained. The image processing device 600 is configured to generate a 3D geometry 624 related to a set of images 610 depicting a site or a part thereof. The images can be color images, grayscale images, thermal images or any other type of 2D images. This can be achieved in different ways, depending on whether a 3D scan of the site exists or not, as checked in box 612. If a 3D scan 614 exists, the 3D geometry 624 can be generated based on 2D-3D alignment of the 3D scan 614 and the images 610. Else, the 3D geometry 624 can be obtained from a 3D reconstruction 616 of the images 610. In block 626 surface heading vectors are estimated in the local coordinate space for a set of geolocated poses 622 as extracted in a geolocation extraction block 618 from the set of images 610 and based on the 3D geometry 624. The surface heading vectors represent partial process outputs 630 that are stored in a storage 628. The 2D images 610, the set of geolocated poses 622, and the 3D geometry 630 are also stored in the storage 628.In block 636, the surface heading vector in the local coordinate space, as represented by partial process output 632, for selection of one image 638 in the set of images 610 and a user-selected point pu634, is then converted to the geospatial coordinate space. The surface heading vector 640 in the geospatial coordinate space is then provided to the image-based viewer block 642.
[0091] In Fig. 7 is schematically illustrated the relation 700 between a 3D model 710 and a 2D image 720. A 2D image 1720 is located at position t, with rotation Rj, in the same coordinate system as 3D model fl. Dashed lines represent the camera frustum, commonly described by the aforementioned intrinsic matrix K,.
[0092] Fig. 8 schematically illustrates, in terms of a number of structural units, the components of an image processing device 700 according to an embodiment.
[0093] Processing circuitry 810 is provided using any combination of one or more of a suitable central processing unit (CPU), multiprocessor, microcontroller, digital signal processor (DSP), etc., capable of executing software instructions stored in a computer program product 910 (as in Fig. 9), e.g. in the form of a storage medium 830. The processing circuitry 810 may further be provided as at least one application specific integrated circuit (ASIC), or field programmable gate array (FPGA).
[0094] Particularly, the processing circuitry 810 is configured to cause the image processing device 700 to perform a set of operations, or steps, as disclosed above. For example, the storage medium 830 may store the set of operations, and the processing circuitry 810 may be configured to retrieve the set of operations from the storage medium 830 to cause the image processing device 700 to perform the set of operations. The set of operations may be provided as a set of executable instructions.
[0095] Thus the processing circuitry 810 is thereby arranged to execute methods as herein disclosed. The storage medium 830 may also comprise persistent storage, which, for example, can be any single one or combination of magnetic memory, optical memory, solid state memory or even remotely mounted memory. The image processing device 700 may further comprise a communications (comm.) interface 820 at least configured for communications with other entities, functions, nodes, and devices. As such the communications interface 820 may comprise one or more transmitters and receivers, comprising analogue and digital components. The processing circuitry 810controls the general operation of the image processing device 700 e.g. by sending data and control signals to the communications interface 820 and the storage medium 830, by receiving data and reports from the communications interface 820, and by retrieving data and instructions from the storage medium 830. Other components, as well as the related functionality, of the image processing device 700 are omitted in order not to obscure the concepts presented herein.
[0096] The image processing device 700 may be provided as a standalone device or as a part of at least one further device. A first portion of the instructions performed by the image processing device 700 may be executed in a first device, and a second portion of the of the instructions performed by the image processing device 700 may be executed in a second device; the herein disclosed embodiments are not limited to any particular number of devices on which the instructions performed by the image processing device 700 may be executed. Hence, the methods according to the herein disclosed embodiments are suitable to be performed by an image processing device 700 residing in a cloud computational environment. Therefore, although a single processing circuitry 810 is illustrated in Fig. 8 the processing circuitry 810 may be distributed among a plurality of devices, or nodes. The same applies to the computer program 920 of Fig. 9.
[0097] Fig. 9 shows one example of a computer program product 910 comprising computer readable storage medium 930. On this computer readable storage medium 930, a computer program 920 can be stored, which computer program 920 can cause the processing circuitry 810 and thereto operatively coupled entities and devices, such as the communications interface 820 and the storage medium 830, to execute methods according to embodiments described herein. The computer program 920 and / or computer program product 910 may thus provide means for performing any steps as herein disclosed.
[0098] In the example of Fig. 9, the computer program product 910 is illustrated as an optical disc, such as a CD (compact disc) or a DVD (digital versatile disc) or a Blu-Ray disc. The computer program product 910 could also be embodied as a memory, such as a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), or an electrically erasable programmable read-only memory (EEPROM) and more particularly as a non-volatilestorage medium of a device in an external memory such as a USB (Universal Serial Bus) memory or a Flash memory, such as a compact Flash memory. Thus, while the computer program 920 is here schematically shown as a track on the depicted optical disk, the computer program 920 can be stored in any way which is suitable for the computer program product 910.
[0099] The inventive concept has mainly been described above with reference to a few embodiments. However, as is readily appreciated by a person skilled in the art, other embodiments than the ones disclosed above are equally possible within the scope of the inventive concept, as defined by the appended patent claims.
Claims
CLAIMS1. A method for estimating surface heading of a user-selected point within a two-dimensional, 2D, image, the method being performed by an image processing device (100, 400:700), the method comprising:obtaining (S102) the user-selected point puin the 2D image, wherein the 2D image is depicting a scanned scene, is associated with a 2D image plane, and is provided in a geospatial coordinate space;obtaining (S104) a three-dimensional, 3D, model of the scanned scene, wherein the 3D model is composed of 3D points and is provided in a local coordinate space;projecting (S106) the 3D model onto the 2D image plane;estimating (S108) a surface heading vector in the local coordinate space for the user-selected point puas a normal to a local surface given by depth values of the user-selected point and 3D points neighbouring the user-selected point; andestimating (S110) the surface heading by converting the surface heading vector from the local coordinate space to the geospatial coordinate space.
2. The method according to claim 1, wherein the 3D model is projected onto the 2D image plane as a depthmap D, of same width and height as the 2D image.
3. The method according to claim 1 or 2, wherein the 3D model fl is projected onto the 2D image plane by a 3D-to-2D projection of each 3D point P[X, Y, Z] of the 3D model to a 2D pixel coordinate p = [x,y].
4. The method according to claim 3, wherein a depth value d = ||P - t|| is stored for the position P with 2D pixel coordinate p = [x, y], where the depth value represents Euclidean distance between the position P and image position t in the local coordinate space.
5. The method according to any preceding claim, wherein the 3D points neighbouring the user-selected point are estimated by applying a reverse 2D-to-3D projection to 2D points in the 2D image that are neighbouring the user-selected point.
6. The method according to claim 2, wherein the user-selected point puhas pixel coordinates [xu, yu] and an associated depth duin the depthmapand wherein the depth values of the 3D points neighbouring the user-selected point are retrieved from the depthmap D, in a window ranging in width from xu— m to xu+ m, and ranging in height from yu— n to yu+ n, where m, n > 0 are integers.
7. The method according to any preceding claim, wherein the 3D points neighbouring the user-selected point approximate a surface of an object placed at the user-selected point pu.
8. The method according to claim 7, wherein the surface heading vector nuin the local coordinate space is estimated as a normal to a plane fitting the surface.
9. The method according to any preceding claim, wherein converting the surface heading vector from the local coordinate space to the geospatial coordinate space comprises (i) obtaining a reverted surface heading vector by reverting local image rotation and position of the surface heading vector from the local coordinate space, and (ii) applying a global image rotation and position to the reverted surface heading vector.
10. The method according to any preceding claim, wherein the 3D model is projected onto the 2D image plane upon the user-selected point puhaving been obtained.
11. The method according to any of claims 1 to 9, wherein the 2D image is part of a set of 2D images depicting the scanned scene, wherein each 2D image in the set of 2D images is associated with a respective 2D image plane, and wherein the 3D model is projected onto each of the 2D image planes ahead of the user-selected point pubeing obtained.
12. The method according to claim 11, wherein a respective surface heading vector in the geospatial coordinate space is estimated for a plurality of points in the 2D images in the set of 2D images ahead of the user-selected point pubeing obtained, and wherein estimating the surface heading of the user-selected point is based on the estimated surface heading vector for at least one of the plurality of points in the 2D images in the set of 2D images.13- The method according to any preceding claim, wherein the method further comprises:providing (S112) the surface heading vector huin the geospatial coordinate space to an image processing application.
14. The method according to claim 13, wherein the image processing application is any of: a digital twin application, a 3D modelling application, a Business Information Modelling, BIM, application, a computer aided design, CAD, application.
15. An image processing device (100, 400:700) for estimating surface heading of a user-selected point within a two-dimensional, 2D, image, the image processing device (100, 400:700) comprising processing circuitry (810), the processing circuitry being configured to cause the image processing device (100, 400:700) to:obtain the user-selected point puin the 2D image, wherein the 2D image is depicting a scanned scene, is associated with a 2D image plane, and is provided in a geospatial coordinate space;obtain a three-dimensional, 3D, model of the scanned scene, wherein the 3D model is composed of 3D points and is provided in a local coordinate space;project the 3D model onto the 2D image plane;estimate a surface heading vector in the local coordinate space for the user-selected point puas a normal to a local surface given by depth values of the user-selected point and 3D points neighbouring the user-selected point; andestimate the surface heading by converting the surface heading vector from the local coordinate space to the geospatial coordinate space.
16. The image processing device (100, 400:700) according to claim 15, further being configured to perform the method according to any of claims 2 to 14.
17. A computer program (920) for estimating surface heading of a user-selected point within a two-dimensional, 2D, image, the computer program comprising computer code which, when run on processing circuitry (810) of an image processing device (100, 400:700), causes the image processing device (100, 400:700) to:obtain (S102) the user-selected point puin the 2D image, wherein the 2D image is depicting a scanned scene, is associated with a 2D image plane, and is provided in a geospatial coordinate space;obtain (S104) a three-dimensional, 3D, model of the scanned scene, wherein the 3D model is composed of 3D points and is provided in a local coordinate space;project (S106) the 3D model onto the 2D image plane;estimate (S108) a surface heading vector in the local coordinate space for the user-selected point puas a normal to a local surface given by depth values of the user-selected point and 3D points neighbouring the user-selected point; andestimate (S110) the surface heading by converting the surface heading vector from the local coordinate space to the geospatial coordinate space.
18. A computer program product (910) comprising a computer program (920) according to claim 17, and a computer readable storage medium (930) on which the computer program is stored.