A method for monitoring construction sites and a system implementing the method
The method corrects trajectory and georeferences smartphone-acquired images to generate precise three-dimensional models of construction sites, addressing the issue of inconsistent geolocation in existing technologies and enabling real-time parameter display.
Patent Information
- Application Number
- PCT/EP2025/064101
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-23
- Filing Date
- 2025-05-22
- Publication Date
- 2025-11-27
AI Technical Summary
Existing methods for monitoring construction sites using smartphones fail to provide high-precision three-dimensional models with accurate georeferencing, leading to fragmented and inconsistent representations.
A method involving image acquisition, precise trajectory correction using API Geospatial, depth map generation with ArCore, and computational geometry techniques to create a georeferenced three-dimensional model, followed by real-time correction and transformation into textured meshes, enabling accurate geometric parameter display.
Enables remote, high-precision monitoring of construction sites with exact geographical localization, allowing real-time display of geometric parameters like width, length, and depth.
Smart Images

Figure EP2025064101_27112025_PF_FP_ABST
Abstract
Description
[0001] Title: A method for monitoring construction sites and a system implementing the method
[0002] DESCRIPTION
[0003] Field of application
[0004] The present invention refers to a method of monitoring, namely acquiring images of excavations, construction sites, archeological sites or generally of construction geographical sites, for which it is necessary to monitor the progress of construction and / or the development of the state over time.
[0005] The present invention also refers to a system implementing said method.
[0006] More in detail, the invention concerns a method of deriving a high- definition three-dimensional model of a construction site, which is also georeferenced through a communication device, such as a smartphone and the like.
[0007] Hereinafter, the description will be directed to a method and a system for remotely monitoring construction sites through smartphone, but it is clear that it should not be considered limited to this specific use.
[0008] Prior art
[0009] Currently, to monitor the progress of an excavation or of a construction site, it is necessary that an operator inspects the site, detects the state of the site, and fill out a field notebook, which is a tool to note all the operations carried out in the field.
[0010] To this end, methods and systems for monitoring sites that allow to obtain high-definition three-dimensional models through smartphone, for example “iPhone Pro” smartphones, equipped with time-of-flight (TOF) cameras are known. Methods and systems that allow to obtain photogrammetric models from surveys made with smartphones provided with accessories, such as antennas to improve both the relative and the absolute geographical positioning, are also known.
[0011] Methods and systems that allow to obtain high-definition three- dimensional models through smartphone are also known, which, however, are not reliable with regard to measurement precision and georeferencing.
[0012] All the methods and systems known in the state of the art that use smartphones do not allow to make three-dimensional models of construction-site excavations that can be obtained remotely, with a high level of measurement precision and an exact georeferencing or geographical localization.
[0013] In the light of the disadvantages and problems unsolved by the state of the art, an object of the present invention is to provide a method capable of creating a three-dimensional model of a construction site remotely, with high precision and with exact geographical localization.
[0014] A further object of the present invention is to provide a system implementing the monitoring method.
[0015] The technical problem underlying the present invention is thus to provide a method and a system for remotely monitoring a construction site.
[0016] Summary of the invention
[0017] The idea underlying the present invention is to provide a method and a system that allow to detect the state of a geographical site, in particular of a construction site, and to remotely monitor the same.
[0018] An object of the present invention is a method for remotely monitoring a geographical site, comprising the following steps:
[0019] 110. acquiring a plurality of images and / or sequences of photograms of said site, by means of an acquisition device in shots along an acquisition trajectory;
[0020] 120. correcting the position and the orientation of each single shot, so as to obtain a correct, oriented and georeferenced trajectory, along which said shots were taken;
[0021] 130. calculating the three-dimensional model of said site, starting from depth maps of said site calculated on said plurality of images and / or sequences of photograms;
[0022] 140. sending said plurality of images and / or sequences of photograms, correct trajectory and three-dimensional model to a remote server for the calculation of a mesh, so as to associate said pluralities of images and / or sequences of photograms with said three-dimensional model, in order to display on said three-dimensional model the geometric parameters of said site like minimum, maximum and average width, length and depth.
[0023] Further according to the invention, said step 120. is carried out by means of a programming interface of the application or application programming interface (API) which is able to provide the absolute positioning of said acquisition device.
[0024] Again according to the invention, in said step 130. a plurality of depth maps are used, connected to each other by computational geometry techniques and robotics.
[0025] Preferably according to the invention, said step 130. comprises a substep 131. of correcting said three-dimensional model in real time, where, given said trajectory, deviation vectors are generated in order to identify a discontinuity between said shots and the logical and temporal order thereof and to connect the discontinuity points by means of a roto- translation matrix.
[0026] Again according to the invention, said step 140. comprises a first substep 141. where said three-dimensional model is transformed into a textured mesh provided with position, orientation, scale, opacity and color.
[0027] Further according to the invention, said step 140. comprises a second sub-step 142. wherein several textured meshes associated with three- dimensional models of images acquired on different days are joined.
[0028] A further object of the present invention is a system for remotely monitoring a geographical site which is able to carry out the method according to the preceding claims, comprising one or more devices for the acquisition of pluralities of images and / or sequences of photograms of said site and for a first processing thereof, so as to obtain a three- dimensional model of said site, a remote server or cloud which is able to receive the data from said one or more devices, being able to associate said three-dimensional model with said plurality of images and / or sequences of photograms, in order to display on said three-dimensional model geometric parameters of the site like minimum, maximum and average width, length and depth.
[0029] Further according to the invention, said one or more devices is a smartphone comprising a camera.
[0030] Again according to the invention, on each device an operating environment or application is installed, which an operator can access by an authentication with username and password and select the target construction site, from an available list.
[0031] Brief description of the drawings
[0032] In the drawings:
[0033] Figure 1 shows an outline of a block diagram of the operation steps of the method, object of the present invention;
[0034] Figure 2 shows a geographical top view of the construction site to be monitored; Figure 3 shows a schematic view of a shot-acquisition trajectory;
[0035] Figure 4 shows a graphic view of a non-corrected three- dimensional model;
[0036] Figures 5a, 5b, and 5c show a graphic view of the acquisition trajectory, of the trajectory with an interruption point, and of the reconstructed trajectory, respectively;
[0037] Figure 6 shows a point cloud of the three-dimensional model;
[0038] Figure 7 shows a schematic top view of the joined meshes of the three-dimensional model;
[0039] Figure 8 shows a schematic view of a first operation sub-step of the method;
[0040] Figure 9 shows a schematic view of a second operation sub-step of the method;
[0041] Figure 10 shows a schematic view of third operation sub-step of the method;
[0042] Figure 11 shows a graphic representation of the top view of the excavation;
[0043] Figure 12 shows a graphic representation of the top view of the excavation;
[0044] Figure 13 shows a further graphic representation of the top view of the excavation, with which the geometric parameters are associated;
[0045] Figure 14 shows a schematic representation of the system implementing the method according to the present invention.
[0046] Detailed description
[0047] With reference to Figure 1, the method 100 for monitoring construction sites C, which is object of the present invention, comprises a plurality of steps.
[0048] In the acquisition step 110, a series of images or a video as a series of photograms I are taken through an acquisition device D or a camera T comprised in an acquisition device D, framing the part of interest of construction site C, shown in Figure 2, carrying out shots S, also called camera-poses.
[0049] The device D is preferably a smartphone.
[0050] However, without departing from the scope of protection of the present invention, the device D may also be any other device provided with a camera T.
[0051] As shown in Figure 3, the positions of each single shot S originate the trajectory R, which is a polyline made of segments, traveled by the camera T and the various orientations of the camera T or of the acquisition device D comprising the camera T, during the acquisition step 110.
[0052] In Figure 3, each numbered point represents a position of each single shot S taken by the camera T or by the device D, with which the orientation is associated. The points may be acquired with even high frequencies n / s.
[0053] The more this trajectory R is accurate, the better the quality of the three- dimensional model M deriving therefrom is, considering that the error variance inside the trajectory R affects the relative accuracy (size of the three-dimensional model M), the error of the totality of positions with respect to space affects instead the absolute accuracy.
[0054] It is thus necessary to generate a precise trajectory to correct the position and the orientation of each single shot or photogram I.
[0055] In a correction step 120, the correction of the position and of the orientation of each single shot S is carried out, thereby obtaining a precise correct trajectory Rc.
[0056] The advantage associated with the construction of a precise trajectory Rcconsists in having shots S that are robust in terms of relative precision, in case of operation outside of the covering areas of the Streetview module, and relative precision and absolute accuracy, in relation to a geographical system, in case of operation in areas in which the Streetview module is present. Thus, the shots S or camera-poses will always be consistent, namely relatively robust, as well as being well oriented, and, where the Streetview module is present, also correctly georeferenced.
[0057] In particular, in said correction step 120, a known interface, such as the application API Geospatial, is used, where API means a programming interface of the application or application programming interface.
[0058] The application API Geospatial is very accurate, having at disposal a very large amount of registered data, and allows a more efficient and effective interaction with augmented reality for a user using it, in particular, comparing the images of the environment in which an artificial intelligence has extracted trillions of three-dimensional points from panoramic views of the known Streetview module. All the points are then analyzed and used to calculate the position and the orientation of a device D.
[0059] Thus, the API Geospatial allows, where the Streetview module is present, to considerably improve the absolute positioning of the camera T, which would otherwise be provided only by GPS which entails errors of several meters, up to several tens of meters.
[0060] In a step 130 of three-dimensionally modeling the construction site C, the three-dimensional model M is generated taking into account the data of depth, width, length and thickness of the asphalt and the like.
[0061] For the generation of the three-dimensional model M, the depth map generated through the known application API ArCore is generated. A depth map, in applications of three-dimensional computer graphics or computer vision, is an image or an image channel containing information related to the distance of the surfaces of the objects in a scene, from a standpoint.
[0062] In this invention, the depth maps may also be generated with softwares and procedures that are already available in the prior art, and may be connected to each other by computational geometry techniques and robotics, such as Simultaneous localization and mapping - SLAM.
[0063] Whichever of the two technologies is used for building the depth map, the three-dimensional models generated are affected by discontinuity between the series of images or photograms I taken during the acquisition step 110, since errors of hardware or software acquisition might occur.
[0064] This is mainly due to poor GPS quality in smartphones.
[0065] As a consequence, the reconstructed three-dimensional model M is not representative of the reality acquired, because it is fragmented and not consistent, as it can be seen in Figure 4.
[0066] Thus, to make the three-dimensional model M consistent, a sub-step 131 of correcting the three-dimensional model M in real time is carried out.
[0067] In particular, using the correct trajectory Rcof the shots S that ensures the actual course of the shooting by consistently identifying the distances between the various take centers or shots S; comparing the positions of the depth maps, deviation vectors are generated chronologically based on the time stamp, namely a sequence of characters representing a date and / or a time to ascertain that a certain event has actually occurred, and this allows to understand when a discontinuity between the shots S occurred, besides allowing to identify the logical and temporal order.
[0068] With reference to Figures 5a, 5b and 5c, Figure 5a shows the exact trajectory R of the shots S taken, Figure 5b shows the discontinuity occurred in a shot and Figure 5c shows the correct trajectory Rc. When an inconsistency between the points of the trajectory R and the local orientations is found, as it is the case in Figure 5b in the point 5-5’, the connection of the points is forced by means of the roto-translation matrix based on the matrices related to the points.
[0069] Considering that the GPS position and the orientation undergo successive refinements, of the single depth maps, coinciding with the shots S or with the frames of a video, in Figure 5, the points 6 and 7 are more precise than the points 1-2-3-4. Thus, in the correction sub-step 131, it is possible to attribute a weighted average weight that increases as a function of the acquisition time.
[0070] Figure 6 shows the point cloud representing the three-dimensional model M generated by the single compensated depth maps, which is already visible on the device D of the operator in real time.
[0071] After step 130 of three-dimensional modeling, all information is uploaded to the cloud in a step 140 of data uploading and management, during which a mesh to associate the real images I with the three-dimensional model M is obtained.
[0072] In a first sub-step 141, in a server, the three-dimensional model M is transformed into textured meshes or Gaussian-splatting representation, which represents a three-dimensional scene by means of millions of Gaussian particles provided with position, orientation, scale, opacity and color.
[0073] This occurs through the generation of preparatory files, of the type bundle-adjustment shot calibration, which is a way of adjusting the bundle, i.e. it is a simultaneous improvement of the three-dimensional coordinates describing some characteristics of a scene based on the optical characteristics of the cameras T used, of creation of undistorted images, so as to remove lens distortion, and subsequent typical known processes of creation of a textured mesh.
[0074] The same elements such as the three-dimensional model M, the undistorted images and the shots S may also be used for the generation of a point cloud of the Gaussian-splatting type.
[0075] The depth map obtained as described above can be used together with shots S or camera poses consistent with the depth map itself, to modify, by means of the Structure From Motion - SFM process, the obtained mesh which will be affected by GPS limitations, despite the use of RANdom Sample Consensus - RANSAC algorithms, geometric analysis and least squares weighting.
[0076] In particular, the distances between the shots S, consistent with the depth map, are measured and the sequence of points considered most robust is used to check and correct the same points adjusted with the GPS during the SFM process.
[0077] In this way, the shots S and the depth map, corrected in relative coordinates, i.e. not positioned in the geographical and cartographic context as they are not positioned according to a global position, will constitute the basis for scaling the shots S and making the 3D model scaled precisely.
[0078] If the GPS is of poor quality, it is possible to use the center of gravity of the various SFM positions for positioning, in order to place the model in coordinates, averaging the error.
[0079] The georeferenced model thus obtained will then be improved on the basis of known points or automatically positioned on existing maps or orthophotos, through scale-invariant feature transform - SIFT and Oriented FAST and rotated BRIEF - ORB artificial intelligence algorithms.
[0080] As regards the orientation of the model itself, if the GPS is not reliable, the orientation of the shots S will be used, independent of the GPS.
[0081] If the 3D model is a Gaussian-splatting cloud, from the depth map and the shots S that recall the images, the new Gaussian-splatting point cloud will be generated, which will be georeferenced and oriented, as described above.
[0082] In short, on the device D a depth map is generated using the known application API ArCore, mentioned above, or other similar ones. This depth map and shots S are to be considered metrically correct, even if in a relative sense without having GPS data, because it will be adjusted, as described above.
[0083] Based on the 3D point cloud, i.e. the depth map and the sequence of shots S, the relative distances will be sampled. These will be used in the case of a georeferenced 3D model, generated through the SFM process in which the relative positions become absolute through the use of GPS. In the case where the GPS is not accurate, which would affect the metric accuracy of the model, the metric unit of the model is corrected based on the correction of the shots S of the ArCore acquisition and the georeferenced shots S coming from the SFM process.
[0084] These relative distances will also be used when going directly from the depth map to the textured mesh. In this case, the model is already metrically correct. The center of gravity of the GPS points will be used, which will be the basis for georeferencing the model itself in space, as described above. The ArCore vector will be used for orientation.
[0085] In the case of Gaussian-splatting, the point cloud generated by the depth map will be transformed into Gaussian-splatting and will be georeferenced and oriented as described above.
[0086] With reference to Figure 7, a second sub-step 142 of joining different meshes associated with three-dimensional models M of images I acquired by the operator on different days is carried out, so as to display the state of a construction site C in a single three-dimensional model M, even though the acquisition was carried out on different days.
[0087] Thus, it will be possible to see in a single three-dimensional model different areas of the same construction site C that have the same working state. In Figure 7, the meshes or the point clouds are visible, comprising the Gaussian-splatting cloud, represented through web app via browser based on the calendar, i.e. based on the day of acquisition of the images I, and univocally according to the attributed working state.
[0088] The second sub-step 142 of joining or matching the meshes is carried out through the extraction of characteristics and matching thereof, using known computer-vision and photogrammetry algorithms.
[0089] The measurements can also be carried out automatically.
[0090] The second sub-step 142 in turn comprises further sub-steps.
[0091] With reference to Figure 8, in a first further sub-step 142a, starting from one mesh, a spatial grid G of the excavation of the construction site C is automatically created and the nadiral angle, which is the angle formed by a direction with the straight line intersecting the nadir in a vertical plane.
[0092] In a second further sub-step 142b, for each single quadrant of the grid G, the point at the lowest depth is identified.
[0093] As shown in Figure 9, in a third further sub-step 142c, all the quadrants having the lowest depth being constant or contained within a certain range are identified, if the depth is constant, this indicates the presence of an excavation.
[0094] As shown in Figure 10, in a fourth further sub-step 142a, the excavationcenter line is identified, which is obtained by clustering, i.e. grouping the points at the desired depth and extracting the central points, and sections perpendicular to the profile generated by the points are made, thereby highlighting the discontinuity points, such as slope variations, from which the following information are extracted: maximum depth of the excavation; width of the excavation at 50% depth, so as to avoid wrong measurements due to landslide, minimum depth of the central section, from -25% to +25% of the width of the section of the lower base, so as to identify the artifacts.
[0095] Subsequently, the widths of the sections are compared to identify anomalies in comparison with the average width, the depths are compared with each other and with the expected depths to highlight anomalies, and the endpoints of the excavation are identified, which typically have a sloped trend being separately signaled, possible abrupt anomalies of width or depth that affect one or two sections, are analyzed and possibly automatically filtered.
[0096] In this way, the excavation representations shown in Figures 11 and 12 are obtained, where the minimum, maximum and average excavation width, excavation length, namely the distance between the outer points, and depth are automatically determined, on a single mesh or on a set of meshes, as shown in Figure 7.
[0097] As shown in Figure 13, it is possible to display said parameters on the three-dimensional model M of the excavation C.
[0098] With reference now to Figure 14, the system A implementing the method 100 comprises one or more acquisition devices D comprising a camera T, such as a smartphone, assigned to one or more operators who carry out the acquisition of the images or of the photograms I, in a given day.
[0099] On each device D, an operating environment or application is installed, which the operator can access by an authentication with username and password and select the target construction site C, from an available list, wherein specific colors are used to indicate the state of the construction site.
[0100] Subsequently, the operator carries out the acquisition of the sequence of images I through shots S, walking along the excavation of the construction site C. The straight line connecting the positions in which the shots S were taken defines a trajectory R.
[0101] The acquired images I are processed in the device D, thereby producing the three-dimensional model M of the construction site. Subsequently, the processing of the images I is sent to a remote server or cloud B which is able to associate the three-dimensional model M with the images I, in order to display on the three-dimensional model M the geometric parameters of the excavation C like minimum, maximum and average width, length and depth.
[0102] As it is evident from the above description, the advantage of the method and the system that are objects of the present invention is to remotely monitor the geometric parameters of an excavation using a commonly used device such as a smartphone used by an operator, which can also display the three-dimensional model of the excavation and its geometric parameters in real time.
[0103] Obviously a person skilled in the art, in order to satisfy any specific requirements which might arise, may make numerous modifications and variations to the system described above, all of which are however contained within the scope of protection of the invention, as defined by the following claims.
Claims
CLAIMS1. A method (100) for remotely monitoring a geographical site (C), comprising the following steps:
110. acquiring a plurality of images and / or sequences of photograms (I) of said site (C), by means of an acquisition device (D, T) in shots (S) along an acquisition trajectory (R);120. correcting the position and the orientation of each single shot (S), so as to obtain a correct, oriented and georeferenced trajectory (Rc), along which said shots (S) were taken;130. calculating the three-dimensional model (M) of said site (C), starting from depth maps of said site (C) calculated on said plurality of images and / or sequences of photograms (I);140. sending said plurality of images and / or sequences of photograms (I), correct trajectory (Rc) and three-dimensional model (M) to a remote server for the calculation of a mesh, so as to associate said pluralities of images and / or sequences of photograms (I), with said three-dimensional model (M), in order to display on said three-dimensional model (M) the geometric parameters of said site (C) like minimum, maximum and average width, length and depth.
2. The method (100) according to the preceding claim, characterized in that said step 120. is carried out by means of a programming interface of the application or application programming interface (API) which is able to provide the absolute positioning of said acquisition device (D, T).
3. The method (100) according to any one of the preceding claims, characterized in that in said step 130. a plurality of depth maps are used, connected to each other by computational geometry techniques and robotics.
4. The method (100) according to any one of the preceding claims,characterized in that said step 130. comprises a sub-step 131. of correcting said three-dimensional model (M) in real time, wherein, given said trajectory (R), deviation vectors are generated in order to identify a discontinuity between said shots (S) and the logical and temporal order thereof and to connect the discontinuity points by means of a roto- translation matrix.
5. The method (100) according to any one of the preceding claims, characterized in that said step 140. comprises a first sub-step 141. wherein said three-dimensional model (M) is transformed into a textured mesh provided with position, orientation, scale, opacity and color.
6. The method (100) according to the preceding claim, characterized in that said step 140. comprises a second sub-step 142. wherein several textured meshes associated with three-dimensional models (M) of images (I) acquired on different days are joined.
7. A system (A) for remotely monitoring a geographical site (C) which is able to carry out the method according to the preceding claims 1-6, comprising: one or more devices (D, T) for the acquisition of pluralities of images and / or sequences of photograms (I) of said site (C) and for a first processing thereof, so as to obtain a three-dimensional model (M) of said site (C); a remote server or cloud (B) which is able to receive the data from said one or more devices (D, T), being able to associate said three-dimensional model (M) with said plurality of images and / or sequences of photograms (I), in order to display on said three-dimensional model (M) geometric parameters of the site (C) like minimum, maximum and average width, length and depth.
8. The system (A) according to the previous claim, characterized in that said one or more devices (D) is a smartphone comprising a camera (T) .
9. The system (A) according to any one of claims 7 and / or 8, characterized in that on each device (D, T) an operating environment or application is installed, which an operator can access by an authentication with username and password and select the target construction site (C), from an available list.
Citation Information
Patent Citations
Performing object modeling by combining visual data from images with motion data of the image acquisition device
EP4064192A1