Method for capturing images and for corresponding three-dimensional reconstruction and system that implements the method
A smartphone-based method and system for image capture and three-dimensional reconstruction address the challenges of cost and complexity in existing technologies, providing accurate and integrated data for urban planning and maintenance.
Patent Information
- Application Number
- PCT/EP2025/072458
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-09
- Filing Date
- 2025-08-05
- Publication Date
- 2026-02-12
AI Technical Summary
Current methods for capturing images and generating three-dimensional models of urban and non-urban areas are expensive, impractical, and require specialized personnel and equipment, leading to outdated data and integration challenges with industrial processes.
A method and system using a smartphone to capture images, associate cartographic data, and perform three-dimensional reconstruction through a combination of image processing, AI, and SfM techniques, enabling automatic data collection and integration with existing software systems.
Enables cost-effective, precise, and up-to-date data capture and reconstruction without specialized skills, facilitating integration with industrial processes like fiber optic infrastructure design.
Smart Images

Figure EP2025072458_12022026_PF_FP_ABST
Abstract
Description
[0001]
[0002] Title: Method for capturing images and for corresponding three- dimensional reconstruction and system that implements the method
[0003] DESCRIPTION
[0004] Field of application
[0005] The present invention relates to a method for capturing images, in particular images of urban areas and generally landscapes, and for a subsequent three-dimensional reconstruction of the images themselves.
[0006] The present invention also relates to a system that implements the method.
[0007] More in detail, the invention relates to a method able to identify the objects present in a captured image, in addition to deriving a corresponding three- dimensional model, which is also georeferenced by means of an acquisition device, such as a smartphone and the like.
[0008] Hereinafter, the description will be directed to a method and a system for capturing images of urban, non-urban areas and of landscapes by means of a smartphone, but it is well clear that the same shall not be considered limited to this specific use.
[0009] Prior art
[0010] Currently, to represent the surrounding reality and its evolution it is necessary to use data collection and analysis methods that are often expensive, impractical, and not immediately easy to use.
[0011] Currently, there are several methods and databases available to private and / or public entities to keep their real estate assets up to date or to map urban and non-urban areas.
[0012] In the case of fiber optic design, for instance, it is essential to adequately estimate the number of property units, commercial buildings and offices that will fall within the service coverage area, in order to subscribe to a plan.
[0013] A widespread problem in almost all countries is the lack of precise knowledge of the location of buildings, their state of maintenance, the house numbers associated with buildings, the property units along a street, and the like.
[0014] Currently, public and / or private bodies that must carry out the checks described above use well-known technologies, such as “Google Street View” or “Apple Look Around”, containing more or less up-to-date data for Western countries and obsolete ones for other countries.
[0015] The lack of precise information makes it impossible to make decisions that involve significant budgets without adequately estimating the catchment area; to fill this impossibility, attempts are often made to integrate the missing information through field inspections.
[0016] Field inspections, however, involve both a very high cost for pedestrian surveys, but also a choice of appropriate hardware and software to keep the databases synchronized and updated.
[0017] On the one hand, therefore, information should be extracted from the above-mentioned maps, when available, with specific procedures that involve the development of specialized software; on the other hand, these data should be integrated with other data coming from field surveys, with all the problems and costs that come with it.
[0018] Currently, the Mobile Mapping System - MMS survey methodology is very widespread, which involves the use of a LiDAR system combined with a spherical camera and an inertial system, mounted on a vehicle or a backpack, to be able to reach restricted traffic areas or areas inaccessible by vehicles.
[0019] These systems, in addition to being extremely expensive, require highly specialized personnel to carry out the surveys and a very specific and professional data processing procedure. Once the processed data has been obtained, the known systems do not yet provide useful information for budget estimation and cannot be integrated with other software used in industrial processes, such as those related to the design and construction of fiber optic infrastructures.
[0020] Moreover, in most countries, customs procedures for sending such equipment involve very high costs and lengthy procedures, both in terms of money and time.
[0021] A further problem is data obsolescence; in fact, once this has been obtained, after having organized the purchase or rental of equipment, shipping, customs procedures, installations, staff training, acquisitions, post-processing, and extraction of data important for design, it is evident that the significant data are already obsolete.
[0022] In light of the prior art drawbacks and unsolved problems, an object of the present invention is to provide a method that enables anyone possessing an acquisition device, such a smartphone, to carry out surveys both by car and on foot without the need for particular skills.
[0023] A further object of the present invention is to provide an automatic method that allows obtaining all of the data necessary and preparatory to specific purposes such as, for instance, the installation of optical fibre.
[0024] Still another object of the present invention is to provide a method for generating a three-dimensional model that is reliable and accurate.
[0025] A further object of the present invention is to make a system that implements the method of geographic image detection.
[0026] Therefore, the technical problem underling the present invention is to provide a method and a system for detecting images of urban areas and the like and for the three-dimensional reconstruction.
[0027] Summary of the invention
[0028] The idea underling the present invention is to provide a method and a system that allow capturing images of urban and non-urban areas, to recognise the objects present in the images and to carry out a three- dimensional reconstruction.
[0029] An object of the present invention is a method for capturing one or more images by an operator moving along a path, and corresponding three- dimensional reconstruction, comprising the following steps:
[0030] 110. identifying the type of an acquisition device that accesses a server, by means of a program;
[0031] 130. by means of said acquisition device, capturing one or more images according to an acquisition time interval t precalculated by said program, along said path;
[0032] 140. by means of said device, assigning a name to the plurality of images, sending said one or more images to said server and associating cartographic data with said one or more images, by means of said program;
[0033] 150. by means of said device and said program, performing specific requests or queries relating to at least one image of said one or more images to said server and processing a reply by means of said server;
[0034] 160. by means of said server, processing a three-dimensional model associated with the specific image object of the query; and
[0035] 170. exporting said processed queries and said three-dimensional model by means of said program.
[0036] Furthermore according to the invention, said method comprises a further step 120. between step 110. and 130. of printing a three-dimensional support for installing said device in a pre-established position, inside a means of transport.
[0037] Still according to the invention, said step 150. comprises the following further sub-steps:
[0038] 151. optimising said queries by means of a Mobile VLM type vision model; and
[0039] 152. performing the prompting to identify and analyse the road elements in said one or more images.
[0040] Preferably according to the invention, said step 130. comprises a sub-step 131. of calculating said time interval t of acquisition of said images, as a function of the distance d of the device from the photographed object and the speed V of movement of said operator along said path, according to the formula: where SD is the sample distance obtained through the following formula: where s is the size of the camera sensor of said device; F is the focal length and R is the resolution of the camera of said device in pixels.
[0041] Still according to the invention, the estimation of said distance d is done by means of a depth map, generated by means of the API ArCore application for Android mobile or AR Kit for iOS mobile.
[0042] Furthermore according to the invention, said speed V is also detected by the GPS system present in said device.
[0043] Still according to the invention, in said step 130., said program processes the path followed by the user and associates a corresponding cartography.
[0044] Preferably according to the invention, said three-dimensional processing of said step 160. occurs by means of the Structure from Motion - SfM technique.
[0045] Still according to the invention, said step 150. is performed by a pretrained Al artificial intelligence module.
[0046] Furthermore according to the invention, said Al artificial intelligence module is trained according to the method, comprising the following steps:
[0047] 210. preparing the dataset by converting images and data into a Comma Separated Values - CSV file format, wherein each data includes a unique identifier (id), the path of the image and the associated metadata for each road element;
[0048] 220. performing a first fine tuning of the data, by using Low-Rank Adaptation - LoRa type method;
[0049] 230. performing a complete fine tuning of the model, thus optimising all parameters, ensuring accuracy in condition detection and analysis; and
[0050] 240. adjusting the hyperparameters, by means of the Bayesian statistics, to fit the specific dataset and hardware constraints.
[0051] A further object of the present invention is a system able to perform the previously described method, comprising one or more acquisition devices for capturing one or more images of urban or non-urban environments, a data processing program installable on said one or more acquisition devices, at least one remote server operatively connected to said one or more devices by means of said program, able to process said one or more images by means of an Al artificial intelligence module, and to generate a three-dimensional model corresponding to at least one image of said one or more images and to generate data associated with the objects present in said at least one image of said one or more images.
[0052] Brief description of the drawings
[0053] In the drawings: figure 1 shows a schematic view of a block diagram of the operating steps of the method object of the present invention; figure 2 shows a schematic view of the components of the system that implements the method object of the present invention; figure 3 shows a schematic view of graphic icons associated with requests that a user may express; figure 4 shows a view of a processed image in which an object has been highlighted, in particular a tree; figure 5 shows a view of a processed image in which an object has been highlighted, in particular a pit; figure 6 shows a view of a processed image in which an object has been highlighted, in particular a pothole in the asphalt; figure 7 shows a schematic view of a three-dimensional model associated with an image captured along a path; and figure 8 shows a block diagram of a training method used by the method object of the present invention.
[0054] Detailed description
[0055] With reference to figure 1, the method 100 for detecting geographic images and three-dimensional reconstruction, object of the present invention, comprises a plurality of steps.
[0056] The method 100 is implemented by a system E comprising at least one acquisition device 1 for capturing a plurality of images I of urban or non- urban environments, a program A or mobile application installable both on at least one acquisition device 1 and on other remote devices, to connect to a remote server or cloud C, operatively connected to said at least one acquisition device 1 and / or to said remote devices by means of said program A.
[0057] The acquisition device 1 is preferably a smartphone.
[0058] However, without departing from the scope of protection of the present invention, the acquisition device 1 may also be any other device provided with a camera.
[0059] The server C is able to process said plurality of images I by means of an artificial intelligence module, and is able to generate a three-dimensional model corresponding to at least one image I of said plurality of images I, with which associating data relating to objects present in said at least one image I of said plurality of images I, upon request of an operator or of a remote user by means of said program A.
[0060] In an identification step 1 10, a user has an acquisition device 1 at his / her disposal, through which accessing the remote server C.
[0061] In the identification step 110, the program A recognises the type of the acquisition device 1 and selects a STL. type file, i.e. the well-known data transmission format in the rapid prototyping industry.
[0062] In a second step 120., the user may decide to download the .STL file and to 3D print the support 2 for his / her acquisition device 1 , and installs it in a means of transport such as a car or the like, so as to be able to insert the device 1 in a unique and recognisable position, to perform the shooting.
[0063] In a third step 130., the acquisition device 1 performs a survey, i.e. an acquisition of images I, according to an image I acquisition time interval t, while the user is travelling along a route R on board a means of transport or on foot, as shown in figure 5.
[0064] In a sub-step 131., the image I acquisition time interval t is then calculated.
[0065] The I image acquisition time interval t, which may also be indicated as a predefined number of shots for a predefined amount of meters, or acquisition interval, or distance between shots, is calculated automatically, therefore not requiring any intervention or adaptation of the operating mode by the user, taking into account two fundamental parameters such as the distance D of the object to be photographed with respect to the position of the device 1 and the user’s travel speed V by the means of transport or on foot, of the means on which the user is travelling, as it is now described in detail.
[0066] The I image acquisition time interval t must be such as to obtain a suitable overlapping of images I, in order to generate three-dimensional models and such as to obtain the image of an object repeated on several visual cones, in order to increase the possibility of having at least one image that frames a certain element, so that it can then be analysed by the Artificial Intelligence - Al. Not only does the overlapping between images I offer the possibility of seeing the same element from multiple images, thus excluding or drastically reducing shadow cones or obstacles observable in a particular image, for instance, a parked vehicle obstructing the view of a building, but it also ensures, or otherwise excludes, the possibility of obtaining a three-dimensional model, as shown in figure 5, from the frames using the Structure from Motion - SfM technique. In particular, the 3D model is generated by means of the SfM methodology if it concerns multiple images I, i.e. more than one. Instead, if the 3D model is only used to perform measurements within a single image I, the overlapping of the depth map on the image I of interest is used. In this way, the operator, having two perfectly overlapped images, the RGB image with its depth map behind it, visually takes measurements on the image I, whereas actually the program A takes them in background on the depth map.
[0067] Moreover, image overlapping also helps training the Al module, which will be described hereinafter, thus offering the possibility of solving some ambiguities.
[0068] By using the speed V at which the vehicle travels or at which the operator walks on foot, also detected by the Global Positioning System - GPS present in the device 1, and the distance D from objects or buildings detected by the depth map, obtained in real time, the optimal time interval t for capturing images I is estimated, in order to have a sufficient overlapping between the subsequent shots and at the same time to avoid an excessive number of shots, which would lead to an overload in terms of space, occupied by the acquired data, and of processing time. The time interval t is calculated in meters travelled from one shot to another, according to the following formula: where t = shot speed, i.e. the time between 2 shots
[0069] R = frame resolution in pixels
[0070] V= speed at which the operator moves with the mobile device D
[0071] SD = sample distance or size of each pixel calculated according to the following formula where d = distance of the smartphone from the object which one wishes to recognize, such as trees, asphalt, buildings and the like, as preset by the operator or user s = camera sensor size
[0072] F = focal length
[0073] R = camera resolution in pixels
[0074] The depth map is generated via the well-known application API ArCore for Android mobile or AR Kit for iOS mobile.
[0075] A depth map, in three-dimensional graphics computer applications or computer vision, is an image or image channel that contains information relating to the distance of the surfaces of objects in a scene, from a given point of view.
[0076] Therefore, the use of the depth map ensures the overlapping between images.
[0077] The time interval t will be greater for objects that the depth map estimates to be further away from the device 1 and vice versa.
[0078] As above reported, another considered factor is the speed of the user’s means of transport. Still considering the objective of obtaining a fixed overlapping, by knowing the distance D of the object to be captured from device 1, the result will be: as the distance D increases, the field of view increases with an equal field of view - FOV and therefore the focal length F decreases the image definition; when the distance D decreases, the field of view decreases and the image definition increases.
[0079] It follows that the closer the distance D measured by the depth map, the closer the photographs of the device 1 and therefore the smaller the time interval t.
[0080] Since the images are always the same in terms of size and dimension, it is also possible to automatically calculate the remaining acquisition autonomy, i.e. the distance in meters that can still be travelled along the path P, taking into account the memory space still available in the device 1.
[0081] Prior to starting the acquisition step 130., the user may set the type of survey, i.e. the object he / she intends to capture, as in the camera options of the device D, for instance houses, sidewalks, house fences, and the like, using icons 3, as shown in figure 3, each associated with the most frequent road elements present in urban and non-urban environments.
[0082] Considering that the depth map is generated over the entire scene captured by the image I, it becomes extremely important to understand which data set is of interest to the user.
[0083] For instance, if a user was interested in mapping and obtaining information relating to buildings, and these are surrounded by fences, there would be a risk of calculating a distance closer to that of the building, also averaging the distance from the fences in the calculation.
[0084] This problem is solved by exposing the user to a series of options, for instance buildings in the presence of fences which, once set, will allow only using the top part of the depth map in this case.
[0085] This can be done by generating a regular grid on the depth map and deciding a priori not to use all of the parts of the grid equal to the lower 50%.
[0086] Another aspect to consider is the presence of obstructing objects such as street trees or parked cars.
[0087] In this case as well, special filtering techniques are used for data outside the average range.
[0088] The average interval is determined based on the average of the shots previously taken.
[0089] The images captured by the acquisition device 1 are georeferenced, since they are linked to the path followed by the user.
[0090] At the end of the image I acquisition, in a fourth step 140., the user attributes a name or order code to the image I package, and sends the images I to the server or cloud C.
[0091] The image I package is associated with the date, name of the city and name of the street by the program P, deriving them from the Web Map Service - WMS type cartography that can be connected to the images, using Google Map, Opens Street Map or others as a cartographic base.
[0092] At the same time, the program A processes the path P followed by the user and associates a map.
[0093] From the images I saved in the server C and from the geographic and cartographic information associated with the captured images I, it is possible to process different requests to obtain information, both through the user’s device 1 and through other remote devices by other users able to connect, through said program A, to a Software-as-a-Service - SaaS type website, i.e. a cloud computing service that allows users to access a platform through a cloud application. In the concerned invention, the SaaS website allows accessing the georeferenced information saved on the server C, through pre-established queries that are pre-established questions and are summarised in an icon interface 3, and downloading them in standard format.
[0094] In particular, in a fifth step 150. of the method, the user may authenticate himself / herself using his / her device 1 to the program A and select a portion of the path P.
[0095] In a sixth step 151., the user may send a specific request or query to the server C, through the SaaS site, such as retrieving a subset of the images I containing a building, or a specific request about a building, such as the street number, the number of floors, the type of building and the number of property units present in each building, using the icons 3.
[0096] In a seventh step 152., queries related to asphalt, trees, buildings, road signs, road products, traffic lights are optimized by launching only the expected one(s), thus allowing the use of this Al on the user’s mobile device D as well, through a vision model such as, for example, the Vision
[0097] Language Model for Mobile Devices - Mobile VLM.
[0098] MobileVLM is used to optimize queries and filter visual data, allowing users to capture only relevant data based on object type.
[0099] SfM (Structure from Motion) is used to filter out irrelevant objects, such as cars blocking roads. SfM is also employed to automatically control image overlap based on depth data.
[0100] Moreover, prompting is also used, that is, supports are provided to the Al module to obtain the desired result.
[0101] Various prompting techniques are employed. In this step, specific prompts that guide the model to accurately identify and analyse road elements in the input images are created.
[0102] In this step there is also an identification process through queries.
[0103] For data collection, images are uploaded by users or come from a database of street-level images, usually selected by the user for streets, cities, or specific areas.
[0104] For unique identification, each image is assigned a unique ID connected to its geographic coordinates and timestamps. This ensures that the data produced by the analysis is accurately mapped to the specific location where the photo was taken, providing a reliable reference for urban planning and maintenance.
[0105] Since all questions are designed to have answers of the type: yes / no / how much = 0 / 1 / n, each request regarding a specific dataset will correspond to precise answers regarding each image.
[0106] Each image has its own coordinates, so a specific question will correspond to a geospatial response that can be exported in the GIS sharing standard formats (OCG).
[0107] In a sixth step 160., the server C processes these requests by means of an Artificial Intelligence - Al module, and moreover processes a three- dimensional model and returns the requested data to the user who can also export them to his / her device 1, in a seventh step 170.
[0108] During post-processing, the software automatically detects and masks the portion of the image containing the support, so it does not interfere with 3D reconstruction or Al-based visual analysis.
[0109] The Al module uses the well-known Large Language and Vision Assistant - LLaVA engine, namely a large, end-to-end trained multimodal model that combines a vision encoder and Vicuna for general visual and language understanding.
[0110] LLaVA employs the well-known open-source type Large Language and Model Assistant - LLaMA, which is a language decoder, capable of performing language-only instruction optimization projects. It is based on the pre-trained CLIP visual encoder ViT-L / 14 for visual content processing.
[0111] The vision encoder and the language decoder are connected to each other through an attention mechanism.
[0112] The vision encoder extracts visual features from input images and connects them to language embeddings, i.e. to distributed representations of words, through a trainable projection matrix, translating the visual features into language embedding tokens, i.e. a set of digital information, thus bridging the gap between text and images.
[0113] As for the automatic generation of three-dimensional models associated with each two-dimensional image I, by selecting the specific image I of the path P, a methodology known as Structure from Motion - SfM is used.
[0114] This methodology is a computational technique that allows reconstructing the shape of objects through the automatic collimation of points from a set of photos, or photogrammetry, based on computer vision algorithms. The SFM extracts the notable points from the individual photos, infers the photographic parameters and cross-references the recognizable points on multiple photos, thus finding the coordinates in the space of the points themselves. At the output of the first calculation step, the SFM builds a cloud of points, with the original colour, forming a spatial model of a “solid” photo of the captured objects.
[0115] Once the cloud of points has been generated and knowing the angles and positions of the image capture centers, it is possible to apply a rasterization technique such as Gaussian Splatting and / or Mip Splatting.
[0116] The final result of the three-dimensional model processing will then be the dense cloud of points, or the normal textured mesh, or the Gaussian Splatting and / or Mip Splatting cloud.
[0117] Finally, it is also possible to generate reports. These reports, in addition to the information generated by LLaVA, also include georeferencing information.
[0118] Georeferencing is ensured by the sensors of the device 1, by the GPS and by the inertial system, or even by advanced technologies such as Augmented Reality (AR) and Computer Vision to further improve the precision of localization and orientation. For instance, Google’s ARCore and Apple’s ARKit use the device’s 1 camera and the inertial movement unit (IMU) type motion sensors to map the surrounding environment and improve localization accuracy.
[0119] In order to be able to reply to the requests sent via the SaaS site by the user, the Al module must be trained.
[0120] For the training method 200, a fine tuning of the data is performed, in order to refine the recognition of the objects of interest, in the captured images I.
[0121] In a first step 210., the Dataset is prepared.
[0122] In particular, the data is converted into a Comma Separated Values - CSV file format where each sample includes a unique identifier (id), the image path and the associated metadata for each road element.
[0123] In a second step 220., a first fine tuning of the data is performed, using Low-Rank Adaptation - LoRa methods.
[0124] This method is effective for limited task-specific data. It optimizes the checkpoints of LLaVA, by tuning a subset of parameters, making it suitable for the specific purposes of the invention.
[0125] In a third step 230., a complete fine tuning of the model is performed, that is, when the specific data is deemed sufficient, the entire model is optimized by LLaVA checkpoints. This process involves the optimization of all parameters, ensuring accuracy in the detection and analysis of conditions such as asphalt, types of tree, features of buildings, road signs, various road products and traffic lights.
[0126] In particular, for asphalt, the Al module detects and analyses cracks, potholes and repairs in asphalt surfaces.
[0127] For trees, the Al module counts the number of trees and identifies their type (broadleaf / deciduous).
[0128] For buildings, the Al module analyses buildings, focusing on common features such as structure, condition, and use.
[0129] For road signs and billboards, the Al module detects and analyses both vertical and horizontal road signs to check the presence, status, and type thereof.
[0130] For the various road products, the Al module identifies various products at the street level, including electrical cabinets, hollow pits, poles, sewers, sidewalks, and fiber optic connection elements.
[0131] For traffic lights, the Al module detects and analyses traffic lights to check the presence, functionality, and status thereof.
[0132] For buildings, the Al module analyses buildings for street number, number of floors, residential status, presence of commercial floors, and estimated number of families living therein.
[0133] In a fourth step 240., the hyperparameters are adjusted by means of Bayesian statistics.
[0134] In particular, hyperparameters are tuned to fit the specific dataset and hardware constraints. This step ensures that the model achieves optimal performance for recognition and analysis tasks, providing accurate data on road elements that are crucial for urban planning and maintenance activities.
[0135] As it is evident from the above description, the advantage of the method and system object of the present invention is to continuously capture images of scenarios or road entities during the execution of a path and both identify the objects within the images and reconstruct a three- dimensional model starting from the two-dimensional images.
[0136] Obviously, a person skilled in the art, in order to meet contingent and specific needs, may make numerous modifications and variations to the support described above, all of which are however included within the scope of protection of the invention as defined by the following claims.
Claims
CLAIMS1. Method (100) for capturing one or more images (I) by an operator moving along a path (P), and corresponding three-dimensional reconstruction, comprising the following steps:
110. identifying the type of an acquisition device (1) that accesses a server (C) , by means of a program (A) ;130. by means of said acquisition device (1), capturing one or more images (I) according to an acquisition time interval t, pre-calculated by said program (A), along said path (P);140. by means of said device (1), assigning a name to the plurality of images (I), sending said one or more images (I) to said server (C), and by means of said program (A) associating cartographic data with said one or more images (I);150. by means of said device (1) and said program (A), performing specific requests or queries relating to at least one image (I) of said one or more images (I) to said server (C) and processing a reply by means of said server (C);160. by means of said server (C), processing a three-dimensional model associated with the specific image (I) object of the query; and170. exporting said processed queries and said three-dimensional model by means of said program (A) .
2. Method (100) according to the preceding claim, characterized by comprising a further step 120. between said steps 110. and 130. of printing a three-dimensional support (2) for installing said device (1) in a pre-established position, inside a means of transport.
3. Method (100) according to any one of claims 1 or 2, characterized in that said step 150. comprises the following further sub-steps:
151. optimizing said query by means of a Mobile VLM type vision model; and152. performing the prompting, to identify and analyse the road elements in said one or more images (I).
4. Method (100) according to any one of the preceding claims, characterized in that said step 130. comprises a sub-step 131. of calculation of a time interval t of acquisition of said images (I), as a function of the distance d of the device (1) from the photographed object and the speed V of movement of said operator along said path (P), according to the formula:where SD is the sample distance obtained through the following formula:where s is the size of the camera sensor of said device (1); F is the focal length and R is the resolution of the camera of said device (1) in pixels.
5. Method (100) according to any one of the preceding claims, characterized in that the estimation of said distance d is done by means of a depth map, generated by means of the API ArCore application for Android mobile or AR Kit for iOS mobile.
6. Method (100) according to any one of the preceding claims, characterized in that said speed V is also detected by the GPS system present in said device (1).
7. Method (100) according to any one of the preceding claims, characterized in that in said step 130., said program (A) processes the path (P) followed by the user and associates a corresponding cartography.
8. Method (100) according to any one of the preceding claims, characterized in that said three-dimensional processing of said step 160. occurs by means of the Structure from Motion - SfM technique.
9. Method (100) according to any one of the preceding claims, characterized in that said step 150. is performed by a pre-trained Al artificial intelligence module.
10. Method (100) according to the preceding claim, characterized in that said Al artificial intelligence module is trained according to the method (200), comprising the following steps:
210. preparing the dataset by converting images and data into a Comma Separated Values - CSV file format, wherein each data includes a unique identifier (id), the path of the image and the associated metadata for each road element;220. performing a first fine tuning of the data, by using Low-Rank Adaptation - LoRa type methods;230. performing a complete fine tuning of the model, thus optimizing all parameters, ensuring accuracy in condition detection and analysis; and240. adjusting the hyperparameters, by means of the Bayesian statistics, to fit the specific dataset and hardware constraints.
11. System (E) able to perform the method (100) according to the preceding claims 1-10, comprising: one or more acquisition devices (1) for capturing one or more images (I) of urban or non-urban environments; a data processing program (P) installable on said one or more acquisition devices (1); at least one remote server (C) operatively connected to said one or more devices (1) by means of said program (A), able to process said one or more images (I) by means of an Al artificial intelligence module, and to generate a three-dimensional model corresponding to at least one image (I) of said one or more images (I) and to generate data associated with the objects present in said at least one image (I) of said one or more images (I).
Citation Information
Patent Citations
Image processing apparatus for generating virtual viewpoint image and method therefor
US20180204381A1
Three-dimensional reconstruction method, system and apparatus based on aerial photography by unmanned aerial vehicle
US20200255143A1
Localization processing service and observed scene reconstruction service
US20230400327A1