Method and system for identification and cataloguing of objects in a scene using 2d and 3D data modalities
The method and system generate a 3D model from drone images, using AI-based 3D semantic segmentation and patch processing to efficiently identify and catalog objects, addressing integration challenges and enhancing remote monitoring capabilities.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-11
- Publication Date
- 2026-03-26
AI Technical Summary
Existing drone-based image processing methods struggle with integrating 2D and 3D information, requiring manual data processing and lacking suitability for accurate detection and cataloguing of objects, especially in complex scenes with variations in viewpoint and occlusion.
A method and system that generates a 3D model from drone-captured images, converts it into a point cloud representation, and uses AI-based 3D semantic segmentation to identify and catalog objects, with overlapping patch processing and instance identification to enhance accuracy and reduce manual effort.
Efficiently and accurately identifies and catalogs objects in 3D models, enabling remote monitoring and detection of defects with reduced manpower and time, while handling large-scale scenes with high accuracy and scalability.
Smart Images

Figure MY2025050060_26032026_PF_FP_ABST
Abstract
Description
[0001] METHOD AND SYSTEM FOR IDENTIFICATION AND CATALOGUING OF
[0002] OBJECTS IN A SCENE USING 2D AND 3D DATA MODALITIES
[0003] FIELD OF INVENTION
[0004] The present invention generally relates to the identification of objects using computer vision. More particularly, the invention relates to an artificial intelligence-assisted method for the identification, cataloguing, and monitoring of objects present in captured scene or images and reconstructed 3D data using computer vision.
[0005] BACKGROUND
[0006] Computer vision is widely employed for analysing images to detect objects present within the images. Object recognition involves extracting information from images or 3D models, such as meshes and point clouds to detect and identify objects within a scene. This task, which humans perform effortlessly, poses challenges for computer-based methods. These challenges include variations in viewpoint, size, and occlusion of objects, especially in images, which can affect the accuracy of recognition algorithms.
[0007] Advancements in image capture technologies, particularly through the use of drones, also known as unmanned aerial systems or unmanned aerial vehicles, as well as other manned or unmanned vehicles / robots equipped with cameras and sensors, offer various solutions to overcome the limitations of traditional inspection methods. By capturing image data remotely, drones enable inspectors to assess objects and infrastructure from a safe distance, reducing the need for physical presence at the inspection site. Real-time viewing of drone-captured images enhances safety and efficiency, allowing inspectors to identify potential issues without being onsite.
[0008] However, existing methodologies for reviewing drone-captured data have limitations, including difficulties in integrating 2D and 3D information and challenges in data analysis. Current approaches require users to toggle between different views and / or data modalities and manually process the data, leading to inefficiencies and potential oversight of critical information. There is a growing need for improved methodologies that streamline the review process and enhance the usability of drone-captured data for object inspection and analysis. One method, disclosed in US Patent No. US 11,216,663 Bl, extracts 2D and 3D information from images captured by unmanned aerial vehicles. This method enables real-time visualization of objects and scenes during drone navigation, benefiting activities such as inspection and asset management. However, this prior art method primarily focuses on user interface display and lacks suitability for detecting and cataloguing objects in 3D.
[0009] Another method, as described in Chinese Patent Application No. CN116091778A, involves performing semantic segmentation processing on target scene data. This method includes steps such as data acquisition, semantic segmentation, input into a segmentation model, and splicing of segmentation results. While effective for large-scale natural scenes, the disclosed method identifies a target object by scanning the surface of the target object using a laser beam. However, this method does not address 3D data processing or asset identification for the accurate detection, cataloguing and monitoring of objects.
[0010] Hence, there is a need for an improved system and method capable of processing data from a captured scene or from one or more images, including drone images, for the effective detection, cataloguing and monitoring of assets. Such a system should incorporate artificial-intelligence (Al) based processing to enhance prediction accuracy and minimize manual effort, along with scalability to accommodate increased workloads. Further, the needed system and method would be capable of being used in a wide variety of detections, such as for monitoring equipment installations, including those on communication towers, transmission and distribution towers, indoor equipment, and the detection of other desired objects in a captured scene or image. Additionally, the system should enable the processing of more towers or identify objects in a wide area or scene at lower costs compared to manual inspection methods.
[0011] SUMMARY
[0012] In one aspect, the present invention provides a method for identifying one or more objects in a scene, the method comprising the steps of: receiving and storing an input scene containing the one or more objects; generating a three-dimensional (3D) model of the input scene, including the one or more objects therein; creating a first point cloud representation from the 3D model; characterized by: converting the first point cloud representation into a plurality of overlapping point cloud patches; identifying a type of the one or more obj ects in the overlapping point cloud patches using a trained 3D semantic segmentation method; combining the type-identified objects in the overlapping point cloud patches to create a second point cloud representation; identifying instances of each of the type-identified objects in the second point cloud representation using an instance identification method; and presenting the type-identified objects, including each instance thereof, within the 3D model and / or an output scene.
[0013] The image processing method mentioned above efficiently and accurately identifies multiple instances of objects from the 3D model generated from the captured input scene or images and determines the type of these objects, requiring less manual effort. Furthermore, in one instance, the method of dividing the entire point cloud into multiple overlapping 3D patches mitigates the limitation of working with less dense versions of the original high-density point cloud due to memory constraints. This approach is more practical as it requires less memory, enables processing of arbitrary amounts of dense 3D point cloud data, maximizes accuracy in identifying the objects from the 3D model generated from the captured images or scene, along with scalability to accommodate increased workloads.
[0014] In one embodiment, the method for identifying one or more objects in the scene further includes cataloguing the type-identified objects, including each instance, in the second point cloud representation. This is advantageous as it allows users to remotely compare the catalogued objects with a predefined equipment library. This comparison facilitates the collection of equipment information and helps to identify unauthorized or missing equipment.
[0015] In one embodiment of the invention, the 3D model created is a digital twin of the captured images or, more appropriately, a digital twin of the input scene or images captured by a UAV. Advantageously, the digital twin accurately presents the objects in the 3D model generated from the captured images or input scene in a virtual environment and facilitates further processing to identify the types of objects, remote monitoring and detect any associated defects. Another advantage of detecting objects in 3D is that it provides the user with accurate 3D information about the detected objects, such as precise location and dimensions of the detected objects in the real -world, which cannot be computed or inferred from the captured 2D images.
[0016] In one embodiment of the invention, the 3D model is represented in a mesh format. This is advantageous as it provides a highly detailed surface representation of the objects, allowing for precise modelling of complex geometries.
[0017] In another embodiment of the invention, the 3D model is represented in a point cloud format. Advantageously, the point cloud format provides the ability to capture and represent vast amounts of data points with high accuracy, leading to precise 3D model representations. Additionally, this format is particularly effective for capturing detailed spatial information, such as the exact locations and dimensions of objects within a scene, which enhances the detection and analysis of objects in 3D space. Another advantage of using point clouds for the proposed artificial intelligence (Al) enabled image processing and object detection is their high compatibility with downstream tasks, such as segmentation and patch generation. Point clouds also simplify the training of highly accurate Al models, leading to more efficient and precise object detection from the captured scene or images, as well as improved 3D model creation.
[0018] In another embodiment, the method of creating the first point cloud representation includes: extracting a plurality of points from a surface of the 3D model represented in the mesh format to create a point cloud, each point representing a geometric coordinate of each object; and eliminating at least a portion of the plurality of points of the point cloud using a downsampling method.
[0019] In another embodiment, the method of creating the first point cloud representation includes: eliminating at least a portion of the plurality of points of the 3D model represented in the point cloud format using a downsampling method.
[0020] If the 3D model is in the mesh format, it is converted into a first point cloud representation by first extracting a plurality of points from the surface of the 3D model and eliminating certain redundant points while keeping all the important contextual details. If the 3D model is already in the point cloud format, certain redundant points are eliminated while keeping all the important contextual details. This approach is advantageous as this eliminates the necessity for high-end hardware to process the dense point cloud dataset while preserving all the necessary domain context and / or information and hence maintaining accuracy of the object identification process. In one embodiment, the method of identifying the one or more objects in the overlapping point cloud patches further includes the steps of: predicting on a plurality of points on one or more objects in each of the overlapping point cloud patches using the trained 3D semantic segmentation method; aggregating the predictions on the plurality of points in the overlapping point cloud patches; and refining the aggregated predictions using a local neighbourhood consensus method.
[0021] The method of identifying the objects includes the steps of comparing the points and their associated predictions on the objects that are common to the overlapping point cloud patches and comparing them in the local neighbourhood of the objects in the overlapping point cloud patches using the local neighbourhood consensus methodT This approach maximizes accuracy in identifying the object instance and cataloguing the objects in the captured images, 3D model and / or scene. Moreover, converting the initial point cloud representation into multiple overlapping point cloud patches is advantageous as this enables the processing of objects across expansive areas, 3D scenes, meshes, or point cloud datasets.
[0022] In another embodiment, the method of presenting the type-identified objects within the 3D model, output scene or captured images further includes the step of mapping a plurality of points on the type-identified objects onto at least one 3D mesh surface. This feature is beneficial as this enables users to directly visualize the object predictions on the 3D mesh, which is a 3D representation of the input scene or the captured images.
[0023] In another embodiment, creating a visual representation of the 3D model further includes the step of projecting the 3D model and / or the type-identified objects into at least one captured 2D image or input scene. This feature is advantageous as this enables users to easily identify the objects within the captured 2D images or input scene with optimal or unobstructed views of the objects.
[0024] In another embodiment, the method of identifying the type of an object and cataloguing the type- identified objects include identifying, segregating and annotating a plurality of objects in close proximity which are difficult to distinguish from one-another. This feature is advantageous as this enables the detection of small objects, which are installed on close proximity on a communication or transmission tower, as separate instances or objects. This feature also facilitates the detection of overlapping objects in the captured images, which is otherwise not possible during manual visual inspection of the images.
[0025] In another aspect, the present invention includes a system for identifying and cataloguing the one or more objects in the image. The system comprising: a processor; and at least one storage device in communication with the processor, the storage device being configured to store instructions, which when executed by the processor enable the system to perform the steps of: receiving and storing an input scene containing the one or more objects; generating a three-dimensional (3D) model of the input scene, including the one or more objects therein; creating a first point cloud representation from the 3D model; characterized by: converting the first point cloud representation into a plurality of overlapping point cloud patches; identifying a type of the one or more obj ects in the overlapping point cloud patches using a trained 3D semantic segmentation method; combining the type-identified objects in the overlapping point cloud patches to create a second point cloud representation; identifying instances of each of the type-identified objects in the second point cloud representation using an instance identification method; and presenting the type-identified objects, including each instance thereof, within the 3D model and / or an output scene.
[0026] The system configured to perform the above steps offers efficient and accurate identification of objects in the 3D model and / or input scene as well as the captured images, facilitating cataloguing and maintenance of an object database. Simultaneously, the system aids in monitoring the status of each object, such as health of telecommunication equipment installed at remote locations, with minimal manual intervention, thereby enhancing operational efficiency.
[0027] In one embodiment, the system comprises an unmanned aerial vehicle (UAV) such as a drone to capture one or more images of an area or scene containing the objects. Advantageously, this helps to obtain multiple or optimal aerial views of the objects, facilitating accurate and rapid identification of the objects present in the captured images or scene.
[0028] In one embodiment, the system is capable of identifying objects, including a variety of telecommunication equipment installed on a communication tower. This helps in the remote identification, cataloguing and monitoring of the telecommunication equipment installed on remote communication towers, offering advantages such as reduced manpower and time.
[0029] In an alternate embodiment, the system is capable of identifying various types of electrical equipment installed on or having an electrical connection with an electrical transmission or distribution tower. This enables the remote identification, cataloguing, and monitoring of the electrical equipment, including the power lines supported by the towers. This is advantageous as it helps in the remote monitoring of the electrical equipment, reducing manpower and time.
[0030] In an alternate embodiment, the system is capable of identifying vegetation in proximity to one or more objects, including electrical equipment, transmission or distribution towers, and / or power lines supported by the towers. This capability is advantageous as it facilitates remote monitoring and detection of vegetation near power lines or equipment, thereby preventing power outages caused by overgrown vegetation with minimal manpower and time.
[0031] In one embodiment, the processor is configured to select one or more images from the captured images or input scene which provides optimal views of the identified objects. This is advantageous as it helps the users to visually identify the objects including any defects in the identified objects by presenting them from multiple optimal angles.
[0032] In another embodiment, the processor is configured to obtain a plurality of information related to the one or more objects from a predefined equipment library stored in the storage device. This is advantageous as it allows for the collection of additional details or information related to the identified the objects that may not be available from the captured images alone. Additionally, comparing the information related to the identified objects from the captured images or scene with the data from the predefined equipment library helps in detecting missing or unauthorized equipment installations on the towers. In another embodiment, the system is configured for easier identification of defects on the objects automatically or facilitates the detection of defects by allowing visual inspection of the objects from multiple optimal viewpoints in the selected images or scene.
[0033] Furthermore, the system can generate at least one alert upon identification of the defects and / or presence of vegetation in proximity to the electrical equipment, transmission or distribution towers and / or power lines and upon identifying the missing or unauthorized objects in the scene and defects in the identified objects. This is advantageous as this accelerates the detection of missing equipment or installation of unauthorized equipment on the towers as well as the defect detection process, facilitating prompt repair or replacement of defective object or equipment or removal of the overgrown vegetation to prevent power outages or the unauthorized equipment.
[0034] BRIEF DESCRIPTION OF DRAWINGS
[0035] The invention will be more understood by reference to the description below taken in conjunction with the accompanying drawings herein:
[0036] FIG. 1 illustrates a flow chart showing the steps of a method for identifying one or more objects in an input scene or captured images, according to an embodiment of the present invention;
[0037] FIG. 2 illustrates a block diagram of a system for identifying and cataloguing the one or more objects in the 3D model, scene and / or 2D images, according to an embodiment of the present invention;
[0038] FIG. 3 depicts a flow diagram illustrating the steps for identifying and cataloguing the one or more objects in the images or input scene received by the system of FIG. 2, in accordance with an embodiment of the present invention;
[0039] FIG. 4 is a flow diagram illustrating the steps for identifying and cataloguing the one or more objects in a digital twin generated from the images or input scene, in accordance with an embodiment of the present invention;
[0040] FIG. 5 presents photographs illustrating the received images, the generated digital twin, and an output 3D model displaying the identified objects produced by the method shown in FIG. 1, in accordance with an exemplary embodiment of the present invention; and FIG. 6 depicts a flow diagram illustrating the steps for identifying and cataloguing the one or more objects, in accordance with the alternate embodiment.
[0041] DETAILED DESCRIPTION
[0042] In line with the above summary, the following description of a number of specific and alternative embodiments is provided to understand the inventive features of the present invention. It shall be apparent to one skilled in the art, however that this invention may be practiced without such specific details. Some of the details may not be described at length so as not to obscure the invention. For ease of reference, same reference numerals will be used and applied throughout the description and figures when referring to the same or similar features in the disclosure of preferred embodiments of the present invention.
[0043] Embodiments of the invention are described by way of illustration. As will be realized, the invention is capable of other and different embodiments and its several details are capable of modifications in various respects, all without departing from the scope of the present invention. It should be noted that the standard equipment or components may have not been illustrated since they are known in the art.
[0044] Furthermore, embodiments of the invention utilize standard image -capturing equipment for capturing outdoor images, especially images of equipment installed on outdoor communication or electrical transmission or distribution towers. However, the invention can also be utilized for the identification of a variety of equipment installed outdoors or even indoor.
[0045] In this description, the term "scene" refers to a visual area captured in an image or video by an image-capturing device, such as a camera or sensors on an unmanned aerial vehicle (UAV) or drone. A "scene" may include images or videos obtained through aerial photogrammetry from an elevated position and at various angles, as captured by a UAV or drone. Additionally, scenes captured by a UAV may contain metadata, such as three-dimensional (3D) coordinates and, in some cases, the camera's angle of view relative to a specific plane or reference. This metadata can include spatial orientation and location data, such as latitude, longitude, and altitude, which are useful for processing the captured scene or image. In certain instances, 3D coordinates refer to a spatial location defined by latitude, longitude, and altitude or elevation, as well as "pose" in the context of 3D coordinates. The term "pose" encompasses both the position and orientation of an object in space. In certain instances, "pose" includes the position, determined by latitude, longitude, and altitude, and the orientation, defined by angles such as roll, pitch, and yaw, which describe the object's rotation or orientation, such as that of a UAV or camera in 3D space. Thus, while 3D coordinates focus on position, incorporating pose provides a more comprehensive understanding by also considering the object's orientation.
[0046] In some instances, the scene captured by the UAV, which is sent to a computer system comprising a processor and memory for storage or further processing, is referred to as an "input scene." The "output scene," on the other hand, refers to the scene generated by the computer system after identifying the objects in the input scene through various steps, as described in the method below. The input scene can be a 2D image or a collection of 2D images, or data received from sensors such as LIDAR, which provides depth images or other forms of 3D data. Additionally, the input may include data combining both 2D and 3D modalities. Similarly, the output scene may be a 2D image or a group of 2D images that provide optimal views of the objects in the scene. Alternatively, the output scene could be a 3D model offering optimal views of the identified objects.
[0047] The present invention introduces a method for the identification and cataloguing of objects within a 3D model, scene and / or images. As illustrated in FIG. 1, the method utilizes computer vision to process an input scene captured by a UAV and converts it into 3D reconstructed data through a series of steps for the automatic identification and cataloguing of objects in the input scene. In one instance, the input scene captured by the UAV is a 2D image or multiple 2D images, and this image or images is processed through a series of steps for the automatic identification and cataloguing of objects within the captured 2D image. Initially, the input scene or one or more images of an area containing the objects is received and stored. Subsequently, a three- dimensional (3D) model, often referred to as a digital twin, is generated from the input scene or the captured images, encompassing the one or more objects therein. This 3D model is then utilized to uniformly sample or extract or generate a multitude of points from the surface of the 3D model, thereby creating a first point cloud representation where each point signifies a geometric coordinate of an object within the 3D model or the input scene. In one instance, the plurality of points is uniformly generated on the surface of the 3D model based on the structural context and domain information of each point.
[0048] In one embodiment, the digital twin is in the form of a mesh or mesh format and one or more artificial intelligence based 3D segmentation methods are utilized to generate the type and instance predictions directly on this mesh surface or the digital twin. Initially, the method accepts the digital twin in the mesh format as input and generates the required number of points directly from its surface to create a point cloud data. In one or more instances, the mesh is converted into point cloud data through pre-processing steps such as downsampling. In one instance, a domain- aware 3D downsampling method, such as voxel downsampling, is used to eliminate certain redundant information or points while keeping all the important contextual details. Advantageously, this technique reduces the need for high computational resources, and the proposed solution can extend to very large 3D scenes.
[0049] Subsequently, in a post processing step, the predictions from the point cloud are mapped onto the mesh surface leading to direct mesh predictions. This approach is advantageous as it allows predictions to be embedded directly within the mesh, significantly enhancing visualizations where different objects can be highlighted, such as by applying distinct colour coding directly on the mesh, improving clarity and interpretability. Additionally, this approach requires less memory, enables processing of arbitrary amounts of dense 3D point cloud data and maximizes accuracy in identifying and cataloguing the objects in the captured image with less computing power.
[0050] In another embodiment, the digital twin is in the form of a point cloud or point cloud representation, and artificial intelligence-based 3D segmentation methods are used to generate type and instance predictions directly on this point cloud representation or digital twin. In one or more instances, redundant information or points are eliminated through pre-processing steps such as downsampling. In one instance, a domain-aware 3D downsampling method, such as voxel downsampling, is employed to eliminate selected points while retaining all important contextual details. This technique advantageously reduces the need for high computational resources to process the high-density point cloud of the digital twin, allowing the proposed solution to extend to very large 3D scenes.
[0051] The image processing methods mentioned above efficiently and accurately identifies multiple objects in the captured images, the 3D model and / or scene and assists to determine the type of the objects, and cataloguing of the objects, requiring less manual effort.
[0052] In a next step, the first point cloud representation thus generated is converted into multiple overlapping point cloud patches to facilitate further processing. The step of dividing the entire point cloud or the first point cloud representation into multiple overlapping 3D patches mitigates the limitation of working with less dense versions of the original high-density point clouds due to memory constraints. The next step involves identifying the one or more objects including a type of the objects present in these overlapping patches utilizing a supervised or trained 3D semantic segmentation method.
[0053] In certain embodiments, a deterministic algorithm or method is utilized to form instances from the semantic segmentation predictions. In some cases, a trained 3D point cloud semantic segmentation model is used for semantic segmentation and to facilitate identification of the one or more objects present in these overlapping patches. In one instance, this step involves training the semantic segmentation method or model using a large set of curated data. In some instances, the trained semantic segmentation method or model does not necessarily get updated over time. This model can be used for identifying one or more objects including the types of the identified objects in the overlapping point cloud patches. In other embodiments, an instance identification method is also utilized for identifying the various instances of the type-identified objects present in the overlapping patches. In various embodiments as discussed below, the 3D semantic segmentation and instance identification methods are used in a number of stages for classification and identification of the objects present in the captured images including each instances of the objects present in the 3D model and / or the input scene.
[0054] Typically, the semantic segmentation and instance identification methods utilize artificial intelligence based models for refinement and improved accuracy during the identification of the objects. At first, training dataset, i.e. 2D images of an area containing the objects are collected using a drone. The drone camera sensors capture a diverse array of assets and objects from multiple viewpoints. The drone is further operated to capture images from different angles and distances to ensure a comprehensive 3D scene of the area is represented, covering all potential occlusion scenarios. The dataset thus received encompasses a wide variety of assets, objects, backgrounds, and geographic locations to adequately represent the entire distribution on which the trained model is intended to be applied.
[0055] Once the images are captured, 3D models or digital twins of assets and objects are meticulously reconstructed in the pre-processing step, and precise annotations are manually conducted on these reconstructed 3D digital twins. Subsequently, the semantic segmentation model is trained to effectively segment objects by minimizing the loss between model predictions and the groundtruth annotations.
[0056] In an alternate embodiment, multiple models are trained and the model exhibiting optimal performance on the validation set is selected based on various accuracy metrics. This selected model is then subjected to evaluation on an unseen test split to further validate its effectiveness and performance.
[0057] Once the objects and its types are identified, these overlapping patches are combined to form a second point cloud representation. In the subsequent phase, the method involves determining the instance of each object and cataloguing the type-identified objects, including each instance thereof, within the second point cloud representation using an instance identification method, often based on domain knowledge to create instances. The type identified objects including each instance thereof are then visually presented in the 3D model and / or scene.
[0058] In an alternate embodiment, instead of presenting the type identified objects, including each instance thereof, on the point cloud model, a plurality of points of the segmented objects in the point cloud model are mapped onto a surface of the respective objects in the 3D model or an output scene or the corresponding geometry within the model, creating a 3D model or the output scene. Additionally, in certain instances, the method includes projecting this 3D model or output scene into a 2D model or one or more 2D images captured by the UAV to facilitate visual representation of the identified objects.
[0059] Additionally, the method includes various steps to enhance accuracy throughout the segmentation and identification processes. These include eliminating redundant points through domain-aware downsampling methods and achieving consensus of model predictions in a local neighbourhood. These techniques contribute to refining the segmentation and object identification methods, ultimately improving the overall accuracy of the proposed method.
[0060] In one embodiment, the accuracy of the segmentation and object identification steps is enhanced by employing a neighbouring points consensus method or step. This step leverages consensus over predictions from overlapping point cloud patches and consensus over predictions in a local neighbourhood of the objects present in the overlapping point cloud patches to improve point cloud semantic segmentation predictions. This additional processing step updates a predicted confidence score of each of the object predictions and, consequently, the predicted class labels of each point in the point cloud representation by aggregating the confidence scores of points in a local neighbourhood. By aggregating predictions based on neighbouring points, incorrect predictions are mitigated, and correct predictions are strengthened by utilizing correct high- confidence predictions from neighbouring points. In certain instances, the neighbouring point consensus step is repeated at least 3-5 times to thoroughly eliminate incorrect predictions and improve correct predictions.
[0061] In another embodiment, the image processing steps disclosed in FIG. 1 are performed using a system 100 shown in FIG. 2. The system 100 comprises an application running on a cloud server 110 having at least one processor 102 and a storage device 104 in communication with the processor 102. In one instance, the application configured to run on the cloud server 110 is a web-based application. In one configuration, the storage device 104 holds instructions, enabling the processor 102 to execute tasks related to object identification within images and cataloguing of the identified objects. Typically, the processor 102 receives and stores images containing objects, often captured by drones like UAV 106.
[0062] In an alternative configuration, the drone 106 captures visual image data and other information using a range of sensors, including LIDAR, IR, or thermal cameras attached to the drone. In one instance, the captured data is retrieved from the storage of the drone 106 and is further processed in the cloud server through a series of steps disclosed in FIG. 1 to identify and catalogue the objects in the image. In some other instances, the data is transmitted over a network for processing or for monitoring or for storage, which can be retrieved later by the processor 102 from the storage device 104 for further processing.
[0063] In an embodiment, the drone 106 is operated to capture multiple images of the area or scene, providing diverse views of the objects from multiple angles. The system 100 processes these images or scene to achieve precise object identification of the objects and cataloguing of the identified objects in the area or scene.
[0064] Typically, the system 100 features a display device 108, presenting 2D visual representations of both the images and the identified objects. In certain instances, the processor 102 is set to display 3D representations of the objects via the display device 108. Alternatively, users can view the captured images or scene and the identified objects through the display device 108. FIG. 3 illustrates a flow diagram detailing the steps for identifying and cataloguing the objects in the captured input scene or the 2D image or images received by the system 100, as depicted in FIG. 2, in accordance with the present embodiment. Initially, the processor 102 undertakes processing of the captured input scene or 2D image or images to generate a three-dimensional (3D) model encompassing the objects within the scene or images. In a specific embodiment, the 3D model generated by the processor 102 serves as the digital twin of the objects featured in the captured input scene or images. This digital twin offers the advantage of accurately representing the objects in a virtual environment, thereby facilitating subsequent processing for the identification and cataloguing of the objects.
[0065] In one embodiment, the digital twin generated from processing the input scene or 2D input images takes the form of a mesh or mesh format or mesh representation. The proposed method is configured to predict on the mesh based 3D representations in an end-to-end manner via generating an intermediate point cloud representation. Initially, the method accepts the mesh representation as input and the mesh representation is converted into point cloud data. Subsequently, the predictions from the point cloud are mapped onto the mesh surface and smoothed to mimic direct mesh predictions. The method is not limited to only 3D meshes but is equally capable of handling 3D point clouds in which case some of the steps pertaining to conversion to and from the mesh may not be needed, according to an alternate embodiment. Thus, the proposed method can operate with meshes directly by integrating pre-processing steps like downsampling and post-processing steps such as mapping point cloud predictions onto the original mesh surface or digital twin.
[0066] In one embodiment, the processor 102 processes the data obtained from the digital twin to perform an artificial intelligence (Al) based 3D semantic segmentation of the data to identify or predict the objects present in the digital twin. FIG. 4 illustrates the steps of the Al based 3D semantic segmentation method. The processor 102 is configured to massage meshes into point cloud data in the input 3D model by uniformly sampling the plurality of points on the surface of the input 3D model. The sampling process creates the first point cloud representation, wherein each sampled point represents a geometric coordinate of an object in the input 3D model. In an embodiment, the processor 102 is configured to uniformly sample the plurality of points on the surface of the input 3D model by leveraging the structural context and domain information of each point. This approach transforms the process from mesh segmentation to a more accurate and detailed point cloud segmentation, which facilitates the usage of highly accurate 3D semantic segmentation models or techniques to maximize the accuracy of identification and cataloguing the objects in the digital twin.
[0067] In an alternate embodiment, the processor 102 is additionally configured to employ a voxel downsampling method to eliminate a subset of points from the first point cloud representation. In one instance, the voxel downsampling is performed by determining the voxel size, taking into account variations in the physical dimensions of objects. This ensures that the geometry and significant features of the objects remain unchanged post-downsampling. The process selectively eliminates redundant points that do not provide unique information about object features. This technique offers a distinct advantage by effectively reducing the number of points or the density of the point cloud dataset while preserving crucial details and minimizes the memory needed for point cloud processing thus mitigating the need for high-end hardware for processing. Importantly, this downsampling process facilitates streamlined processing without compromising the accuracy of the object identification procedure.
[0068] In one instance, voxel downsampling is performed by determining the voxel size, considering variations in the physical dimensions of objects. This ensures that the geometry and significant features of the objects remain unchanged after downsampling. The process selectively eliminates redundant points that do not provide unique information about the object's features. This technique effectively reduces the number of points or the density of the point cloud dataset while preserving crucial details, thereby minimizing the memory required for point cloud processing and reducing the need for high-end hardware. Importantly, this downsampling process facilitates streamlined processing without compromising the accuracy of the object identification procedure.
[0069] Further, the present system 100 implements a patch-based inference scheme, accomplished by dividing the entire point cloud dataset into multiple overlapping point cloud patches. Specifically, the processor 102 is configured to generate these overlapping patches by leveraging structural context and domain information, thereby alleviating the need to work with less dense versions of the original high-density point clouds due to memory constraints. This strategy effectively enhances accuracy in identifying and cataloguing objects within the captured input scene or images. Additionally, the partitioning of the point cloud dataset enables the system 100 to efficiently process large-scale 3D scenes or images encompassing a vast number of objects within the input scene or area.
[0070] Furthermore, during the formation of point cloud patches, the overlap between patches is chosen to ensure complete coverage of the objects within at least one of the overlapping regions. This is advantageous as it guarantees the availability of high-quality object predictions, thereby enhancing the accuracy of the subsequent inference process. In a patch-based inference scheme, the model conducts inference on individual patches, assigning semantic segmentation predictions to each point within the patch. Subsequently, these predictions are aggregated by merging the overlapping point cloud patches into a single point cloud, resulting in duplicate points and predictions in the overlapping areas. To refine the dataset, deduplication is performed, retaining only the most confident predictions while eliminating redundant points and their associated predictions. This process of selecting unique points or predictions based on high confidence scores ensures the retention of high-quality predictions without the presence of duplicate data, thereby improving the overall accuracy of the model.
[0071] In one embodiment, the overlapping point cloud patches undergo processing using a trained or supervised 3D semantic segmentation method to identify objects within them. Subsequently, the processor 102 combines the point cloud data associated with the identified objects from the overlapping patches to create a second point cloud representation.
[0072] In one embodiment, the processor 102 employs a refinement method to identify objects by comparing the positions of a plurality of points on respective objects common to overlapping point cloud patches and in the local neighbourhoods of objects within the overlapping patches. This two stage predictions refinement process enhances semantic segmentation predictions by aggregating confidence scores from neighbouring points within overlapping patches and the local neighbourhoods of the objects. The neighbouring points consensus step leverages consensus over predictions from overlapping point cloud patches and consensus over predictions in a local neighbourhood of the objects present in the overlapping point cloud patches to improve point cloud semantic segmentation predictions. This step aggregates the confidence scores of points in a local neighbourhood and in the overlapping point cloud patches. By aggregating predictions based on neighbouring points and overlapping point cloud patches, incorrect predictions are mitigated. In certain instances, the above consensus steps are repeated at least 3-5 times to thoroughly eliminate incorrect predictions and improve correct predictions. In another embodiment, the processor 102 is configured to perform semantic segmentation predictions on overlapping point cloud patches generated from the first point cloud representation to identify objects, including determining the type of each object. This step is followed by an instance identification method applied to the second point cloud representation created from the point cloud patches. This method identifies individual instances of each type-identified object and allows for the cataloguing of the identified objects within the second point cloud representation. In one scenario, the catalogued objects are stored in a database to monitor the health of each object.
[0073] In an exemplary embodiment, the processor 102 runs the 3D semantic segmentation model to generate predictions for each point in the first point cloud representation. This prediction also provides information such as whether any particular point belongs to an object of interest, e.g., an antenna. However, this prediction does not give details such as whether any two points predicted as belonging to the antenna actually belong to the same antenna or different antennas.
[0074] Thus, in the next step, the processor 102 runs the post-processing steps, e.g., point cloud prediction aggregation and local neighbourhood consensus, to refine these predictions. Then, in the next stage, the processor 102 runs the instance identification method to provide information about whether the points predicted as an antenna actually belong to the same antenna or different antennas, thus forming the instances. This is advantageous as it enables the identification of points on the same or different objects or equipment e.g., antennas, which further improves the accuracy of predictions.
[0075] In one embodiment, the step of cataloguing the identified objects in a 3D model or scene includes storing data related to each identified object in a data repository. This stored data can be linked to items in a predefined equipment library or repository for further analysis. For example, in applications such as monitoring telecommunication towers, the processor 102 is configured to connect the objects stored in the data repository to corresponding objects in the predefined equipment library to retrieve detailed information such as manufacturer name, model, make, operating frequency, voltage rating, etc. This helps in tasks such as structural loading assessments by providing complete and accurate information about the detected equipment. Furthermore, the catalogued objects stored in the data repository can be utilized for auditing equipment installed on telecommunication towers, electrical transmission towers, or distribution towers. By comparing the catalogued objects, whose details are stored in the data repository, with the information related to the installed equipment stored in the equipment library, it is possible to remotely identify any unauthorized installations on the towers. This process also facilitates the remote detection of stolen or damaged equipment, enhancing the ability to maintain and secure the assets located on these towers.
[0076] In one or more instances, the identified objects are visually depicted on the display device by mapping multiple points of the identified objects onto a surface of the corresponding objects in the 3D model or output scene or the corresponding geometry within the 3D output scene. In a particular embodiment, the processor 102 is programmed to visually represent the recognized objects in the 3D model and / or scene by projecting the point cloud points of the identified objects onto a surface of the corresponding objects in the 3D model or output scene or the corresponding geometry within the 3D scene. This is advantageous as this allows users to directly visualize the object predictions on the 3D model or output scene, enhancing the interpretation of the identified objects.
[0077] In an alternate embodiment, the processor 102 is configured to convert the 3D model or output scene into a 2D model within the captured 2D images through 3D to 2D projection. This is advantageous as this allows users to easily identify the objects within the captured 2D images or input scene, enhancing the accessibility and interpretability of the object identification results. In some instances, the processor 102 is configured to facilitate selection of one or more images which provide optimal views of the one or more identified objects. In certain instances, the processor 102 is configured to automatically select one or more images which provide optimal views of the one or more identified objects. Here, "automatically" refers to the ability of the processor 102 to perform this selection without manual intervention, relying on predefined algorithms, machine learning models, or programmed criteria. In certain instances, the processor 102 is configured to analyse factors such as the clarity, angle, lighting, and relevance of each captured image to determine which views best showcase the objects of interest. Typically, the images selected by the processor 102 offer distinct views of the identified object or objects of interest. These images are chosen to capture different perspectives, angles, or features, ensuring that critical details of the objects are clearly visible. By providing varied viewpoints, the processor 102 enhances the user's ability to analyse and interpret the objects, improving the accuracy and comprehensiveness of the visual representation. Additionally, the automatic selection process enhances efficiency by reducing the need for user input and ensuring that the most informative images are chosen, facilitating a more accurate and consistent representation of the objects. Further, the identified objects are visually represented using colour coding or visual markings, such as overlays or shapes within one or more of the captured images of input scene which offers optimal views of the identified objects. The approach maximizes accuracy in identifying and cataloguing objects while requiring less computational power. Additionally, providing clear and interpretable visual representations helps to quickly and efficiently highlight the identified objects, reducing the time and resources needed for manual analysis.
[0078] In one embodiment, the processor 102 is designed to identify, segregate, and annotate objects in close proximity. This functionality is beneficial as this allows for the detection of small objects, such as transceiver junction devices and other in-close proximity objects installed on the communication tower, treating them as separate instances or objects. This capability enhances the accuracy of the system 100 in identifying and cataloguing objects, particularly in scenarios where precise differentiation of closely positioned items is required.
[0079] FIG. 5 presents photographs illustrating the received image or input scene, the generated digital twin, and the 3D model or output scene displaying the identified objects produced using the method of FIG. 1, in accordance with an exemplary embodiment of the present invention. According to the exemplary embodiment, the method and system 100 is trained on point cloud datasets of telecommunication towers for the identification of telecommunication equipment installed therein. This capability enables efficient identification and cataloguing of specific equipment types, contributing to streamlined maintenance, monitoring, and management of communication infrastructure.
[0080] In another embodiment, the method and system 100 is capable of identifying various objects, encompassing a range of unauthorized equipment installed on telecommunication towers. This functionality facilitates remote identification, cataloguing, and monitoring of such equipment situated on distant towers, thereby reducing the need for extensive manpower and time resources.
[0081] Additionally, in another embodiment, the processor 102 is configured to assist in detection of numerous external defects associated with the identified objects. In one instance, the processor 102 automatically identifies structural anomalies such as unexpected tilting in the antennas mounted on the towers and alerts users for further examination of the 3D or 2D output images or scene. Alternatively, in a different scenario, the processor 102 enables downstream processing to assist in predicting defects, such as structural deformations in the antennas. This capability expedites the defect detection process, enabling timely repair or replacement of faulty equipment, thus contributing to enhanced operational efficiency and reliability.
[0082] In yet another embodiment, the method and system 100 can be used to assist in detecting numerous external defects such as rust, exterior physical damages, disorientation, etc., associated with the identified objects. In such scenarios, the method and system 100 are configured to automatically generate alerts in the form of audio, visual signals, or notifications such as email, SMS or push notifications on selected devices. This capability further expedites the identification, repair, or replacement of faulty equipment.
[0083] In another exemplary embodiment, the method and system 100 is trained on point cloud datasets of electrical transmission or distribution networks, including transmission or distribution towers, associated electrical equipment, and power lines or cables supported by the towers. In this embodiment, the method leverages point cloud data received from sensors such as LIDAR to detect objects including electrical equipment, transmission or distribution towers, power lines or cables and vegetation near them. In some instances, this method helps to identify and catalogue the types of vegetation surrounding the tower, which helps to prevent electrical faults in the power lines.
[0084] Further, the method and system 100 can be used to detect various components on transmission or distribution towers including detection of types of equipment such as insulators, cables, and transformers on the towers and associated defects. The high quality point-cloud or mesh data obtained using the present method facilitates capturing small equipment with high fidelity, ensuring accurate Al detection.
[0085] In yet another exemplary embodiment, the method and system 100 is trained to detect various objects placed indoors, such as in a warehouse. In this embodiment, the method can be utilized to detect various objects placed in the warehouse, providing the warehouse owner with a comprehensive list of objects present in the warehouse remotely, which can be used to update or compare with the warehouse's inventory management system. In another exemplary embodiment, the method and system 100 is trained to identify objects in a public area, such as a street. The model can be used to identify objects including traffic signs around the scene of an accident, etc. For example, the method can be used to assist in forensic analysis by detecting objects of interest in a crime or accident scene. For instance, in the case of a traffic accident, authorities might be interested in analysing traffic signs around the scene. In such a scenario, the present method with its highly detailed mesh or point cloud representation can be used to detect and catalogue these objects, fully capturing all relevant details, thereby aiding authorities in their investigations.
[0086] In another embodiment of the invention, the method is configured to work directly with point cloud data as shown in FIG. 6, This eliminate the step of converting the mesh to a first point cloud representation if the 3D model is in the mesh format. When the 3D model is already in the point cloud format, the step of converting the mesh to a first point cloud representation if the 3D model is in the mesh format is not needed. The first point cloud representation can be directly obtained by elimination of certain points in the point cloud format through downsampling techniques during pre-processing. Further, the mapping of point cloud data onto the original mesh surface during post-processing step is also not needed as the 3D model is already presented in the point cloud format.
[0087] The alternate method shown in FIG. 6 starts with receiving and storing the images or input scene containing the objects of interest. The images or the input scene is further processed to create a point cloud data. Certain points in the point cloud are eliminated in the next step using a downsampling method to create a first point cloud representation forming a corresponding 3D model, or digital twin to accurately represent the objects in a three-dimensional space. This point cloud representation is then converted into overlapping patches, dividing the point cloud into smaller sections to facilitate object identification and analysis using a supervised 3D semantic segmentation model. The identified objects are combined to form a second point cloud representation in a number of steps, as previously disclosed. Subsequently, an instance identification method is applied to determine the instances of each object present in the second point cloud representation. Finally, the plurality of points of the type-identified objects in the 3D model in the second point cloud representation is projected into an output scene or the captured images for easier analysis. This method mitigates the limitation of working with less dense versions of the original high-density point clouds due to memory constraints. This approach is more practical as it requires less memory, enables processing of dense 3D point cloud data, maximizes accuracy in identifying and cataloguing the objects in the captured image.
[0088] In summary, the present invention presents a method for efficiently identifying and cataloguing objects within images, the 3D model and / or scene. By employing a combination of steps, including the generation of a 3D model or the digital twin, point cloud sampling, supervised semantic segmentation and instance identification methods, and mesh surface mapping, the method enables accurate identification and categorization of objects. The resulting 3D predictions provide a comprehensive representation of the identified objects, with optimal views for further analysis. It is understood that variations and modifications to the method described herein may be made without departing from the scope of the invention, as defined by the appended claims.
Claims
CLAIMS1. A method for identifying one or more objects in a scene comprising the steps of: receiving and storing an input scene containing the one or more objects; generating a three-dimensional (3D) model of the input scene, including the one or more objects therein; creating a first point cloud representation from the 3D model; characterized by: converting the first point cloud representation into a plurality of overlapping point cloud patches; identifying a type of the one or more obj ects in the overlapping point cloud patches using a trained 3D semantic segmentation method; combining the type-identified objects in the overlapping point cloud patches to create a second point cloud representation; identifying instances of each of the type-identified objects in the second point cloud representation using an instance identification method; and presenting the type-identified objects, including each instance thereof, within the 3D model and / or an output scene.
2. The method according to claim 1 further includes the step of cataloguing the type- identified objects, including each instance thereof, in the second point cloud representation.
3. The method according to claim 1 wherein the 3D model is a digital twin of one or more captured images or input scene.
4. The method according to claim 1 wherein the 3D model is represented in a mesh format or a point cloud format.
5. The method according to claim 1 wherein creating the first point cloud representation includes: extracting a plurality of points from a surface of the 3D model represented in the mesh format to create a point cloud, each point representing a geometric coordinate of each object; and eliminating at least a portion of the plurality of points of the point cloud using a downsampling method.
6. The method according to claim 1 wherein creating the first point cloud representation includes: eliminating at least a portion of the plurality of points of the 3D model represented in the point cloud format using a downsampling method.
7. The method according to claim 1 wherein identifying the one or more objects in the overlapping point cloud patches further includes the steps of: predicting on a plurality of points on one or more objects in each of the overlapping point cloud patches using the trained 3D semantic segmentation method; aggregating the predictions on the plurality of points in the overlapping point cloud patches; and refining the aggregated predictions using a local neighbourhood consensus method.
8. The method according to claim 1 wherein presenting the type-identified objects within the 3D model and / or output scene further includes the step of mapping a plurality of points on the type-identified objects onto at least one 3D mesh surface.
9. The method according to claim 1 further includes creating a visual representation from the 3D model by projecting the 3D model and / or the type -identified objects into the one or more captured images and / or input scene.
10. The method according to claim 2 wherein identifying the type of the object and cataloguing the type -identified objects includes identifying, segregating and annotating a plurality of objects in close proximity.
11. A system (100) for identifying one or more objects in a scene comprising: a processor (102); and at least one storage device (104) in communication with the processor (102), the storage device (104) being configured to store instructions, which when executed by the processor (102) enable the system (100) to perform the method as claimed in the claim 1.
12. The system (100) according to claim 11 further comprises an unmanned aerial vehicle (UAV) to capture the input scene.
13. The system (100) according to claim 11 wherein the one or more identified objects includes at least one telecommunication equipment installed on at least one communication tower.
14. The system (100) according to claim 11 wherein the one or more identified objects include at least one electrical equipment installed on or in connection with an electrical transmission or distribution tower.
15. The system (100) according to claim 14 wherein the processor (102) is configured to identify vegetation in proximity to the one or more objects including the electrical equipment, transmission or distribution towers and / or power lines supported by the towers.
16. The system (100) according to claim 11 wherein the processor (102) is configured to select one or more images from the captured images or input scene with optimal views of the identified objects.
17. The system (100) according to claim 11 wherein the processor (102) is configured to obtain a plurality of information related to the one or more objects from a predefined equipment library stored in the storage device (104).
18. The system (100) according to any one of claim 11 wherein the processor (102) is configured to facilitate identification of at least one missing or unauthorized object and defects to the identified objects in the scene.
19. The system (100) according to claim 15 wherein the processor (102) is configured to generate at least one alert upon identification of the presence of vegetation in proximity to electrical equipment, transmission or distribution towers and / or power lines or upon identifying missing or unauthorized objects and defects in the identified objects.
Citation Information
Patent Citations
Systems and methods for analyzing remote sensing imagery
US20170076438A1
Point cloud compression using non-orthogonal projection
US20190139266A1
Video based point cloud compression-patch alignment and size determination in bounding box
US20200314435A1
Apparatus and method for generating lightweight three-dimensional model based on image
US20230186565A1
Instance segmentation systems and methods for SPAD lidar
US20240029271A1