Method for detection of public safety elements in real world enviornment using a 2d image
The computer vision-based method addresses the limitation of existing systems by incorporating 3D spatial awareness and an adjacency matrix to enhance the detection of public safety elements, improving accuracy and efficiency in real-time hazard assessment.
Patent Information
- Application Number
- PCT/IN2025/051093
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-18
- Filing Date
- 2025-07-18
- Publication Date
- 2026-01-22
AI Technical Summary
Existing image analysis systems lack the capability to understand and interpret the full spatial context of objects and interactions within an environment, limiting their effectiveness in real-world hazard detection and assessment.
A computer vision-based method using deep learning and neural networks for detecting public safety elements, incorporating 3D spatial awareness and an adjacency matrix to capture object interactions, which includes image quality evaluation, scene segmentation, depth estimation, and object property extraction.
Enhances the accuracy and efficiency of detecting public safety risks in real-time, reducing manual effort and providing timely alerts for potential hazards, thereby creating a safer environment.
Smart Images

Figure IN2025051093_22012026_PF_FP_ABST
Abstract
Description
METHOD FOR DETECTION OF PUBLIC SAFETY ELEMENTS IN REAL WORLD ENVIORNMENT USING A 2D IMAGE FIELD OF INVENTION
[0001] The present invention pertains to the field of image processing, with a particular focus on computer vision utilizing deep learning, neural networks, and mathematical analysis. Specifically, this invention relates to a method for identifying and detecting elements within real-world environments. More precisely, the invention addresses a method for recognizing various elements in such environments and detecting those that may pose potential public safety risks. BACKGROUND OF THE INVENTION
[0002] The increasing prevalence of digital imagery and advancements in computational power have significantly enhanced the capabilities of image processing and computer vision technologies. These developments have paved the way for sophisticated methods to analyse and interpret visual data from the real world.
[0003] Deep learning and neural networks have emerged as powerful tools for image analysis, offering the ability to learn and identify patterns with high accuracy. These technologies have found applications in various fields, including autonomous driving, medical imaging, and surveillance systems. However, there remains a substantial need for improved methods to detect and identify elements that pose public safety risks in real-world environments.
[0004] Public safety is a critical concern in urban, semi-urban, residential and commercial areas where hazards can arise from various sources such as structural damages, environmental hazards such as debris, uncleared waste, obstacles in public spaces etc.
[0005] This also refers to the disruptive occurrence that limits the movability of the people on the public roads such as excavation barriers, potholes, and dilapidated sidewalks. This not only impacts the aesthetic appeal of the environment but also poses hazards including health hazards to human beings in the affected areas.
[0006] The non patent literature Visual Pollution Prediction Framework Based on a Deep Active Learning Approach Using Public Road Images by Mohammad AlElaiwi et. al. presents a deep learning-based framework to detect and classify visual pollution (VP) in public road images, focusing on three main types: excavation barriers, potholes, anddilapidated sidewalks. Using a large dataset of 34,460 images from Saudi Arabia and a deep active learning approach for automatic annotation, the system trains and evaluates several AI models to accurately identify and localize these pollutants.
[0007] Chinese patent application CN110022466A- Video analysis platform based on smart big data and control method thereof, describes a video analysis platform that uses smart big data to analyze and detect violations in video streams, such as intrusion or safety hazards, by applying intelligent algorithms for object detection and classification, with real- time alerts and management through an integrated system.
[0008] Another patent application CN109166293A- Remote auxiliary early warning method based on substation stereoscopic detection- describes a remote early warning system for substations that uses stereoscopic detection and optical character recognition (OCR) to identify and label power equipment, enabling effective monitoring and prediction of potential violations or operational issues by analyzing image features and motion detection based on the optical flow field.
[0009] These patent applications, primarily focus on the detection of safety violations and the analysis of video data streams in a controlled environment like substations. These systems leverage 2D image analysis and object detection techniques, primarily aimed at identifying specific violations such as intrusion, personnel safety hazards, or improper equipment usage. However, these solutions are generally limited in their ability to understand and interpret the full spatial context of objects and interactions within an environment. They lack the capability to assess the 3D spatial relationships between detected objects, which is crucial for more accurate and context-aware hazard detection.
[0010] The existing prior arts are limited to detecting the types of visual pollution and lacks scene understanding, depth estimation, and inter-object relationship modeling, reducing its effectiveness in real-world hazard assessment. These gaps may be overcome by expanding the detection scope, incorporating 3D spatial awareness, and using an adjacency matrix to capture object interactions for a more comprehensive and actionable analysis.
[0011] The ability to automatically identify and assess these risks through visual data analysis can greatly enhance safety measures, providing timely alerts and preventing potential incidents.
[0012] Governments and municipalities recognize the importance of minimizing public safety risks to create a safe and appealing environment for citizens. Many have launched campaigns involving manual inspections across cities to identify potential safety violation and notify violators for corrective actions. However, this method is non-automatic, time-consuming, economically unfeasible, and imposes considerable mental and physical strain on employees.
[0013] The present invention aims to address this by leveraging advancements in deep learning, neural networks, and statistical analysis to develop a robust method for detecting and identifying elements that could pose public safety risks. OBJECTIVE OF THE INVENTION
[0014] The primary objective of the present invention is to provide a computer vision based method for identifying elements causing potential public safety risk.
[0015] Another objective of this invention is to develop a computer vision-based system for public safety monitoring. The system aims to: 1. Automate Detection: Utilize advanced image processing and deep learning techniques to automatically detect various elements causing public safety issues in real-time 2. Improve Accuracy: Enhance the precision of identifying potential risks by leveraging neural networks and mathematical analysis 3. Efficiency: Reduce the time and resources required for manual inspections, making the process economically viable and less burdensome for employees 4. Real-Time Monitoring: Provide continuous, real-time monitoring of urban environments to promptly identify and report safety violations 5. Enhanced Safety: Ensure a safer environment for citizens / residents by swiftly addressing and mitigating risks related to public infrastructure
[0016] Another objective of the invention is to provide a foundation towards a comprehensive solution for real-time monitoring and analysis of environments, ultimately contributing to enhanced public safety and risk management. Its applications include: a) Integrating with city surveillance systems to provide real-time alerts for public safety incidents b) Identifying public safety elements and taking corrective actions to create a safe and appealing environment for citizens. Public safety elements encompasses the visible deterioration and bad aesthetic quality of natural and human-made landscapes. c) Compliance monitoring: Ensuring adherence to safety protocols in construction sites, factories, and other high-risk areas d) Environment monitoring: Detecting environmental hazards such as spills, debris, or obstacles in public spaces. Monitoring weather conditions and their impact on safety (e.g., flooding, snow accumulation).SUMMARY OF THE INVENTION
[0017] The present invention relates to a computer vision-based method for detecting public safety elements in the real-world environment.
[0018] In accordance with the present invention, the method comprises the steps of: capturing a 2D image of a real-world environment using an image-capturing device, image quality check, understanding the captured image by segmenting the scene in the image, depth estimation of each object in the scene, identifying and extracting properties of objects in the scene, generating an adjacency matrix representing spatial and contextual relationship and detecting the presence of public safety risk elements in the segmented scene of the image. To summarize, the proposed framework provides the following non-limiting features: a. Understanding the scene in the 2D image by detecting the objects present b. Mapping each of the objects in relation to each other including the location in the 2D scene c. Detecting public safety elements and extracting their properties
[0019] The method disclosed in the present invention encompasses a wide range of categories contributing to public safety risk, including, but not limited to, roads, manholes, construction waste, non-construction waste, concrete barriers, construction fencing, signages, light poles, and any other elements that may poses environmental and health hazards.
[0020] The present invention is a computer vision-based system, for real time detection of public safety risk elements. The system comprises at least one image capturing device configured to collect 2D images of the environment and a computing unit.
[0021] The computing unit comprises an image quality evaluation module, a scene segmentation module, a depth estimation module, an object property extraction engine, and an adjacency matrix generator.
[0022] The image quality evaluation module configured to access the image resolution and detect blurriness of an image using a CNN classifier. Only those images that qualify a predefined threshold will be used for further processing. The scene segmentation module uses a deep learning model, trained via self-supervised contrastive learning and partial manual annotation, to segment objects within the image. The depth estimation module reconstructs 3D coordinates from 2D images, and understands the spatial relationship ofobject and its positions in the detected scene. The object property extraction engine identifies key visual and spatial attributes such as geometry, color, texture, and relative placement. The adjacency matrix generator models spatial relationships and potential interactions between detected objects. These modules work jointly to transform raw visual input into actionable safety insights.
[0023] It can be extended to assess public safety risks, take corrective actions and monitors the effectiveness of efforts to minimize public safety incidents over time.
[0024] These objectives and advantages of the present invention will become more evident from the following detailed description when taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The objective of the present invention will now be described in more detail with reference to the accompanying drawing, wherein: FIGS.1A, 1B & 1C show a few examples of various elements posing public safety risk towards humans in the real-world environment; FIG.2 shows the overall flow of the method proposed in the present invention; FIG.3 shows the flow chart representing the steps followed in the method proposed in the present invention; and FIG.4 shows the pictorial flow of the generic framework of the method proposed in the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0026] The present invention discloses a computer vision-based method for detecting public safety elements in the real-world environment.
[0027] According to the present invention, the method 400 comprises the steps of: capturing an 2D image of a real-world environment using an image-capturing device401, validating the image 402, understanding the captured image by segmenting the scene in the image 404, depth estimation of each object in the scene 403, identifying and extracting properties of objects in the scene 405, and generating an adjacency matrix representing spatial and contextual relationship 406.
[0028] Before starting the process of detecting public safety elements, the framework is trained using deep learning, specifically Convolutional Neural Networks (CNNs), for efficient scene classification. Input images are captured from various sources. An open- source image segmentation model is employed to train the CNNs for specific public safety elements. This involves training the CNNs with multiple layers of artificial neurons, focusing on image recognition and vision computing.
[0029] Furthermore, achieving accurate training for deep learning models in the context of public safety requires a substantial dataset encompassing various public safety elements in a real-world environment. However, the manual annotation process required to create such datasets can be time-consuming.
[0030] The innovative based approach proposed in the method of the present invention incorporates a self-supervised method to create a foundational model. This training approach involves taking the entire dataset and using contrastive learning to extract features and the correlations between them. The model learns a meaningful representation of data by leveraging the inherent patterns and structure within the data itself. Using this as a base model, 10% of the dataset is then annotated manually by humans to create an initial segmentation model using the foundational model as a backbone.
[0031] This image segmentation model is then employed for fully automatic and prompt-able data annotation. This AI-based data annotation significantly reduces the time otherwise required for manual labelling.
[0032] Consequently, annotators can now interact with the model using interactive points and boxes, prompting the image segmentation model to automatically segment and mask public safety objects with a single click. Over time, the image segmentation model is trained with a diverse dataset, encompassing different public safety categories and general outdoor scene recognition. As a result, the model becomes generalized and capable of performing auto-data annotation with minimal to no manual intervention.
[0033] After training the framework, the images captured from the end devices / image capturing devices are transmitted for inferencing the trained model. The captured image is processed in the scene understanding model using the trained model. The imageoutput from the scene understanding model is transmitted for classification. The classifier gives the AI model result of the respective public safety element.
[0034] After training the neural network pipeline, the images captured from the end devices / image capturing devices are transmitted for inferencing the trained model. The captured image is processed via a series of steps to provide a final outcome which is the detection of public safety elements. The detailed process which takes place in each step is described below: Step 1: Capturing a 2D scene - a scene is captured 401, representing a view of a real-world environment that includes multiple surfaces and objects, organized in a meaningful way. The scene is obtained through various means, including but not limited to, employing a drone to acquire an image of the environment, utilizing an inspector employed by a government or municipality to capture an image using his / her camera, integrating dashcams with vehicles like cars to capture images, or using any other possible image acquisition unit. The captured image is subsequently fed to the framework for further processing. Step 2: Image Quality Detection 402 - To ensure optimal performance of computer vision systems, it is crucial to assess the quality of input images based on resolution and blur. Images that do not meet predefined quality standards are not utilized for further processing. Images are evaluated to ensure they meet a minimum resolution threshold. The minimum threshold for image resolution is set as 800x800 pixels to ensure optimal performance. Low- resolution images, which lack sufficient detail, are flagged. CNN based image classifier is used to perform a binary classification indicating if the image is blurred or not. Images classified as blurred, indicate significant loss of detail, and are deemed unsuitable for further processing. Each image is evaluated against these checks. Images that do not meet the required quality standards are rejected and not considered for further processing. By implementing these quality checks, the framework ensures that only images with sufficient clarity and detail are used, thereby improving the accuracy and reliability of the computer vision process. Step 3: Scene Detection and Understanding 404– This step helps us to understand the image and derive the context of different objects in the image. The captured scene is processed using a robust CNN based instance segmentation model to detect all objects in the image. Generally, humans are capable of recognizing and classifying scenes in a tenth of a second or less, as they possess the ability to capture the gist of the scene, even though this usually means having missed many of its details. Similarly, the proposed method aims at understanding the context of the scene prior to further detection and assessment.a) The computer vision model examines each and every pixel of the image which is fed to it and maps the entire image into a scene with multiple objects b) Additional business logic is implemented that determines if the identified scene is relevant to the task being performed, i.e. detection of public safety elements. If the scene is not relevant, the image is not considered for further processing thereby optimizing computational resources. The primary objective of this step is to perceive, interpret, and make sense of the visual environment, allowing further steps to respond appropriately based on the analysis of the scene. The output of this step provides the following information: a. Objects present in the image b. Number of instances of each unique object c.2D location of the objects in relation to each other Step 4: Depth Estimation and Mapping 403 – Encoder – Decoder based depth estimation model is used to generate a depth map, which assigns depth values to each pixel in the image via x, y and z coordinates. Depth estimation provides distance of each pixel relative to the camera. A custom algorithm is used to combine the output from scene understanding with depth values to estimate the distance between the camera and various objects detected in the scene. The resulting depth mapping information helps in understanding the spatial relationships and dimensions of objects. Step 5: Property extraction for objects of interest 405- This step primarily identifies and further analyses "objects of interest" i.e. specific items, entities, or features within an image that are relevant or related to public safety. From all the objects identified within the scene, specific ones that have the potential to cause public safety risk are detected using CNN based objection detection models. Furthermore, these specific “objects of interest” are put through a pipeline of traditional image processing, image classifiers and additional rule engine for extracting relevant properties. This involves identifying and isolating specific attributes or characteristics of visual data that are relevant for public safety. These properties are essential for understanding the content and context of the visual data and are used for tasks such as object recognition, scene understanding, and image classification. Properties extracted include and not limited to Geometric Properties such as shape, size, orientation of objects within the image, Colour Properties such as dominant colours of the
[0035] The adjacency matrix A combined with the property matrix “P” fully represents the objects in a scene, their 3D positions, and their properties. This formalization allows for comprehensive analysis and interaction modeling within the scene.
[0036] Using this method, public safety elements can be detected from a 2D image captured via various sources. This detection system can form the core component of a robust solution for real-time monitoring and analysis of environments and provide real- time alerts for public safety incidents; not limited to potholes, excavations, debris, uncleared public waste etc.
[0037] In one embodiment, the present invention provides a system 200 for the real- time detection of public safety hazards using computer vision. The system comprises at least one image capturing device 201 configured to collect two-dimensional (2D) images of an environment. The image capturing device may be deployed in various operational settings and can be selected from a group including, but not limited to, drones, dashcams, handheld cameras, or mobile surveillance systems. These devices continuously acquire visual data from public or monitored environments and transmit the captured image streams to a computing unit 210 for further processing.
[0038] The computing unit 210 is equipped with a plurality of specialized modules for end-to-end analysis of visual data. First, an image quality evaluation module 212 processes incoming images to assess resolution adequacy and identify visual distortions such as blurriness. This is achieved using a convolutional neural network (CNN) classifier trained to differentiate between high-quality and suboptimal images, so that only images meeting predefined quality thresholds are further processed.
[0039] Once the image quality is verified, a scene segmentation module 214 utilizes a trained deep learning model to segment the image into semantically meaningful regions by segmenting objects of interest from the background and from one another. The segmentation model is initially trained using a self-supervised contrastive learning technique, which allows the model to learn effective feature representations fromunlabeled data. To refine the model’s accuracy, manual annotation is performed on a subset comprising only ten percent of the total dataset. This semi-supervised approach significantly reduces labeling overhead and improves segmentation performance.
[0040] In certain embodiments, the scene segmentation model 214 is further enhanced to perform automatic annotation of public safety-relevant elements through interactive prompting mechanisms. This includes the use of user-provided inputs such as spatial points and bounding boxes to guide the model in identifying and labeling specific objects. Thus the system adapts dynamically to new environments and evolving public safety contexts without extensive retraining.
[0041] The processed image is fed into a depth estimation module 216, which generates three-dimensional (3D) spatial coordinates from the two-dimensional (2D) images. This module employs techniques such as monocular depth prediction or stereo matching algorithms, depending on the camera setup, to reconstruct the 3D layout of the scene and accurately position objects in a spatial context.
[0042] The depth information, along with the segmented image, is then analyzed by an object property extraction engine 218. This engine identifies and extracts critical attributes necessary for understanding the content and context of the visual scene. These attributes include, but are not limited to, the geometric structure, color distribution, surface texture, and spatial properties such as orientation, scale, and relative position of each object of interest.
[0043] An adjacency matrix generator 220 constructs a spatial interaction model representing the relationships between detected objects. This matrix captures proximity, orientation, and potential interaction pathways among entities within the visual frame. The adjacency matrix, in conjunction with the extracted object properties and spatial data, is used for identifying elements or object configurations that may pose public safety hazards.
[0044] In various embodiments, the entire computing architecture may be deployed on a cloud-based platform, integrating with municipal infrastructure or smart city surveillance systems. Such a deployment facilitates centralized monitoring, large-scale data aggregation, and coordinated response management by public safety authorities.The cloud-based configuration also supports real-time alerts and dashboard visualizations, thereby enabling prompt detection and mitigation of public safety hazards in dynamic urban environments. EXAMPLE Model Architecture, Training, and Optimization:
[0045] At the core of the system is a deep learning model that leverages a pre-trained Masked Autoencoder (MAE) encoder as its foundational backbone. This encoder is initially trained in a self-supervised manner on a vast dataset of general images, typically comprising millions of diverse, unlabeled images. These images may come from a proprietary scene dataset containing various environments, objects, and lighting conditions. During this pre-training phase, the MAE learns meaningful visual representations by reconstructing masked portions of input images. The masking process forces the encoder to learn complex visual patterns, structural relationships, and contextual dependencies across the scene without the need for human-provided labels. This pre-training provides an exceptionally powerful and generalizable foundation
[0046] Attached to this pre-trained MAE encoder is a segmentation head based on the Mask R-CNN architecture. This segmentation head refines the encoded features to produce pixel-level segmentation masks, allowing for precise object and region identification within a given image. The combined architecture of the MAE backbone and the Mask R-CNN head is subsequently fine-tuned using a large-scale, custom- curated dataset consisting of over 500,000 images. Each image in this dataset is meticulously pixel-annotated to define every relevant object and scene category.
[0047]
[0048] During this fine-tuning, the model is optimized using a multi-component loss function. This comprehensive loss guides the learning process to ensure accuracy in several critical areas simultaneously. It includes components that ensure accurate object classification, precise localization of bounding boxes, and the generation of detailed instance maskshighly accurate, pixel-level masks for each segmented instance.
[0049] In addition to the primary architecture, the system is modular, allowing for alternative model embodiments where needed. For instance, variations may include lightweight versions optimized for edge computing or the integration of differentencoder-decoder combinations for improved performance in specific domains such as aerial imagery or low-light conditions. Scene Understanding and Output Representation:
[0050] The output of the system consists of a highly structured, semantically rich representation of the captured scene. This output includes two main data components: a property matrix and an adjacency matrix. Together, these matrices provides a comprehensive understanding of the scene. The property matrix catalogs all objects identified within the scene and encodes their key characteristics. These characteristics may include geometric properties such as shape, size, and spatial coordinates (x, y, z), as well as appearance-based attributes such as color, texture, and observable damage conditions.
[0051] The adjacency matrix supplements this by capturing the spatial relationships between all objects in the scene. Specifically, it encodes the distances and orientations between every pair of identified objects in the scene. From this comprehensive representation, the system is able to assess potential public safety hazards. Depending on the object properties and their contextual relationships, the system can generate automated alerts for incidents of interest, such as road hazards, structural damage, or improper object placement. Real-World Example: Pothole Detection
[0052] To illustrate the system’s functionality in a real-world scenario, consider the detection of potholes on urban roadways.
[0053] 1. Image capture: In this use case, a municipal vehicle equipped with a standard dashcam records video footage of city streets. As the vehicle moves, individual 2D frames are extracted and assessed by the system.
[0054] 2. Quality check: The system first assesses the captured image to ensure it is not blurry and meets a minimum resolution threshold. If the image is clear enough, it proceeds to the next step. The minimum threshold for image resolution is set at 800x800 pixels to ensure optimal performance
[0055] 3. Scene Understanding: Once an acceptable image is identified, it is processed by the trained segmentation model. The image is processed by the segmentation model which segments the scene and identifies all objects within it, such as 'road,' 'car,''sidewalk,' and an anomaly on the road surface. The system recognizes the context as relevant for public safety analysis. If deemed relevant, the system continues to depth estimation and property extraction.
[0056] 4. Depth Estimation & Property Extraction: A dedicated encoder-decoder model then generates a depth map from the 2D image, allowing the system to estimate the relative distance of the detected anomaly from the vehicle-mounted camera. The anomaly is classified as a pothole, an “object of interest” for public safety purposes. The system extracts specific properties of the pothole, including its shape, depth, surface roughness, and estimated width and length.
[0057] 5. Matrix construction: These properties are stored in the property matrix under a structured format, such as P_pothole = (x, y, z, color, texture, damage_state). Simultaneously, the adjacency matrix is populated with spatial relationships between the pothole and surrounding elements, such as nearby lane markings, the road edge, and pedestrian sidewalks. This contextual spatial data enables the system to determine not only the presence of the pothole but also its proximity to areas of pedestrian or vehicular traffic, influencing the urgency and type of alert generated. System Deployment and Infrastructure:
[0058] The system is designed for broad hardware compatibility, allowing image capture from multiple sources. These may include dashcams integrated with municipal or commercial vehicles, drones conducting aerial inspections, or handheld imaging devices used by on-ground inspectors. On the software side, the system is deployable both in cloud-based infrastructures and on-premise environments, facilitating seamless integration with existing municipal monitoring platforms, web portals, or mobile inspection apps. Performance Evaluation:
[0059] Model performance has been rigorously assessed on a dedicated test dataset containing 10,000 annotated samples. These samples represent a wide array of urban topologies, lighting conditions, and seasonal variations. The model's evaluation metrics include a precision of 84%, a recall of 76%, and an F1 score of 79%. These metrics confirm the model’s reliability and robustness across diverse real-world deployment scenarios.
[0060] The present invention offers several advantages as set forth herein. ^ Automation and Efficiency: The framework operates automatically, reducing manual effort and improving the speed of detecting public safety elements ^ General Purpose Applicability: Its versatility allows for use in various areas, both urban and rural, making it accessible to anyone ^ Configurability and Extensibility: The framework is customizable, enabling adaptation to different use cases such as smart city surveillance, visual pollution detection etc. ^ Cloud-Based and Compatible: Built on a cloud-based architecture, it is easily integrated into various systems and devices, compatible with software and web applications.
Claims
AMENDED CLAIMS received by the International Bureau on 31 December 2025 (31.12.2025)We claim1. A computer vision-based method (400) for detecting public safety elements in a real -world environment, comprising: a. capturing a 2D image (401) of the environment using an image-capturing device; b. performing image quality assessment (402) for the selection of images for further processing; c. segmenting the selected image to detect and understand the scene (404), and to classify objects using a convolutional neural network (CNN)-based segmentation model; d. estimating depth of each object (403) using an encoder-decoder based depth estimation model to determine 3D spatial relationships between objects; e. identifying and extracting properties (405) of objects relevant to public safety from the segmented image; f. generating an adjacency matrix (406) representing spatial and contextual relationships among detected objects; g. analyzing the adjacency matrix in conjunction with extracted object properties to identify object configurations indicative of a public safety risk based on spatial proximity, interaction, or contextual relationships among the objects; and h. generating real-time alerts or notifications when objects posing a potential public safety risk are detected.
2. The method (400) as claimed in claim 1, wherein the image quality assessment (402) comprises: a) checking image resolution against a predefined threshold; and b) performing binary classification for blur detection using a CNN.
3. The method (400) as claimed in claim 1, wherein object properties include at least one of geometric properties, color characteristics, surface texture patterns, spatial positioning in the 3D environment, and damage indicators.
4. The method (400) as claimed in claim 1, wherein the adjacency matrix is a square matrix of size N x TV, where TV is the number of objects detected in the image, and each matrix element comprises the distance between object z and object j; and a relationship score based on spatial or contextual proximity.
5. The method (400) as claimed in claim 1, further comprising generating a separate property matrix P, wherein each row Pi represents the object’s 3D position and its extracted features.
6. The method (400) as claimed in claim 1, further comprising filtering out irrelevant scenes before depth estimation and object property analysis using pre-defined rules.
7. A system for real-time detection of public safety hazards (200) using computer vision, comprising: a) at least one image capturing device (201) configured to collect 2D images of an environment; b) a computing unit (210) compri sing : i. an image quality evaluation module (212) configured to assess image resolution and detect blurriness using a CNN classifier; ii. a scene segmentation module (214) using a trained deep learning model to segment objects in the image; iii. a depth estimation module (216) for generating 3D coordinates of objects from 2D images; iv. an object property extraction engine (218) to extract properties essential for understanding the content and context of the visual data which include but not limited to geometry, colour, texture, and spatial properties of objects of interest; v. an adjacency matrix generator (220) to model spatial relationships and interaction between detected objects; and vi. a public safety analysis module configured to analyze the adjacency matrix in combination with the properties extracted by the property extraction engine to identify object configurations that pose public safety hazards and to generate real-time alerts or notifications.
8. The system (200) as claimed in claim 8, wherein the image capturing device (201) is selected from the group of a drone, a dashcam, a handheld camera, or a mobile surveillance system.
9. The system (200) as claimed in claim 8, wherein the scene segmentation module (214) is trained using a self-supervised contrastive learning approach, followed by manual annotation of only a subset (10%) of the dataset.
10. The system (200) as claimed in claim 8, wherein the scene segmentation model (214) is used for automatic annotation of public safety elements through interactive prompting using points and bounding boxes.
11. The system (200) as claimed in claim 8, wherein the computing unit (210) is deployed on a cloud-based architecture for integration with municipal or smart city surveillance systems.
Citation Information
Patent Citations
Hazard detection through computer vision
US11176383B2
Image assessment using deep convolutional neural networks
US20160035078A1
Three-dimensional object detection based on image data
US20230029900A1
Systems and methods for extracting information about objects from scene information
US9904867B2