An artificial intelligence (AI) based system and method of estimating traffic count
Patent Information
- Application Number
- PCT/IN2025/050255
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-20
- Filing Date
- 2025-02-20
- Publication Date
- 2025-10-30
AI Technical Summary
Existing traffic count estimation methods require manual data collection, which is time-consuming, labor-intensive, and prone to human error, necessitating a more efficient and accurate automated solution.
An AI-based system utilizing knowledge distillation and deep learning on edge devices for real-time traffic surveillance, capturing traffic scenes, detecting objects, tracking them until they exit the field of view, and estimating traffic count.
Automates traffic count estimation, providing continuous, accurate, and reliable results with reduced computational complexity, capable of handling high traffic volumes and complex scenes.
Smart Images

Figure IN2025050255_30102025_PF_FP_ABST
Abstract
Description
AN ARTIFICIAL INTELLIGENCE (Al) BASED SYSTEM AND METHOD OF ESTIMATING TRAFFIC COUNTTECHNICAL FIELD
[0001] The present disclosure relates to traffic monitoring and / or traffic surveillance systems. Particularly, the present disclosure relates to an Artificial Intelligence (Al) based system and method of estimating traffic count.BACKGROUND OF THE DISCLOSURE
[0002] The following description includes information that may be useful in understanding the present disclosure It is not an admission that any of the information provided herein is prior art or relevant to the presently claimed disclosure, or that any publication specifically or implicitly referenced is prior art.
[0003] The drastic growth in the number of vehicles in the last few decades has necessitated significantly better traffic management and planning. To manage the traffic efficiently, traffic volume is an essential parameter. Most existing methods solve the vehicle counting problem under the assumption of state-of-the-art computation power. With the recent growth in cost-effective Internet of Things (loT) devices and edge computing, several machine learning models are being tailored for such loT devices. Further, with the recent advancements in Artificial Intelligence and computer vision, objection detection techniques are traction which uses various traffic cameras or other vision-based sensors for capturing live feed of traffic scenes. However, existing solutions estimate traffic counts with the need for manual data collection, which can be time-consuming, labor-intensive, and prone to human error.
[0004] Thus, there exists a need for a mechanism for efficiently, accurately, and effectively estimating traffic counts, thereby overcoming the above-mentioned limitations of the conventional methods.
[0005] The above-mentioned drawbacks / difficulties / disadvantages of the conventional techniques are explained just for exemplary purpose and this disclosure and description mentioned below would never limit its scope only to such problems. A person skilled in the art may understand that this disclosure and below mentioneddescription may also solve other problems or overcome the other drawbacks / disadvantages of the conventional arts, which are not explicitly captured above.SUMMARY OF THE DISCLOSURE
[0006] The present disclosure overcomes one or more shortcomings of the prior art and provides additional advantages discussed throughout the present disclosure. Additional features and advantages are realized through the techniques of the present disclosure. Other embodiments and aspects of the disclosure are described in detail herein and are considered a part of the claimed disclosure.
[0007] In one non-limiting embodiment of the present disclosure, an Artificial Intelligence (Al) based method of estimating traffic count is disclosed. The method comprises capturing at least one traffic scene in a Field of View (FOV) of an edge device. Further, the method detects presence of one or more objects within the at least one traffic scene by analyzing at least one frame of the at least one traffic scene, using an object detection process. Furthermore, the method comprises tracking the one or more of detected objects by analyzing at least one consecutive frame with respect to the one or more detected objects until the one or more of detected objects exits the FoV. Thereafter, the method estimates the traffic count based on the tracking of the one or more detected objects.
[0008] In one non-limiting embodiment of the present disclosure, for detecting the presence of one or more objects within the at least one traffic scene, the method further comprises detecting presence of one or more objects using a pre-trained student network.
[0009] In one non-limiting embodiment of the present disclosure, wherein the student network is trained using knowledge distillation.
[0010] In one non-limiting embodiment of the present disclosure, for tracking the one or more of detected objects, the method further comprises tracking the one or more of detected objects using a pre-trained student network. The student network is trained using knowledge distillation.
[0011] In one non-limiting embodiment of the present disclosure, for estimating the traffic count of the one or more detected objects, the method further comprises counting the one or more detected objects passing through a road segment captured within the at least one traffic scene.
[0012] In another non-limiting embodiment of the present disclosure, for counting the one or more detected objects based on the tracking of the one or more detected objects, the method comprises defining at least one virtual line within the at least one traffic scene and tallying the one or more detected objects crossing the at least one virtual line to count the one or more detected objects.
[0013] In another non-limiting embodiment of the present disclosure, the method further comprises processing the estimated traffic count to generate one or more traffic insights and displaying the one or more traffic insights on a user interface.
[0014] In another non-limiting embodiment of the present disclosure, the method further comprises detecting speed of the one or more detected objects. Further, the detecting comprises performing an optical mapping of the at least one traffic scene to real -world coordinates visible in the FoV of the edge device.
[0015] In another non-limiting embodiment of the present disclosure, the method further comprises detecting illegal parking of the one or more detected objects. Further, the determining comprises determining whether the one or more detected objects are stationary in the at least one consecutive frame for a predetermined time.
[0016] In another non-limiting embodiment of the present disclosure, the method further comprises detecting movement of the one or more detected objects in inappropriate lanes. Further, the detecting comprises extracting one or more trajectories of the one or more detected objects within the at least one traffic scene and clustering the one or more trajectories to identify the movement of the one or more detected objects in the inappropriate lanes.
[0017] In another non-limiting embodiment, the present disclosure discloses an Artificial Intelligence (Al) based system for estimating traffic count. The Al basedsystem comprises an edge device and at least one image capturing device. The image capturing device is electronically coupled to the edge device. The at least image capturing device configured to capture at least one traffic scene in a a Field of View (FoV) of the edge device. The edge device is configured to detect presence of one or more objects within the at least one traffic scene by analyzing at least one frame of the at least one traffic scene, using an object detection process. Thereafter, the edge device is configured to track the one or more of detected objects by analyzing at least one consecutive frame with respect to the one or more detected objects until the one or more of detected objects exits the FoV. Finally, the edge device is configured to estimate the traffic count based on the tracking of the one or more detected objects.
[0018] The foregoing summary is illustrative only and is not intended to be in any way limiting. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features will become apparent by reference to the drawings and the following detailed description.BRIEF DESCRIPTION OF DRAWINGS
[0019] The embodiments of the disclosure itself, as well as a preferred mode of use, further objectives and advantages thereof, will best be understood by reference to the following detailed description of an illustrative embodiment when read in conjunction with the accompanying drawings. One or more embodiments are now described, by way of example only, with reference to the accompanying drawings in which:
[0020] Fig. 1 illustrates an exemplary environment indicating an Artificial Intelligence (Al) based system for estimating traffic count, in accordance with an embodiment of the present disclosure.
[0021] Fig. 2 depicts a detailed block diagram of an edge device of the Al based system of Fig. 1 for estimating traffic count, in accordance with an embodiment of the present disclosure.
[0022] Fig. 3a depicts a knowledge distillation technique for training student network from a teacher network, in accordance with an embodiment of the present disclosure.
[0023] Fig. 3b depicts deployment of trained student network from training phase to inference phase, in accordance with an embodiment of the present disclosure.
[0024] Fig. 4 shows a flowchart depicting an Al based method of estimating traffic count, in accordance with an embodiment of the present disclosure.
[0025] The figures depict embodiments of the disclosure for purposes of illustration only. One skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods illustrated herein may be employed without departing from the principles of the disclosure described herein.DETAILED DESCRIPTION
[0026] The foregoing has broadly outlined the features and technical advantages of the present disclosure in order that the detailed description of the disclosure that follows may be better understood. It should be appreciated by those skilled in the art that the conception and specific embodiment disclosed may be readily utilized as a basis for modifying or designing other structures for carrying out the same purposes of the present disclosure.
[0027] The novel features which are believed to be characteristic of the disclosure, both as to its organization and method of operation, together with further objects and advantages will be better understood from the following description when considered in connection with the accompanying Figures. It is to be expressly understood, however, that each of the Figures is provided for the purpose of illustration and description only and is not intended as a definition of the limits of the present disclosure.
[0028] Existing solutions estimate traffic counts with the need for manual data collection, which can be time-consuming, labor-intensive, and prone to human error. The present disclosure recognizes this problem and provides a knowledge distillation and deep learning-based computer vision system for real-time heterogeneous road traffic surveillance on edge devices. Particularly, the present disclosure discloses an Artificial Intelligence (Al) based system and method of estimating traffic count. To do so, the method initially acquires video frames / feeds or traffic scenes through high- resolution image capturing devices placed at suitable traffic locations. For example, theimage capturing devices may be installed on traffic signals, overpasses or a vantage point which may provide a clear road view.
[0029] Further, the method captures traffic scenes in a Field of View (FoV) of the edge device and detects presence of objects within the at least one traffic scene by analyzing at least one frame of the at least one traffic scene. The method detects the presence of the objects using an object detection process such as a pre-trained student network which may be trained using knowledge distillation. Thereafter, the method tracks the detected objects by analyzing at least one consecutive frame with respect to the detected objects until the detected objects exit the FoV. Finally, the method estimates the traffic count based on the tracking of detected objects.
[0030] Thus, the present disclosure estimates the traffic counts automatically without manual data collection which can be time-consuming, labor-intensive, and prone to human error. By automating the process of estimating the traffic count, traffic data may be collected continuously and consistently, allowing for more accurate and reliable results. Moreover, the present disclosure utilizes object detection process for efficiently handling high traffic volumes and complex scenes, providing a comprehensive view of traffic patterns. Additionally, the present disclosure utilizes only one teacher student network for object detection process which makes the computation relatively simple and easy for estimating the traffic count.
[0031] Fig. 1 illustrates an exemplary environment 100 indicating an Artificial Intelligence (Al) based system 102 (hereinafter referred to as system) for estimating traffic count, in accordance with an embodiment of the present disclosure. The architecture 100 depicts the system 102 installed on a traffic signal 112 for estimating traffic density or traffic count, tracking traffic violations, detecting lane changes, detecting speed of vehicles, etc., In some embodiments, the system 102 may be installed on overpasses, vantage points, etc., which may provide a clear road view. In some embodiments, the system 102 may be defined as the traffic control device, or a sensing device installed on traffic signals. Further, the system 102 may use knowledge distillation and deep learning-based computer vision techniques for real-time heterogeneous road traffic surveillance.
[0032] The system 102 may comprise an image capturing device 104 and an edge device 106. The image capturing device 104 may be electronically coupled to the edge device 106. In some embodiments, the image capturing device 104 may comprise but not limited to high resolution camera or vision-based sensors for capturing traffic scenes. In other words, the image capturing device 104 may capture the traffic scenes within a Field of View (FoV) 114 of the edge device 106. In some embodiments, the FoV 114 of the edge device 106 may be defined as an area of the at least one traffic scene captured by the image capturing device 104. The captured traffic scenes may comprise video frames or image frames of moving objects such as vehicles 116, pedestrians, etc., Accordingly, in one embodiment, the traffic scene may indicate movement of vehicles 116 in virtual lanes (118a, 118b, 118c). In another embodiment, the traffic scene may indicate movement of pedestrians crossing a road segment. It may be worth noting that the image capturing device 104 may continuously monitor and capture the traffic scenes of the moving objects until the objects have left the FOV 114 of the edge device 106.
[0033] Further, in one embodiment, the edge device 106 may be configured internally within the image capturing device 104. In an alternative embodiment, the edge device 106 may be configured external to the image capturing device 104. In some embodiments, the edge device 106 configured to perform the functionality of the present disclosure. In a non-limiting embodiment, the edge device 106 may be a chipset which may be installed within the system 102. In another non-limiting embodiment, the edge device 106 may be installed in vehicles 116 for analyzing the traffic data captured by image capturing device (such as 360° camera) or sensors placed on the front end and / or rear end of the vehicles 116. The edge device 106 receives the traffic scenes captured by the image capturing device 104 and analyzes the traffic for estimating the traffic count. Further, the edge device 106 may comprise, but not limited to, a processor 108 and a memory 110. In some embodiments, the processor 108 may perform the method disclosed in the present disclosure. In some embodiments, the memory 110 may store the traffic scenes, image or video frames, captured during different time intervals. The detailed explanation of estimating traffic count is provided here below with reference to Fig. 2.
[0034] Fig. 2 depicts a detailed block diagram of an edge device 106 for estimating traffic count, in accordance with an embodiment of the present disclosure. The edge device 106 may be electronically coupled to the image capturing device 104. The image capturing device 104 may capture at least one traffic scene in a Field of View (FoV) 114 of the edge device 106. The edge device 106 may receive the at least one captured traffic scene and analyze the at least one traffic scene for traffic count estimation. In an embodiment, the edge device 106 may be a dedicated control unit configured to perform the method of the present disclosure and comprise a processor 108, an Input / Output (I / O) interface 202, one or more modules 204, and a memory 110, which may be communicatively coupled with each other. However, the edge device 106 may comprise any other additional components, which may be required for performing functionalities in accordance with the present disclosure.
[0035] As used herein, the processor 108 may refer to an Application Specific Integrated Circuit (ASIC), an electronic circuit, a hardware processor (shared, dedicated, or group) and memory that execute one or more software or firmware programs, a combinational logic circuit, and / or other suitable components that provide the described functionality. In another implementation, the processor 108 may comprise various modules 204 including an image capturing module 206, a detection module 208, a tracking module 210, an estimation module 212 and other modules 214. In an embodiment, one or more other modules apart from the above-mentioned modules 204 may be used to perform various miscellaneous functionalities of the edge device 106. It shall be appreciated that such modules may be represented as a single module or a combination of different modules. In some embodiments, the processor 108 may perform one or more functions of the edge device 106 for traffic count estimation.
[0036] In an embodiment, the I / O interface 202 may include any input or output interface that may be configured for receiving inputs such as, without limiting to, traffic scenes comprising frames of captured vehicles, traffic count, etc. from various sources. Additionally, the I / O interface 202 may include a variety of software and hardware interfaces, for example, a web interface, a graphical user interface, an input device, an output device, and the like, for receiving the traffic scenes or data captured by image capturing device 104 or any other information from various sources.
[0037] In a non-limiting embodiment, the memory 110 may be an external memory chip or an inbuilt Electrically Erasable Programmable Read-Only Memory (EEPROM) memory. In an embodiment, the memory 110 may be a computer-readable medium known in the art including, for example, volatile memory, such as static random-access memory (SRAM) and dynamic random-access memory (DRAM), and / or synchronous dynamic random-access memory (SDRAM) and / or non-volatile memory, such as read only memory (ROM), erasable programmable ROM, flash memories, hard disks, optical disks, and magnetic tapes. In some implementations, the memory 110 may store the traffic count 216 related to the traffic, historical data 218 comprising traffic scenes captured during different time intervals in a day or week or month or years, frames 220 of the traffic scenes, and other data 222 in the form of various data structures. Additionally, the historical data 218 may be organized using data models, such as relational or hierarchical data models. The other data 222 may include various temporary data and files generated by the processor 108 while performing various functions of the edge device 106. In some embodiments, the memory 110 may be communicatively coupled to the processor 108 and may store various data processed by the processor 108.
[0038] In accordance with the present disclosure, the image capturing module 206 may capture at least one traffic scene in a Field of View (FoV) of the edge device 106. The at least one traffic scene indicates movement of number of vehicles in real-time. In some embodiments, the image capturing module 206 may capture traffic scenes continuously. In some embodiments, the image capturing module 206 may capture traffic scenes during different time intervals. Further, the memory 110 may store the captured traffic scenes. Further, the image capturing module 206 may transmit the captured traffic scenes to the edge device 106 or detection module 208 for estimating the traffic count.
[0039] In some embodiments, the detection module 208 may detect presence of one or more objects within the at least one traffic scene by analyzing at least one frame of the at least one traffic scene. In some embodiments, the one or more objects may include but not limited to, vehicles, pedestrians, etc., The detection module 208 may detect presence of the one or more objects using an object detection process. The goal of theobject detection process is to identify the objects present in the at least one frame and give boundary boxes to the identified objects. The detection module 208 may detect the presence of the one or more objects using a pre-student network. In some embodiments, the student network may be trained using the knowledge distillation technique. In some embodiments, the object detection technique for object detection may include but not limited to Histogram of Oriented Gradients (HOG), Scale Invariant Feature Transform (SIFT), Support Vector Machines (SVM) classification etc., In a non-limiting embodiment, the deep learning models for object detection may include, but not limited to, knowledge distillation, Faster Region-based Convolutional Neural Networks (RCNN), You Only Look Once (YOLO), and Single Shot Detector (SSD) etc.,
[0040] In some embodiments, the detection module 208 may detect speed of the one or more detected objects by performing an optical mapping of at least one traffic scene to real-world coordinates visible in the FoV 114 of the edge device 106. Optical mapping is a well-known technique for detecting speed of the objects, however a skilled person may appreciate that any other suitable technique may be used in the present disclosure to detect the speed of the one or more detected objects.
[0041] In some embodiments, the detection module 208 may be further configured to detect illegal parking of the one or more detected objects. To detect illegal parking of the one or more detected objects, the detection module 208 may determine whether the one or more detected objects are stationary in the at least one consecutive frame for a predetermined time. In an exemplary embodiment, if a vehicle is stationary in one or more consecutive frames for at least 10-15 minutes, then the vehicle is said to be illegally parked on the road segment. In other words, if a vehicle is not moving in consecutive frames continuously, then the vehicle is said to be parked illegally. It may be worth noting this is merely an example to understand the concept and should not be taken into limiting sense.
[0042] In some embodiments, the detection module 208 may detect movement of the one or more detected objects in inappropriate lanes. To do so, the detection module 208 may extract one or more trajectories of the one or more detected objects within the at least one traffic scene . Thereafter, the detection module 208 may cluster the one or moretrajectories to identify the movement of the one or more detected objects in the inappropriate lanes or wrong lanes. In some embodiments, the detection module 208 may send alerts upon identifying that the object is moving in the wrong lane.
[0043] In some embodiments, the tracking module 210 may track the one or more detected objects by analyzing at least one consecutive frame with respect to the one or more detected objects until the one or more detected objects exits the FoV 114. In an exemplary embodiment, once the vehicle 116 is detected, the tracking module 210 may track location or movement of the vehicle 116 in next particular frame or successive frame or consecutive frames with respect to the detected vehicle 116 until the vehicle 116 leaves or exits the FoV 114 of the edge device 106. In some embodiments, the tracking module 210 may track the one or more detected objects using a pre-trained student network. The student network may be trained using the knowledge distillation technique explained in Fig. 3a.
[0044] Once the one or more detected objects has left the FoV, the estimation module 212 may estimate the traffic count based on the tracking of the one or more detected objects. For estimating the traffic count, the estimation module 212 may define at least one virtual line within the at least one traffic scene. Further, the estimation module 212 may tally the one or more detected objects crossing the at least one virtual line to count the one or more detected objects passing through a road segment. In some embodiments, the traffic count may be stored in the memory 110 and a dataset is prepared in dashboard for indicating real-time statistics of the traffic such as traffic hours, congestion in particular region or area, etc.,
[0045] In some embodiments, the estimation module 212 may process the estimated traffic count to generate one or more traffic insights. In a non -limiting embodiment, the one or more traffic insights may be related to each of the one or more detected objects for determining traffic violations performed by the detected object in past hours. For example, the one or more traffic insights may indicate whether the speed of the one or more detected object is within a predefined threshold. In another embodiment, the one or more traffic insights may indicate whether the detected object is illegally parked on the road. In another embodiment, the one or more traffic insights may indicate whetherthe detected object is moving in wrong lane. Further, the estimation module 212 may display the one or more traffic insights on a user interface. In an embodiment, the one or more traffic insights may play a crucial role in determining traffic density during peak hours in a day. In another embodiment, the one or more traffic insights may be useful for identifying congestion in a particular area.
[0046] Thus, the present disclosure estimates the traffic counts automatically without manual data collection which can be time-consuming, labor-intensive, and prone to human error. By automating the process of estimating the traffic count, traffic data may be collected continuously and consistently, allowing for more accurate and reliable results. Moreover, the present disclosure utilizes object detection process for efficiently handling high traffic volumes and complex scenes, providing a comprehensive view of traffic patterns. Additionally, the present disclosure utilizes only one teacher student network for object detection process which makes the computation relatively simple and easy for estimating the traffic count.
[0047] Fig. 3a depicts a knowledge distillation technique for training a student network from a teacher network, in accordance with an embodiment of the present disclosure. The knowledge distillation technique may be defined as a process of knowledge transfer 302 from a teacher (large) network 304 to a student (smaller) network 306. In some embodiments, the teacher network 304 may be a relatively complex network trained on the data 308 (i.e., dataset) of interest. Thereafter, the teacher network 304 may be used as source of knowledge which may be transferred to the student network 306. Accordingly, the teacher network 304 transfers the knowledge to the student network 306 which is lightweight and easier to use. This trained student network may be deployed in the edge device 106 for estimating traffic count in the aspects of the present disclosure. It may be worth noting that the present disclosure utilizes only one teacher - student network for estimating traffic count which makes the computation of the traffic data efficient and easier.
[0048] Fig. 3b depicts deployment of trained network (student network) for estimating traffic count from training phase to inference phase, in accordance with an embodiment of the present disclosure. In the training phase, initially, the image capturing device 104implemented on traffic signal captures traffic scenes within the FoV of the edge device 106. The traffic scenes may include but not be limited to, frames 310 (i.e., image or video frames). The traffic scenes may be captured during different time intervals in a day in real-time. The captured traffic scenes may be stored in memory and may be used as a dataset for training the teacher-student network. In an exemplary scenario, the dataset prepared for training may comprise the traffic scenes captured in 2 months, 4 months, 7 months, 1 year, 2years, etc., Further, during process of data augmentation 312, the dataset is to be made ready for the deep learning model to be trained using knowledge distillation 304.
[0049] During the inference phase, after the student network is fully trained, it is deployed on the edge device 106. Thereafter, the input traffic stream 314 is passed through the trained student network 306 and vehicle’s trajectories may be estimated. In some embodiments, a Simple Online Real -time Tracker (SORT) may be used to estimate the vehicle’s trajectories for estimating the traffic count. Further, a Kalman filter is applied on the estimated vehicle’s trajectories. In an exemplary embodiment, the Kalman filter 316 applies 8 variables; (u, v, a, h, u’ , v’, a’ , h’) on the vehicle’s trajectories, where (u, v) are centres of the bounding boxes, a is the aspect ratio and h, the height of the image. The other variables are the respective velocities of the variables. In this manner, a “Track” for each detection with all necessary state information of vehicles may be created. Finally, unique Identification (IDs) may be created and assigned to each “Tracks” using a Hungarian Algorithm 316 for differentiating the vehicles in the FoV 114 of the edge device. However, the description should not be taken into a limiting sense.
[0050] Fig. 4 depicts an Artificial Intelligence (Al)-based method 400 of estimating traffic count, in accordance with an embodiment of the present disclosure. The method 400 may be described in the general context of computer executable instructions. Generally, computer executable instructions may include routines, programs, objects, components, data structures, procedures, modules, and functions, which perform specific functions or implement specific abstract data types.
[0051] The order in which the method 400 is described is not intended to be construed as a limitation, and any number of the described method blocks may be combined inany order to implement the method. Additionally, individual blocks may be deleted from the methods without departing from the spirit and scope of the subject matter described. Further, it may be noted that the method 400 may be performed by the processor 108 or any of the modules 204 of the edge device 106, as shown in Fig. 2. The description of Fig. 4 is provided with reference to Figs. 1-2.
[0052] At block 402, the method 400 may comprise capturing at least one traffic scene in a Field of View (FoV) 114 of the edge device 106. The at least one traffic scene may be captured by an image capturing device 104 which may include, but not limited to, analog camera, digital camera, vision based sensors, etc. In some embodiments, the at least one traffic scene may include images or video frames of moving objects. In some embodiments, the objects may include, but not limited to, vehicles 116, pedestrians, etc., The method 400 at block 402 may be performed by the processor 108 or the image capturing module 206.
[0053] Further, the method 400 at block 404 may comprise detecting presence of one or more objects within the at least one traffic scene by analyzing at least one frame of the at least one traffic scene, using an object detection process. Accordingly, the method detects the presence of one or more vehicles or pedestrians captured in the at least one frame of the at least one traffic scene. The method may detect the presence of the one or more vehicles or pedestrians using any suitable object detection process. The method 400 at block 404 may be performed by the processor 108 or a detection module 208.
[0054] In some embodiments, for detecting the presence of the one or more objects within the at least one traffic scene, the method may comprise detecting presence of one or more objects using a pre-trained student network 306. The student network 306 may be trained using the knowledge distillation 302 explained with reference to Fig. 3a.
[0055] The method 400 at block 406 comprises tracking the one or more detected objects by analyzing at least one consecutive frame with respect to the one or more detected objects until the one or more detected objects exits the FoV 114. In some embodiments, the method for tracking the one or more detected objects further comprises tracking one or more detected objects using a pre-trained student network 306. The student network 306 may be trained using the knowledge distillation processexplained with reference to Fig. 3a. The method 400 at block 406 may be performed by the processor 108 or a tracking module 210.
[0056] The method 400 at block 408 comprises estimating the traffic count based on the tracking of the one or more detected objects. In some embodiments, the method for estimating the traffic count of the one or more detected objects further comprises counting the one or more detected objects passing through a road segment captured within the at least one traffic scene. In some embodiments, the method step 408 may be the processor 108 or an estimation module 212. In some embodiments, the method counts the one or more detected objects by defining at least one virtual line within the at least one traffic scene and tallying the one or more detected objects crossing the at least one virtual line to count the one or more detected objects.
[0057] In some embodiments, the method may comprise processing the estimated traffic count to generate one or more traffic insights and displaying the one or more traffic insights on a user interface. The method may be performed by the processor 108.
[0058] In some embodiments, the method may comprise detecting speed of the one or more detected objects by performing an optical mapping of the at least one traffic scene to real-world coordinates visible in the FoV 114 of the edge device 106. The method 400 may be performed by the processor 108.
[0059] In some embodiments, the method may further comprise detecting movement of the one or more detected objects in inappropriate lanes by extracting one or more trajectories of the one or more detected objects within the at least one traffic scene. Thereafter, the method may perform clustering of the one or more trajectories to identify the movement of the one or more detected objects in the inappropriate lanes. The method 400 may be performed by the processor 108.
[0060] A description of an embodiment with several components in communication with each other does not imply that all such components are required. On the contrary, a variety of optional components are described to illustrate the wide variety of possible embodiments of the disclosure.
[0061] When a single device or article is described herein, it will be clear that more than one device / article (whether they cooperate) may be used in place of a single device / article. Similarly, where more than one device or article is described herein (whether they cooperate), it will be clear that a single device / article may be used in place of the more than one device or article or a different number of devices / articles may be used instead of the shown number of devices or programs. The functionality and / or the features of a device may be alternatively embodied by one or more other devices which are not explicitly described as having such functionality / features. Thus, other embodiments of the disclosure need not include the device itself.
[0062] Finally, the language used in the specification has been principally selected for readability and instructional purposes, and it may not have been selected to delineate or circumscribe the inventive subject matter. It is therefore intended that the scope of the disclosure be limited not by this detailed description, but rather by any claims that issue on an application based here on. Accordingly, the embodiments of the present disclosure are intended to be illustrative, but not limiting, of the scope of the disclosure, which is set forth in the following claims.
[0063] While various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for purposes of illustration and are not intended to be limiting, with the true scope and spirit being indicated by the following claims.
[0064] Reference Numerals:
Claims
We claim:
1. An Artificial Intelligence (Al)-based method of estimating traffic count, comprising: capturing at least one traffic scene in a Field of View (FoV) of an edge device; detecting presence of one or more objects within the at least one traffic scene by analyzing at least one frame of the at least one traffic scene, using an object detection process; tracking the one or more of detected objects by analyzing at least one consecutive frame with respect to the one or more detected objects until the one or more of detected objects exits the FoV; and estimating the traffic count based on the tracking of the one or more detected objects.
2. The method as claimed in claim 1 , wherein detecting the presence of one or more objects within the at least one traffic scene, comprising: detecting presence of one or more objects using a pre-trained student network, wherein the student network is trained using knowledge distillation.
3. The method as claimed in claim 1 , wherein tracking the one or more of detected objects, comprising: tracking the one or more of detected objects using a pre-trained student network, wherein the student network is trained using knowledge distillation.
4. The method as claimed in claim 1, wherein estimating the traffic count of the one or more detected object, comprising: counting the one or more detected objects passing through a road segment captured within the at least one traffic scene.
5. The method as claimed in claim 4, wherein counting the one or more detected objects based on the tracking of the one or more detected objects, comprising: defining at least one virtual line within the at least one traffic scene; and tallying the one or more detected objects crossing the at least one virtual line to count the one or more detected objects.
6. The method as claimed in claim 1, further comprising: processing the estimated traffic count to generate one or more traffic insights; and displaying the one or more traffic insights on a user interface.
7. The method as claimed in claim 1, further comprising detecting speed of the one or more detected objects, wherein the detecting comprising: performing an optical mapping of the at least one traffic scene to real-world coordinates visible in the FoV of the edge device.
8. The method as claimed in claim 1, further comprising detecting illegal parking of the one or more detected objects, wherein the determining comprising: determining whether the one or more detected objects are stationary in the at least one consecutive frame for a predetermined time.
9. The method as claimed in claim 1, further comprising detecting movement of the one or more detected objects in inappropriate lanes, wherein the detecting comprising: extracting one or more trajectories of the one or more detected objects within the at least one traffic scene; and clustering the one or more trajectories to identify the movement of the one or more detected objects in the inappropriate lanes.
10. An Artificial Intelligence (Al) based system for estimating traffic count, comprising: an edge device; and at least one image capturing device electronically coupled to the edge device, the at least one image capturing device configured to capture at least one traffic scene in a Field of View (FoV) of the edge device; the edge device is configured to: detect presence of one or more objects within the at least one traffic scene by analyzing at least one frame of the at least one traffic scene, using an object detection process;track the one or more of detected objects by analyzing at least one consecutive frame with respect to the one or more detected objects until the one or more of detected objects exits the FoV; and estimate the traffic count based on the tracking of the one or more detected objects.
11. The Al based system as claimed in claim 10, wherein to detect the presence of one or more objects within the at least one traffic scene, the edge device is configured to: detect presence of one or more objects using a pre-trained student network, wherein the student network is trained using knowledge distillation.
12. The Al based system as claimed in claim 10, wherein to track the one or more of detected objects, the edge device is further configured to: track the one or more of detected objects using a pre-trained student network, wherein the student network is trained using knowledge distillation.
13. The Al based system as claimed in claim 10, wherein to estimate the traffic count of the one or more detected objects, the edge device is further configured to: count the one or more detected objects passing through a road segment captured within the at least one traffic scene.
14. The Al based system as claimed in claim 13, wherein to count the one or more detected objects, the edge device is further configured to: define at least one virtual line within the at least one traffic scene; and tally the one or more detected objects crossing the at least one virtual line to count the one or more detected objects.
15. The Al based system as claimed in claim 10, wherein the edge device is further configured to: process the estimated traffic count to generate one or more traffic insights; and display the one or more traffic insights on a user interface.
16. The Al based system as claimed in claim 10, wherein the edge device is configured to detect speed of the one or more detected objects, and wherein to detect the speed the edge device is further configured to: perform an optical mapping of the at least one traffic scene to real-world coordinates visible in the FoV of the edge device.
17. The Al based system as claimed in claim 10, wherein the edge device is configured to detect illegal parking of the one or more detected objects, wherein to detect the illegal parking, the edge device is further configured to: determine whether the one or more detected objects are stationary in the at least one consecutive frame for a predetermined time.
18. The Al based system as claimed in claim 10, wherein the edge is configured to detect movement of the one or more detected objects in inappropriate lanes, wherein to detect the movement the edge is further configured to: extract one or more trajectories of the one or more detected objects within the at least one traffic scene; and cluster the one or more trajectories to identify the movement of the one or more detected objects in the inappropriate lanes.
Citation Information
Patent Citations
Methods and systems for interpreting traffic scenes
US20210183238A1
Method and apparatus for detecting traffic anomaly
US20230036864A1