Image marking system and method thereof
By automatically generating labeled datasets using proximity sensors and training machine learning models, the tediousness and instability of manual labeling have been solved, enabling automated object recognition and safety prediction, and improving the service quality and safety of transportation systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-18
- Publication Date
- 2026-03-20
AI Technical Summary
Existing machine learning models rely on manual labeling during the training phase, which makes the labeling process cumbersome and of inconsistent quality, affecting the service quality and safety of the transportation system.
By automatically generating labeled datasets using proximity sensor data, machine learning models are trained to identify and track objects of interest, reducing reliance on human operators and enabling predictive decision-making and potential collision avoidance through machine learning models.
It enables automated generation of tagged data and object recognition, improving the service consistency and security of the transportation system and reducing the occurrence of human error.
Smart Images

Figure CN116635302B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates to a system and method for labeling images to identify and monitor objects of interest. Furthermore, the present invention relates to image processing and machine learning methods and systems. In particular, but not exclusively, to identifying and labeling objects or entities of interest captured in a series of images. The present invention also relates to the training of machine learning models. The trained machine learning models can uniquely identify objects in videos and images and track their positions. Furthermore, the machine learning models can detect anomalies to prevent damages or accidents. Moreover, the trained models can be used to remotely control mobile objects to autonomously perform their tasks. BACKGROUND
[0002] Current machine learning models use artificially annotated objects in the training phase. The labeling process is cumbersome and the quality of the human operation and labeling is closely related to the individual’s knowledge, bias in performing the task and often decreases due to distraction or fatigue of the expert.
[0003] Almost all service operations in transportation systems heavily rely on human and the experience of the staff directly influences the quality of the service. Variations in decision making and providing the service lead to uncontrolled and inconsistent quality of the service. In complex environments, such as in the transportation industry, distraction or fatigue of the operator can lead to errors with catastrophic consequences in the aftermath. SUMMARY
[0004] Embodiments of the present invention seek to solve the above-mentioned problems by providing a system that labels data of mobile devices and vehicles from cameras or other monitoring sensors. In particular, embodiments of the present invention utilize data from one or more proximity sensors, or in other words proximity detectors, to identify or detect objects of interest in raw data from one or more monitoring sensors. Advantageously, embodiments of the present invention do not require a human operator with specific domain knowledge to perform a manual annotation of the data. Instead, data from one or more proximity sensors is used to annotate data from one or more monitoring sensors. Embodiments of the present invention are thereby able to automatically generate a labeled dataset.
[0005] In further embodiments of the present invention, the labeled dataset is used for training of a machine learning model. The machine learning model is trained using the labeled dataset to identify specific objects of interest, such as mobile devices or vehicles that can be captured in the monitoring sensor data.
[0006] In a further embodiment of the invention, the trained machine learning model is used to identify objects of interest, such as mobile devices or vehicles captured in the monitoring sensor data. The identified objects of interest are localized in order to provide tracking, generate alerts in response to predicted collisions, improve service and provide guidance.
[0007] Embodiments of the invention can predict the subsequent effects of decisions and predict the outcome of scenarios to avoid undesirable actions.
[0008] Also disclosed is a method for generating a labeled dataset or for training a machine learning model or for detecting one or more objects of interest, and a computer program product that, when executed, performs the method for generating a labeled dataset or for training a machine learning model or for detecting one or more objects of interest. BRIEF DESCRIPTION OF DRAWINGS
[0009] Figure 1 An aircraft in situ on a typical airport apron is shown within the field of view of one or more monitoring sensors, such as cameras;
[0010] Figure 2 The position of one or more proximity sensors mounted at the end of a boarding bridge on an airport apron, and the position of one or more proximity devices mounted on one or more associated vehicles are shown;
[0011] Figure 3 An example of a list of device IP connections is shown;
[0012] Figure 4 is a schematic diagram showing the different functional components of an embodiment of the invention;
[0013] Figure 5 is a flowchart showing the different steps of a method of annotating monitoring sensor data according to an embodiment of the invention;
[0014] Figure 6 Annotated monitoring sensor data output by a trained machine learning model of the invention is shown;
[0015] Figure 7 Some examples of system applications for improving aircraft tracking are shown;
[0016] Figure 8 A method of automatic collision detection according to an embodiment of the invention is shown;
[0017] Figure 9 A method for providing autonomous service vehicles according to an embodiment of the invention is shown; and
[0018] Figure 10A flowchart of a method for providing autonomous service vehicles according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0019] The following exemplary description is based on systems, devices and methods for the aviation industry. However, it should be understood that the present application can find application outside the aviation industry, including in other transportation industries, or in the delivery industry that transports items between different locations, or in industries that involve coordination of multiple vehicles. For example, embodiments of the present application can also find application in the shipping, railway or highway industries.
[0020] The following described embodiments can be implemented using the Python programming language, for example using the OpenCV TM , Tensorflow TM and Keras TM libraries.
[0021] Embodiments of the present application have two main stages:
[0022] 1 - Label data in order to train a machine learning model for the unique identification of objects.
[0023] 2 - Monitor and control devices and vehicles in the environment and analyze the decision results to achieve optimal operation.
[0024] Dataset creation phase
[0025] The monitoring data of the objects of interest is captured by one or more monitoring sensors, such as cameras, LiDAR or time-of-flight cameras. The monitoring sensors are also referred to as first sensors. The one or more monitoring sensors generate monitoring sensor data. The monitoring sensor data is also referred to as first data. The monitoring sensor data can comprise one or more frames. Each frame can comprise an image, point cloud or other sensor data captured at a certain moment in time, and an associated timestamp indicating the time at which the image, point cloud or other sensor data was captured.
[0026] Proximity devices with associated unique identifiers are installed on each of the plurality of objects of interest. For example, on each vehicle of a fleet of vehicles.
[0027] One or more proximity sensors or in other words proximity detectors are installed at a location of interest. The proximity sensors or proximity detectors are also referred to as second sensors. The second sensors detect second data. The proximity detectors detect the presence of transmitters or proximity devices installed on, attached to or coupled to the objects of interest. Typically, the proximity sensors are installed at one end of a passenger boarding bridge that allows passengers to deplane or board a plane from an airport terminal. Other locations are also possible.
[0028] Each bridge is typically movable on the tarmac so that it can be positioned in a stationary position very close to the aircraft. Because each proximity sensor can be mounted at an end of the movable bridge, the specific location of each proximity sensor can vary depending on the location of each aircraft at the gate. Although the following description is with reference to identifying and monitoring objects of interest in the vicinity of an aircraft using marker images, this is exemplary and embodiments of the present invention can be applied to identifying and monitoring objects of interest in the vicinity of other modes of transportation or indeed in the vicinity of any point in space.
[0029] The one or more proximity sensors can be any suitable kind of sensor capable of detecting the presence of a proximity device within the range of the one or more proximity sensors. Illustrative examples of proximity sensors are WiFi TM sensors, Bluetooth sensors, inductive sensors, weight sensors, optical sensors, and radio frequency identifiers.
[0030] The coverage of the one or more proximity sensors is aligned with the field of view of the one or more monitoring sensors. In this way, the three-dimensional space corresponding to the coverage of the one or more proximity sensors is captured within or corresponds to the field of view of the one or more monitoring sensors.
[0031] For example, the range or coverage of the proximity sensor can be substantially circular. The field of view of the camera or one or more monitoring sensors is trained or directed to the range or area of the coverage of the proximity sensor.
[0032] The one or more proximity sensors generate proximity sensor data. In some embodiments, the proximity sensor data can comprise one or more entries. Each entry comprises a unique identifier, such as an IP address or other device identifier corresponding to a particular proximity device, and a timestamp indicating the time at which the unique identifier entered or exited the coverage of the proximity sensor.
[0033] When an object of interest enters the coverage of the one or more proximity sensors, a proximity device mounted on the object of interest is automatically detected by the one or more proximity sensors.
[0034] Automatic detection can be performed as follows. Each proximity sensor, such as a wireless network interface controller (WNIC), has a unique ID (e.g., MAC address, IP), and can connect to a radio-based wireless network using an antenna to communicate via microwave radiation. The WNIC can operate in infrastructure mode to directly interface with all other wireless nodes on the same channel. A wireless access point (WAP) provides a SSID and wireless security (e.g., WEP or WPA). The SSID is broadcast by the station in beacon packets to announce the presence of the network. The wireless network interface controller (WNIC) and wireless access point (WAP) must share the same key or other authentication parameters.
[0035] The system provides a dedicated hotspot (network sharing) at each operating station. For example, the 802.11n standard operates in 2.4 GHz and 5 GHz bands. Most newer routers are capable of utilizing both wireless bands, referred to as dual-band. This allows data communications to avoid the crowded 2.4 GHz band, which is also shared with Bluetooth devices. The 5 GHz band is also wider than the 2.4 GHz band, with more channels, allowing a larger number of devices to share the space.
[0036] WiFi or Bluetooth access points or similar GPS sensors can provide location data that can be used to train the machine learning model. Optionally, this data can be used in conjunction with the training model to provide higher performance.
[0037] The one or more proximity sensors capture a unique identifier of the proximate device and a timestamp corresponding to a time at which the proximate device was detected.
[0038] The one or more proximity sensors detect the departure of the proximate device when the object of interest leaves the coverage area of the one or more proximity sensors. The one or more proximity sensors capture a unique identifier of the proximate device and a timestamp corresponding to a time at which the departure of the proximate device was detected.
[0039] In other embodiments, the one or more proximity sensors capture proximity sensor data comprising one or more frames. Each frame of the proximity sensor data can include a list of proximate devices that are currently within the coverage area of the one or more proximity sensors, and a timestamp indicating a time at which the frame of the proximity sensor data was captured.
[0040] The system receives monitoring sensor data from the one or more monitoring sensors and receives proximity sensor data from the one or more proximity sensors.
[0041] In some embodiments, the system stores proximity sensor data including a timestamp and a unique identifier in a device IP connection list. Each entry in the list includes the unique identifier, a timestamp corresponding to a time when proximity to the device was detected or when proximity to the device was detected to have left, an object name, and one or more pre-processed videos associated with the object name.
[0042] The object name can be identified via a lookup table. The lookup table contains a list of unique identifiers and the names of the objects of interest that they are installed on. The system queries the lookup table to determine the object name associated with the unique identifier of any particular proximity device. The object name can be added as an additional field to the relevant entry in the device IP connection list.
[0043] In some embodiments, the system processes the data stored in the device IP connection list to calculate one or more time intervals during which any particular object of interest was within the coverage of one or more proximity sensors. The time intervals can correspond to the time between detecting that a proximity device installed on the object of interest entered the coverage of the one or more proximity sensors and detecting that the proximity device left the coverage of the one or more proximity sensors. Thus, the calculated time intervals represent the times when the object of interest was present within the coverage of the one or more proximity sensors. The one or more calculated time intervals can be stored as additional fields in the relevant entry in the device IP connection list.
[0044] The system processes the monitoring sensor data from the one or more monitoring sensors to automatically annotate the monitoring sensor data. The system selects each frame of the monitoring sensor data and reads the timestamp associated with that frame. The system then compares the timestamp to one or more timestamps or time intervals of the proximity sensor data. The system determines whether the monitoring sensor timestamp matches a timestamp of the proximity sensor data or falls within a calculated time interval of the proximity sensor data.
[0045] If the system determines that the timestamp of the selected frame falls within a time interval of the proximity sensor data, the system annotates the selected frame with the unique identifier associated with that time interval. Thus, the system annotates the selected frame with the unique identifier of any proximity device that was within the coverage of the one or more proximity sensors at the time the selected frame was captured.
[0046] If the timestamp of the selected frame falls within multiple time intervals of the proximity sensor data, the system annotates the selected frame with the unique identifier associated with each of the multiple time intervals. Thus, the system annotates the selected frame with the unique identifier of each proximity device that was within the coverage of the one or more proximity sensors at the time the selected frame was captured.
[0047] The process is repeated for each frame of the monitoring sensor data until all frames have been annotated with one or more unique identifiers representing the objects of interest present in each frame.
[0048] In some embodiments, the system can utilize one or more image processing techniques to assist in annotating the selected frames. The one or more image processing techniques can include segmentation algorithms, noise reduction techniques, and other object recognition techniques. The system can also apply an optical character recognition algorithm to the selected frames in order to identify distinctive textual markings on one or more objects of interest. If humans are not included in the objects of interest during the dataset generation phase, the system can utilize additional object recognition algorithms to annotate additional features of interest, such as humans.
[0049] When annotating the selected frames, the system can utilize annotations from previously processed frames as prior information.
[0050] According to some embodiments, the system can utilize positioning data from one or more positioning sensors installed on the objects of interest. For example, certain objects of interest, such as vehicles, can include pre-installed global positioning system (GPS) sensors. The system can use data from the one or more positioning sensors to assist in annotating the selected frames of the monitoring sensor data. For example, positioning data indicating the latitude and longitude corresponding to the field of view of one or more monitoring sensors at a certain time can be used as prior when annotating the selected frame corresponding to that time.
[0051] Training phase
[0052] The annotated monitoring sensor data can be used to train a machine learning model. In preferred embodiments, the machine learning model can be a neural network classifier, such as a convolutional neural network. The trained neural network classifier can be configured to take a single frame of monitoring sensor data as input and provide an annotated frame of the monitoring sensor data as output. For example, the neural network can take a single video frame from a CCTV camera (monitoring sensor) as input and output a labeled frame identifying one or more objects of interest in the frame.
[0053] In the case of a camera, the monitoring sensor data consists of frames or successive images from the camera sensor. In the case of a LiDAR, the data is a point cloud, and in RGB-D or time-of-flight cameras, the data is a combination of images and point clouds.
[0054] The machine learning model can be trained by a machine learning training module. Machine learning model training with various possible methods is well known to the skilled person. In one specific implementation, the machine learning model is a deep learning method and is a convolutional neural network based method. Implementation in Python can be done using TensorFlow or PyTorch modules.
[0055] Thus, it will be appreciated that in order to train the machine learning model, labelled data (e.g. images of vehicles in the field of view and their names) is required. In order to obtain the names (labels), a sensor (e.g. a wireless interface card) can be installed on the vehicle and an access point installing a data acquisition (e.g. a camera). Once the vehicle is in the coverage of the access point, its wireless interface card detects the access point SSID and connects to it. We use the access point’s timestamp and the list of connected devices to label the images (or point clouds).
[0056] In the specific example of WiFi communication, the proximity sensor data is WiFi connection data. However, it will be appreciated that Bluetooth or GPS data can be used for the same purpose in addition to or instead of WiFi connection data.
[0057] Thus, embodiments of the present application comprise a system capable of learning. During the training process, the machine learning training module iteratively adjusts one or more parameters of the machine learning model to reduce a “cost function”. The value of the cost function is representative of the performance of the machine learning model. For example, the cost function can depend on the accuracy of the machine learning model in predicting one or more unique identifiers associated with a given frame of monitored sensor data. One well-known algorithm for training neural network models is the backpropagation gradient descent algorithm. Gradient descent is an optimization algorithm used to find a local minimum of a function by taking steps proportional to the negative of the function gradient at the point. For each input, the backpropagation algorithm computes the gradient of the loss function with respect to the output and to the network weights. Instead of computing each weight directly and inefficiently individually, the backpropagation algorithm computes the gradient of the loss function with respect to each weight using the chain rule. This algorithm computes the gradient of one layer at a time and iterates backward from the last layer to avoid redundant computation of intermediate terms in the chain rule. Using this approach, backpropagation makes it feasible to use gradients for multi-layer networks such as multi-layer perceptrons (MLPs).
[0058] Application of trained machine learning model
[0059] The trained machine learning model can be applied to unlabeled monitoring sensor data to automatically label any objects of interest present in any given frame of the monitoring sensor data. Once the machine learning model has been trained, no proximity sensors need to be installed on the objects of interest, and the trained machine learning model can receive monitoring sensor data as input. In some embodiments, the trained machine learning model can not receive further proximity sensor data. The trained machine learning model can output labeled monitoring sensor data that identifies one or more objects of interest present in one or more frames of the monitoring sensor data.
[0060] The system can be configured to perform real-time analysis of monitoring sensor data. A real-time feed of monitoring sensor data can be fed to the system to be used as input to the trained machine learning model. In this way, the system can provide labeled monitoring sensor data in substantially real-time.
[0061] In some embodiments, the performance of the trained machine learning model can be continually improved during use. Machine learning model parameters can be fine-tuned using detected monitoring sensor data and / or other sensor data such as GPS sensor data.
[0062] The system can use the labeled data output from the machine learning model as part of an object tracking process. For example, the system can track the location of one or more labeled objects of interest on one or more subsequent frames. In further embodiments, the system can use the computed location of one or more labeled objects of interest over time to predict a likely future location of the one or more labeled objects of interest. The system can use the predicted location to determine that an impending collision between the one or more objects of interest is about to occur.
[0063] The system can provide tracking information for one or more labeled objects of interest as output. In some embodiments, the system can use the tracking information to perform automated guidance of the one or more objects of interest. If the system determines that a collision between the one or more objects of interest is about to occur, it can automatically take action to prevent the collision, for example by issuing a command to stop movement of the one or more objects of interest.
[0064] Particular embodiments in an aviation industry context
[0065] Specific embodiments of the application applied in an aviation industry environment will now be further described with reference to the drawings.
[0066] Figure 1A typical ramp is shown, which can be found at any airport or airfield. Also known as an airport ramp, taxiway, or apron, the ramp is the area of the airport where aircraft are parked between flights. While an aircraft is parked on the ramp, it can be loaded, unloaded, refueled, deplaned, and boarded. Passengers and crew can board / deboard via a boarding bridge 105. The boarding bridge 105 forms a bridge between the entrance door of the aircraft and the passenger terminal. Alternatively, passengers can board / deboard via other means, such as a mobile stairway, or in the case of smaller aircraft, no stairway at all.
[0067] Figure 1 Also shown in the center are several vehicles 101, which are typically found on the ramp of any airport or airfield. The vehicles 101 can include baggage carts, tanker trucks, passenger vehicles, and other vehicles. Any vehicle can be present on the ramp at any time, whether or not an aircraft is parked on the ramp. In addition, one or more emergency vehicles 102 can be called to the ramp to provide emergency response.
[0068] To facilitate and track the arrival and departure of aircraft, and to coordinate the numerous vehicles on the ramp at any given time, various monitoring sensors can be used.
[0069] Figure 1 A surveillance camera 104 is shown, which is configured to continuously monitor the ramp. The surveillance camera 104 is preferably positioned such that the entire ramp is within the field of view of the camera(s) 104. The camera 104 outputs video data, which includes a number of frames. Each frame of the video data includes a still image captured by the camera’s sensor, and a timestamp indicating the time at which the frame was captured. In addition to, or instead of, the surveillance camera 104, other types of monitoring sensors can be used. One alternative type of monitoring sensor is a Light Detection and Ranging (LiDAR) sensor. The LiDAR sensor can output LiDAR data, which includes a number of frames. Each frame of the LiDAR data includes a point cloud representing computed distance measurements distributed over the field of view, and a timestamp indicating the time at which the frame was captured. Another alternative type of monitoring sensor is a time-of-flight camera. A time-of-flight camera is a range imaging camera that can resolve distances between a sensor and an object. A laser-based time-of-flight camera captures distance measurements over the entire field of view with each pulse of a light (LASER) source. Those skilled in the art will appreciate that the present invention need not be limited to these particular types of monitoring sensors, and that other types of monitoring sensors can likewise be utilized.
[0070] During the dataset creation phase, one or more proximity sensors can be installed within the ramp. Figure 2A proximity sensor 201 is shown that can be located within the apron. In some embodiments, the proximity sensor can be a WiFi proximity sensor, such as a WiFi or wireless communication router. WiFi routers are well known to those skilled in the art and are inexpensive and widely available. However, those skilled in the art will appreciate that the proximity sensor need not be limited to a WiFi router and any suitable proximity sensor can be used.
[0071] In fact, the proximity sensor can be any receiver that detects the presence of a transmitter that is within the detection range of the receiver.
[0072] The WiFi router 201 can be installed at a central location of the apron. The WiFi router 201 can be positioned such that the coverage range of the WiFi router 201 extends to completely cover the apron. In some embodiments, the WiFi router 201 can be positioned such that the coverage range of the WiFi router 201 extends to cover at least a portion of the apron. Figure 2 In the embodiment shown, the WiFi router 201 is installed at the end of the boarding bridge 205. In this case, the router 201 is installed at the end of the boarding bridge that is closest to the aircraft. In some embodiments, multiple WiFi routers 201 can be installed within the apron such that the combined coverage range of the multiple WiFi routers 201 extends to completely cover the apron. It will be appreciated that in certain embodiments, the coverage range of one or more WiFi routers 201 need not extend to cover the entire apron, but can cover only a portion of the apron. The coverage range of the one or more WiFi routers 201 is aligned with the field of view of the one or more surveillance cameras. In some embodiments, the one or more surveillance cameras are positioned such that the apron fills the field of view of the one or more surveillance cameras and the one or more WiFi routers 201 are installed such that the coverage range of the one or more WiFi routers 201 extends to completely cover the apron. In other embodiments, the one or more surveillance cameras can be positioned such that the apron is within the field of view and the one or more WiFi routers 201 can be positioned such that the coverage range of the one or more WiFi routers 201 extends to cover at least the apron. The coverage range of the one or more WiFi routers 201 and the field of view of the one or more surveillance cameras 204 are aligned such that objects present in the field of view of the one or more surveillance cameras are within the coverage range of the one or more WiFi routers 201.
[0073] During the dataset creation phase, one or more proximity devices can be installed on one or more vehicles or equipment within the apron. Figure 2One such proximity device 203 is shown mounted on a vehicle on a tarmac. The one or more proximity devices 203 are detectable by the one or more proximity sensors 201. In embodiments where the one or more proximity sensors 201 are WiFi routers, the one or more proximity devices 203 can be WiFi enabled devices such as mobile phones, tablets, smart watches and other devices. In cases where the proximity sensors 201 are other types of sensors, the proximity devices 203 can be devices that are detectable by the proximity sensors 201. The one or more proximity devices 203 are each associated with a unique identification number for identifying each proximity device 203. In embodiments of the present application where the one or more proximity devices 203 are WiFi enabled devices, the unique identification number can be the IP address of the device. It will be appreciated that the proximity devices can be any suitable device that is detectable by the one or more proximity sensors 201. The advantage of using WiFi enabled devices such as mobile phones is that they are relatively inexpensive and very readily available.
[0074] The unique identification numbers and the vehicles or devices on which the associated proximity devices 203 are mounted can be recorded in a lookup table. The lookup table can include a list of unique identification numbers corresponding to each of the one or more proximity sensors 201 and an indication of the vehicle or device on which the proximity device 203 is mounted. In embodiments where the one or more proximity sensors are WiFi routers and the one or more proximity devices are WiFi enabled devices, the lookup table includes a list of IP addresses of the one or more WiFi enabled devices and an indication of the vehicle or device on which each WiFi enabled device is mounted.
[0075] When a vehicle or device equipped with a WiFi-enabled device 203 is within the coverage area of one or more WiFi routers 201, the one or more WiFi routers 201 detect the WiFi-enabled device 203. In some embodiments, when the vehicle or device and its associated WiFi device 203 first enter the coverage area of one or more WiFi routers 201, the one or more WiFi routers 201 detect the WiFi device 203 and record the IP address of the WiFi device 203 and a timestamp corresponding to the time when the WiFi device 203 enters the coverage area of one or more WiFi routers 201. Subsequently, while the WiFi device 203 remains within the coverage area of one or more WiFi routers 201, the one or more WiFi routers 201 continue to detect the WiFi device 203. When the vehicle or device and its associated WiFi device 203 leave the coverage area of one or more WiFi routers 201, the one or more WiFi routers 201 detect the departure of the WiFi device 203 and record the IP address and a timestamp corresponding to the time when the departure of the WiFi device 203 is detected. One or more WiFi routers 201 can then output proximity sensor data, which includes a series of timestamps indicating when any of the one or more WiFi-enabled devices 203 enters or leaves the coverage area of the one or more WiFi routers 201, and the IP address associated with each detected WiFi-enabled device 203.
[0076] like Figure 3 As shown, proximity sensor data (including a series of timestamps representing the time when any of one or more WiFi-enabled devices 203 enters or leaves the coverage area of one or more WiFi routers 201, and the IP address associated with each detected WiFi-enabled device 203) can be stored in a device IP connection list 301. Each entry in the list includes the IP address of the WiFi-enabled device 203, a timestamp corresponding to the time when the WiFi-enabled device 203 was detected or when the WiFi-enabled device was detected leaving, and a vehicle name or device name. Entries in the device IP connection list 301 may also include one or more pre-processed videos associated with the vehicle name or device name. In other embodiments, the vehicle name or device name associated with each IP address can be stored and accessed via a lookup table.
[0077] In other embodiments, the one or more WiFi routers 201 can continuously detect any WiFi enabled devices 203 present within the coverage area of the one or more WiFi routers. The one or more WiFi routers 201 can output proximity sensor data comprising one or more frames. Each frame can comprise a list of IP addresses corresponding to any WiFi enabled devices detected within the coverage area of the one or more WiFi routers 201, and a timestamp corresponding to the time at which the devices were detected.
[0078] Figure 4 A schematic of a system in accordance with an embodiment of the present application is shown. The system comprises an input module 401, a processing module 402, a machine learning model 403, a machine learning training module 404, and an output module 405. The input module 401 is configured to receive monitoring sensor data and proximity sensor data. In accordance with the preferred embodiment described above, the input module 401 is configured to receive monitoring sensor data from one or more CCTV cameras and proximity sensor data from one or more WiFi routers 201. The input module 401 passes the monitoring sensor data and the proximity sensor data to the processing module 402.
[0079] In some embodiments, the processing module processes the received proximity sensor data to calculate one or more time intervals during which each vehicle or device was present within the coverage area of the one or more WiFi routers 201. In embodiments in which the proximity sensor data comprises one or more timestamps representing the time at which any of the one or more WiFi enabled devices 203 entered or exited the coverage area of the one or more WiFi routers 201, the processing module 402 can process the proximity sensor data to generate one or more frames of proximity sensor data, each frame having an associated timestamp and a list of unique identifiers of any WiFi enabled devices 203 present within the range of the one or more WiFi routers at the time corresponding to the timestamp. The time interval between consecutively generated frames of proximity sensor data is preferably equal to the time interval between consecutive frames of monitoring sensor data. In some embodiments, the processing module processes the proximity sensor data to generate one or more frames of proximity sensor data having timestamps matching the timestamps associated with one or more frames of received monitoring sensor data.
[0080] The processing module 402 processes the monitoring sensor data to create a training dataset. Figure 5 A method of automatically labelling monitoring sensor data is shown, the method being performed by the processing module 402 to create a labelled training dataset.
[0081] First, at step 501, the processing module 402 selects a first frame of the received monitoring sensor data and reads the timestamp associated with the selected frame. Second, at 502, the processing module 402 compares the timestamp associated with the selected frame to one or more timestamps or time intervals of the proximity sensor data. At 503, the processing module 402 determines whether the selected timestamp of the monitoring sensor data matches one or more timestamps of the proximity sensor data or falls within one or more time intervals of the proximity sensor data. At 504, if the determination is positive, the processing module 402 reads one or more IP addresses associated with the one or more timestamps or time intervals of the proximity sensor data. The processing module 402 then determines the vehicle or device name associated with each of the one or more IP addresses, for example, via a lookup table or a list of device IP connections. At 505, the processing module 402 labels the selected frame of the monitoring sensor data with the one or more device names.
[0082] The processing module 402 then repeats steps 501-505 for each frame of the monitoring sensor data. In this way, each frame of the monitoring sensor data is labeled with the name of any vehicle or object of interest that appeared in the field of view of one or more of the surveillance cameras 104 during that frame.
[0083] The processing module 402 can utilize one or more other models, such as computer vision, image processing, and machine learning methods, to assist in labeling the frames of the monitoring sensor data. For example, the processing module can apply segmentation algorithms, noise reduction algorithms, edge detection filters, and other object recognition techniques to the selected frames of the monitoring sensor data to assist in labeling the monitoring sensor data. In some embodiments, the processing module 402 can use previously labeled frames as prior information when labeling a currently selected frame. In some embodiments, other models can be utilized to generate embedded feature vectors, where each embedded feature vector is associated with an object of interest labeled in one or more frames of the monitoring sensor data. The embedded feature vectors can also include characteristic information related to the object of interest associated therewith. For example, one or more other models can be used to extract the color, model, make, and license ID of a vehicle appearing in a frame of the monitoring sensor data.
[0084] The labeling of the monitoring sensor data can include one or more identification images associated with each identified object of interest. For example, the labeling of a vehicle image identified within a frame of the monitoring sensor data can include several images showing the vehicle from different angles and / or under different lighting conditions. The identification images can be stored in the local memory of the system and used to improve and / or enhance the identification of objects of interest in unlabeled monitoring sensor data.
[0085] In some embodiments, processing module 402 is configured to utilize optical character recognition algorithms to detect distinctive text markers, such as tail fin numbers or vehicle license plates, to identify vehicles or objects of interest. In some embodiments, input module 401 is also configured to receive positioning data from one or more positioning sensors mounted on the vehicle or equipment. Specifically, positioning sensors may include GPS sensors or other suitable positioning sensors. Input module 401 passes the positioning data to processing module 402 to aid in monitoring the labeling of sensor data. These methods can be applied during the initial dataset creation phase or during the application phase as a method to continuously improve system performance by providing additionally labeled monitoring sensor data.
[0086] The machine learning training module 404 receives labeled monitoring sensor data and uses it to train the machine learning model 403. In a preferred embodiment, the machine learning model 403 is a neural network classifier model, and more preferably a deep neural network classifier. The machine learning model 403 takes frames of monitoring sensor data as input and outputs labeled frames that indicate the machine learning model's prediction of the presence of one or more vehicles or devices within the frame. In a preferred embodiment, the machine learning model 403 outputs labeled frames that indicate the predicted location of one or more vehicles or devices within the frame. In other embodiments, the labeled frames may include one or more labels indicating that the model 403 predicts the presence of one or more vehicles or devices at a certain location within the frame, wherein the labels may be stored as metadata along with the frames.
[0087] During the training process familiar to technicians, the machine learning training module 404 adjusts the weights and biases of the neural network model to reduce the cost function value. The cost function value is calculated based on the accuracy of predicting vehicles or devices present within a given frame on a labeled monitoring sensor dataset. The weights and biases are updated in a manner that increases prediction accuracy. The machine learning training module 404 uses the backpropagation gradient descent algorithm to calculate the necessary changes to the weights and biases and update the network.
[0088] Once the machine learning model 403 has been trained on labeled surveillance sensor data, the trained machine learning model 403 is applied to unlabeled surveillance sensor data to automatically label any vehicles or devices present in the received frames of surveillance sensor data. The input module 401 receives unlabeled surveillance sensor data from one or more surveillance cameras or LiDAR sensors 104. The unlabeled surveillance sensor data is input into the machine learning model 403. The machine learning model 403 outputs labeled surveillance sensor data. Figure 6The annotated monitoring sensor data output by the trained machine learning model 403 is shown. One or more monitoring cameras or other monitoring sensors 601 capture monitoring sensor data which is then labeled by the machine learning model to detect vehicles present on the airport tarmac.
[0089] Figure 7 Some examples of system applications for improving aircraft tracking are shown.
[0090] To detect the aircraft 702, its type, and its unique ID, embodiments of the invention can use three data sources (or combination of data sources) and combine them to accurately identify the aircraft:
[0091] 1 - Aircraft schedule data
[0092] 2 - OCR reading of the tail number 701
[0093] 3 - Aircraft type detection machine learning
[0094] Knowing the type, location, and schedule of the aircraft, a machine learning model can optimize services. Moreover, when the radar of the landing aircraft is off (e.g. at night), the exact location of each aircraft can be identified (e.g. in a digital twin). AI-based schedule can be used to optimize refueling, thawing, baggage loading, or other services, which can globally optimize all processes. Optimal operations can be used to train a machine learning model to learn optimal decisions and present these decisions to operators. Once the accuracy and robustness of the model are tested, the machine learning model can provide optimal decisions to ground crew and pilots, while expert operators can only need to monitor and validate the optimized operations for double check, ensuring no conflicts or anomalies.
[0095] According to illustrative embodiments, the system can utilize multiple data sources to track aircraft in an airport environment. The system receives schedule data related to an airport via input module 401. The schedule data includes expected arrival and departure times for one or more aircraft, assigned gate numbers, and aircraft information such as tail number and aircraft type. The system identifies aircraft within the field of view of one or more monitoring sensors. The system can use optical character recognition (OCR) to read the tail number of an aircraft. In other embodiments, the system can use a machine learning model 403 and / or other location data from the aircraft, such as GPS data and radio frequency identification. The machine learning model can be trained to identify different aircraft types using labeled training data as previously described. It should be appreciated that the system can utilize any combination of one or more of the inputs described above, and can utilize other data sources than those described in the illustrative embodiments above. In identifying aircraft within the field of view of one or more monitoring sensors, the system uses the identified aircraft type, location, and schedule information to optimize aircraft services. Aircraft services that can be optimized using the improved aircraft tracking of the present invention include refueling, thawing, baggage loading, provisioning, and aircraft maintenance. In some embodiments, the improved aircraft tracking is provided to a human operator who optimizes one or more aircraft services based on the tracking information. In other embodiments, the system can use the optimized aircraft services to train a machine learning model. The machine learning model can thereby learn how to optimize one or more aircraft services based on aircraft tracking information. In some embodiments, an expert operator monitors the decision process of the trained machine learning model to ensure there are no errors or anomalies. In other embodiments, the system automatically provides optimized aircraft services based on the improved aircraft tracking without the need for human supervision or input.
[0096] Figure 8An automated collision detection method that can be implemented using the system of the present application is shown. According to the illustrative embodiment shown, the system identifies two vehicles-aircraft 801 and an emergency vehicle 802 within the field of view of one or more monitoring sensors. The machine learning model 403 outputs labeled monitoring sensor data that indicates the location of the vehicles within the field of view of the one or more monitoring sensors. In some embodiments, the system utilizes additional location information from location sensors mounted on the vehicles, such as GPS sensors. The processing module 402 stores the location of each vehicle over a plurality of time frames and uses the time-varying location of each vehicle to calculate a predicted trajectory of each vehicle over time. The processing module 402 performs a check as to whether the predicted trajectories of the two vehicles intersect at a given future point in time, indicating a possible collision between the two vehicles. If a possible collision 803 is detected, the processing module outputs a collision detection alert to the output module 405. The output module 405 outputs the collision detection alert via any suitable communication protocol. For example, in some embodiments, the output module outputs the collision detection alert to both vehicles via radio frequency communication, alerting the drivers of the detection of a possible collision 803. In other embodiments, the collision detection alert can be output to an autonomous vehicle system as an instruction to take evasive action to prevent a collision.
[0097] Figure 9 and 10 A method of providing autonomous service vehicles using the present application is shown. Figure 9 An autonomous service vehicle 901 is shown. In certain embodiments, a number of collision detection sensors, such as ultrasonic ranging sensors, can be mounted on the autonomous service vehicle 901. In addition, the autonomous service vehicle 901 can be mounted with monitoring sensors, such as cameras and LiDAR systems.
[0098] At step 1001, one or more monitoring sensors 902 receive monitoring sensor data of an aircraft 903 located on a tarmac. For example, an aircraft has arrived and parked at the end of a boarding bridge 904. The input module 401 receives the monitoring sensor data and passes it to the processing module 402. At step 1002, the processing module 402 applies an OCR algorithm to the received monitoring sensor data to identify the tail number of the aircraft 903. In other embodiments, the input module 401 can receive the tail number or flight identification number of the aircraft located on the tarmac from an external source. At step 1003, once the tail number of the aircraft is identified, the communication module can query an external source to send the maintenance schedule of the aircraft. In other embodiments, the input module 401 can receive the maintenance schedule from an external source.
[0099] At step 1004, the processing module reads a first maintenance item from the maintenance schedule. For example, the maintenance item can be to replenish food or perishable items for the next scheduled flight of the aircraft. Based on this, the processing module determines that the self-driving service vehicle 901 should navigate to the aircraft to deliver the required supplies. At step 1005, the processing module 402 sends instructions to the self-driving vehicle 901 via the communication module 405 to start navigating to the aircraft.
[0100] At 1006, the self-driving vehicle 901 starts navigating to the aircraft. One or more monitoring sensors capture the self-driving vehicle 901 within the field of view and output monitoring sensor data to the input module 401 substantially in real-time. The machine learning model 403 analyses the monitoring sensor data and outputs labelled monitoring sensor data indicative of the location of the self-driving vehicle 901. Based on the detected location of the self-driving vehicle 901, the processing module 402 determines navigation instructions 905 that should be sent to the vehicle 901 to assist in navigating to the aircraft. The system repeats this process for all planned maintenance items in the maintenance schedule.
[0101] It will be appreciated that the present application need not be limited to application within the aviation transport industry, but can be applied to other industries, such as shipping and other modes of public transport, plant equipment management, parcel tracking and traffic management.
[0102] Application to marine vessel tracking
[0103] One such alternative application of the present application is in the shipping industry, in particular the automated tracking of ships within a port or harbour.
[0104] According to the present application as described above, a labelled training dataset for training a machine learning model can be generated. In particular, monitoring sensor data of one or more objects of interest, including marine vessels such as small boats and ships, is collected using one or more monitoring sensors installed to monitor a region of interest in a marine vessel tracking environment. For example, the region of interest in a marine vessel tracking environment can be a harbour or port, but other regions of interest can also be considered. The monitoring sensors can include, for example, CCTV cameras, cameras attached to one or more unmanned aerial vehicles, including alternative types of video cameras such as LiDAR sensors, thermal imaging cameras and the like.
[0105] At the same time, proximity sensor data of the one or more objects of interest can be collected via one or more proximity sensors installed in the region of interest and one or more proximity devices installed on the one or more objects of interest.
[0106] The proximity sensor data and the monitored sensor data are used to automatically label the monitored sensor data, with each frame being labeled to show any images of objects of interest present in that frame.
[0107] As noted above, the system can utilize position data from one or more positioning sensors. For example, a marine vessel can utilize a GPS system, a differential GPS system, a RADAR system, a global navigation satellite system, and / or an acoustic transponder system. This is not an exhaustive list of positioning systems for marine vessels, and other suitable positioning systems can likewise be used. The system can use the position data to augment the labeled training data for use during the training phase.
[0108] The system can also utilize additional image processing techniques to assist in labeling the monitored sensor data. These techniques can include object tracking and extraction, pattern matching, human and face detection, object recognition, gesture recognition, and the like. As noted above, the system can utilize an object recognition algorithm to identify and label images of humans appearing in selected frames of the monitored sensor data.
[0109] In preferred embodiments, one or more other models, such as computer vision, data mining, and machine learning methods, can be utilized to assist in labeling frames of the monitored sensor data. In some embodiments, one or more other models can be used to generate an embedded feature vector associated with one or more identified objects of interest present in the monitored sensor data. Such models are known to the skilled person. The embedded feature vector can include further characteristic information associated with the identified object of interest. For example, the one or more other models can be used to extract one or more of the following: a color, a model, a brand, an ID number, a registered owner, a designated location, and route information of a marine vessel appearing in a frame of the monitored sensor data. The feature vector can include position data relating to the object of interest captured from one or more position sensors as described above.
[0110] The embedded feature vector can be embedded in the labeled monitored sensor data prior to training the first machine learning model. Alternatively or additionally, a second machine learning model can be used online to generate an embedded feature vector associated with an object of interest labeled by the first machine learning model in real-time or substantially real-time monitored sensor data.
[0111] As described above, in accordance with the present application, annotated monitoring sensor data is used to train a machine learning model. The machine learning model can utilize a deep learning approach, and can be a neural network classifier implemented as a convolutional neural network. The machine learning model can be trained on annotated data by a machine learning training module. The machine learning model can be configured to take as input a single frame of unannotated monitoring sensor data, and output an annotated frame of the monitoring sensor data annotated to reflect the location of any object of interest present in the frame of monitoring sensor data. As described above, in some embodiments, the annotated frame output by the first machine learning model can also include embedded feature vectors provided by one or more other models.
[0112] Once the machine learning model has been trained, the system including the trained machine learning model is applied to unannotated monitoring sensor data to automatically annotate any marine vessels, equipment, or people appearing in received frames of monitoring sensor data.
[0113] One example application of the system in a marine vessel tracking environment is to alert a registered owner if their vessel leaves its assigned location. In this example, the system receives monitoring sensor data from one or more CCTV cameras positioned to monitor a dock or port. The system processes the monitoring sensor data and identifies one or more vessels present in the monitoring sensor data. For each of the one or more identified vessels, the system can perform a further identification process to generate an embedded feature vector associated with each identified vessel. The embedded feature vector can include one or more of a color, a brand, a model, a registered owner, an assigned mooring location, route information, and other identifying characteristics. For example, the system can use optical character recognition (OCR) to identify a name or ID of the vessel and include that information in the embedded feature vector. The system can receive identifying information, including expected arrival and departure times, assigned mooring locations, registered owners, etc., from external sources, such as one or more servers of a dock booking system, for example.
[0114] The system can determine that an identified vessel has left its assigned mooring location by tracking the location of the vessel in the monitoring sensor data, or by monitoring location data received from one or more location sensors of the vessel. The system can determine that the departure from the assigned mooring location is abnormal based on a comparison between a detected vessel departure time and an expected departure time based on route information associated with the vessel. In response to determining that the departure is an abnormal departure, the system can alert the registered owner of the abnormal departure by generating an electronic notification, an SMS message, an alert notification, etc.
[0115] Other applications of the system in a vessel tracking environment include anomaly detection, fire detection, and theft detection.
[0116] Generally, when an object of interest enters the coverage range of one or more proximity sensors, a proximity device mounted on the object of interest is detected by the one or more proximity sensors. An entry timestamp is generated corresponding to the time the object of interest entered the coverage range or range of the one or more proximity sensors.
[0117] Further, when the object of interest leaves the coverage range of the one or more proximity sensors, the one or more proximity sensors detect the departure of the proximity device. The one or more proximity sensors capture the unique identifier of the proximity device and a timestamp corresponding to the time the departure of the proximity device was detected. A departure timestamp is generated corresponding to the time the object of interest left the coverage range or range of the one or more proximity sensors.
[0118] The entry and departure timestamps define a time period during which the object of interest within the field of view of the monitoring sensor is tagged. In some embodiments, the tagging operation is performed for about 10 minutes to 1 hour. This is a typical time period during which the object of interest is located within the field of view of the monitoring sensor and, thus, within the range of the proximity sensor.
[0119] Of course, there can typically be multiple objects of interest within the coverage range or range of the one or more proximity sensors. Typically, each object of interest has an associated entry timestamp and an associated departure timestamp. Typically, the entry or departure timestamps associated with each object of interest are different because the objects of interest typically enter or leave the range of the one or more proximity sensors at different times. However, it should be appreciated that in some instances, multiple objects of interest can enter the coverage range or range of the one or more proximity sensors substantially simultaneously. Thus, different entry and departure timestamps can be generated for each object of interest.
[0120] In one particular application for tagging images to identify and monitor objects of interest in the vicinity of large objects, the proximity sensor 201 has a range of up to about 100 meters. This means that a particular object of interest can be identified and tagged anywhere in the vicinity of a large aircraft or other object of interest. In some applications, the range of the proximity sensor can be greater. For example, embodiments of the present invention find application in tagging images to identify and monitor objects of interest in the vicinity of vehicles, such as highway vehicles, ships, cruise ships, and aircraft carriers, among others. In this and other applications, the proximity sensor can have a range of up to 300 meters to 400 meters. In some embodiments, long-range Wi-Fi can be used.
[0121] In one specific example, the proximity sensor can include a software or hardware module configured to adjust the range of the proximity sensor in response to a range adjustment command.
[0122] Embodiments of the invention can detect the size of an object in the proximity of the proximity sensor, such as a vehicle.
[0123] In one specific example, this can be performed by using known image processing techniques to detect and read a unique identifier, such as a tail number (e.g. JA8089) or a ship registration number. A lookup table of unique identifiers and corresponding aircraft types or / and sizes can be used. Embodiments of the invention can determine the unique identifier and use the lookup table to determine the size of the object, such as a vehicle in the proximity of the proximity sensor.
[0124] Typically, a hardware or software module configured to adjust the range of the proximity sensor receives a command to adjust its range depending on the size of an object, such as an aircraft, located in the proximity of the proximity sensor.
[0125] Instead of using a lookup table, LIDAR or features based on the aircraft can be used to determine the size of an aircraft or stationary object located in the proximity of the proximity sensor.
[0126] Advantageously, some embodiments determine the time period between an entry timestamp and an exit timestamp for each object entering the coverage range of the proximity sensor.
[0127] This allows certain objects to be in the coverage range of the proximity sensor for a short period of time to not be considered, or to not be processed by the detection algorithm or the algorithm that marks each image to identify and monitor objects of interest.
[0128] For example, a threshold can be defined where if an object is in the range of the proximity sensor for, for example, 5 seconds, then the object is not processed. Other time periods can advantageously be used, such as 10 minutes. This allows the system to not process or ignore different objects entering the range of the proximity sensor for a short period of time.
[0129] Reference is made to the drawings Figure 2 The predetermined area 202 shown, which area (and range of the proximity sensor 201) is typically substantially circular.
[0130] However, certain sub-sectors that are fully contained within the predetermined area can be excluded from processing by the detection algorithm or the algorithm that marks each image to identify and monitor objects of interest.
[0131] This can be achieved by placing two or more additional proximity sensors (in addition to proximity sensor 201) or Wi-Fi routers at different locations within area 202. One or more sub-sectors of interest within the area can be defined using known triangulation techniques. Sub-sectors of any shape can be defined. For example, a square sub-sector can be defined using four longitude and latitude coordinates. Alternatively, a circular sub-sector can be defined using a single longitude and a single latitude coordinate, along with the radius or diameter of the circular sub-sector.
[0132] In some embodiments, the system or method can be configured to uniquely identify objects of interest and to distinguish multiple objects of the same type within the range of one or more proximity sensors. For example, suppose two baggage carriers serve an aircraft. Each baggage carrier can be assigned a different tag. For example, the first baggage carrier could be labeled "Baggage_Truck_1" and the second baggage carrier could be labeled "Baggage_Truck_2". This allows for analysis of the performance of different baggage carriers. It also allows for the identification of specific faults in each object of interest. For example, a problem can be identified based on the average speed of objects or carriers serving the aircraft or other points in space. Each of these carriers or objects can be uniquely identified, tracked, tagged, etc.
[0133] In alternative applications of embodiments of the invention, objects are identified, monitored, tracked, and marked, typically vessels arriving at or departing from a port. Vessels may have unique characteristics, such as the specific size and color of their sails.
[0134] Embodiments of the present invention typically process large amounts of image data or frames of video feeds within a predetermined time period. As mentioned above, this could be on the order of 10 to 20 minutes.
[0135] Note that in this invention, from Figure 2 The same video or image feed for each camera 204 shown is used for both the training phase and the labeling phase according to an embodiment of the invention. Since the training environment is exactly the same as the deployment environment, this provides better labeled images for identifying and monitoring objects of interest.
[0136] An embodiment of the present invention determines whether an object is within a predetermined region 202 aligned with a first sensor.
[0137] Therefore, it should be understood that embodiments of the present invention typically include remote proximity sensors with a range greater than 0.5 m. Due to the use of wireless network protocols such as ASS 802.11, the range of the proximity sensors is greater than that of optical scanner systems. Embodiments of the present invention can operate in a frequency range of 2.4 GHz to 6 GHz.
[0138] Typically, the proximity sensor range and the CCTV field of view are aligned or substantially aligned.
[0139] In some embodiments of the present application, an initial detection of objects of interest is performed prior to performing automatic labelling. For example, embodiments of the present application can detect objects in a frame and then use the above described method to assign labels based on proximity data. It will be appreciated that some objects of interest, such as baggage carriers, can already be fitted with proximity devices as described above. However, objects can be retrofitted with proximity devices by installing proximity devices on or within one or more objects of interest.
[0140] From the foregoing it will be appreciated that embodiments of the present application advantageously:
[0141] • Identify relatively large objects (e.g. airport carriers, indoor robots) for training machine learning models. Embodiments of the present application use long-range sensors (on the order of tens of meters) for data collection. In contrast, bar code laser scanners have much smaller range (on the order of tens of centimeters)
[0142] • Configured to run continuously day and night, without interruption, to collect data in all types of outdoor environmental conditions
[0143] • Accurately label images for training machine learning models. This can be achieved by detecting when an object appears in the camera view and when the object leaves the view.
[0144] • Precisely determine the time when the proximity sensor enters the reader coverage. This can be achieved by aligning and calibrating the camera view (or the region of interest of the camera) with the sensor detectability coverage area and the timestamp of when the object leaves the view.
[0145] Furthermore, embodiments of the present application have the following advantages:
[0146] • Use of long-range detection that enables a variety of applications (such as the goods and parcel shipping industry, sea and rail transportation, carrier detection, etc.) to label images for training machine learning models in real-world environments.
[0147] • Does not add any manual scanning process. This makes embodiments of the present application useful for industries and operations that cannot tolerate interruptions
[0148] • Generates two timestamps (start and end timestamps of the object in the view).
[0149] Manual scanning introduces noise and the machine learning model will disadvantageously learn these features as the target object, as the human will be observed in all labelled images, which confuses the machine learning model. Furthermore, manual scanning is technically impossible to detect large objects as the range of manual scanning is very short.
[0150] From the foregoing, it will be appreciated that the system can include a computer processor running one or more server processes for communicating with client devices. The server processes include computer readable program instructions for performing the operations of the present application. The computer readable program instructions can be source or object code in assembly or any other programming language, including a procedural programming language, an object oriented programming language, a scripting language, an assembly language, machine code instructions, instruction set architecture (ISA) instructions, and state setting data, written in a suitable programming language, or any combination of the above.
[0151] The wired or wireless communication network described above can be a public, private, wired, or wireless network. The communication network can include one or more of a local area network (LAN), a wide area network (WAN), the Internet, a mobile telephone communication system, or a satellite communication system. The communication network can include any suitable infrastructure, including copper cables, optical cables or fibers, routers, firewalls, switches, gateway computers, and edge servers.
[0152] The system described above can include a graphical user interface. Embodiments of the present application can include a screen graphical user interface. For example, the user interface can be provided in the form of a widget embedded in a website, as an application of a device, or on a dedicated login web page. Computer readable program instructions for implementing the graphical user interface can be downloaded from a computer readable storage medium to a client device via a network, such as the Internet, a local area network (LAN), a wide area network (WAN), and / or a wireless network. The instructions can be stored in a computer readable storage medium within the client device.
[0153] As will be appreciated by those skilled in the art, the application described herein can be embodied in whole or in part as a method, system, or computer program product including computer readable instructions. Thus, the present application can take the form of an entirely hardware embodiment, or an embodiment combining software, hardware, and any other suitable method or apparatus.
[0154] The computer readable program instructions can be stored on a non-transitory, tangible computer readable medium. The computer readable storage medium can include one or more of an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, a portable computer disk, a hard disk, a random access memory (RAM), a read only memory (ROM), an erasable programmable read only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk.
[0155] Exemplary embodiments of the present application can be implemented as a circuit board that can include a CPU, a bus, a RAM, a flash memory, one or more ports for operation of I / O devices such as a printer, a display, a keyboard, a sensor, and a camera for connection, a ROM, a communication subsystem such as a modem, and a communication medium.
[0156] Furthermore, the above detailed description of embodiments of the application is not intended to be exhaustive or to limit the application to the precise form disclosed. For example, while processes or blocks are presented in a given order, alternative embodiments can perform routines having steps, or employ systems having blocks, in a different order, or employ a different arrangement of steps or blocks, and some processes or blocks can be deleted, moved, added, subdivided, combined, and / or modified. Each of these processes or blocks can be implemented in a variety of different ways. Also, while processes or blocks are at times shown as being performed in series, these processes or blocks can instead be performed or implemented in parallel, or can be performed at different times.
[0157] The teachings of the application provided herein can be applied to other systems, not necessarily the system described above. The elements and acts of the various embodiments described above can be combined to provide further embodiments.
[0158] While some embodiments of the application have been described, these embodiments have been presented by way of example only, and are not intended to limit the scope of the disclosure. Indeed, the novel methods and systems described herein can be embodied in a variety of other forms; furthermore, various omissions, substitutions and changes in the form of the methods and systems described herein can be made without departing from the spirit of the disclosure.
[0159] Embodiments of the application can be described by the following numbered clauses.
[0160] 1. A system for generating a labeled dataset, the system comprising:
[0161] a processing component configured to:
[0162] receive monitoring sensor data comprising one or more frames, wherein the image data comprises data defining an object of interest;
[0163] receive proximity sensor data;
[0164] analyze the one or more frames of monitoring sensor data to identify one or more objects of interest present in the one or more frames of monitoring sensor data based on the proximity sensor data;
[0165] label the one or more frames of monitoring sensor data based on the analysis to generate a labeled dataset;
[0166] output the labeled dataset.
[0167] 2. The system of clause 1, wherein monitoring each of the one or more frames of sensor data comprises monitoring an instance of sensor data and a timestamp.
[0168] 3. The system of clause 2, wherein the instance of sensor data comprises a single frame of video data and the timestamp corresponds to a time at which the single frame of video data was detected.
[0169] 4. The system of clause 2, wherein the instance of sensor data comprises a single instance of point cloud depth data and the timestamp corresponds to a time at which the single instance of point cloud depth data was detected.
[0170] 5. The system of clause 1, wherein the proximity sensor data comprises one or more timestamps and one or more unique object identifiers.
[0171] 6. The system of clause 5, wherein the one or more unique identifiers are Internet Protocol (IP) addresses.
[0172] 5. The system of clause 1, wherein the processing component is further configured to receive positioning data indicative of a location of the one or more objects of interest and analyze the one or more frames of monitored sensor data to identify the one or more objects of interest present in the one or more frames of monitored sensor data based on the proximity sensor data and the positioning data.
[0173] 6. The system of clause 5, wherein the positioning data is Global Positioning System (GPS) data.
[0174] 7. The system of clause 1, further comprising one or more monitoring sensors and one or more proximity sensors, wherein each monitoring sensor and each proximity sensor is configured such that a field of view or coverage of the one or more monitoring sensors and the one or more proximity sensors are aligned.
[0175] 8. The system of clause 7, wherein the field of view or coverage of the one or more monitoring sensors and the one or more proximity sensors are uniform.
[0176] 9. The system of clause 7, wherein the proximity sensor data comprises one or more timestamps and one or more unique identifiers; and each of the one or more timestamps indicates a time at which the one or more objects of interest entered or exited a coverage of the one or more proximity sensors.
[0177] 9. A system for training a machine learning model, the system comprising:
[0178] a processing component configured to:
[0179] receiving data tagged using the system of any preceding clause;
[0180] training a machine learning model using the received data.
[0181] 10. A detection system for detecting one or more objects of interest, the system comprising:
[0182] an input module configured to receive monitoring sensor data comprising one or more frames;
[0183] a machine learning model trained using the system of clause 9;
[0184] a processing component configured to:
[0185] determine, using the machine learning model, whether any object of interest is present in each frame of the one or more frames of monitoring sensor data.
[0186] 11. The system of clause 10, wherein the system is further configured to track one or more vehicles within a field of view of the monitoring sensor data based on the determination.
[0187] 12. The system of clause 11, wherein the system is further configured to determine a likelihood of a collision between the one or more tracked vehicles.
Claims
1. A system for generating labeled datasets, the system comprising: The processing unit (402) is configured as follows: First data is received from one or more first sensors (204), wherein the first data includes one or more frames, each frame including an image or point cloud, and also includes an associated timestamp indicating the time when the image or point cloud was captured; and wherein the first data includes data defining an object of interest (203) within a predetermined region (202); Second data is received from one or more second sensors (201), wherein the second data is associated with the object of interest within the predetermined area, and wherein the second data includes one or more unique object identifiers and a plurality of timestamps, at least one of the plurality of timestamps representing the time when the object of interest enters the coverage area of the one or more second sensors, and at least one of the plurality of timestamps representing the time when the object of interest leaves the coverage area of the one or more second sensors; Analyze the one or more frames of the first data to identify the objects of interest present in the first data based on the second data; Based on the analysis of the first data, one or more frames are labeled to generate the labeled dataset; Output the dataset with the specified tags.
2. The system according to claim 1, wherein, At least one of the one or more first sensors is a camera configured to capture an image of the area, and at least one of the one or more second sensors is configured to detect radio frequency signals.
3. The system according to claim 1 or 2, wherein, The first data is associated with two-dimensional or three-dimensional image data of the predetermined area, and the second data is associated with radio frequency data.
4. The system according to claim 1, wherein, The first data includes a single frame of video data, the timestamp corresponding to the time when the single frame of the video data was detected, and / or, wherein the first data includes a single instance of point cloud depth data, the timestamp corresponding to the time when the single instance of the point cloud depth data was detected.
5. The system according to claim 1, wherein, One or more unique identifiers are Internet Protocol (IP) addresses.
6. The system according to claim 1 or 2, wherein, The processing unit (402) is also configured to receive location data indicating the location of the object of interest, and analyze the one or more frames of the first data to identify the object of interest present in the one or more frames of the first data based on the second data and the location data.
7. The system according to claim 6, wherein, The location data is GPS data.
8. The system according to claim 1 or 2, wherein, Each of the one or more first sensors (204) and each of the one or more second sensors (201) are configured such that the field of view or coverage of the one or more first sensors and the one or more second sensors are aligned.
9. The system according to claim 1 or 2, wherein, The system is also configured to determine the time period of an object of interest (203) within a predetermined area (202).
10. The system according to claim 9, wherein, The system is also configured to generate the labeled dataset only when the time period is greater than a predetermined threshold.
11. The system of claim 1 or 2, further comprising a module configured to adjust the range of the one or more second sensors (201) in response to a range adjustment command.
12. The system according to claim 1 or 2, wherein, The system is also configured to determine the size of the other object (206) by comparing a detected identifier associated with the other object (206) with a lookup table, wherein the lookup table is a lookup table of identifiers and associated sizes of the other object (206).
13. The system of claim 12 further includes a module configured to adjust the range of the one or more second sensors (201) based on a determined size, length or width of the other object (206).
14. The system according to claim 1 or 2, wherein, The system is also configured to determine one or more sub-sectors within a predetermined region (202) based on triangulation of radio frequency signals from multiple signal detectors.
15. The system according to claim 14, wherein, The labeled dataset is generated only for objects that are not located within the one or more sub-sectors.
16. A training system for training a machine learning model, the training system comprising: The processing unit is configured as follows: Receive data using the system tag described in any of the preceding claims; The machine learning model is trained using the received data.
17. A detection system for detecting one or more objects of interest, the detection system comprising: The input module is configured to receive first sensor data comprising one or more frames; A machine learning model trained using the training system of claim 16; The processing unit is configured as follows: The machine learning model is used to determine whether any object of interest exists in each of the one or more frames of the first sensor data.
18. The detection system according to claim 17, wherein, The detection system is also configured to be based on one or more vehicles within the field of view that are tracking the data from the first sensor.
19. The detection system according to claim 18, wherein, The detection system is also configured to determine the likelihood of a collision between one or more tracked vehicles.
20. The detection system according to any one of claims 17-19, wherein, The detection system is also configured to detect when an object of interest (203) enters a predetermined area (202) and generate an entry timestamp corresponding to the time when the object of interest enters the predetermined area (202).
21. The detection system according to any one of claims 17-19, wherein, The detection system is also configured to detect when an object of interest (203) leaves a predetermined area (202) and generate a departure timestamp corresponding to the time when the object of interest leaves the predetermined area (202).
22. A method for generating a labeled dataset, the method comprising: First data is received from one or more first sensors, wherein the first data includes one or more frames, each frame including an image or point cloud, and also includes an associated timestamp indicating the time when the image or point cloud was captured; and wherein the first data includes data defining an object of interest (203) within a predetermined region (202); Second data is received from one or more second sensors, wherein the second data is associated with the object of interest within the predetermined area, and wherein the second data includes one or more unique object identifiers and one or more timestamps, the one or more timestamps representing the time when the object of interest leaves the coverage area of the one or more second sensors; Analyze the one or more frames of the first data to identify the objects of interest present in the first data based on the second data; as well as Based on the analysis, one or more frames of the first data are labeled to generate the labeled dataset.
23. The method according to claim 22, wherein, At least one of the one or more first sensors is a camera configured to capture an image of the area, and at least one of the one or more second sensors is configured to detect radio frequency signals.
24. The method according to claim 22 or 23, wherein, The first data is associated with two-dimensional or three-dimensional image data of the predetermined area, and the second data is associated with radio frequency data.
25. The method according to claim 22 or 23, wherein, The first data includes a single frame of video data, the timestamp corresponding to the time when the single frame of the video data was detected, and / or, wherein the first data includes a single instance of point cloud depth data, the timestamp corresponding to the time when the single instance of the point cloud depth data was detected.
26. The method according to claim 22, wherein, One or more unique identifiers are Internet Protocol (IP) addresses.
27. The method of claim 22 or 23, further comprising receiving location data indicating the location of the object of interest, and analyzing the one or more frames of the first data to identify the object of interest present in the one or more frames of the first data based on the second data and the location data.
28. The method according to claim 27, wherein, The location data is GPS data.
29. The method according to claim 22 or 23, wherein, Each of the one or more first sensors (204) and each of the one or more second sensors (201) are configured such that the field of view or coverage of the one or more first sensors and the one or more second sensors are aligned.
30. The method according to claim 22 or 23, wherein, The method also includes determining the time period of the object of interest (203) within the predetermined area (202).
31. The method according to claim 30, wherein, The method further includes generating the labeled dataset only when the time period is greater than a predetermined threshold.
32. The method according to claim 22 or 23, wherein, The method further includes adjusting the range of the one or more second sensors (201) in response to a range adjustment command.
33. The method according to claim 22 or 23, wherein, The method further includes determining the size of the other object (206) by comparing a detected identifier associated with the other object (206) with a lookup table, wherein the lookup table is a lookup table of identifiers and associated sizes of the other object (206).
34. The method of claim 33 further includes adjusting the range of the one or more second sensors (201) based on a determined size, length or width of the other object (206).
35. The method according to claim 22 or 23, wherein, The method further includes determining one or more sub-sectors within a predetermined region (202) based on triangulation of radio frequency signals from multiple signal detectors.
36. The method according to claim 35, wherein, The labeled dataset is generated only for objects that are not located within the one or more sub-sectors.
37. A training method for training a machine learning model, the training method comprising: Receive data marked using the method described in any of the preceding claims; The machine learning model is trained using the received data.
38. A detection method for detecting one or more objects of interest, the detection method comprising: The input module receives first sensor data, including one or more frames. The machine learning model is trained using the training method described in claim 37; The machine learning model is used to determine whether any object of interest exists in each of the one or more frames of the first sensor data.
39. The detection method according to claim 38, wherein, The detection method also includes one or more vehicles within the field of view that are determined based on the data from the first sensor.
40. The detection method according to claim 39, wherein, The detection method also includes determining the probability of a collision between one or more tracked vehicles.
41. The detection method according to any one of claims 38 to 40, wherein, The detection method further includes detecting when an object of interest (203) enters a predetermined area (202) and generating an entry timestamp corresponding to the time when the object of interest enters the predetermined area (202).
42. The detection method according to any one of claims 38 to 40, wherein, The detection method further includes detecting when the object of interest (203) leaves the predetermined area (202) and generating a departure timestamp corresponding to the time when the object of interest leaves the predetermined area (202).
Citation Information
Patent Citations
Automatically tagging images to create labeled dataset for training supervised machine learning models
US10607116B1