System and method of detecting curved mirror in image
Patent Information
- Application Number
- JP2023195726
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-22
- Filing Date
- 2023-11-17
- Publication Date
- 2025-06-25
- Estimated Expiration
- 2043-11-17
AI Technical Summary
Existing Advanced Driver Assistance Systems (ADAS) struggle with accurate, real-time detection of curved mirrors due to their reflective nature and lack of fixed patterns, leading to false positives and slow computational speeds.
A computer-implemented method using a Deep Autoencoder Gaussian Mixture Model (DAGMM) for anomaly detection, combined with a machine learning algorithm like YOLOv3, to identify and confirm curved mirrors by analyzing image data and transforming it into a domain where points are compared to traffic sign clusters for verification.
The method significantly reduces false positives and enhances detection speed, achieving over 95% accuracy and reducing detection time to 21.5 milliseconds, supporting collision prediction and avoidance systems.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates generally to a system and computer-implemented method for detecting convex mirrors in an image. [Background technology]
[0002] Road traffic accidents tend to occur at and around road junctions and at points on roads, such as at sharp bends, where it is difficult for vehicles to see traffic coming from different directions. In particular, the number of accidents occurring at junctions without traffic lights is generally higher than at junctions with traffic lights.
[0003] In order to reduce the number of accidents occurring at intersections, especially at intersections without traffic lights, road safety mirrors, such as convex mirrors, are typically installed at intersections and at points on roads, such as at points with sharp curves. Such convex mirrors allow vehicles to see blind spots and the like that exist at these intersections and points, thereby allowing road users to feel safe and avoid or reduce the occurrence of accidents. Naturally, the convex mirrors are located outside the vehicle.
[0004] It has been proposed to provide vehicles with image capture devices, e.g. video cameras, in order to detect convex mirrors at certain points on the road, e.g. at intersections, sharp turns, etc. The detection of convex mirrors plays an important role in blind spot detection.
[0005] It has been recognized that improvements to existing Advanced Driver Assistance Systems (ADAS) are desirable. In particular, existing object recognition technology can be improved to accommodate real-time convex mirror recognition for use in ADAS.
[0006] For example, existing ADAS cannot quickly detect convex mirrors in real time. Furthermore, the accuracy of existing systems is also problematic because many convex mirrors that are detected are false positive detections. Accurate detection of convex mirrors is difficult because convex mirrors simply reflect an image of the surroundings where they are located and do not have any fixed pattern.
[0007] Furthermore, the inventors have noticed that the computation speed of convex mirror detection can also affect collision risk prediction. The computation load increases with the complexity of the detection system, so the computation speed of convex mirror detection tends to be slower. One problem that may occur is that with current convex mirror detection, it may take a significant amount of time before the convex mirror is detected. [Prior art documents] [Non-patent literature]
[0008] [Non-Patent Document 1] Redmon, Joseph & Farhadi, Ali. (2018). Yolov3: An Incremental Improvement [Non-Patent Document 2] Ammar A,Koubaa A,Ahmed M,Saad A,Benjdira B.Vehicle Detection from Aerial Images Using Deep Learning:A Comparative Study.Electronics.2021:10(7):820 [Non-Patent Document 3] Zong.B.,Song.Q.,Min.M.,Cheng.W.,Lumezanu,C.,Cho.D.,&Chen.H.(2018).Deep Autoencoding Gaussian Mixture Model for Unsupervised Anomaly Detection.ICLR Summary of the Invention [Problem to be solved by the invention]
[0009] Therefore, there is a need for a system and computer-implemented method for detecting convex mirrors in an image that attempts to address, or at least ameliorate, one of the problems discussed above. [Means for solving the problem]
[0010] According to one aspect of the present disclosure, a computer-implemented method for detecting convex mirrors in an image is provided, the method comprising the steps of applying a machine learning algorithm to the image to identify potential convex mirrors in the image, and applying an anomaly detection algorithm to the identified potential convex mirrors to determine whether the potential convex mirrors are convex mirrors.
[0011] The anomaly detection algorithm of the methods disclosed herein may be based on the Deep Autoencoder Gaussian Mixture Model (DAGMM).
[0012] A convex mirror candidate of the method disclosed herein may be confirmed as a convex mirror if it is not a traffic sign.
[0013] The anomaly detection algorithm of the method disclosed herein may include the steps of applying an autoencoder function to image data associated with a convex mirror candidate to obtain a point on a transform map, said point representing the convex mirror candidate on the transform map, and confirming the convex mirror candidate as a convex mirror if the point is beyond a threshold distance from a cluster representing a traffic sign on the transform map.
[0014] There may be multiple clusters representing traffic signs, and a candidate convex mirror is confirmed to be a convex mirror if its point is beyond a threshold distance from each of the clusters.
[0015] A machine learning algorithm can be trained using a bounding box with parameters optimized for the convex mirror.
[0016] The optimized parameters for the convex mirror may include the bounding box height, the bounding box width, the abscissa of the bounding box center point, the ordinate of the bounding box center point, and the bounding box height-to-weight ratio.
[0017] The method may further include acquiring the image using an image capture device.
[0018] Acquiring the images may include acquiring a plurality of images over different time instances.
[0019] The image may be a real-time image.
[0020] According to another aspect of the present disclosure, there is provided a system for detecting convex mirrors in an image, the system including an image capture device and an electronic control unit coupled to the image capture device, the image capture device configured to acquire the image, the electronic control unit configured to apply a machine learning algorithm to the image to identify potential convex mirrors in the image, and apply an anomaly detection algorithm to the identified potential convex mirrors to confirm whether the potential convex mirrors are convex mirrors.
[0021] The anomaly detection algorithm of the system disclosed herein may be based on the Deep Autoencoder Gaussian Mixture Model (DAGMM).
[0022] A convex mirror candidate in the system disclosed herein may be identified as a convex mirror when it may not be a traffic sign.
[0023] The anomaly detection algorithm of the system disclosed in this specification may include the steps of applying an autoencoder function to image data associated with the convex mirror candidate to obtain a point on a transformation map, the point representing the convex mirror candidate on the transformation map, and confirming the convex mirror candidate as a convex mirror if the point is beyond a threshold distance from a cluster representing a traffic sign on the transformation map.
[0024] According to another aspect of the present disclosure, there is provided a computer-readable storage medium having stored thereon instructions for instructing a processing unit of the system to execute a computer-implemented method for detecting convex mirrors in an image, the method comprising the steps of applying a machine learning algorithm to the image to identify potential convex mirrors in the image, and applying an anomaly detection algorithm to the identified potential convex mirrors to determine whether the potential convex mirrors are convex mirrors.
[0025] Exemplary embodiments of the invention will be better understood and readily apparent to those skilled in the art from the following description, which is given by way of example only, and which refers to the drawings in which: [Brief description of the drawings]
[0026] [Figure 1] FIG. 1 is a schematic block diagram illustrating a system for detecting convex mirrors in an image in an exemplary embodiment. [Diagram 2] 1 is a schematic flow chart illustrating a computer-implemented method for detecting convex mirrors in an image in an exemplary embodiment; [Diagram 3] 1 shows a series of images to illustrate dataset creation for detecting road safety mirrors, e.g., convex mirrors, using a machine learning algorithm in an exemplary embodiment. [Figure 4] 13 is a frame shot of a verification result of road safety mirror detection in an exemplary embodiment. [Diagram 5]1 shows a sequence of images depicting examples of true positive images in the form of road safety mirrors detected by an illustrative embodiment; [Figure 6] 5 shows a sequence of images depicting examples of false positive images in the form of road traffic signs detected by an illustrative embodiment; [Figure 7] 1 is a transformation map including data points for convex mirror candidates identified in an image in an exemplary embodiment; [Figure 8] 1 is a diagram illustrating an improved anomaly detection method in an exemplary embodiment. [Figure 9] 11 is a photograph illustrating the effectiveness of an improved anomaly detection method in an example embodiment; [Figure 10] 11A-11C are diagrams and photographs illustrating road safety mirror detection at a T-junction in an exemplary embodiment; [Figure 11] 11 is a chart showing dimensions of a convex mirror relative to dimensions of a default box for model training in an exemplary embodiment. [Figure 12] 1 is a schematic diagram of a computer system suitable for implementing the exemplary embodiments. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0027] Exemplary non-limiting embodiments may provide a system and computer-implemented method for detecting road safety mirrors, such as convex mirrors, in an image.
[0028] In various embodiments, the term "convex mirror," as used herein, refers broadly to a mirror having a curved reflective surface. In various embodiments, the convex mirror includes a curved mirror portion. In various embodiments, the surface is convex (i.e., convex outward). In various embodiments, the convex mirror is a convex mirror.
[0029] In various embodiments, the terms "image" or "image data," as used herein, broadly refer to any content or data that can be rendered for viewing by a user. In various embodiments, the image or image data includes computer-readable data. In various embodiments, the image or image data may be converted to computer-readable data. In various embodiments, the image or image data may be converted to a format suitable or compatible for use by components of the systems and methods disclosed herein. For example, the image or image data may be a frame from a video, a portion of a frame from a video, a still image, a portion of a still image, or otherwise.
[0030] 1 is a schematic block diagram illustrating a system 100 for detecting convex mirrors in an image in an exemplary embodiment. The system 100 is mounted on a vehicle. The system 100 includes an image capture device 102 and a processing unit 104 coupled to the image capture device 102.
[0031] In an exemplary embodiment, the image capture device 102 is configured to capture images. The vehicle image capture device 102 may be a camera or video camera located on the front of the vehicle body near the rearview mirror or on the front grill of the vehicle. The image capture device 102 is positioned such that its image capture area is at a predetermined angle toward the front of the vehicle.
[0032] In an exemplary embodiment, the processing unit 104 may be configured to control the functions of the components of the system 100. In an exemplary embodiment, the processing unit 104, e.g., an electronic control unit (ECU) of a vehicle, is configured to apply machine learning algorithms to the images to identify potential convex mirrors in the images, and to apply anomaly detection algorithms to the identified potential convex mirrors to determine whether the potential convex mirrors are convex mirrors.
[0033] In an exemplary embodiment, the anomaly detection algorithm may be based on a Deep Autoencoder Gaussian Mixture Model (DAGMM). The anomaly detection algorithm may be configured to identify a convex mirror candidate as a convex mirror when it cannot be a road sign. For example, the anomaly detection algorithm may include applying an autoencoder function to image data associated with the convex mirror candidate to obtain a point, e.g., a data point, on a transformation map, the point representing the convex mirror candidate on the transformation map, and identifying the convex mirror candidate as a convex mirror if the point is beyond a threshold distance from a cluster representing a traffic sign on the transformation map. In an exemplary embodiment, there may be multiple clusters of points representing traffic signs. The convex mirror candidate is identified as a convex mirror if the point is beyond a threshold distance from each of the clusters.
[0034] In an exemplary embodiment, the processing unit 104 may be further coupled to an object detection unit 106, which may be further coupled to an action unit 108. The object detection unit 106 may be configured to perform object detection on the convex mirror detected by the processing unit 104. The action unit 108 may be configured to receive one or more commands from the object detection unit 106, such as, but not limited to, to activate a braking function of the vehicle, and / or activate a warning system for a vehicle user, and / or perform steering control, or to not perform any action.
[0035] During operation, a crossroads or intersections may be identified by image capture device 102. In some example embodiments, map information may be provided to processing unit 104, which may use the map information to identify crossroads or intersections, for example by matching the map with Global Positioning System (GPS) coordinates. At such crossroads or intersections, processing unit 104 applies machine learning algorithms to images captured by image capture device 102 to identify potential convex mirrors in the images, and then applies anomaly detection algorithms to the identified potential convex mirrors to confirm whether the potential convex mirrors are convex mirrors.
[0036] 2 is a schematic flow chart 200 illustrating a computer-implemented method for detecting convex mirrors in an image in an exemplary embodiment. In step 202, a machine learning algorithm is applied to the image to identify convex mirror candidates in the image. In step 204, an anomaly detection algorithm is applied to the identified convex mirror candidates to confirm whether the convex mirror candidates are convex mirrors. The method for detecting convex mirrors in an image may be implemented using the system 100 of FIG. 1. For example, the identification and confirmation of convex mirror candidates in the image may be performed by the processing unit 104.
[0037] In an exemplary embodiment, the method may further include capturing / obtaining the images using an image capture device. The image capture device may be a camera or a video camera. The image capture device may be mounted on the vehicle and positioned in front of the vehicle body near a rearview mirror or on a front grill of the vehicle. Capturing the images may include capturing a forward looking image toward the front of the vehicle. Capturing the images may include capturing multiple images across different time instances. For example, the multiple images across different time instances may be a time-continuous series of frames from a video, with each frame representing an image captured at a particular time instance. The images may be real-time images. The images may be frames obtained from a real-time video.
[0038] In an exemplary embodiment, no convex mirror candidates may be identified in the image. In an exemplary embodiment, one or more convex mirror candidates may be identified in the image. In an exemplary embodiment, no convex mirror may be identified from one or more convex mirror candidates identified in the image. In an exemplary embodiment, one or more convex mirrors may be identified from one or more convex mirror candidates identified in the image.
[0039] In an exemplary embodiment, the machine learning algorithm may be a real-time object detection system. The real-time object detection system may be a one-pass object detection system in which an image is analyzed only once. The real-time object detection system may be a You Only Look Once (YOLO) v3 system. In an exemplary embodiment, the machine learning algorithm may be trained using a bounding box of optimized parameters for the convex mirror. The optimized parameters for the convex mirror may include a bounding box height, a bounding box width, an abscissa of a center point of the bounding box, an ordinate of a center point of the bounding box, and a height-to-weight ratio of the bounding box.
[0040] In an exemplary embodiment, the anomaly detection algorithm may be based on the likelihood that the convex mirror candidate is not a traffic sign. The anomaly detection algorithm may be based on a Deep Autoencoder Gaussian Mixture Model (DAGMM). In an exemplary embodiment, the anomaly detection algorithm may include applying an autoencoder function to image data associated with the convex mirror candidate to transform the candidate into a different domain to obtain a point on a transformation map, the point representing the convex mirror candidate on the transformation map, and confirming the convex mirror candidate as a convex mirror if the point is beyond a threshold distance from a cluster representing a traffic sign on the transformation map. In an exemplary embodiment, there may be multiple clusters of points representing traffic signs. The convex mirror candidate is confirmed as a convex mirror if the point is beyond a threshold distance from each of the clusters.
[0041] 3A-3C show a series of images to illustrate the creation of a dataset for road safety mirror, e.g., convex mirror detection, using a machine learning algorithm in an exemplary embodiment. In an exemplary embodiment, the machine learning algorithm used for road safety mirror detection is based on a real-time object detection system known as You Only Look Once (YOLO) v3. In an exemplary embodiment, the latest version as of 2019 has been selected.
[0042] FIG. 3A shows an original image captured by an image capture device in its original pixel size. For the training data, a high-resolution pixel image (approximately about 12 million pixels) can be used. As shown in FIG. 3A, the Region of Interest (ROI) information of the road safety mirror is included in the original image. The ROI means the region of interest of the target object, i.e., the area of the road safety mirror, which is surrounded by a box 302 in FIG. 3A. FIG. 3B shows a predetermined reduced pixel size image scaled down from the original image of FIG. 3A. FIG. 3C shows the L image to be used for the training dataset. * a * b *FIG. 3C shows an image in color space converted from the reduced pixel size image of FIG. 3B.
[0043] In an exemplary embodiment, the original image of FIG. 3A is scaled to a predetermined reduced pixel size of FIG. 3B and the L of FIG. * a * b * As those skilled in the art will appreciate, the RGB color space is based on the colors red, green, and blue. * a * b * In color space, capital L * means brightness, and lowercase a * , b * means complementary color. * a * b * I noticed that the color space seems to be close to human vision. In the RGB color space, it can be difficult to distinguish the color gamut depending on the brightness situation. On the other hand, L * a * b * In color space, it may be possible to distinguish color regions like the human eye. Therefore, in an exemplary embodiment, the original image is converted to a L-pixel image for use as an input for the YOLOv3 deep learning method. * a * b * converted to color space.
[0044] FIG. 4 is a frame shot of the result of the confirmation of road safety mirror detection in the exemplary embodiment. A public road test was conducted using the exemplary embodiment of the system and method disclosed herein. In the exemplary implementation, a low-resolution camera (approximately, about 320,000 pixels) was used as the image capture device, i.e., not a high-resolution camera used for data collection. In the exemplary implementation, an NVIDIA Jetson Xavier module was used as the processing unit, e.g., controller. The road safety mirror detected by the exemplary implementation is surrounded by the frames 402 and 404 shown in FIG. 4.
[0045] Figures 5A and 5B show a sequence of images representing examples of true positive images in the form of road safety mirrors detected by an exemplary embodiment. Figures 6A-6D show a sequence of images representing examples of false positive images in the form of road traffic signs detected by an exemplary embodiment.
[0046] Despite using a low-resolution camera (approximately 320,000 pixels), the results showed a high accuracy of over 95%, as shown in Table 1 and equation (1).
number
[0047] [Table 1]
[0048] As can be seen from the results in Table 1, there is some FP in the results of road safety mirror detection.
[0049] Exemplary embodiments of the systems and methods disclosed herein further provide a solution to the problem of false positive results, as will be described below with reference to FIGS.
[0050] FIG. 7 is a transformation map 700 that includes data points, such as 702a and 704a, for potential convex mirrors identified in an image in an exemplary embodiment.
[0051] 7, the transformation map 700 includes a first data point 702a representing a first traffic sign 702b, a second data point 704a representing a second traffic sign 704b, and a cluster 706 of data points representing, for example, a first convex mirror 708 shown as a circle on the image, and a second convex mirror 710 shown as an oval. Each data point on the transformation map 700 is obtained by applying an anomaly detection method to image data associated with a candidate convex mirror. The first data point 702a and the second data point 704a are examples of false positive results that do not represent a convex mirror.
[0052] Anomaly detection methods can be used as a countermeasure to overcome the problem of false positive results. An example of an anomaly detection method is the Deep Autoencoder Gaussian Mixture Model (DAGMM). By using the Gaussian Mixture Model, one data point can be obtained for each object, and clusters of data points can be obtained for multiple objects on the transformation map 700.
[0053] However, the inventors have realized that even with existing anomaly detection methods such as DAGMM, it is difficult to define the distribution of road safety mirrors compared to traffic signs due to the dynamic and inconsistent images reflected from the road safety mirrors. Unlike traffic signs that typically display a consistent pattern, convex mirrors do not have a fixed pattern but reflect images of the surroundings where they are located. Therefore, there is no consistent pattern displayed by the convex mirrors.
[0054] To overcome the difficulty in distinguishing between true and false positive results, the inventors devised an improved anomaly detection algorithm.
[0055] 8 is a schematic diagram illustrating an improved anomaly detection method in an exemplary implementation. In the exemplary implementation, the improved anomaly detection method is based on DAGMM.
[0056] In an exemplary implementation, a convex mirror candidate 800 is identified from an image 802. The convex mirror candidate 800 may be detected by applying a machine learning algorithm, for example, a deep learning method YOLOv3, to the image 802. The convex mirror candidate 800 may be either a road traffic sign (i.e., a false positive) or a convex mirror (i.e., a true positive). One or more convex mirror candidates 800 may be identified from the image 802. For example, there are two convex mirror candidates in the image 802, which are represented by two frames 804 and 806.
[0057] The improved anomaly detection method then involves applying an autoencoder function to image data associated with the convex mirror candidate 800 to obtain data points (e.g., 808, 810) on a transformation map 812, where the data points (e.g., 808, 810) represent the convex mirror candidate 800 on the transformation map 812, and confirming the convex mirror candidate 800 as a convex mirror 814 if the data point (e.g., 808) is beyond a threshold distance from a cluster representing the traffic sign on the transformation map 812.
[0058] As shown in FIG. 8, there are multiple clusters (e.g., 816, 818) representing traffic signs 820. If the data points in a cluster (e.g., 816, 818) represent a traffic sign, the data points close to the cluster may be traffic signs, while the data points far from the cluster may be road safety mirrors, e.g., convex mirrors. Thus, in an exemplary implementation, the distance of the data points to the cluster representing the road traffic sign is calculated. If the distance of the data points to the cluster representing the traffic sign exceeds a predetermined threshold, the convex mirror candidate is deemed to be a road safety mirror. If the distance of the data points to the cluster representing the traffic sign is within a predetermined threshold, the convex mirror candidate is deemed not to be a road safety mirror. In other words, the convex mirror candidate 800 is confirmed as a convex mirror 814 if the data points, e.g., 808, are beyond the threshold distance from each of the clusters (e.g., 816, 818). In general, the threshold distance used to confirm the convex mirror candidate may depend on factors such as the resolution of the camera used. Thus, in various exemplary implementations, the threshold distance for identifying convex mirror candidates can be adjusted and optimized to maximize true positive results and minimize false positive results of convex mirror detection.
[0059] 9A and 9B are pictures showing the effectiveness of the improved anomaly detection method in an exemplary implementation. FIG. 9A shows an image taken from a portion of a video (i.e., a sequence of frames / images) obtained without an improved anomaly detection method, e.g., DAGMM. Since no anomaly detection was applied, there were 31 false positive detections in this video. On the other hand, FIG. 9B shows an image taken from the same portion of a video obtained with an improved anomaly detection method, e.g., DAGMM. As shown in FIG. 9B, there were no false positive detections in the same video. Thus, the effectiveness of the improved anomaly detection method disclosed herein was proven by this video.
[0060] 10 is a diagram and a photograph showing the detection of road safety mirrors at a T-junction in an exemplary embodiment. In an exemplary embodiment, road safety mirror detection by the method of applying the improved anomaly detection algorithm DAGMM is confirmed after the subject vehicle 1000 approaches the T-junction. Two road safety mirrors are identified, which are indicated by boxes 1002 and 1004. Therefore, it is demonstrated that the system and method for detecting convex mirrors disclosed herein are effective in reducing false positive results in road safety mirror detection and can successfully detect road safety mirrors.
[0061] 11 is a chart showing the dimensions of a convex mirror relative to the dimensions of a default box for model training in an exemplary embodiment. In an exemplary embodiment, a machine learning algorithm can be trained using a bounding box with parameters optimized for the convex mirror. The parameters optimized for the convex mirror can include the height of the bounding box, the width of the bounding box, the abscissa of the center point of the bounding box, the ordinate of the center point of the bounding box, and the height-to-weight ratio of the bounding box.
[0062] In an exemplary embodiment, a deep learning method is utilized for the recognition of convex mirrors. Instead of using a normal default box for model training, the machine learning model may use an optimized convex mirror size from a normal default box. The optimal mirror size is identified based on test results in which an image capture device, e.g., a camera, detects convex mirrors on public roads. The optimal mirror size may depend on factors such as the resolution of the camera used. Thus, in various exemplary embodiments, the mirror size may be adjusted and optimized accordingly. In the exemplary embodiment shown in FIG. 11, the optimal mirror size is about 0.2 with respect to the overall image size. Compared with the use of a conventional normal default box, the adoption of the optimized convex mirror size allows for convex mirror detection from a longer distance. In an exemplary embodiment, the use of the optimized convex mirror size achieves a faster detection time of 21.5 milliseconds (ms) while maintaining a comparable accuracy level. An example of the machine learning model may be a Single Shot Multibox detector (SSD).
[0063] [Table 2]
[0064] In an exemplary implementation using the optimized convex mirror size, first, a crossroad is detected by an image capture device, a camera. Optionally, map information may be useful. For example, map information may be provided and matched against Global Positioning System (GPS) coordinates to identify a crossroad or intersection. Then, a machine learning algorithm is applied at this crossroad to identify / detect a convex mirror area.
[0065] In an exemplary embodiment, YOLOv3 is described as being used as a real-time detection system for convex mirror candidate detection. YOLOv3 is a real-time object detection system that applies a single neural network to the entire image. The network divides the image into regions and predicts bounding boxes and probabilities for each region. These bounding boxes are weighted by the predicted probability. The YOLOv3 system has several advantages over classifier-based systems. Because it looks at the entire image at test time, predictions are usually based on the global context in the image. It also makes predictions in a single network evaluation, unlike systems like R-CNN, which require thousands for a single image. This makes it extremely fast, over 1000 times faster than R-CNN and 100 times faster than Fast R-CNN. Additional information about YOLOv3 can be found in (Non-Patent Document 1), which is incorporated herein by reference. An example of YOLOv3 can be found in (Non-Patent Document 2), which is incorporated by reference herein, an example of the architecture of YOLOv3 can be found at least in Section 3.2.1 and FIG. 2, and an example of training YOLOv3 can be found at least in Section 4.2.
[0066] A first example of a dataset that may be used to train the YOLOv3 machine learning model is the COCO dataset available at https: / / cocodataset.org / #home, which is a large-scale object detection, segmentation, and captioning dataset. A second example of a dataset that may be used to train the YOLOv3 machine learning model is the CIFAR-10 or CIFAR-100 dataset available at http: / / www.cs.toronto.edu / ~kriz / cifar.html, which is a labeled subset of a small 80 million image dataset. It is contemplated that any other suitable dataset may be used.
[0067] In an exemplary embodiment, it is noted that DAGMM is used as an anomaly detection algorithm to determine whether a convex mirror candidate is a convex mirror or not. DAGMM utilizes a deep autoencoder to generate a low-dimensional representation and reconstruction error for each input data point, which is then fed into a Gaussian Mixture Model (GMM). Instead of using a separate two-stage training and a standard Expectation-Maximization (EM) algorithm, DAGMM jointly optimizes the parameters of the deep autoencoder and the mixture model end-to-end at the same time, and uses a separate estimation network to facilitate parameter learning of the mixture model. The joint optimization, which balances the autoencoder reconstruction, density estimation of the latent representation, and regularization, helps the autoencoder avoid less attractive local optima and further reduce the reconstruction error, eliminating the need for pre-training. Additional information on DAGMM can be found in (Non-Patent Document 3), which is incorporated herein by reference.
[0068] One example of a dataset that may be used to train a DAGMM machine learning model is the KDDCUP99 10 percent dataset from the UCI repository, available at https: / / archive.ics.uci.edu / ml / datasets / KDD+Cup+1999+Data. It is contemplated that any other suitable dataset may be used.
[0069] The exemplary embodiment described above has been verified using a front camera mounted on a vehicle in a feasibility study, and the effectiveness of the method has been demonstrated by experimental results on public roads.
[0070] The above-described exemplary embodiments can advantageously detect convex mirrors that are depicted as perfect circles, as well as convex mirrors that are depicted as ellipses from the viewing angle of the image capture device, in contrast to existing ADAS that typically cannot detect convex mirrors when they appear as ellipses from the viewing angle of the image capture device.
[0071] The above-described exemplary embodiments can advantageously reduce the occurrence of false positive results and improve the accuracy of convex mirror detection compared to approaches that use only known deep learning models, in contrast to existing detection systems that incorrectly identify road traffic signs as convex mirrors.
[0072] The exemplary embodiments described above can advantageously provide faster computations and beneficially save time compared to approaches that use only known deep learning models.
[0073] The exemplary embodiments described above may be used for collision prediction and avoidance, for example, at road locations with blind spots and at intersections without traffic lights.
[0074] The exemplary embodiments described above may advantageously reduce the investment required to build new facilities at intersections and low visibility locations by utilizing a conventional piece of infrastructure already present on roadways: road safety mirrors.
[0075] The exemplary embodiments described above may be advantageously used to support Autonomous Driving (AD) systems at levels 1 and 2, such as Advanced Driver Assistance Systems (ADAS), and may be further extended to support AD systems above level 3. The exemplary embodiments described above may also be advantageously applied in other forms of vehicles, such as unmanned ground vehicles (uGVs), automated guided vehicles (AGVs), autonomous vehicles, drones, etc. The exemplary embodiments described above may also be advantageously used in the field of robotics.
[0076] The terms "coupled" or "connected," as used in this description, are intended to cover both direct connection and connection through one or more intermediary means, unless expressly stated otherwise.
[0077] Terms such as "configured to perform (a task / operation)", "configured for performing (a task / operation)", etc., as used in this description, include being programmable, programmed, connectable, wired, or otherwise configured to have the capability to perform the task / operation when arranged or installed as described herein. Terms such as "configured to perform (a task / operation)", "configured to perform (a task / operation)", etc. are intended to cover "when used, the task / operation is performed", e.g., specifically performing or performing the task / operation and / or being specifically configured and / or being specifically arranged and / or being specifically made to be so.
[0078] The term "and / or," e.g., "X and / or Y," should be understood to mean either "X and Y" or "X or Y," and should not be understood as expressly endorsing either or both meanings.
[0079] The terms "associated with," "related to," and the like, when used herein in reference to two elements, refer to a broad relationship between the two elements. This relationship includes, but is not limited to, a physical, chemical, or biological relationship. For example, if element A is associated with element B, elements A and B may be directly or indirectly attached to each other, or element A contains element B, or vice versa.
[0080] Terms such as "exemplary embodiment," "exemplary implementation," "exemplary," and the like, as used herein, are intended to provide an example of the subject matter described in the present disclosure. Such an example may relate to one or more features defined in the claims, and is not necessarily intended to highlight the best example or any fundamental importance of any feature.
[0081] Certain portions of the description herein may be explicitly or implicitly described in terms of algorithms and / or functional operations performed on data within a computer memory or electronic circuitry. These algorithmic descriptions and / or functional operations are typically used by those skilled in the information processing arts for efficient description. An algorithm generally relates to a self-consistent sequence of steps leading to a desired result. The steps of an algorithm may include physical manipulations of physical quantities, such as electrical, magnetic, or optical signals that can be stored, sent, transmitted, combined, compared, and otherwise manipulated.
[0082] Furthermore, unless expressly stated otherwise, and as will generally be apparent from the following, those skilled in the art will understand that discussions throughout this specification utilizing terms such as "scan," "calculate," "identify," "replace," "generate," "initialize," "output," and the like refer to operations and processes of an instruction processor / computer system or similar electronic circuits / devices / components that manipulate / process data represented as physical quantities within the described system and convert it into other data similarly represented as physical quantities in that system or other information storage, transmission, or display device, and the like.
[0083] This description also discloses devices / apparatus associated with performing the method steps described above. Such apparatus may be specially configured for this purpose, or may include a general-purpose computer / processor or other device selectively activated or reconfigured by a computer program stored in a memory member. The algorithms and displays described herein are not inherently related to any particular computer or other apparatus. It is to be understood that a general-purpose device / machine may be used in accordance with the teachings herein. Alternatively, the construction of a specialized device / apparatus to perform the method steps may be desired.
[0084] In addition, the description is also proposed to implicitly cover computer programs, in that it is clear that the steps of the methods described herein may be embodied by computer code. It will be appreciated that a variety of programming languages and encodings can be used to implement the teachings of the description herein. Furthermore, the computer program is not limited to any particular control flow, if applicable, and may use different control flows without departing from the scope of the present invention.
[0085] Furthermore, one or more of the steps of the computer program may be executed in parallel and / or sequentially, if applicable. Such a computer program may be stored on any computer-readable medium, if applicable. The computer-readable medium may include storage devices such as magnetic or optical disks, memory chips, or other storage devices suitable for interfacing with a suitable reader / general-purpose computer. In such a case, the computer-readable storage medium is non-transitory. Such storage media also covers any computer-readable medium, such as register memory, processor cache, and Random Access Memory (RAM), and others, if data is stored only for a short time and / or only when there is power. The computer-readable medium may also include wired media, such as exemplified in the Internet system, or wireless media, such as exemplified in Bluetooth technology. The computer-readable medium may be, for example, a cloud storage on the Internet or in an intranet. The computer program, when loaded into a suitable reader and executed therein, effectively results in an apparatus capable of implementing the steps of the above-mentioned method, for example in a physical embodiment. The computer readable medium is intended to be portable and reproducible in the sense that the computer program, if applicable, is reproducible.
[0086] Exemplary embodiments may also be implemented as hardware modules. A module is a functional hardware unit designed for use with other components or modules. For example, a module may be implemented using digital or discrete electronic components, or may form part of an entire electronic circuit, such as an Application Specific Integrated Circuit (ASIC). Those skilled in the art will appreciate that exemplary embodiments may also be implemented as a combination of hardware and software modules.
[0087] Additionally, in describing some embodiments, the present disclosure may disclose methods and / or processes as a particular sequence of steps. However, unless otherwise required, it should be understood that the method or process should not be limited to the particular sequence of steps disclosed. Other sequences of steps may be possible. The particular order of steps disclosed herein should not be construed as unduly limiting. Unless otherwise required, the methods and / or processes disclosed herein should not be limited to the steps being performed in the order described. The sequence of steps may be varied and still remain within the scope of the present disclosure.
[0088] Further, in the description herein, whenever the term "substantially" is used, it should be understood to include, but not be limited to, "entirely", or "completely", etc. Additionally, whenever used, terms such as "comprising" are intended to be open-ended descriptive language in that they broadly include the elements / components listed following the term, as well as other components not explicitly listed. For example, when "comprising" is used, a reference to "a" feature is also intended to be a reference to "at least one" of that feature. Terms such as "consisting" may be considered subsets of terms such as "comprising" in the appropriate context. Thus, in embodiments disclosed herein using terms such as "comprising", it should be understood that these embodiments also provide teachings regarding corresponding embodiments using terms such as "consisting". Additionally, terms like "about," "approximately," and the like, whenever used, typically refer to a reasonable variation, such as a + / - 5% variation from the disclosed value, or a 4% variation from the disclosed value, or a 3% variation from the disclosed value, or a 2% variation from the disclosed value, or a 1% variation from the disclosed value.
[0089] Furthermore, in the description herein, certain values may be disclosed in ranges. The endpoints of a range are intended to exemplify the preferred range. Whenever a range is described, the range is intended to cover and teach all the individual values within the range, as well as all possible subdivisions. That is, the endpoints of a range should not be interpreted as inflexible limits. For example, a description of a range of 1% to 5% is intended to specifically disclose the individual values within the range, e.g., 1%, 2%, 3%, 4%, 5%, as well as subdivisions such as 1% to 2%, 1% to 3%, 1% to 4%, 2% to 3%, etc. The specific disclosure intent of the recitation applies to any depth / width of a range.
[0090] Various exemplary embodiments may be implemented in terms of data structures, program modules, programs, and computer instructions executed within a computer-implemented environment. A specially configured general-purpose computing environment is briefly disclosed herein. One or more exemplary embodiments may be embodied in one or more computer systems, such as, for example, the one illustrated generally in FIG.
[0091] One or more exemplary embodiments may be implemented as software, such as a computer program executing within computer system 1200 and instructing computer system 1200 to perform methods of the exemplary embodiments.
[0092] The computer system 1200 includes a computer unit 1202, an input module, such as a keyboard 1204 and a pointing device 1206, and a number of output devices, such as a display 1208 and a printer 1210. A user can interact with the computer unit 1202 using the above-mentioned devices. The pointing device can be implemented as a mouse, a trackball, a pen device, or any similar device. One or more other input devices (not shown), such as a joystick, a game pad, a satellite dish, a scanner, a touch-sensitive screen, etc., can also be connected to the computer unit 1202. The display 1208 can include a cathode ray tube (CRT), a liquid crystal display (LCD), a field emission display (FED), a plasma display, or any other device that generates an image viewable by a user.
[0093] The computer unit 1202 can be connected to a computer network 1212 via a suitable transceiver device 1214, which can provide access to, for example, the Internet or other network systems such as a Local Area Network (LAN) or a Wide Area Network (WAN) or a personal network. The network 1212 can include servers, routers, networked personal computers, peer devices or other common network nodes, wireless telephones or wireless personal digital assistants. Networking environments can be found in offices, company-wide computer networks, home computer systems, and the like. The transceiver device 1214 can be a modem / router unit within the computer unit 1202 or external to it, and can be any type of modem / router, such as a cable modem or a satellite modem.
[0094] It is to be understood that the illustrated network connections are exemplary and other methods of establishing a communications link between the computers may be used. Any of a variety of protocols may be assumed to be present, such as TCP / IP, Frame Relay, Ethernet, FTP, HTTP, etc., and the computer unit 1202 may operate in a client-server configuration to allow a user to retrieve web pages from a web-based server. Additionally, any of a variety of web browsers may be used to display and manipulate data on the web pages.
[0095] The computer unit 1202 in this example includes a processor 1218, a Random Access Memory (RAM) 1220, and a Read Only Memory (ROM). The ROM 1222 may be a system memory that stores Basic Input / Output System (BIOS) information. The RAM 1220 may store one or more program modules, such as an operating system, application programs, and program data.
[0096] The computer unit 1202 further includes a number of input / output (I / O) interface units, such as an I / O interface unit 1224 with the display 1208 and an interface unit 1226 with the keyboard 1204. The components of the computer unit 1202 are typically connected and communicate and interfaced / coupled via an interconnected system bus 1228 and in a manner known to those skilled in the relevant art. The bus 1228 can be any of several types of bus structures including a memory bus or memory controller, a peripheral bus, and a local bus using any of a variety of bus architectures.
[0097] It should be understood that other devices can also be connected to the system bus 1228. For example, a video or digital camera can be coupled to the system bus 1228 using a Universal Serial Bus (USB) interface. An IEEE 1394 interface can be used to couple additional devices to the computer unit 1202. Other manufacturer interfaces can also be used, such as FireWire developed by Apple Computer and i.Link developed by Sony. Coupling of devices to the system bus 1228 can also be done through a parallel port, a game port, a PCI board, or any other interface used to couple input devices to a computer. It should also be understood that a microphone and a speaker can be used to record and play back voice / audio, although these components are not shown. A sound card can be used to couple the microphone and the speaker to the system bus 1228. It should be understood that several peripheral devices can be simultaneously coupled to the system bus 1228 through alternative interfaces.
[0098] The application program can be provided to a user of the computer system 1200 in a state encoded / stored on a data storage medium, such as a CD ROM or a flash memory carrier. The application program can be read using a corresponding data storage medium drive of the data storage device 1230. The data storage medium is not limited to being portable and can include an example embedded within the computer unit 1202. The data storage device 1230 can include a hard disk interface unit and / or a removable memory interface unit (neither of which are shown in detail), which respectively couples a hard disk drive and / or a removable memory drive to the system bus 1228. This allows data to be read and written. Examples of removable memory drives include magnetic disk drives and optical disk drives. The drives and their associated computer readable media, such as floppy disks, provide non-volatile storage of computer readable instructions, data structures, program modules, and other data for the computer unit 1202. It is to be understood that the computer unit 1202 can include more than one of such drives. In addition, the computer unit 1202 can include drives for interfacing with other types of computer readable media.
[0099] Application programs are read and controlled by execution by the processor 1218. Intermediate storage of program data may be provided using RAM 1220. The methods of the exemplary embodiments may be implemented as computer readable instructions, computer executable components, or software modules. One or more software modules may alternatively be used. These may include executable programs, data link libraries, configuration files, databases, graphic images, binary data files, text data files, object files, source code files, or the like. When one or more computer processors execute one or more of the software modules, the software modules interact to cause one or more computer systems to operate in accordance with the teachings herein.
[0100] The operation of the computer unit 1202 can be controlled by various program modules. Examples of program modules are routines, programs, objects, components, data structures, libraries, etc. that perform particular tasks or implement particular abstract data types. The exemplary embodiments can also be implemented with other computer system configurations including handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, personal digital assistants, mobile phones, etc. Moreover, the exemplary embodiments can also be implemented in distributed computing environments where tasks are performed by remote processing devices that are linked through a wireless or wired communications network. In a distributed computing environment, program modules can be located in both local and remote memory storage devices.
[0101] The exemplary embodiments may also be implemented with other computer system configurations including handheld devices, multiprocessor systems / servers, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, personal digital assistants, mobile phones, etc. Moreover, the exemplary embodiments may also be implemented in distributed computing environments where tasks are performed by remote processing devices that are linked through a wireless or wired communications network. In a distributed computing environment, program modules may be located in both local and remote memory storage devices.
[0102] In the above exemplary embodiment, the vehicle may be described as an automobile driven by a user. It should be understood that the exemplary embodiment is not so limited. For example, the vehicle may include any movable object capable of detecting a convex mirror.
[0103] Those skilled in the art will appreciate that various modifications and / or improvements may be made to the specific embodiments without departing from the scope of the invention as broadly described. For example, within the description herein, features of different exemplary embodiments may be mixed, combined, exchanged, incorporated, adopted, improved, included, etc., among different exemplary embodiments. For example, the exemplary embodiments are not necessarily mutually exclusive, since some may be combined with one or more embodiments to form new exemplary embodiments. Furthermore, although the present disclosure provides embodiments having one or more of the features / features discussed herein, one or more of these features / features may also be waived in other alternative embodiments, and the present disclosure supports such waiver and alternative embodiments related thereto. The present embodiments are therefore to be considered in all respects as illustrative and not restrictive. [Explanation of symbols]
[0104] 100 System for detecting convex mirrors in an image 102 Image capture device 104 Processing Unit 106 Object Detection Unit 108 Operating Unit 302 slots 402 frames 404 frames 700 Conversion Map 702a First Data Point 702b First Traffic Sign 704a Second Data Point 704b Second Traffic Sign A cluster of 706 data points 706 First Convex Mirror 708 Second Convex Mirror 800 Convex mirror candidate 802 images 804 frames 806 frames 808 data points 810 data points 812 Conversion Map 814 Convex Mirror Cluster of 816 data points A cluster of 818 data points 820 traffic sign 1000 vehicles 1002 frames 1004 frames 1200 Computer Systems 1202 Computer Unit 1204 Keyboard 1206 Pointing Device 1208 Display 1210 Printer 1212 Computer Networks 1214 Transceiver Device 1218 Processor 1220 Random Access Memory 1222 Read Only Memory 1224 Input / Output Interface Unit 1226 Interface Unit 1228 System Bus 1230 Data Storage Device
Claims
1. A computer-implemented method for detecting a curve mirror (814) in an image (802), comprising: applying a machine learning algorithm to the image (802) to identify a curve mirror candidate (800) in the image (802); applying an anomaly detection algorithm to the identified curve mirror candidate (800) to confirm whether the curve mirror candidate (800) is a curve mirror (814); and the curve mirror candidate (800) is confirmed as a curve mirror (814) if it cannot be a traffic sign (820); the anomaly detection algorithm comprises: applying a function of an autoencoder to image data associated with the curve mirror candidate (800) to obtain points (808, 810) on a transformation map (812), the points (808, 810) representing the curve mirror candidate (800) on the transformation map (812); if the points (808, 810) are at a threshold distance or more away from clusters (816, 818) representing traffic signs (820) on the transformation map (812), confirming the curve mirror candidate (800) as a curve mirror (814); and a computer-implemented method characterized by the above.
2. The method according to claim 1, wherein the anomaly detection algorithm is based on a Deep Autoencoder Gaussian Mixture Model (DAGMM).
3. There are a plurality of clusters (816, 818) representing traffic signs (820), and the curve mirror candidate (800) is confirmed as a curve mirror (814) if the points (808, 810) are at a threshold distance or more away from each of the clusters (816, 818). The method according to claim 1 or 2.
4. The method according to claim 1 or 2, wherein the machine learning algorithm is trained using a bounding box having parameters optimized for curve mirrors.
5. The parameters optimized for the curve mirror comprise the height of the bounding box, the width of the bounding box, the abscissa of the center point of the bounding box, the ordinate of the center point of the bounding box, and the ratio of the height to the width of the bounding box. The method according to claim 4.
6. The method according to claim 1 or 2, further comprising the step of acquiring the image (802) using the image capturing device (102).
7. The method according to claim 6, wherein the step of acquiring the image (802) includes the step of acquiring a plurality of images over different time instances.
8. The method according to claim 1 or 2, wherein the image is a real-time image.
9. A system (100) for detecting a curved mirror (814) in an image (802), comprising: an image capturing device (102); an electronic control unit (104) coupled to the image capturing device (102); and the image capturing device (102) is configured to acquire the image (802); the electronic control unit (104) is configured to apply a machine learning algorithm to the image (802) to identify a curved mirror candidate (800) in the image (802), and apply an anomaly detection algorithm to the identified curved mirror candidate (800) to confirm whether the curved mirror candidate (800) is a curved mirror (814); the curved mirror candidate (800) is confirmed as a curved mirror (814) if it cannot be a traffic sign (820); the anomaly detection algorithm includes: applying a function of an autoencoder to image data associated with the curved mirror candidate (800) to obtain points (808, 810) on a transformation map (812), wherein the points (808, 810) represent the curved mirror candidate (800) on the transformation map (812); confirming the curved mirror candidate (800) as a curved mirror (814) if the points (808, 810) are at a threshold distance beyond a cluster (816, 818) representing a traffic sign (820) on the transformation map (812); and characterized in that.
10. The system according to claim 9, wherein the anomaly detection algorithm is based on Deep Autoencoder Gaussian Mixture Model (DAGMM).
11. A computer-readable storage medium storing instructions for instructing a processing unit of a system to execute a computer-implemented method for detecting a curved mirror (814) in an image (802), the method comprising: Applying a machine learning algorithm to the image (802) to identify a curve mirror candidate (800) in the image (802); Applying an anomaly detection algorithm to the identified curve mirror candidate (800) to check whether the curve mirror candidate (800) is a curve mirror (814); including the curve mirror candidate (800) is confirmed as a curve mirror (814) when it cannot be a traffic sign (820); the anomaly detection algorithm applying the function of an autoencoder to the image data associated with the curve mirror candidate (800) to obtain points (808, 810) on a conversion map (812), where the points (808, 810) represent the curve mirror candidate (800) on the conversion map (812); if the points (808, 810) are beyond a threshold distance from clusters (816, 818) representing traffic signs (820) on the conversion map (812), confirming the curve mirror candidate (800) as a curve mirror (814); including A computer-readable storage medium characterized by the above.