Coarse and Fine Attention Network for Optical Signal Detection and Recognition
The integration of coarse and fine attention modules into vehicle optical signal detection networks addresses the challenges of accurate bounding box localization and category prediction in autonomous driving, significantly improving detection accuracy and noise tolerance.
Patent Information
- Application Number
- JP2023515033
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-09-04
- Filing Date
- 2021-07-21
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2041-07-21
AI Technical Summary
Existing methods for detecting and recognizing vehicle tail optical signals in autonomous driving face challenges due to the difficulty in accurately localizing bounding boxes and predicting categories, especially in noisy and ambiguous environments.
A novel vehicle optical signal detection network incorporating a coarse attention module and a fine attention module, which promote each other and can be integrated into existing object detection networks, to dynamically extract patterns and improve noise tolerance in region proposals.
The proposed coarse and fine attention mechanism enhances the accuracy of bounding box localization and category prediction, overcoming the limitations of previous methods by effectively filtering noise and extracting semantic information from vehicle tail optical signals.
Smart Images

Figure 0007695018000010 
Figure 0007695018000011 
Figure 0007695018000012
Abstract
Description
Technical Field
[0001] The present invention generally relates to a method for detecting and recognizing vehicle optical signals, and more particularly, but not limited thereto, to a system, method, and computer program product for detecting and recognizing the semantics of vehicle tail optical signals for autonomous driving.
Background Art
[0002] The problem of vehicle tail light detection for autonomous driving has attracted attention because it is essential for a vehicle to accurately grasp the driving intention of other vehicles and support quick driving judgments based on detection outputs. The "attention mechanism" has been proven to be useful for enhancing the discriminative power to obtain a richly expressive model. Detecting self-luminous objects is difficult due to the abstract semantic meaning and ambiguous boundaries.
[0003] Other conventional techniques that use complex image processing and employ prior knowledge of specific datasets ignore the common characteristics of self-luminous targets. Since there is a lot of noise that affects the prediction of ambiguous boundaries and the extraction of semantic information, previous methods generate low-quality proposals.
[0004] Therefore, there are limitations to the accuracy of the position of the bounding box, and the predicted categories that depend on the quality of the region proposals are also greatly influenced by confusing information. Therefore, the prior art is not optimal for autonomous driving applications.
Summary of the Invention
[0005] Accordingly, the inventors have identified a need in the art and discovered a novel vehicle optical signal detection network having two components including a coarse attention module and a fine attention module. These two sub-modules promote each other and can be incorporated into various existing object detection networks based on neural networks. Also, they can be learned end-to-end without additional monitoring. Thus, the inventors propose a coarse and fine attention (CFA) mechanism to address the conventional problem of dynamically localizing patterns with a large amount of information corresponding to abstract regions. Specifically, the feature quantities responsible for bounding box localization and category classification are clustered into expert channels by the newly conceived coarse and fine attention module, and patterns are extracted from noisy proposal regions by the fine attention module.
[0006] That is, the present invention systematically learns the feature quantities of a self-luminous target (i.e., the taillight of a vehicle), dynamically extracts patterns with a large amount of information from region proposals, and can increase the noise in the proposal regions. As a result of technological improvements, it includes those that have been put into practical use. Thereby, it is possible to quickly and accurately obtain bounding boxes and predicted categories that were hardly obtained in the prior art. Furthermore, the coarse and fine attention module of the present invention can be integrated into any two-stage detector while maintaining an end-to-end training paradigm, thereby solving the problem that the existing technology is restricted by the specific structure and appearance of data.
[0007] In one exemplary embodiment, the present invention provides a method for detecting and recognizing vehicle optical signals implemented on a computer. The method includes delimiting one or more regions of an image of a vehicle including at least one of a brake light and a signal light generated by a vehicle signal including an illumination unit using a coarse attention module to generate one or more boundary regions; removing noise from the one or more boundary regions using a fine attention module to generate one or more noise-free boundary regions; and identifying the at least one of the brake light and the signal light from the one or more noise-free boundary regions.
[0008] In another exemplary embodiment, the present invention provides a computer program product including a computer-readable storage medium having program instructions embodied therein. The program instructions are executable by a computer and cause the computer to delimit one or more regions of an image of a vehicle including at least one of a brake light and a signal light generated by a vehicle signal including an illumination unit using a coarse attention module to generate one or more boundary regions; remove noise from the one or more boundary regions using a fine attention module to generate one or more noise-free boundary regions; and identify the at least one of the brake light and the signal light from the one or more noise-free boundary regions.
[0009] In another exemplary embodiment, the present invention provides a vehicle optical signal detection and recognition system, the system including a processor and a memory, the memory causing the processor to: demarcate one or more regions of an image of a vehicle including at least one of a brake light and a signal light generated by an automotive signal including an illumination unit using a coarse attention module to generate one or more boundary regions; remove noise from the one or more boundary regions using a fine attention module to generate one or more noise-free boundary regions; and identify the at least one of the brake light and the signal light from the one or more noise-free boundary regions.
[0010] In an optional embodiment, the coarse attention module receives a small feature map from a region of interest (ROI) pooling module, and the fine attention module receives a large feature map from the ROI pooling module.
[0011] In any other optional embodiment, the coarse attention module includes two branches for calculation including an attention score branch and an expected learning score branch. In the attention score learning branch, the coarse attention module converts the small feature map into an original feature vector. In the expected score learning branch, the coarse attention module calculates the average of all previous features of the same category as the channel expected attention score (C-E score) of each category, which is multiplied by the small feature map to obtain a refined feature map 1.
[0012] In another optional embodiment, the fine-grained attention module uses the large feature map and the coarse attention score obtained from the coarse attention module to generate a fine-grained attention map by calculating the average of all channel scores. To obtain the refined feature map 2, the fine-grained attention map is multiplied by the large feature map, which is used to generate an identification region for the prediction of the bounding box and category.
[0013] Other details and embodiments of the present invention are described below so that the contribution of the present invention to the technology can be better understood. Nevertheless, the present invention is not limited in its application to such details, expressions, terms, illustrations, or arrangements, or combinations thereof, as described in this specification or shown in the drawings. Rather, the present invention allows for additional embodiments beyond those described, and can be practiced and implemented in various ways and should not be regarded as limited.
[0014] Thus, those skilled in the art will understand that the underlying concept of this disclosure can be readily utilized as a basis for the design of other structures, methods, and systems for achieving some of the objectives of the present invention. Therefore, it is important that the claims be considered to include such equivalent structures as long as they do not depart from the spirit and scope of the present invention.
[0015] Aspects of the present invention will be better understood from the following detailed description of exemplary embodiments of the present invention with reference to the drawings.
Brief Description of the Drawings
[0016]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9a
Figure 9b
Figure 9c
Figure 9d
Figure 10
Figure 11
Figure 12
DETAILED DESCRIPTION OF THE INVENTION
[0017] Next, the present invention will be described with reference to FIGS. 1 to 12. In FIGS. 1 to 12, like reference numerals refer to like parts throughout. It is emphasized that, in accordance with a general convention, the various features of the drawings are not necessarily to scale. On the contrary, the dimensions of the various features can be arbitrarily enlarged or reduced for clarity.
[0018] Referring to the example depicted in FIG. 1, an embodiment of a vehicle optical signal detection and recognition method 100 according to the present invention can include various steps for using a computing device to identify a brake light and a signal light (i.e., a tail light) of an automobile with a coarse attention module and a fine attention module implemented in a two-stage detector.
[0019] By introducing the example depicted in FIG. 10, one or more computers of a computer system 12 according to an embodiment of the present invention can include a memory 28 having instructions stored in a storage system for performing the steps of FIG. 1.
[0020] One or more embodiments can be implemented in a cloud environment 50 (e.g., FIG. 12), but nevertheless, it is understood that the present invention can be implemented outside of a cloud environment.
[0021] Generally referring to FIGS. 1 to 9d, detection is a basic task in the computer vision community and has been successful in many real-world scenarios including general object detection, road object detection, and face detection. Despite such progress, self-luminous target detection remains an open research problem. Conventional techniques either use complex image processing or employ prior knowledge regarding a specific dataset and ignore the characteristics common to such targets. In the present invention, a new attention mechanism called a coarse-to-fine attention mechanism is described to address the problems underlying the self-luminous target detection task.
[0022] Self-luminous target detection is different from general object detection for at least the following reasons. First, there is more noise in such scenarios (for example, the bounding box is deceived into including light noise from other vehicles and surrounding road lamps, and the signal category classifier will give incorrect predictions). Second, since all ground truths look the same in terms of general contours and appearances, and the semantic information within the bounding box is crucial for making correct predictions, the semantic information within the bounding box is important. Third, due to the large variations in image brightness and contrast, the boundaries of objects may become ambiguous, or separated light signals may be difficult to distinguish.
[0023] Figures 9a to 9d exemplarily show explanatory diagrams of the self-luminous object detection task. Figures 9a and 9b depict the contrast of the optical signal and the fact that the background has a large variance (that is, it is much easier to recognize the signal in 9a than in 9b). Figure 9c shows the ambiguous state of the boundary of the bounding box characterizing the semantic category, and Figure 9d shows the state of the interference noise of the ambient light.
[0024] In two-stage general object detection, region proposals are generated in the first stage, and the exact positions of the bounding box and object category are predicted in the second stage. This pipeline has been shown to be suitable for general object detection tasks, but when the pipeline is used to detect self-luminous targets with ambiguous boundaries and a lot of noise for extracting semantic information, it will lead to a performance degradation. For example, the boundaries of objects in low-light environments are difficult to judge, and the characteristics of traffic lights in different scenarios (such as during rainy days or at night) have large variations. Therefore, there are limitations in performance for general two-stage detectors. Also, the region proposal generator (for example, the region proposal network (RPN) in Faster R-CNN) will generate low-quality proposals due to surrounding interference. Therefore, the accuracy of the position of the bounding box is limited, and the predicted category, which highly depends on the quality of the region proposals, is also greatly influenced by the confused information.
[0025] To address these limitations, there are two ways: (a) learning a better region proposal network that can generate both the proposed positions and foreground predictions more accurately, or (b) dynamically extracting informative patterns from the region proposals and being more tolerant of the noise in the proposed regions.
[0026] As described below, in the present invention, the method 100 and the architecture follow a second efficient way of inserting an attention module for noise filtering and information transmission. An exemplary motivation is to extract features from low-quality rough region proposals by prompting a set of expert extractors within the network, and then to follow fine-grained attention to discover accurate regions in the spatial regions for guiding detection. This new idea can mitigate both of the above problems. On the one hand, the accurate bounding box can be predicted by considering the spatial positions of specific patterns generated by the expert extractor with rough attention. In particular, for structured objects, it is necessary to consider both the structure and the correlations between them to make accurate predictions. And the important information patterns (e.g., the signal semantics of traffic signals) for classifying such self-luminous targets can be effectively integrated using the attention mechanism of the present invention, named the "Coarse-to-Fine Attention mechanism" (CFA) in the present invention, to improve performance, and can be integrated with a two-stage detector for performance improvement.
[0027] When using the CFA module, since accurate sub-parts are detected in high correlation with the core information, it is expected that both the position of the bounding box and the category prediction will be more accurate in the detection task of processing non-rigid objects or the semantic information of objects that need to be extracted from noisy proposed regions.
[0028] Specifically, the coarse attention can be interpreted as an attention mechanism in the channel dimension that encourages the expert channel to extract information with coarse proposals. The modal feature vector serves to distinguish sub-parts that do not have the specific appearance semantics of the object and is generated to represent specific patterns in the basic data in an unsupervised manner. Since the coarse attention bridges between the channel dimension and the spatial dimension, the expert feature extractor is expected to focus on spatial patterns with a specific amount of information. Benefiting from the information transmission to the modal feature vector in the expert channel, fine attention is followed to predict an accurate attention map in the spatial region. In this way, the region proposal network is optimized implicitly using backpropagation to include desirable information (i.e., regions characterizing semantic meaning and boundaries).
[0029] As described below, the Coarse and Fine Attention (CFA) module is detailed and can be directly inserted into Faster R-CNN and other conventional two-stage detectors for the task of self-luminous object detection. A mathematical proof and intuitive inspiration of the proposed CFA mechanism are provided, and its effectiveness is verified. Also, as shown in FIG. 8, experiments on the CIFAR-10 dataset are conducted to prove the correctness of the coarse attention module and report the performance of the attention mechanism on the dataset, significantly exceeding the baseline.
[0030] FIG. 2 provides the backbone model of the present invention. The input image 201 of the vehicle taillight passes through a convolutional neural network (CNN) module 202 for feature extraction and a region proposal network (RPN) module 203 for original region proposal.
[0031] As shown in FIG. 3, the CFA module 301 in the present invention supplies the classification and localization module 205 and provides an output 206 in place of the region of interest (ROI) pooling section 204 of FIG. 2. The CFA includes two combined attention modules: a coarse attention module 301a and a fine attention module 301b. The ROI pooling is supplied to the coarse attention module and the fine attention module. The coarse attention module clusters high-information features in the expert channels as the central feature vector and repeatedly assigns attention weights to each channel. The fine attention module utilizes the information from the coarse attention module and extracts a spatial attention region from the coarse proposals. By generating the central feature vector from the coarse attention, the fine attention is followed. In the present invention, the central feature vector and the modal feature vector are used interchangeably, both meaning the feature pattern responsible for bounding box localization and category classification.
[0032] Coarse attention bridges the connection between the channel dimension and the spatial dimension by considering the high-information patterns underlying the data each channel has. Fine attention fuses the central feature vector to provide a spatial attention map for bounding box regression and class prediction. The CFA can be inserted into any two-stage detector and the entire network can be optimized in an end-to-end manner. An example of the CFA module integrated with Faster R-CNN is shown in FIG. 3. Since the fine attention is responsible for generating the spatial attention map, a large spatial dimension size is desired. By assigning attention weights to each channel and promoting the central features to the channels with high attention weights, (for example, in the expert channels, the pattern corresponding to the self-luminous object of the coarse proposal is extracted from the noise).
[0033] For example, FIG. 4 shows the ROI pooling input in more detail. By using ROI pooling with different kernels on the original feature map from the RPN, two different feature maps, namely, a "large feature map" and a "small feature map", are generated. Since the fine attention module emphasizes refined information extraction, the feature map used in the fine attention module is, for example, twice the size of the coarse attention module (e.g., H2 = 2H3, W2 = 2W3). The present invention can generate a plurality of CAs and FAs with different scales of the feature map of the ROI. Therefore, the size of the feature map is not limited.
[0034] Next, as shown in FIG. 4, the CFA module extracts a pattern with a large amount of information from the original region proposal in order to predict a more accurate boundary and category. With the help of the CFA module, the accuracy of the output of the vehicle optical signal detection will be significantly improved. As described above, the CFA module composed of the coarse attention module and the fine attention module can be incorporated into any two-stage detector for self-luminous object detection.
[0035] FIG. 5 shows a coarse attention module (CA) having an input of a small feature map from ROI pooling. The CA module uses a set of expert feature extractors to extract an expressive modal vector that plays a role in locating separated spatial patterns. In order to filter out interference noise and assign attention weights to the feature learning channels, the CA module clusters the features with a large amount of information into the expert channels as the central feature vector and repeatedly assigns weights to each channel.
[0036] Since incorrect region proposals contain confusing information, the present invention aims to train a set of feature extractors that carry desirable information. However, directly training a high-dimensional neural network would non-selectively consider noisy patterns outside the informative sub-regions. For this motivation, the present invention discovers desirable parts within a rough region by transmitting important features to a specific channel and learning the modal feature vectors of each part.
[0037] It is recommended that the expert feature extractor focus on specific patterns by forcing the generation of similar modal vectors in the same category containing similar semantic information. The spatial position is expected to be useful for localizing the bounding box symmetrically (i.e., in a repetitive manner). To update the modal vector and transmit useful information bearing such patterns in a dynamic way, backpropagation can transmit correlated distinguishable features to the clustering of specific patterns in a manner similar to the mixture model. To extract informative features from the rough proposed region, the motivation is to optimize the set of expert feature extractors by assigning channels with high attention weights, and to regularize the informative features for clustering in the corresponding expert channels by prompting the expert to extract the central features of each class.
[0038] Each channel of the feature quantity is aware of the position, and all instances of the same category share the central feature quantity using multiple complete pattern extractors. The above characteristics do not apply to the feature quantities lying across each channel, but it encourages clustering important information into multiple channels and implicitly separating feature quantities in a sparse representation by backpropagation. That is, the CA module adds regularization by dynamically weighting the importance of channels to model the desired modal vector for each class. Channels with high attention scores embed specific patterns while avoiding the integration of noisy feature quantities.
[0039] The CA module implements the above idea of generating candidate modal feature quantities for each category across all channels by averaging the samples previously trained in a specific channel. Next, to regularize the clustering of discriminative feature quantities for a set of experts, the weights of the importance of each channel are dynamically assigned by calculating the similarity to the estimated modal feature quantities of the same category and the dissimilarity to other categories. The set of experts shares similarities with the mixture model to repeatedly update the modal vector and weight assignment by the expectation-maximization (EM) algorithm. As the training time lengthens, the central feature quantities of important patterns can be obtained by channels with high attention scores optimized in an end-to-end manner.
[0040] As shown in FIG. 5, to learn the expert feature quantity extractor, the central feature quantity of each class is approximated by averaging all the feature quantities generated in the same channel of the same class. The score in each channel of the current sample is necessary to ensure important feature quantities extracted by channels with high scores.
[0041] By prompting the feature quantity in the expert channel to share similarity with the central feature quantity of the same class while distinguishing it from the central feature quantities of different classes, the channel score S of sample i having category c can be calculated as follows.
[0042] TIFF0007695018000001.tif13133
[0043] Here, F i is the feature quantity of the current sample i, and A c is the average central feature quantity of class c calculated by averaging the central feature quantity from the buffer and F i , C is the total number of classes, and 1 / C - 1 controls the weight of the contribution rate for samples with different labels. The above processing shares similarity with the EM algorithm. First, as the "expectation step", the weighted channel attention score before optimization is generated, and then, as the "maximization step", the feature quantity with a large amount of information is transmitted to the channel with a high score to optimize the parameters of the RPN. The network is optimized to adapt to the channel attention score so that the training process is stable and an average feature quantity buffer does not occur during testing.
[0044] The CA module includes two branches: an attention score branch (input to the FA module described later) and an expectation score learning branch.
[0045] In the attention score learning branch, the rough attention module converts the small feature map into a value called the "original feature vector" through global average pooling (GAP), and uses this to calculate the rough attention score (C-A score) of the feature map by two fully connected layers (FC). The C-A score is used as the input to the fine attention module.
[0046] In the expected score learning branch with 8 categories and 512 channels, the coarse attention module calculates the average of all previous feature maps of the same category as the channel expected attention score (C-E score) for each category, and it multiplies with a small feature map to obtain the refined feature map 1. The refined feature map 1 is output to the classification and localization modules described later.
[0047] Next, how the CA module operates mathematically is shown. Starting with global average pooling and generating a class activation map (CAM) for classifying a specific category, distinguishable image regions are depicted. More specifically, the feature of sample I with class c after GAP is as follows.
[0048] TIFF0007695018000002.tif13138
[0049] Here, d represents the index of the channel dimension, and f d (x,y) represents the feature map at channel d and position (x,y). And following the FC layer, the original classification output score O org is calculated.
[0050] TIFF0007695018000003.tif10138
[0051] Here, W d is the weight of the FC layer for class c in the channel dimension d. In the proposed coarse attention mechanism, an attention weight S is assigned to each channel as in Equation (1), and the corresponding output score is as follows.
[0052] TIFF0007695018000004.tif11147
[0053] TIFF0007695018000005.tif7162
[0054] TIFF0007695018000006.tif34146
[0055] TIFF0007695018000007.tif53168
[0056] TIFF0007695018000008.tif70168
[0057] Referring to FIG. 6, in order to refine position determination, the fine attention (FA) module generates a refined feature map to localize the exact discrimination region for the accurate prediction of the bounding box and category. By calculating the average of all channel scores, the large feature map obtained from the previous CA module and the coarse attention score are adopted to generate a fine attention map. Then, the fine attention map is multiplied by the large feature map to obtain the refined feature map 2, which is used to generate the discrimination region for the prediction of the bounding box and category. The refined feature map 2 is input to the classification and position determination modules described later.
[0058] That is, the coarse attention constructs the connection between the channel dimension and the spatial dimension by clustering distinguishable spatial features into a set of channels. However, in order to better understand the semantic meaning, for example, in order to more accurately predict the bounding box and category, a fine attention module is required to localize the exact spatial region responsible for the semantic category and the ground truth bounding box. Since the present invention already has a plurality of experts of the feature extractor that focus on specific central features of interest (i.e., via the CA module), the present invention can generate a spatial attention mask by the weighted sum of the spatial regions on which each expert focuses to give attention including the sub-regions that interpret the positions of the semantic category and the bounding box. In fact, the fine attention module generates a spatial attention map to find the discrimination part responsible for the prediction of the bounding box and the semantic category.
[0059] As described above, the architecture of the fine-grained attention module is shown in FIG. 6. The input feature map is generated from the RPN via ROI pooling. Next, to generate the spatial attention mask, the predicted channel attention scores obtained from the previous CA module are adopted, and the feature map is refined accordingly.
[0060] TIFF0007695018000009.tif10151
[0061] Here, f(x, y) is the input feature, d is the channel index, and S d is the channel attention score in channel d generated by the coarse attention module, and σ is the sigmoid activation function.
[0062] As shown in FIG. 7, in the classification and localization module 205, the two refined feature maps from the coarse attention module and the fine-grained attention module are converted into values by the fully connected layer and can be used to calculate the coordinates of the bounding box and the category of the object in the bounding box.
[0063] The total loss of the self-luminous object detection task is the weighted sum of the classification loss and the regression loss. For the classification loss, a multi-task loss function and the non-maximum suppression method (NMS) are used to calculate the classification loss. For the regression loss, a residual fitting method for calculating the regression loss is used. Specifically, the deviation between the actual value of the bounding box and the ground truth of the bounding box is calculated by regression learning.
[0064] Turning to FIG. 1, FIG. 1 illustratively shows the method flow of the coarse and fine attention mechanism of the present invention for the task of self-luminous object detection. The CFA is designed to address the issue of characterizing objects including bounding boxes and semantic meanings, and alleviates problems caused by noise in the rough proposal regions (for example, when surrounding road lights are also included in the proposal region, it causes confusion).
[0065] In step 101, the present invention includes receiving, by a computing device, an image of an automobile including brake lights and / or signal lights (i.e., tail lights include brake lights and signal lights) generated by the signals of the automobile.
[0066] In step 102, the present invention includes bounding, by a computing device using a coarse attention module, one or more regions of the image including the lighting part to generate one or more boundary regions.
[0067] In step 103, the present invention includes removing noise from one or more boundary regions by a computing device using a fine attention module to generate one or more noise-free boundary regions. The output from the coarse attention module is utilized in the fine attention module.
[0068] In step 104, the present invention includes identifying, by a computing device, brake lights and / or signal lights or both from one or more noise-free boundary regions.
[0069] Accordingly, the present invention includes a coarse attention module that filters out interference noise in the environment and assigns attention weights to all feature learning channels, which is useful for an end-to-end learning network by combining a fine attention module for localizing an accurate discrimination region for the next prediction of the bounding box and category, the accurate prediction of the bounding box and category, and a new coarse-fine attention module for classical object detection network and self-luminous object detection.
[0070] In fact, as shown in FIG. 8, the inventors adopted the Vehicle Light Signal (VLS) dataset for result analysis due to the issues of localizing optical signals and extracting semantic information. The VLS dataset includes four general vehicle behaviors: forward driving, braking, left turning, and right turning. Each behavior signal is classified into two scenarios: day and night because the lighting signals are not the same during the day and at night. The data is collected from a drive recorder by evenly sampling 15 frames from a 15-minute video. The VLS dataset contains 7,720 images, 8 categories, and 10,571 instances. The bounding boxes are classified into forward driving, braking, left turning, and right turning during the day and at night, respectively. In the experiment, 60% of the samples are randomly selected as training data, 20% as validation data, and 20% as test data, and this ratio can be fixed for different models in other experiments.
[0071] The model is implemented using, for example, Caffe, a deep learning framework. The model is trained on a single NVIDIA GTX 1080Ti. In all models, the initial learning rate is set to 10 -3 and the momentum is set to 0.9, and the weight decay is 5×10 -4 . In the task, the input to the system is the original image, and the output includes the position of the vehicle and the classification result of the vehicle light signal in this image. Thus, as a performance evaluation criterion, the mean average precision (MAP) of object detection, which is the average of the average precision (AP) of each category, is used.
[0072] Most of the state-of-the-art techniques used in the optical signal detection task employ general object detectors. Using prior research, first, the performance of Faster R-CNN, a general two-stage general object detector, is evaluated with different backbones, and the performance when the coarse-to-fine attention mechanism of the present invention is integrated is evaluated. The results are shown in FIG. 8. When trained with CFA using the same backbone network, it can be seen that the detector outperforms the original model, especially when the performance of the backbone network is low. That is, the CFA module of the present invention described herein has the ability to extract informative features from low-quality proposals with more noise or interference.
[0073] <Exemplary embodiments using a cloud computing environment> This detailed description includes exemplary embodiments of the present invention in a cloud computing environment, but the implementation of the teachings described herein is not limited to a cloud computing environment. Rather, embodiments of the present invention can be implemented with any other type of computer environment now known or developed in the future.
[0074] Cloud computing is a service delivery model that enables convenient and on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage devices, applications, virtual machines, and services), where the resources can be rapidly provisioned and released with minimal management effort or interaction with a service provider. This cloud model may include at least five characteristics, at least three service models, and at least four implementation models.
[0075] The characteristics are as follows.
[0076] On-Demand Self-Service: Cloud consumers can unilaterally provision computing capabilities such as server time and network storage automatically as needed, without the need for human interaction with a service provider.
[0077] Broad Network Access: Computing capabilities are available over the network and can be accessed through standard mechanisms, thereby facilitating use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, PDAs).
[0078] Resource Pooling: The provider's computing resources are pooled and offered to multiple consumers using a multi-tenant model. Various physical and virtual resources are dynamically assigned and reassigned according to demand. In general, consumers have a sense of location independence since they do not manage or know the exact location of the provided resources. However, consumers may be able to specify location at a higher level of abstraction (e.g., country, state, data center).
[0079] Rapid Elasticity: Computing capabilities can be provisioned quickly and elastically, enabling them to scale out automatically in some cases and immediately, and to be released quickly and scale in immediately. To the consumer, the available computing capabilities often appear to be unlimited and can be purchased in any quantity at any time.
[0080] Measured Service: Cloud systems leverage measurement capabilities at an appropriate level of abstraction for the type of service (e.g., storage, processing, bandwidth, active user accounts) to automatically control and optimize resource use. It can monitor, control, and report resource usage to provide transparency to both the provider and the consumer of the utilized service.
[0081] The service model is as follows.
[0082] Software as a Service (SaaS): The function provided to the consumer is that the consumer can use the provider's application running on the cloud infrastructure. The application can be accessed from various client lines via a client interface such as a web browser (e.g., webmail). The consumer does not manage or control the underlying cloud infrastructure, including the network, server, operating system, storage, and even individual application functions. However, this does not apply to limited settings of user-specific application configurations.
[0083] Platform as a Service (PaaS): The function provided to the consumer is to deploy the application created or obtained by the consumer to the cloud infrastructure using the programming languages and tools supported by the provider. The consumer does not manage or control the underlying cloud infrastructure, including the network, server, operating system, and storage, but can control the deployed application and, in some cases, also control the configuration of its hosting environment.
[0084] Infrastructure as a Service (IaaS): The function provided to the consumer is to prepare processors, storage, networks, and other basic computing resources that allow the consumer to deploy and run any software that may include an operating system and applications. The consumer does not manage or control the underlying cloud infrastructure, but can control the operating system, storage, and deployed applications, and in some cases, can partially control some network components (e.g., host firewall).
[0085] The deployment model is as follows.
[0086] Private Cloud: This cloud infrastructure is operated exclusively for a specific organization. This cloud infrastructure can be managed by the organization or a third party and can exist on-premises or off-premises.
[0087] Community Cloud: This cloud infrastructure is shared by multiple organizations and supports a specific community with common concerns (e.g., mission, security requirements, policies, and compliance). This cloud infrastructure can be managed by the organization or a third party and can exist on-premises or off-premises.
[0088] Public Cloud: This cloud infrastructure is provided to an unspecified number of people or large industry groups and is owned by an organization that sells cloud services.
[0089] Hybrid Cloud: This cloud infrastructure combines two or more cloud models (private, community, or public). It retains the entities specific to each model but is bound by standard or individual technologies to achieve data and application portability (e.g., cloud bursting for load distribution between clouds).
[0090] The cloud computing environment is a service-oriented environment that emphasizes statelessness, low coupling, modularity, and semantic interoperability. At the core of cloud computing is an infrastructure that includes a network of interconnected nodes.
[0091] Referring now to FIG. 10, a schematic diagram of an example of a cloud computing node is shown. Cloud computing node 10 is merely an example of a suitable node and is not intended to suggest any limitation as to the scope of use or functionality of the embodiments of the invention described herein. Nevertheless, cloud computing node 10 is capable of implementing, executing, or both, any of the functionality described herein.
[0092] Cloud computing node 10 is depicted as a computer system / server 12, but it is understood to be operable in numerous other general purpose or special purpose computing system environments or configurations. Examples of well-known computing systems, environments, or configurations, or combinations thereof, that may be suitable for use with computer system / server 12 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, hand-held or laptop circuitry, multiprocessor systems, microprocessor-based systems, set top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments including any of the above systems or circuitry.
[0093] Computer system / server 12 may be described in the general context of computer system executable instructions, such as program modules, being executed by a computer system. Generally, program modules can include routines, programs, objects, components, logic, data structures, etc. that perform particular tasks or implement particular abstract data types. Computer system / server 12 may be implemented in a distributed cloud computing environment where tasks are performed by remote processing circuits linked through a communications network. In a distributed cloud computing environment, program modules may be located in both local and remote computer system storage media including memory storage circuits.
[0094] Referring now to FIG. 10, computer system / server 12 is shown in the form of a general-purpose computing circuit. The components of computer system / server 12 can include, but are not limited to, one or more processors or processing units 16, a system memory 28, and a bus 18 that couples various system components including system memory 28 to the processor 16.
[0095] Bus 18 represents any one or more of a plurality of types of bus structures including a memory bus or memory controller using any of a variety of bus architectures, a peripheral bus, an accelerated graphics port, and a processor or local bus. By way of example, such architectures can include Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Extended ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, Peripheral Component Interconnect (PCI) bus, and PCI Express (PCIe) bus.
[0096] The computer system / server 12 generally includes various computer system-readable media. Such media can be any available media accessible by the computer system / server 12 and can include both volatile and non-volatile media, as well as both removable and non-removable media.
[0097] System memory 28 can include computer system-readable media as volatile memory, such as RAM 30 or cache memory 32 or both. The computer system / server 12 can further include other removable / non-removable computer system storage media and volatile / non-volatile computer system storage media. As an example, storage system 34 can be provided for reading and writing to a non-removable non-volatile magnetic medium (not shown, generally referred to as a "hard drive"). Also, although not shown, a magnetic disk drive for reading and writing to a removable non-volatile magnetic disk (e.g., a floppy disk), and an optical disk drive for reading and writing to a removable non-volatile optical disk (such as a CD-ROM, DVD-ROM, or other optical media) can be provided. In these examples, each can be connected to bus 18 by one or more data media interfaces. As further illustrated and described below, memory 28 can include a computer program product that stores one or more program modules having computer-readable instructions configured to execute one or more functions of the present invention.
[0098] As an example, a program / utility 40 having a set (at least one) of program modules 42 can be stored in memory 28, similar to an operating system, one or more application programs, other program modules, and program data. Each of the operating system, one or more application programs, other program modules, and program data, or some combination thereof, can include an implementation form of a network environment. In some embodiments, program modules 42 are generally adapted to perform one or more functions or methods of the present invention, or both.
[0099] Further, computer system / server 12 can communicate with one or more external devices 14, such as other peripheral devices like a keyboard, pointing circuitry, display 24, etc., and one or more components that facilitate interaction with computer system / server 12. Such communication can occur via an input / output (I / O) interface 22, or any circuitry (such as a network card, modem, etc.) that enables computer system / server 12 to communicate with one or more other computing circuits, or both. Additionally, computer system / server 12 can communicate with one or more networks (such as a local area network (LAN), a general wide area network (WAN), or a public network (e.g., the Internet), or combinations thereof) via network adapter 20. As shown, network adapter 20 can communicate with other components of computer system / server 12 via bus 18. It should be understood that, although not shown, other hardware components or software components, or both, can be used in conjunction with computer system / server 12. Examples of these include microcode, circuit drivers, redundancy processing units, external disk drive arrays, RAID systems, tape drives, data archive storage systems, etc.
[0100] Referring now to FIG. 11, an exemplary cloud computing environment 50 is depicted. As shown, the cloud computing environment 50 includes one or more cloud computing nodes 10. In contrast, local computer circuitry used by cloud consumers (e.g., a PDA or cellular phone 54A, a desktop computer 54B, a laptop computer 54C, or an automotive computer system 54N or combinations thereof, etc.) can communicate. The nodes 10 can communicate with each other. The nodes 10 can be physically or virtually grouped (not shown) in one or more networks, such as, for example, the private, community, public, or hybrid clouds described above or combinations thereof. Thereby, the cloud computing environment 50 can provide infrastructure, platform, software, or combinations thereof as a service, and cloud consumers do not need to maintain resources on local computer circuitry. It should be understood that the types of computer circuitry 54A - N shown in FIG. 11 are merely exemplary, and the computing nodes 10 and the cloud computing environment 50 can communicate with any type of electronic circuitry via any type of network or network addressable connection (e.g., using a web browser) or both.
[0101] Referring now to FIG. 12, an exemplary set of functional abstractions provided by the cloud computing environment 50 (FIG. 11) is shown. It should be understood in advance that the components, layers, and functions shown in FIG. 12 are merely exemplary and embodiments of the present invention are not limited thereto. As illustrated, the following layers and corresponding functions are provided.
[0102] The hardware and software layer 60 includes hardware components and software components. Examples of hardware components include mainframe 61, a server 62 based on a reduced instruction set computer (RISC) architecture, server 63, blade server 64, memory circuit 65, and network and network components 66. In some embodiments, the software components include network application server software 67 and database software 68.
[0103] The virtualization layer 70 provides an abstraction layer. From this layer, for example, the following virtual entities can be provided: virtual server 71, virtual storage 72, virtual network 73 including a virtual private network, virtual applications and operating systems 74, and virtual client 75.
[0104] As an example, the management layer 80 can provide the following functions. Resource provisioning 81 enables the dynamic procurement of computing resources and other resources used to execute tasks within a cloud computing environment. Metering and pricing 82 enables cost tracking when resources are utilized within a cloud computing environment and billing or invoicing for the consumption of these resources. As an example, these resources may include licenses for application software. Security enables not only the protection of data and other resources but also the identification and authentication of cloud consumers and tasks. The user portal 83 provides access to the cloud computing environment for consumers and system administrators. Service level management 84 enables the allocation and management of cloud computing resources such that the required service level is met. Service quality assurance (SLA) planning and fulfillment 85 enables the advance arrangement and procurement of cloud computing resources expected to be needed in the future according to the SLA.
[0105] The workload layer 90 provides examples of the functions available in the cloud computing environment. Examples of the workloads and functions that can be provided from this layer include mapping and navigation 91, software development and lifecycle management 92, delivery of virtual classroom education 93, data analysis processing 94, transaction processing 95, and the vehicle optical signal detection and recognition method 100 according to the present invention.
[0106] The present invention can be a system, method, computer program product, or a combination thereof integrated at any possible technical detail level. The computer program product may include a computer-readable storage medium storing computer-readable program instructions for causing a processor to execute aspects of the present invention.
[0107] The computer-readable storage medium can be a tangible device that can hold and store instructions used by an instruction execution device. As an example, the computer-readable storage medium can be an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or a suitable combination thereof. A more specific example of the computer-readable storage medium can be a portable computer diskette, a hard disk, a RAM, a ROM, an EPROM (or flash memory), an SRAM, a CD-ROM, a DVD, a memory stick, a floppy disk, a punch card, a mechanically encoded device that records instructions in a raised structure within a groove, and suitable combinations thereof. The computer-readable storage medium used herein should not be construed as a transient signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted via a wire.
[0108] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices or to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, or a wireless network network or a combination thereof). The network is composed of copper wire transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers or a combination thereof. The network adapter card or network interface of each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in a computer-readable storage medium within each respective computing / processing device.
[0109] The computer-readable program instructions for carrying out the operations of the present invention may be source code or object code written in any combination of one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or object-oriented programming languages such as Smalltalk, C++, and procedural programming languages such as the "C" programming language and similar programming languages. The computer-readable program instructions may be executable as a stand-alone software package, entirely on the user's computer, or partially on the user's computer. Alternatively, it may be executable partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize for carrying out aspects of the present invention.
[0110] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0111] These computer-readable program instructions can be provided to a general purpose computer, a special purpose computer processor, or other programmable data processing apparatus to generate a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus implement the functions / acts specified in one or more blocks of a flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can be connected to a computer, a programmable data processing apparatus, or other device that functions in a particular manner, such that the computer-readable program instructions stored therein constitute one of the manufactured articles including instructions for implementing the aspects of the functions / acts specified in one or more blocks of a flowchart and / or block diagram.
[0112] Like instructions that execute functions / acts specified in one or more blocks of a flowchart and / or block diagram on a computer, other programmable apparatus, or other device, the computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to execute a series of operational steps on the computer, other programmable apparatus, or other device to generate a computer-implemented process.
[0113] The flowcharts and block diagrams in the figures illustrate the configuration, functionality, and operation of implementations executable by systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions, which constitutes one or more executable instructions for implementing the specified logical function. In some alternative embodiments, the functions shown in the blocks may occur out of the order shown in the figures. For example, two blocks shown in succession may, in fact, be executed substantially simultaneously, or the blocks may be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams or flowchart diagrams, or combinations of blocks of the block diagrams or flowchart diagrams or both, can be implemented by a special purpose hardware-based system that performs the specified functions or operations, or combinations of special purpose hardware and computer instructions.
[0114] The description of the various embodiments of the present invention is presented for purposes of illustration, but is not intended to be exhaustive or limited to the disclosed embodiments. It will be apparent to those skilled in the art that many modifications and variations are possible without departing from the scope and spirit of the disclosed embodiments. The terms used herein are chosen to best explain the principles of the embodiments, the practical application to the technology found in the marketplace, or the technical improvement, or to enable others skilled in the art to understand the embodiments disclosed herein.
[0115] Furthermore, the applicant's intention is to include equivalents of all elements of all claims, and no amendment of any claim of this application should be construed as a waiver of the benefits or rights with respect to the equivalents of any element or feature of the amended claim.
Claims
1. A method for detecting and recognizing vehicle optical signals implemented on a computer, the method comprising: Using a coarse attention module to demarcate one or more regions of an image of a vehicle including at least one of a brake light and a signal light generated by an automotive signal including an illumination unit to generate one or more boundary regions; Using a fine attention module to remove noise from the one or more boundary regions to generate one or more noise-free boundary regions; Identifying the at least one of the brake light and the signal light from the one or more noise-free boundary regions; A method comprising the above.
2. The method according to claim 1, wherein the coarse attention module and the fine attention module use a neural network.
3. The method according to claim 1, wherein the coarse attention module extracts features from low-quality coarse region proposals by prompting a set of expert extractors in the network, and then fine attention follows in the fine attention module to find accurate regions in the spatial region for guiding detection.
4. The method according to claim 3, wherein the bounding box is predicted by the coarse attention module based on the spatial position of a specific pattern generated by the expert extractor.
5. The coarse attention module performs coarse attention in the channel dimension by prompting an expert channel to extract information with a coarse proposal, The modal feature vector is generated by the coarse attention module to represent a specific pattern in the basic data in an unsupervised manner. The method according to claim 1.
6. The coarse-grained attention (CFA) module includes the coarse attention module and the fine attention module, and the method according to claim 1.
7. The coarse attention module clusters the features with a large amount of information as the central feature vector into the expert channels, and repeatedly assigns attention weights to each channel. The fine attention module utilizes the information from the coarse attention module and extracts the spatial attention regions in the coarse proposals. The method according to claim 1.
8. The coarse attention module receives a small feature map from the region of interest (ROI) pooling module. The fine attention module receives a large feature map from the ROI pooling module. The method according to claim 1.
9. The large feature map is twice the size of the small feature map, and the method according to claim 8.
10. The coarse attention module extracts a modal vector with rich expressiveness from the small feature map. The coarse attention module filters out interference noise, clusters the features with a large amount of information as the central feature vector into the expert channels, and assigns attention weights to the feature learning channels by repeatedly assigning weights to each channel. The method according to claim 8.
11. The coarse attention module includes two branches for calculation including an attention score learning branch and an expected score learning branch. In the attention score learning branch, the coarse attention module converts the small feature map into the original feature vector. In the expected score learning branch, the coarse attention module calculates the average of all previous feature quantities of the same category as the channel expected attention score (C-E score) for each category, which is multiplied by the small feature quantity map to obtain the refined feature quantity map 1. The method according to claim 8.
12. The fine attention module generates a refined feature quantity map to localize an accurate discrimination region for the prediction of the bounding box and category. The method according to claim 8.
13. The fine attention module uses the large feature quantity map and the coarse attention score obtained from the coarse attention module to generate a fine attention map by calculating the average of all channel scores. To obtain the refined feature quantity map 2, the fine attention map is multiplied by the large feature quantity map, which is used to generate a discrimination region for the prediction of the bounding box and category. The method according to claim 11.
14. The classification and localization module that receives the refined feature quantity map 1 from the coarse attention module and the refined feature quantity map 2 from the fine attention module, and converts the refined feature quantity map 1 and the refined feature quantity map 2 into values by a fully connected layer, which is used to calculate the coordinates of the bounding box and the category of the object in the bounding box. The method according to claim 13.
15. The coarse attention module processes the small feature quantity map to output the refined feature quantity map 1 and the coarse attention score (C-A score). The fine attention module processes the large feature quantity map and the coarse attention score (C-A score) to output the refined feature quantity map 2. The method according to claim 1.
16. The method according to claim 15, wherein the refined feature maps 1 and 2 are processed by a classification and localization module to calculate the coordinates of the bounding box and the category of the object in the bounding box.
17. The method according to claim 1, implemented in a cloud computing environment.
18. A computer program product, comprising a computer-readable storage medium having program instructions implemented therein, the program instructions being executable by a computer, and causing the computer to demarcate one or more regions of an image of a vehicle including at least one of a brake light and a signal light generated by an automotive signal including an illumination unit using a coarse attention module to generate one or more boundary regions; remove noise from the one or more boundary regions using a fine attention module to generate one or more noise-free boundary regions; identify the at least one of the brake light and the signal light from the one or more noise-free boundary regions; A computer program product that causes the above to be executed.
19. A vehicle optical signal detection and recognition system, the system comprising a processor; a memory, the memory causing the processor to demarcate one or more regions of an image of a vehicle including at least one of a brake light and a signal light generated by an automotive signal including an illumination unit using a coarse attention module to generate one or more boundary regions; To generate one or more noise-free boundary regions, removing noise from the one or more boundary regions using a fine attention module, identifying at least one of the brake light and the signal light from the one or more noise-free boundary regions, A system storing instructions for causing the above to be executed.
20. The system according to claim 19, implemented in a cloud computing environment.
Citation Information
Patent Citations
Object detection device and program
JP2013232023A