Method for object detection in image data
By processing sequences with a multi-level neural network and random dropout technology, the bounding box proposals are refined step by step, which solves the problem of poor proposal quality in the G-FSOD framework and improves the recognition accuracy and stability of new category detection.
Patent Information
- Application Number
- CN202510326424.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-03-19
- Filing Date
- 2025-03-19
- Publication Date
- 2025-09-19
AI Technical Summary
The existing G-FSOD framework suffers from poor proposal quality, scale distribution mismatch and insufficient training data in object detection, which leads to degraded performance in detecting new categories.
A multi-stage neural network is used to process the sequence, combined with random activation and attention modules to refine the bounding box proposals step by step, handle accidental and cognitive uncertainty through random activation, and improve the proposal quality using feature pyramid and keypoint representation.
Improved object detection performance, especially in detecting new categories, reducing forgetting and improving recognition accuracy and stability.
Smart Images

Figure CN120673107A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to methods for object detection in image data. Background Art
[0002] Object detection (especially in images) is a common task in the context of autonomous control of robotic devices such as robotic arms and autonomous vehicles. For example, a controller for a robotic arm should be able to identify the object that the robotic arm is to pick up (e.g., among multiple different objects), and an autonomous vehicle must be able to identify other vehicles, pedestrians, and fixed obstacles.
[0003] One method for object detection in images, particularly for "new" classes with few training examples (in addition to "base classes" with many training examples), is G-FSOD (Generalized few-shot object detection). The G-FSOD framework is typically based on a two-stage Faster-R-CNN (Region-based Convolutional Neural Network) model. One of the biggest bottlenecks in this type of object recognition is often the poor quality of the object proposals generated and processed in the corresponding machine learning models. The quality of the proposals in G-FSOD deteriorates with the introduction of new classes. The main reasons for this can be: (1) the number of training data (training examples) for these new categories is small, and the training data is therefore generally unrepresentative of the actual class distribution; (2) the model may treat the new categories as background due to the small IoU (Intersection of Union) with the ground truth bounding box (i.e., the ground truth information about the bounding box present in the training data); and (3) the scale distribution (Skalenverteilung) of the new objects is different from the scale distribution in the basic training data. In addition, the small number of training examples for the new categories leads to higher epistemic uncertainty, because the true data distribution is not fully captured (erfassen), which can cause the machine model to overfit or underfit the data. Summary of the Invention
[0004] Therefore, it would be desirable to be able to implement methods for object detection in an improved manner, especially in the G-FSOD framework.
[0005] According to various embodiments, there is provided a method for object detection in image data, comprising:
[0006] Extract features from image data (e.g., determining feature maps at different resolutions, e.g., using a neural convolutional network);
[0007] Determine one or more proposals for bounding boxes for the corresponding object based on the extracted features;
[0008] Correcting the bounding box by a sequence of a plurality of processing stages, wherein the processing stages each include a neural network and each receive one or more bounding box proposals as input and, in a pass through the processing stages, determining a corresponding bounding box correction for each input bounding box proposal, wherein for each processing stage;
[0009] o for each input bounding box proposal, determining multiple bounding box corrections (e.g., offsets to bounding box description data, or also completely new description data for a bounding box, wherein a bounding box is specified, for example, by the position of a corner (e.g., the top left corner), its height, and its width) by performing multiple passes on the input one or more bounding box proposals, wherein the input bounding box proposals differ due to different deactivations of neurons of the neural network (i.e., random dropouts are performed during the neural network passes, so that the passes differ),
[0010] o For each input bounding box proposal, determine the output bounding box correction by averaging the bounding box corrections determined for the input bounding box proposals in these passes.
[0011] The above approach enables the consideration of epistemic uncertainty in the G-FSOD framework, thereby improving the performance of object detection. Aleatoric uncertainty can also be considered.
[0012] The rectified bounding box proposals (possibly with associated classifications) can be the result of object detection or further processed (e.g., each rectified bounding box can be further segmented to separate the object from the background).
[0013] These refinement stages can also output a classification for each bounding box correction, i.e., one or more classification values ("scores," e.g., logits) that predict the class of the object contained in the corresponding bounding box. Furthermore, these refinement stages can also output the bounding box correction and, if applicable, the uncertainty (e.g., divergence or variance) of the classification. From this, a probability distribution over the bounding box position or classification can be formed.
[0014] Examples of various implementations are given below.
[0015] Embodiment 1 is the method for object detection as described above.
[0016] Embodiment 2 is a method according to embodiment 1, wherein each processing stage determines a relevant classification for each bounding box correction in each pass, and determines a classification for each input bounding box proposal by averaging the classifications determined for the input bounding box proposals in these passes.
[0017] Therefore, this further improves object detection (including classification) in light of epistemic uncertainty in classification.
[0018] Embodiment 3 is a method according to embodiment 1 or 2, wherein each processing stage further obtains the extracted features as input.
[0019] Each processing stage can therefore exploit the extracted features, which improves the quality of object detection.
[0020] Embodiment 4 is a method according to one of embodiments 1 to 3, comprising: training at least one of the processing stages, which outputs, for each input bounding box proposal, explanatory data of a bounding box probability distribution in view of the position of the corresponding bounding box; determining bounding box samples by sampling multiple times from the bounding box probability distribution; determining the loss between the bounding box samples and the bounding box ground truth information (i.e., for example, determining the loss for each sample (with respect to (e.g., the nearest) ground truth bounding box) and averaging or summing the losses); and training at least one processing stage to reduce the loss (i.e., adapting parameter values, typically weights, of the processing stage in a direction that reduces the loss, i.e., according to the gradient of the loss, typically using backpropagation).
[0021] Therefore, the aleatory uncertainty about the bounding boxes is taken into account during training, which further improves object detection.
[0022] Embodiment 5 is a method according to one of embodiments 1 to 4, comprising: training at least one of the processing stages, which outputs, for each input bounding box proposal, explanatory data of a classification probability distribution in view of the category of the object contained in the corresponding bounding box; determining classified samples by sampling multiple times from the classification probability distribution; determining the loss between the classified samples and the classification reference true value information (i.e., for example, determining the loss for each sample and averaging the losses or summing the losses); and training at least one processing stage to reduce the loss (i.e., in the direction of reducing the loss, i.e., by adapting the parameter values, typically weights, of the processing stage according to the gradient of the loss, typically by using backpropagation).
[0023] Therefore, aleatory uncertainty about classification is taken into account during training, which further improves object detection.
[0024] Embodiment 6 is a method according to one of embodiments 1 to 5, comprising: determining the one or more proposals for the bounding box based on the extracted features through a region proposal network based on key points (also expressed herein by the common English term “Keypoint”).
[0025] Unlike anchor-based region proposal networks (RPNs) that typically provide fixed-size “anchors,” bond-based RPNs can provide more accurate spatial information and improve the alignment of extracted features with proposals, which improves classification.
[0026] Embodiment 7 is a method according to one of embodiments 1 to 6, comprising: training processing stages, wherein each processing stage comprises an attention block, such as a CBAM (Convolutional Block Attention Module), during training, which processes features derived from the extracted features (e.g., by RoI pooling) and assigned to corresponding one or more bounding box proposals, wherein the processing stage determines a bounding box correction (and possibly a classification) by using the processed features.
[0027] Embodiment 8 is a method for controlling a robotic device, comprising: capturing image data of an environment of the robotic device; detecting (e.g., locating and classifying) an object in the image data by a method according to one of embodiments 1 to 7; and controlling the robotic device based on the detection of the object in the image data (i.e., in particular, whether an object of a particular category has been detected or at which position the object has been detected).
[0028] Embodiment 9 is a data processing device (particularly a control device), which is configured to execute the method according to any one of embodiments 1 to 8.
[0029] Embodiment 10 is a computer program comprising instructions, which, when executed by a processor, causes the processor to perform the method according to any one of embodiments 1 to 8.
[0030] Embodiment 11 is a computer-readable medium storing instructions. When the instructions are executed by a processor, the instructions cause the processor to perform the method according to any one of embodiments 1 to 8. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In the accompanying drawings, like reference numerals generally refer to the same parts in the various views of the whole. The accompanying drawings are not necessarily drawn to scale, but emphasis is generally placed on illustrating the principles of the invention. In the following description, various aspects are described with reference to the following drawings.
[0032] Figure 1 A vehicle is shown.
[0033] Figure 2 A machine learning model according to one embodiment is shown.
[0034] Figure 3 Shown in detail Figure 2 The R-CNN (region-based convolutional neural network) level structure of the machine learning model.
[0035] Figure 4 A flow chart is shown which represents a method for object detection in image data according to one embodiment. DETAILED DESCRIPTION
[0036] The following detailed description refers to the accompanying drawings, which, for illustrative purposes, show specific details and aspects of the present disclosure in which the present invention may be implemented. Other aspects may be used and structural, logical, and electrical changes may be performed without departing from the scope of protection of the present invention. Various aspects of the present disclosure are not necessarily mutually exclusive, as certain aspects of the present disclosure may be combined with one or more other aspects of the present disclosure to form new aspects.
[0037] Various examples are described in more detail below.
[0038] Figure 1 A vehicle 101 is shown.
[0039] exist Figure 1 In the example of FIG. 1 , a vehicle 101 (eg, a passenger car or a truck) is equipped with a vehicle control device (also referred to as an electronic control unit, eg, a control apparatus, such as an electronic control unit (ECU)) 102 .
[0040] The vehicle control device 102 has data processing components such as a processor (eg, CPU (Central Processing Unit)) 103 and a memory 104 for storing control software 107 according to which the vehicle control device 102 operates and data processed by the processor 103. The processor 103 executes the control software 107.
[0041] For example, the stored control software (computer program) includes instructions that, when executed by the processor, cause the processor 103 to perform driving assistance functions (i.e., functions of an ADAS (Advanced Driver Assistance System)) or even autonomously control the vehicle (AD (Autonomous Driving)).
[0042] The control software 107 is transferred from the computer system 105 to the vehicle 101, for example, via the network 106 (or by means of a storage medium, such as a memory card). This can also take place during operation (or at least when the vehicle 101 is with the user), because the control software 107 is updated to a new version over time, for example.
[0043] The control software 107 determines control actions for the vehicle (e.g., steering actions, braking actions, etc.) based on available input data, which contain information about the environment or are derived therefrom (e.g., by detecting other road users, such as other vehicles). These include, for example, sensor data from one or more sensor devices 109, such as sensor data from a camera of the vehicle 101, which are connected to the vehicle control unit 102 via a communication system 110 (e.g., a vehicle bus system such as a CAN (Controller Area Network)).
[0044] The control software 107 can be trained, for example, using machine learning (ML), i.e., the control software 107 implements a neural network (NN) 108 that is trained, for example, based on training data, wherein in this example, the neural network is trained by the computer system 105. Thus, the computer system 105 implements an ML training algorithm for training one (or more) ML models 108.
[0045] For example, an ML model (e.g., a neural network) is an ML model used for object recognition (e.g., other vehicles, etc.) Such a model can be trained using supervised training, however, this requires a large number of training data elements (i.e., training examples) that are labeled with labels (i.e., ground-truth or "ground truth" information).
[0046] In a wide range of applications such as autonomous driving and industrial automation, capturing large-scale training data with such labeled training data elements, e.g., required for training (typically data-intensive) object recognition models, can be very time-consuming, labor-intensive, and costly.
[0047] Few-shot object detection (FSOD) methods attempt to obtain convincing representations from a limited number of training examples. Generalized FSOD (G-FSOD) aims to jointly recognize base classes, for which many training examples exist, and novel classes, for which only a limited number of training examples exist. However, this approach ignores the uncertainty that impacts the recognition performance of both classes. However, naively integrating the uncertainty estimation in the two-stage G-FSOD framework with the Region Proposal Network (RPN) and subsequent R-CNN (region-based convolutional neural network) leads to performance losses.
[0048] Prediction uncertainty can be divided into aleatory uncertainty and epistemic uncertainty. The former represents the inherent variability of the data itself, such as sensor noise. Aleatory uncertainty is usually taken into account by explicitly integrating it into the corresponding machine learning model (e.g., a neural network) as a learnable parameter associated with the predicted outcome. In particular, in neural networks used for object recognition, epistemic uncertainty is often taken into account by incorporating random activation (Dropouts) during the training phase of the model, in which a portion of neurons are randomly omitted during training, thereby creating a model ensemble (or "ensemble model"). By examining the differences between the predictions produced by different models in such an ensemble, the degree of epistemic uncertainty in the model can be approximately determined. Monte Carlo random activation (MC-Dropout) extends this approach during inference by performing multiple forward passes with random activation enabled and averaging the resulting predictions.
[0049] According to various embodiments, a machine learning model (in particular, a G-FSOD framework) is provided that progressively refines (i.e., corrects) initially low-quality, highly uncertain (object) proposals (proposals, i.e., e.g., bounding boxes (also referred to herein as the common English term "Bounding-Box(en)"), possibly with associated classification values (or classification "scores") that are determined within the machine learning model but not yet finalized, i.e., do not necessarily correspond to the finalized predictions) to multiple (processing) stages (each with an R-CNN). Each stage exploits predictive aleatory and epistemic uncertainty to output more reliable predictions. According to various embodiments, these stages include attention blocks during training, which enable learning the most convincing spatial features for each class (even when there are only a few training examples).
[0050] According to various embodiments, a method is therefore provided, hereinafter also referred to as UPPR (Uncertainty-based Progressive Proposal Refinement), in which uncertainty estimation is used in conjunction with the FSOD method to improve object proposals, improve the overall performance of the recognition (detection) and reduce forgetting (detection of already learned classes). UPPR specifically focuses on modeling prediction uncertainty within a two-stage G-FSOD framework, which enables refinement of object proposals. This approach (particularly the modeling of prediction uncertainty in G-FSOD) enables improved recognition performance while simultaneously alleviating the forgetting problem by explicitly incorporating uncertainty modeling.
[0051] Figure 2 A machine learning model 200 is shown according to one embodiment.
[0052] In particular, the machine learning model 200 comprises a plurality of R-CNN stages 204 (i.e., a sequence of (R-CNN) stages, three stages in the example shown), wherein aleatoric and epistemic uncertainty are estimated in each R-CNN stage. Each stage is treated as an ensemble model (with the aid of dropout, see above) that refines the proposals based on an IoU (Intersection over Union) threshold and the estimated uncertainty. During training, increasing IoU thresholds are set (as the stage sequence progresses) to make subsequent stages (i.e., stages further down the sequence) more reliable than earlier ones. During training, after each R-CNN stage 204, each proposal is compared to the ground truth and the IoU is calculated. If the IoU falls below the threshold (for the corresponding stage), the proposal is rejected. The IoU thresholds for the three R-CNN stages 204 are, for example, 50%, 60%, and 70%. This not only improves the predicted recognition but also helps reduce forgetting of base classes.
[0053] Figure 3 The structure of the R-CNN stage 300 of the machine learning model 200 is shown in detail. According to one embodiment, each R-CNN stage 204 has this structure.
[0054] The R-CNN stage 300 includes a RoI (Region of Interest) pooling layer 301. This is followed by an attention block 302 (only during training, not during inference). During the training phase, the R-CNN stage 200, including the attention block 302, is trained, for example, based on a balanced set of training data elements for the base class and the new class.
[0055] The feature extractor 201 is followed by a region proposal network 202 (which, according to one embodiment, is not an anchor-based RPN but a keypoint-based (deeper, i.e., having more layers) RPN).
[0056] Using a cascaded R-CNN architecture (i.e., a sequence of R-CNN 204) for the machine learning model in the G-FSOD framework instead of a single R-CNN stage can improve the quality of features at the instance level (i.e., for each proposal) and achieve improved overall performance in object recognition.
[0057] According to G-FSOD, according to various embodiments, the training dataset Divided into two subsets: The base dataset of training examples and have a limited number of A "new" dataset of training examples It is important to note that there is no overlap between these two classes, i.e. In each training data element, the input image and the ground truth pair, where the ground truth contains the class label (for objects shown in the input image) and the corresponding bounding box coordinates Here, i is the index of the training data element. For both the base and the new dataset, the following applies:
[0058]
[0059] or
[0060]
[0061] The G-FSOD training method consists of two stages. In the first stage, the machine learning model is trained on the basic dataset. is trained on the basis of to build transferable (übertragbar) prior knowledge. In the second stage, the machine learning model uses the acquired knowledge in order to combine of (basic) training examples and (small number of) training examples from In contrast to FSOD, the main goal of G-FSOD is to maximize the average overall precision (AP), where the average overall precision is the weighted average of the AP of the base class (bAP) and the AP of the new class (nAP), i.e.
[0062]
[0063] The following describes in more detail Figure 2 and Figure 3 Components of the machine learning model 200 shown in .
[0064] RPN 202 is an RPN based on multi-scale keypoints. An anchor-based RPN in the form of a class-independent module, for example with a three-layer architecture, typically generates low-value proposals for the subsequent R-CNN detector 205 (which is formed by a sequence of R-CNN stages 204). This problem arises from the reliance on fixed-size anchors, which can result in a large number of proposals for the background and low-value proposals for the foreground. In addition, misalignment of anchors and collapsed features makes bounding box classification difficult.
[0065] On the other hand, keypoint-based methods are expected to alleviate the above limitations by using keypoints to represent each object and thus providing more accurate spatial information. Therefore, according to various embodiments, the anchor-based RPN is replaced by the keypoint-based CenterNet (referred to as CenterNet-RPN). In order to explicitly consider the variability of object sizes, The feature extractor 201 includes a feature pyramid neural network (FPN). This allows The refinement of object proposals (e.g., in the form of bounding box proposals for each resolution) becomes easy. Accordingly, the output 203 of the feature extractor 201 is a collection of feature maps for different resolutions, i.e., a feature pyramid F pyr .
[0066] RPN 202 outputs proposals that are refined by the cascaded R-CNN 204 (which includes increasing IoU thresholds during the sequence). Each R-CNN stage 204 (index m) improves the proposals from the previous stage. The quality of the object proposals (or, in the case of the first R-CNN stage 204, the quality of the RPN 202), and thus increase the number of true positive results that are passed to the next stage 204 (or output in the case of the last stage 204). In each R-CNN stage 204, 300, the classification features (indexed "cls") and the localization features (indexed "box") are decoupled by introducing a dual classification and bounding box regressor head, i.e., each R-CNN stage 300 contains a first MLP 303 (multi-layer perceptron) and a first output layer 304 for classification (i.e., generating class scores) and a second MLP 305 and a second output layer 306 for localization (i.e., determining the bounding box, for example, in the form of a bounding box offset, i.e., a bounding box correction).
[0067] ROI pooling 301 uses each R-CNN stage 204 as input to obtain the feature pyramid F pyr The output of ROI pooling 301 for the mth level is used to “fill” each proposal from the previous level. To express.
[0068] During training, each R-CNN stage 300 contains an attention block 302, so that at the instance level (i.e., for Multi-level attention is implemented for each proposal (in English: attention). The motivation for this is that while feeding instance-level features to the cascaded R-CNN stage 204 helps refine the proposal, not all instance-level features are equally important. To place more emphasis on features related to the correct classification, an attention block (or "module") 302 is provided.
[0069] The attention module 302 is, for example, a convolutional block attention module (CBAM) to achieve selective attention to the most important features of the G-FSOD task. In particular, the channel- and spatial-dependent attention components of the CBAM capture the channel- and spatial-dependent relationships between features at the instance level (e.g., there are multiple image channels, such as color or depth channels), thereby enabling the machine learning model 200 to better capture semantically rich information for both the new class and the base class. Another advantage of the CBAM for object detection in the G-FSOD framework is its lightweight design, which is particularly important because it is integrated into each R-CNN stage 204. To prevent the CBAM from favoring the base class over the new class, the multi-level attention block is only added during the training phase for the new class to ensure a balanced representation of features of the base class and the new class.
[0070] As described above, there are inherent data and model uncertainties (i.e., accidental and epistemic uncertainties) that are considered according to various embodiments to reduce forgetting and improve recognition of new categories. To this end, accidental uncertainty and epistemic uncertainty are estimated in each stage 204 of the cascade R-CNN. Ultimately, a stage-by-stage refinement (of object proposals) is performed based on epistemic uncertainty as well as on accidental uncertainty.
[0071] Level-by-level refinement based on epistemic uncertainty: Epistemic uncertainty is modeled during inference by using a dropout layer ( Figure 3 Neurons are represented by dotted lines in the figure).
[0072] The processing process for the training example obtains the feature pyramid Fpyr (i.e., feature maps for multiple different resolutions) generated by the feature extraction network 201 (which can be regarded as the backbone network) and the object proposal generated by the previous level (or, in the case of the first level, by the RPN 202, i.e., the RPN 202 can be regarded as the zeroth level) in the mth R-CNN level 300. Proposal features are then extracted using RoI pooling 301, guided by a CBAM attention block 302 for focusing, and fed into the classification head 303, 304 and the bounding box regression head 305, 306 to obtain category scores and bounding box offsets.
[0073] The process is a single forward pass through the R-CNN stage 300. In testing and inference, the dropout layer is activated and R such forward passes are performed for each stage 204, aggregating the predictions (classification scores and bounding box offsets) into aggregated classification scores. and the aggregate bounding box offset vector See also Figure 3 ) and passes it to the next stage (or outputs it in the case of the last stage).
[0074] Formally, for M levels 204, the classification features for the mth level are expressed as follows:
[0075]
[0076] where a m (·) is the m-th level CBAM attention module. It is the MLP 303 of the classification head in the RoI head (RoI-Head) 307.
[0077] Similarly, the bounding box features are calculated as follows:
[0078]
[0079] in It is the MLP 305 in the bounding box head in the RoI head 307.
[0080] and It is then used by the output layer 304 or 306 respectively, where it is used in the RoI predictor 308 or Represented to compute classification scores and bounding box regression offsets (and associated aleatoric uncertainty (variance) and (or as a covariance matrix representation or ) (These are output by the output layers 304, 306). As described above, during inference, R forward passes with dropout are performed and the classification scores (e.g., classification logits) and bounding box offsets are aggregated. The same is true for the associated random variance, i.e., (After R passes) the aggregation is And will (After R passes) the aggregation is
[0081] Therefore it applies:
[0082]
[0083] and
[0084]
[0085] in Represents the forward pass of the RoI head 307 (ie, MLP 307) for the rth time The output classification features, and represents the bounding box features output by the RoI head 307 (i.e., MLP 308) for the rth forward pass, where, due to random dropout, these features used for classification and bounding box regression may be different in each pass.
[0086] Level-by-level refinement based on aleatoric uncertainty: aleatoric uncertainty is considered not only for classification but also for bounding box regression. To this end, the classification score is modeled as a multivariate Gaussian distribution consisting of the mean s of the predicted classification scores cls and the diagonal correlation covariance matrix To parameterize, according to the predicted category variance Then N are drawn (i.e. sampled) from the generated Gaussian distribution. cls Category ratings The resulting matrix contains all the samples generated in this way, and is denoted by S cls To express:
[0087]
[0088] The classification loss is then these random classification logit values S cls Softmax cross entropy between the ground truth classification labels and the associated ground truth classification labels.
[0089] The classification loss (and regression loss) are compared to the ground truth for each level during training. The loss (classification loss plus regression loss) is calculated for each R-CNN level and then averaged across these R-CNN levels to determine the corresponding training loss.
[0090] The bounding box regression result (i.e., offset) is modeled as a Gaussian distribution in the same way, with mean b box is the predicted box offset b box , and the diagonal covariance matrix ∑ box The predicted box offset variance Samples are again drawn from this distribution, averaged, and the bounding box regression loss (relative to the ground truth) is determined, for example by using the negative log-likelihood.
[0091] In summary, the process is as follows: First, for each training example, the initial proposals from RPN 202 are compared with the features generated by feature extractor 201. Figure 1 are sent to the first R-CNN level. Next, the RoI head 307 of the first level pools the features and extracts classification and bounding box features, which are passed through the RoI predictor 308, which results in classification scores and variances, as well as bounding box offsets and variances. To capture epistemic uncertainty, randomness is introduced during training through random inactivation layers. During inference, R forward passes are performed and the network predictions are aggregated and averaged to obtain the final prediction. The predicted bounding box offsets are then applied to the input proposals, which results in a refined (i.e., corrected) box that is used as input to the next R-CNN level. This gradual refinement produces more reliable boxes by utilizing average epistemic predictions, which are more robust than predictions from a single pass.
[0092] In summary, according to various embodiments, there is provided Figure 4 The method shown.
[0093] Figure 4 A flow chart 400 is shown illustrating a method of object detection in image data according to one embodiment.
[0094] At 401 , features are extracted from image data (e.g., feature maps at different resolutions are determined, e.g., using a neural convolutional network).
[0095] At 402 , one or more proposals for a bounding box of a corresponding object are determined based on the extracted features.
[0096] At 403, the bounding box is (successively) corrected by a sequence of a plurality of processing stages, wherein the processing stages each comprise a neural network (e.g., an MLP) and each take as input one or more bounding box proposals and determine a corresponding bounding box correction for each input bounding box proposal in a pass through the processing stages (i.e., the bounding box proposal corrected according to the bounding box correction of a processing stage is used as input to the next processing stage (unless it is the last in the sequence)), wherein for each processing stage:
[0097] At 404, for each input bounding box proposal, a plurality of bounding box corrections (e.g., offsets to bounding box description data, or also entirely new description data for a bounding box, where a bounding box is specified, for example, by the location of a corner (e.g., the top left corner), its height, and its width) are determined by performing a plurality of passes on the input one or more bounding box proposals, wherein the input bounding box correction proposals differ due to different deactivations of neurons of the neural network (i.e., random dropouts are performed during passes of the neural network such that the passes differ),
[0098] • In 405 , for each input bounding box proposal, determine an output bounding box correction by averaging the bounding box corrections determined for the input bounding box proposals in the passes.
[0099] Figure 4 The method can be performed by one or more computers having one or more data processing units. The term "data processing unit" can be understood as any type of entity capable of processing data or signals. For example, data or signals can be processed according to at least one (i.e., one or more) specific functions performed by the data processing unit. The data processing unit may include or be formed by analog circuits, digital circuits, logic circuits, microprocessors, microcontrollers, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), field programmable gate array (FPGA) integrated circuits, or any combination thereof. Any other means for implementing the corresponding functions described in more detail herein may also be understood as a data processing unit or logic circuit device. One or more method steps described in detail herein may be performed (e.g., implemented) by a data processing unit through one or more specific functions performed by the data processing unit.
[0100] Thus, according to various embodiments, the method is particularly computer-implemented.
[0101] Various embodiments may receive and use image data from various sensors that can output output data in the form of images, such as single images, video, radar, lidar, ultrasonic, motion, thermal imaging, etc. The sensor data may be measured or simulated over a period of time (e.g., to generate training data elements).
[0102] In particular, these sensor data may be classified, for example in order to detect the presence of an object represented in the sensor data. Figure 4 The method can be integrated into various frameworks where new categories emerge. In this way, it can be deployed by various KI-driven perception systems in, for example, robots and self-driving cars. Figure 4 method.
[0103] Figure 4 The method is often used, for example, to generate control signals for robotic devices. The term "robotic device" can be understood to refer to any technical system with mechanical components whose movements are controlled, such as, for example, computer-controlled machines, vehicles, household appliances, power tools, manufacturing machines, personal assistants, or access control systems. Control rules are learned for this technical system and then the technical system is controlled accordingly.
Claims
1. A method for object detection in image data, the method comprising: Extracting features from the image data (203); determining one or more proposals for bounding boxes for the corresponding object based on the extracted features (203); The bounding box is corrected by a sequence of a plurality of processing stages (204, 300), wherein the processing stages each comprise a neural network (303) and each receive one or more bounding box proposals as input and determine a corresponding bounding box correction for each input bounding box proposal in a pass through the processing stages (204, 300), wherein for each processing stage (204, 300): - for each input bounding box proposal, determining a plurality of bounding box corrections by performing a plurality of passes on the input one or more bounding box proposals, said bounding box proposals differing due to different deactivations of neurons of the neural network (303), - For each input bounding box proposal, determine an output bounding box correction by averaging the bounding box corrections determined for the input bounding box proposals in the pass.
2. The method of claim 1 , wherein each processing stage (204, 300) determines a relevant classification for each bounding box correction in each pass, and determines a classification for each input bounding box proposal by averaging the classifications determined for the input bounding box proposals in the pass.
3. The method according to claim 1 or 2, wherein each processing stage (204, 300) also takes the extracted features (203) as input.
4. The method according to any one of claims 1 to 3, comprising: training at least one of the processing stages (204, 300) to output, for each input bounding box proposal, description data of a bounding box probability distribution with respect to a position of the corresponding bounding box; Determine bounding box samples by sampling multiple times from the bounding box probability distribution; Determine the loss between the bounding box sample and the bounding box ground truth information; and training at least one of the processing stages (204, 300) to reduce the loss.
5. The method according to any one of claims 1 to 4, comprising: training at least one of the processing stages (204, 300) to output, for each input bounding box proposal, description data of a classification probability distribution with respect to a class of an object contained in the corresponding bounding box; Determine the classification samples by sampling multiple times from the classification probability distribution; Determine the loss between the classified sample and the classification benchmark truth information; and training at least one of the processing stages (204, 300) to reduce the loss.
6. The method according to any one of claims 1 to 5, comprising: One or more proposals for bounding boxes are determined based on the extracted features (203) by a keypoint-based region proposal network (202).
7. The method according to any one of claims 1 to 6, comprising: The processing stages (204, 300) are trained, wherein during training each processing stage (204, 300) includes an attention block (302) that processes features derived from the extracted features (203) that are assigned to corresponding one or more bounding box proposals, wherein the processing stage (204, 300) determines a bounding box correction by using the processed features.
8. A method for controlling a robotic device, the method comprising: capturing image data of an environment of the robotic device; detecting an object in the image data by a method according to any one of claims 1 to 7; as well as The robotic device is controlled based on the detection of the object in the image data. 9 . A data processing device configured to execute the method according to claim 1 .
10. A computer program comprising instructions which, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 8.
11. A computer-readable medium storing instructions, which, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 8.