Sample Weighted Learning System and Method for Object Detection in Images
By proposing a sample weight generation system and a sample weighted learning system in image object detection, using feature transformation and dynamic adjustment weight prediction, the problems of category imbalance and sample weighting complexity are solved, the balance of classification and regression tasks is achieved, and the detection accuracy is significantly improved.
Patent Information
- Application Number
- CN202010431851.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-20
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2040-05-20
AI Technical Summary
The prior art is difficult to effectively solve the problem of category imbalance in image object detection, which makes it difficult for the detector to find a balance between accuracy and speed in the first stage, and the sample weighting process is complex and dynamic, affecting the detection accuracy.
A sample weight generation system and a sample weighted learning system are proposed to generate sample weights for image object detection through feature transformation, sample feature generation and weight prediction equipment. The system uses exponential functions to predict the sample weights of classification and regression losses, and dynamically adjusts the transformation function of the sample feature generation device through loss function calculation and gradient adjustment.
The balance between classification and regression tasks in image object detection is achieved, and the detection accuracy is improved, especially under the high IoU standard, which significantly improves the performance of object detection.
Smart Images

Figure CN113723395B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and particularly to a sample weight generation system and method for image processing, and a sample weighted learning system and method for object detection of images. Background Art
[0002] Modern region-based object detection in images is a multi-task learning problem, consisting of object classification and localization. It involves region sampling (sliding window or region proposal), region classification and regression, and non-maximum suppression. Using region sampling, it transforms object detection into a classification task, thus classifying and regressing a large number of regions. According to the way of region search, these detectors can be classified into one-stage detectors and two-stage detectors.
[0003] Generally, the object detectors with the highest accuracy are based on a two-stage framework, such as Faster R-CNN (Fast R-CNN), which rapidly narrows down the region range (mostly from the background) in the region proposal stage. In contrast, one-stage detectors, such as SSD and YOLO, achieve faster detection speed but lower accuracy. This is due to the class imbalance problem (i.e., the imbalance between foreground and background regions), which is a classic challenge for object detection.
[0004] Two-stage detectors handle class imbalance through a region proposal mechanism and then adopt various effective sampling strategies, such as selecting samples with a fixed foreground-to-background ratio and hard sample mining. Although similar hard sample mining can be applied to one-stage detectors, it is less efficient due to the existence of a large number of simple negative samples.
[0005] Sample weighting is a very complex and dynamic process. When applied to the loss function of a multi-task problem, there are various uncertainties in each sample. If the detector uses its ability for precise classification and produces poor localization results, the mislocalized detections will damage the average precision, especially under the high IoU criterion, and vice versa. Summary of the Invention
[0006] According to the present invention, sample weighting in the field of image processing is not only related to data but also related to tasks. On the one hand, different from the prior art, the importance of the samples of an image should be determined by its intrinsic attributes compared with the ground truth annotation and its response to the loss function. On the other hand, object detection of an image is a multi-task problem. The weighting of the samples of an image should maintain a balance between different tasks.
[0007] According to one aspect of the present invention, a sample weight generation system for image processing is proposed, and the system includes:
[0008] A feature transformation device for transforming the input first feature, second feature, third feature, and fourth feature into a first dense feature, a second dense feature, a third dense feature, and a fourth dense feature respectively;
[0009] A sample feature generation device for generating a joint sample feature according to the first dense feature, second dense feature, third dense feature, and fourth dense feature using a transformation function;
[0010] A weight prediction device for predicting a sample weight for classification loss and a sample weight for regression loss for a sample according to the generated sample feature.
[0011] A sample weight generation system according to an aspect of the present invention, wherein:
[0012] The sample weight for classification loss and the sample weight for regression loss predicted by the weight prediction device are obtained by a first exponential function and a second exponential function respectively.
[0013] According to an aspect of the present invention, a sample weighted learning system for object detection in images is provided, the system comprising:
[0014] An input device for receiving the input first feature, second feature, third feature, and fourth feature for each sample;
[0015] A feature transformation device for transforming the input first feature, second feature, third feature, and fourth feature into a first dense feature, a second dense feature, a third dense feature, and a fourth dense feature respectively;
[0016] A sample feature generation device for generating a joint sample feature according to the first dense feature, second dense feature, third dense feature, and fourth dense feature using a transformation function;
[0017] A weight prediction device for predicting a sample weight for classification loss and a sample weight for regression loss for each sample according to the generated sample feature;
[0018] A loss function calculation device for calculating a loss function according to the predicted sample weight for classification loss and the sample weight for regression loss;
[0019] And
[0020] A feature transformation adjustment device for adjusting the transformation function based on the calculated loss function.
[0021] A sample weight generation system according to an aspect of the present invention, wherein the system further comprises:
[0022] A gradient calculation device for deriving a gradient according to the calculated loss function to adjust the transformation function used by the sample feature generation device.
[0023] A sample weighted learning system according to an aspect of the present invention, wherein the sample feature generation device further comprises:
[0024] A first feature transformation device for transforming the input first feature into a first dense feature;
[0025] A second feature transformation device for transforming the input second feature into a second dense feature;
[0026] A third feature transformation device for transforming the input third feature into a third dense feature; and
[0027] A fourth feature transformation device for transforming the input fourth feature into a fourth dense feature.
[0028] A sample weighted learning system according to an aspect of the present invention, wherein the weight prediction device further comprises:
[0029] A classification loss weight prediction device for predicting the sample weight of the classification loss; and
[0030] A regression loss weight prediction device for predicting the sample weight of the regression loss.
[0031] A sample weighted learning system according to an aspect of the present invention, wherein the first feature is a classification loss, the second feature is a regression loss, the third feature is an intersection over union, and the fourth feature is a classification probability.
[0032] A sample weighted learning system according to an aspect of the present invention, wherein:
[0033] The input device is further configured to receive a fifth feature; and
[0034] The sample weighted learning system further comprises:
[0035] A fifth feature transformation device for transforming the input fifth feature into a fifth dense feature.
[0036] A sample weighted learning system according to an aspect of the present invention, wherein:
[0037] The fifth feature is a mask loss.
[0038] A sample weighted learning system according to an aspect of the present invention, wherein:
[0039] The predicted sample weights of the classification loss and the regression loss are obtained by a first exponential function and a second exponential function respectively.
[0040] A sample weighted learning system according to an aspect of the present invention, wherein:
[0041] The sample weights of the classification losses of a set of samples including positive samples and negative samples are averaged to obtain the sample weights of the classification losses of each sample.
[0042] According to one aspect of the present invention, a method for generating sample weights for image processing is provided. The method includes:
[0043] Transforming the input first feature, second feature, third feature, and fourth feature into a first dense feature, a second dense feature, a third dense feature, and a fourth dense feature respectively;
[0044] Using a transformation function to generate a joint sample feature based on the first dense feature, the second dense feature, the third dense feature, and the fourth dense feature;
[0045] Generating sample weights for the classification loss and sample weights for the regression loss of the sample based on the generated sample feature.
[0046] According to the sample weight generation method of one aspect of the present invention, wherein:
[0047] The sample weights of the predicted classification loss and the sample weights of the regression loss are respectively obtained by a first exponential function and a second exponential function.
[0048] According to one aspect of the present invention, a sample weighted learning method for object detection in images is provided. The method includes:
[0049] First step: Receiving the input first feature, second feature, third feature, and fourth feature for each sample;
[0050] Second step: Transforming the input first feature, second feature, third feature, and fourth feature;
[0051] Third step: Using a transformation function to generate a joint sample feature based on the first dense feature, the second dense feature, the third dense feature, and the fourth dense feature;
[0052] Fourth step: Generating sample weights for the classification loss and sample weights for the regression loss of each sample based on the generated sample feature;
[0053] Fifth step: Calculating a loss function based on the predicted sample weights of the classification loss and the sample weights of the regression loss;
[0054] Sixth step: Adjusting the transformation function used by the sample feature generation device according to the calculated loss function.
[0055] According to the sample weighted learning method of one aspect of the present invention, wherein the second step further includes:
[0056] The step of transforming the first feature of the input into a first dense feature;
[0057] The step of transforming the second feature of the input into a second dense feature;
[0058] The step of transforming the third feature of the input into a third dense feature; and
[0059] The step of transforming the fourth feature of the input into a fourth dense feature.
[0060] A sample weighting learning method according to an aspect of the present invention, wherein the fourth step further includes:
[0061] The step of predicting the sample weights of the classification loss; and
[0062] The step of predicting the sample weights of the regression loss.
[0063] A sample weighting learning method according to an aspect of the present invention, wherein the first feature is the classification loss, the second feature is the regression loss, the third feature is the intersection over union, and the fourth feature is the classification probability.
[0064] A sample weighting learning method according to an aspect of the present invention, wherein the first step further includes receiving a fifth feature; and the method further includes the step of transforming the fifth feature of the input into a fifth dense feature.
[0065] A sample weighting learning method according to an aspect of the present invention, wherein the fifth feature is the mask loss.
[0066] A sample weighting learning method according to an aspect of the present invention, wherein:
[0067] The sample weights of the predicted classification loss and the regression loss are obtained by a first exponential function and a second exponential function respectively.
[0068] A sample weighting learning method according to an aspect of the present invention, wherein:
[0069] The sample weights of the classification loss of a set of samples including positive samples and negative samples are averaged to be used as the sample weights of the classification loss of each sample.
[0070] A sample weighting learning method according to an aspect of the present invention, wherein the sixth step further includes:
[0071] Deriving a gradient according to the calculated loss function, and then adjusting the transformation function according to the derived gradient.
[0072] A sample weight generation system or a sample weighting learning system according to an aspect of the present invention can also be applied to a system for object detection in image processing.
[0073] The sample weight generation method or sample weighted learning method according to one aspect of the present invention can also be applied to the method of object detection in image processing.
[0074] The sample weighted learning method for object detection in images proposed by the present invention is simple and effective. By using a sample weighted network to learn weights sample by sample, a balance can be achieved between classification and regression tasks. Specifically, in addition to the basic detection network, the present invention designs a sample weighted network to predict the weights of the classification loss and regression loss of the samples in the image. The sample weighted network takes the classification loss, regression loss, IoU (Intersection over Union) value, and classification probability as inputs. It uses a function that converts the current context features of the sample into sample weights. The sample weighted network according to the present invention has been comprehensively evaluated on the MSCOCO and Pascal VOC datasets, and various one-stage and two-stage detectors have been evaluated.
[0075] In summary, the present invention proposes a general loss function for object detection in images, which covers most region-based object detectors and their sampling strategies, and on this basis designs a unified sample weighted network. Compared with the previous sample weighting methods for images, the method according to the present invention has the following advantages: (1) jointly learning the sample weights of both the classification task and the regression task; (2) being data-dependent, so that the soft weights of each individual sample can be learned from the training data; (3) being applicable to various one-stage and two-stage detectors, and can be easily inserted into most object detectors used, and obtaining a significant performance improvement without affecting the inference time. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] To more fully understand the present invention and its advantages, reference will now be made to the following description in conjunction with the accompanying drawings, in which:
[0077] Figs. 1(a)-1(c) show schematic diagrams of samples with different weights and classification losses during the object detection training process of an image, wherein Fig. 1(a) shows a sample with a large classification loss but a small weight, Fig. 1(b) shows a sample with a small classification loss but a large weight, and Fig. 1(c) shows the inconsistency exhibited before the classification probability and IoU;
[0078] Figs. 2(a)-2(c) show the system architecture diagrams of the sample weighted learning system for object detection in images according to an embodiment of the present invention, wherein Fig. 2(a) shows the structure diagram of a two-stage detector, Fig. 2(c) shows the schematic diagram of the sample weighted network for processing an image according to an embodiment of the present invention, and Fig. 2(b) shows the schematic diagram of how to obtain the loss function in image processing according to Fig. 2(c);
[0079] Figure 3 The block diagram of a sample weighted learning system for object detection in images according to an embodiment of the present invention is shown;
[0080] Figure 4 The flowchart of a sample weighted learning method for object detection in images according to an embodiment of the present invention is shown;
[0081] Figs. 5(a)-5(d) show the schematic comparison effects of applying the sample weighted learning system for object detection in images according to an embodiment of the present invention to object detection.
[0082] Figure 6 The block diagram of an electronic device according to an embodiment of the present disclosure is schematically shown. Detailed implementation manners
[0083] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, many specific details are set forth in order to provide a comprehensive understanding of the embodiments of the present invention. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.
[0084] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0085] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0086] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning that those skilled in the art usually understand this expression (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). In the case of using expressions such as "at least one of A, B, or C, etc.", generally, it should be interpreted according to the meaning that those skilled in the art usually understand this expression (for example, "a system having at least one of A, B, or C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0087] Some block diagrams and / or flowcharts are shown in the drawings. It should be understood that some blocks or combinations of blocks in the block diagrams and / or flowcharts can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, so that when these instructions are executed by the processor, they can create a device for implementing the functions / operations illustrated in these block diagrams and / or flowcharts. The technology of the present invention can be implemented in the form of hardware and / or software (including firmware, microcode, etc.). Additionally, the technology of the present invention can take the form of a computer program product on a computer-readable storage medium storing instructions, and this computer program product can be used by an instruction execution system or in combination with an instruction execution system. The previously mentioned "difficult" samples generally refer to samples with a relatively large classification loss in image processing. However, "difficult" samples are not necessarily important. As shown in Fig. 1(a) (the samples of all images are selected from the training process), the sample has a relatively high classification loss but a small weight ("difficult" but not important). On the contrary, if a "simple" sample captures the key points of the object category shown in Fig. 1(b), it may be very important. Additionally, the assumption that the bounding box regression is accurate when the classification score is high is not always the case as shown in Fig. 1(c). Sometimes, inconsistencies may occur between classification and regression.
[0088] The present invention is applied to object detection in the field of image processing. Object detection refers to the ability of a computer and software system to locate objects in an image (or scene) and identify each object. The samples referred to herein can, for example, be small regions in an object (image). In the object detection process, the input is an image, and the output is the identified (located) object.
[0089] The present invention reconstructs the sample weighting problem in a probabilistic format and measures sample importance by reflecting uncertainty. The probabilistic modeling not only solves the sample weight problem but also solves the balance problem between classification and localization tasks. The present invention makes the sample weighting process flexible and can be learned through deep learning.
[0090] Figures 2(a)-2(c) show the structural diagrams of a sample weighting learning system for object detection in images according to an embodiment of the present invention. Figure 2(a) shows the structural diagram of a two-stage detector (detection network), which can also be replaced by a one-stage detector. Figure 2(c) shows a schematic diagram of a sample weighting network in a sample weighting learning system for object detection in images according to an embodiment of the present invention, and Figure 2(b) shows a schematic diagram of how to obtain a loss function in image processing according to Figure 2(c). In Figure 2(c), each sample is compared with its true annotation to calculate input features. According to an embodiment of the present invention, the input features include four initial features related to the samples of the image and Prob i , which are the classification loss, regression loss, IoU, and classification probability for the samples of each image, respectively. Then, four functions F, G, H, and K corresponding to the four input features related to the samples of the image transform the four features to obtain four dense features. The four functions can all be implemented by an MLP neural network. The sample weights for the classification loss and the sample weights for the regression loss are obtained based on the dense features and provided to Figure 2(b) for further processing. It can be seen from Figure 2(c) that the sample weighting network for obtaining the sample weights for the classification loss and the sample weights for the regression loss can be composed of two levels of multi-layer perceptron (MLP) networks. The losses of all samples can be averaged to optimize the model parameters. Figure 2(b) transmits the generated total loss L i back to the sample weighting network and the detection network respectively, so as to adjust the four functions used in the sample weighting network and the relevant parameters in the detection network respectively. The following will refer to Figure 3 specifically describe the block diagram of a sample weighting learning system for object detection according to an embodiment of the present invention.
[0091] As Figure 3 shown, a sample weighting learning system for object detection in images includes an input device 30 for receiving input image data, and the input image data can be feature-based samples; a sample weighting network 32 for obtaining sample weights through learning according to the input image data such as features; a loss function calculation device 34 for calculating a loss function according to the weights of the samples; and a gradient calculation device 36 for calculating a gradient according to the loss function and providing it to the sample weighting network 32. And a detection network (not shown).
[0092] The sample weighted network 32 includes: a feature transformation adjustment device 301, a feature transformation device 302, a sample feature generation device 303, and a weight prediction device 304.
[0093] The feature transformation device 302 includes a first feature transformation device 3021, a second feature transformation device 3022, a third feature transformation device 3023, and a fourth feature transformation device 3024 that respectively transform the input image data such as features using transformation functions. The feature transformation adjustment device 301 is used to adjust the transformation functions used by the first feature transformation device 3021, the second feature transformation device 3022, the third feature transformation device 3023, and the fourth feature transformation device 3024.
[0094] The weight prediction device 304 includes a classification weight prediction device 3041 for predicting the sample weights of the classification loss based on the generated sample features and a regression weight prediction device 3042 for predicting the sample weights of the regression loss.
[0095] The loss function calculation device 34 is used to calculate the loss function according to the predicted sample weights. The weights of the samples can include the sample weights of the classification loss and the sample weights of the regression loss.
[0096] Before specifically describing the sample weighted learning system for object detection, the algorithm adopted by the present invention is described first. This algorithm can perform object detection in images more effectively.
[0097] The latest research on object detection including one-stage object detectors and two-stage object detectors follows a similar region-based paradigm. Given a set of anchors (samples from the image) i is a natural number., that is, the prior boxes usually placed on the image to densely cover spatial positions, scales, and aspect ratios, the multi-task training objects in the image can be summarized as follows:
[0098]
[0099] where is the classification loss (regression loss), and represents the sampled anchors for classification (regression). N1 and N2 are the numbers of training samples and foreground samples. The relationship applies to most object detectors. Let and be the classification loss weights for sample a i and the regression loss weights for sample a i respectively. The following are the generalized loss functions for two-stage and one-stage detectors with different sampling strategies:
[0100]
[0101] Among them, and are indicator functions that output 1 when the condition is met and 0 otherwise. As a result, and can be used to represent various sample strategies. Here, regional sampling can be interpreted as a special case of sample weighting, enabling soft sampling.
[0102] The present invention jointly learns sample weights for both classification and regression from a data-driven perspective. Previous methods focused on reweighting the classification (such as OHEM and Focal-Loss) or the regression loss (such as KL-Loss), while the present invention jointly reweights the classification and regression losses. Additionally, different from mining "hard" samples (which have a high classification loss) in the OHEM and Focal-Loss methods, the present invention focuses on important samples for object detection in images, which may also be "easy" samples.
[0103] The present invention reconstructs the sample weighting problem in a probabilistic format and measures sample importance by reflecting uncertainty. The present invention makes the sample weighting process flexible and can be learned through deep learning. Probabilistic modeling not only solves the sample weight problem but also solves the balance problem between the classification and localization tasks.
[0104] Next, the sample weighting learning system for object detection in images according to an embodiment of the present invention will be specifically described with reference to FIGS. 2(a)-2(c).
[0105] As shown in FIGS. 2(a)-2(c), the vector gt i represents the true annotation bounding box coordinates. is the estimated bounding box coordinates. By comparing gt i with (related to a i ), four discriminative features for each sample are obtained: and Prob i , which are the classification loss, regression loss, IoU i (Intersection over Union) and Prob i (classification probability), respectively. The four features are used as input data and input into the sample weighting network 32 by the input device 30. Subsequently, how to obtain the classification loss and the regression loss will be specifically described.
[0106] Different from directly using the visual features of samples (which actually loses the information of the true annotation of the object from the corresponding image), the present invention designs four discriminative features from the detector itself, which utilize the interaction between the estimation and the true annotation (i.e., IoU and classification score), because both the classification and regression losses inherently reflect the uncertainty of the prediction to some extent.
[0107] For negative samples, set the feature IoU i and Prob i to 0. Among them, positive samples can be samples including an object (an object in the image). Negative samples can be samples that do not include an object (an object in the image).
[0108] The input device 30 will obtain four different features: and Prob i and input them to the sample weighting network 32. For the four input features and Prob i , they are respectively processed by the first feature transformation device 3021, the second feature transformation device 3022, the third feature transformation device 3023, and the fourth feature transformation device 3024.
[0109] According to an embodiment of the present invention, the first feature transformation device 3021, the second feature transformation device 3022, the third feature transformation device 3023, and the fourth feature transformation device 3024 can respectively be or use four different functions F, G, H, and K for transforming the input into dense features for a more comprehensive representation and providing them to the sample feature generation device 303. According to an embodiment of the present invention, these functions are transformation functions and can all be implemented by an MLP neural network, which can map each one-dimensional value to a higher-dimensional feature through the transformation function. By using the sample feature generation device 303, the transformed features can be encapsulated in the sample-level feature d i :
[0110]
[0111] The sample feature generation device 303 provides the generated joint sample features to the weight prediction device 304. The classification loss weight prediction device 3041 and the regression loss weight prediction device 3042 of the weight prediction device 304 respectively learn the sample weights of the classification loss i and the sample weights of the regression loss from the generated sample features d
[0112]
[0113] and
[0114] Among them, W cls and W reg can respectively represent two separate MLP networks for weight prediction for classification loss and regression loss.
[0115] The weight prediction device 304 provides the predicted classification loss weight and regression loss weight to the loss function calculation device 34.
[0116] How the loss function calculation device 34 calculates the loss function will be specifically described below.
[0117] The object detection target can be decomposed into regression and classification tasks. Given the i-th sample, first, the regression task is modeled as a Gaussian likelihood, where the predicted location offset is used as the mean and the standard deviation
[0118]
[0119] where the vector gt i represents the true annotation bounding box coordinates, and is the estimated bounding box coordinates. To optimize the regression network, maximize the log probability of the likelihood:
[0120]
[0121] By defining (obtained according to gt i and ), multiplying the equation by -1 and ignoring the constant, the loss function calculation device 34 obtains the following regression loss: )
[0122]
[0123] For the training of object detectors for images, there are two opposite sample weighting strategies. On the one hand, some people prefer "difficult" samples, which can effectively accelerate the training process through larger losses and gradients. On the other hand, some people believe that when ranking is more important for evaluation metrics and the class imbalance problem is less important, "easy" examples require more attention. However, it is usually unrealistic to manually judge how difficult or noisy the training samples are. Therefore, the sample-level variance involved in Equation (8) introduces greater flexibility because it allows automatic adjustment of sample weights based on the effectiveness of each sample feature.
[0124] Taking the derivative of Equation (8) with respect to the variance equal to zero and solving (assuming λ2 = 1), the optimal variance value satisfies Reinsert this value into Equation (8) and ignoring the constant, the overall regression object reduces to This function is a concave non-decreasing function, which greatly supports while only applying soft penalties to large values. This makes the algorithm robust to outliers and noisy samples with large gradients, which may degrade performance. This also prevents the algorithm from focusing too much on being very difficult samples. In this way, the regression function of Equation (8) favors selecting samples with large IoU, as this encourages faster speed and drives the loss towards 0. In turn, this motivates the feature learning process to increase the weights on these samples, while samples with relatively small IoU still maintain moderate gradients during training.
[0125] For Equation (8), λ2 is a constant value that absorbs the overall loss scale in object detection of the image. By writing as Equation (8) can be roughly regarded as a weighted version of the regression loss, where a regularization term is used to prevent the loss from becoming a trivial solution. As the deviation increases, the weight on decreases. Intuitively, this weighting strategy places more weight on confident samples and penalizes more the mistakes made by these samples during training. For the classification task, the likelihood is formulated as the softmax function:
[0126]
[0127] where the temperature t i controls the flatness of the distribution. and y i are respectively the logarithm and the true annotation label of. The distribution of is actually the Boltzmann distribution. To make its form consistent with the form of the regression task, define Let (obtained ), the loss function calculation device 34 approximates the classification loss as:
[0128]
[0129] The loss function calculation device 34 combines the weighted classification loss (Equation (10)) with the weighted regression loss (Equation (8)), resulting in the following total loss:
[0130]
[0131] Note the direct prediction bring implementation difficulties because is expected to be a positive number and placing in the denominator has the potential risk of division by zero. According to an embodiment of the present invention, in order to further optimize the equation, prediction is adopted so that the optimization is numerically more stable and allows unconstrained prediction output. The loss function calculation device 34 makes the final total loss function become:
[0132]
[0133] The present invention customizes different weights for each sample and thus allowing adjustment of the multi-task balance weight at the sample level. The loss function calculation device 34 can effectively drive the network to learn useful sample weights through network design.
[0134] After the loss function calculation device 34 calculates the loss function L according to Equation (12) i , it provides the loss function to the gradient calculation device 36. The gradient calculation device 36 calculates the gradient based on the loss function and provides it to the sample weighting network 32. Then, the feature transformation adjustment device 301 adjusts the functions used by the first feature transformation device 3021, the second feature transformation device 3022, the third feature transformation device 3023, and the fourth feature transformation device 3024 based on the gradient. Since the feature transformation device 302 can be dynamically adjusted based on the gradient, the classification loss weight and regression loss weight of each sample can be dynamically learned. The above shows a sample weighting learning system for object detection in images. Obviously, an image sample weight generation system (not shown) composed of the feature transformation device 302, the sample feature generation device 303, and the weight prediction device 304 of the present invention can be used to predict the sample weights of classification loss and regression loss for samples.
[0135] Next, in combination with Figure 3 and Figure 4 describe a sample weighting learning method for object detection according to an embodiment of the present invention.
[0136] In S410, the sample weighting learning system for object detection in images receives the input features related to the sample.
[0137] In S420, the first feature transformation device 3021, the second feature transformation device 3022, the third feature transformation device 3023, and the fourth feature transformation device 3024 respectively transform the input features related to the image sample to obtain dense features.
[0138] In S430, the sample feature generation device 303 generates joint sample features based on the transformed dense features.
[0139] In S440, the weight prediction device 304 generates sample weights for the classification loss and sample weights for the regression loss according to the generated sample features.
[0140] In S450, the loss function calculation device 34 calculates the loss function according to the predicted sample weights for the classification loss and the sample weights for the regression loss.
[0141] In S460, the gradient calculation device 36 calculates the gradient according to the calculated loss function.
[0142] In S470, the feature transformation adjustment device 301 adjusts the first feature transformation device 3021, the second feature transformation device 3022, the third feature transformation device 3023, and the fourth feature transformation device 3024 according to the obtained gradient, so that the first feature transformation device 3021, the second feature transformation device 3022, the third feature transformation device 3023, and the fourth feature transformation device 3024 execute S420 - S470 on the next sample again until all samples are processed. The present invention can execute S410 - S470 based on each sample. Preferably, the present invention executes S410 - S470 according to each batch (each group) of samples. Each batch of samples can include at least four samples, thereby improving the calculation efficiency.
[0143] The method according to the present invention can calculate the loss function by adaptively learning the sample weights for the classification loss and the sample weights for the regression loss in object detection of image processing, thereby more preferably training an accurate object detection model.
[0144] Since the sample - weighted learning system for object detection according to the present invention makes no assumptions about the basic object detector, this means that it can be used with most region - based object detectors, including Faster R - CNN, RetinaNet, and Mask R - CNN. The method according to the present invention is general and only makes minimal modifications to the original framework. Faster R - CNN consists of a Region Proposal Network (RPN) and a Fast R - CNN network. Keep the RPN unchanged and insert the sample - weighted learning method for object detection of images according to the present invention into the Fast R - CNN branch. For each sample, first calculate and Prob i as the input of the Sample - Weighted Network (SWN). Then the predicted weights and Insert into Equation (12) and backpropagate the gradient to the basic detection network and the sample weighting network. For RetinaNet, the present invention follows a similar process to generate classification weights and regression weights for each sample. Since Mask R-CNN has an additional mask branch, the method of the present invention can incorporate another branch into the sample weighting network to generate adaptive weights for the mask loss, where classification, bounding box regression, and mask prediction are jointly estimated. According to an example of the present invention, in order to match the additional mask weights, the mask loss is also used as an input to the sample weighting network, so that the mask loss and the other four inputs and Prob i are used together as inputs to the sample weighting network to calculate the prediction weights and (perform steps S410 - S470). Thus, according to an embodiment of the present invention, the sample weighting network for image processing may further include a fifth feature transformation device (not shown) for transforming the input mask loss using another function to obtain dense features, and the function may be implemented by an MLP neural network. After that, the dense features obtained by the fifth feature transformation device and the four dense features obtained by the four functions F, G, H, and K are provided to the sample feature generation unit 303 to generate sample features. The weight prediction device 304 generates sample weights for the classification loss and sample weights for the regression loss based on the generated sample features.
[0145] According to an embodiment of the present invention, the sample weight generation system (or method) for image processing or the sample weighting learning system (or method) for image processing can be applied to a system (or method) for object detection in image processing (not shown), where the system for object detection in image processing may include an input device for receiving an input image, and the sample weighting learning system according to the present invention is used to perform sample weighting learning on the input image to obtain sample weights for the classification loss and sample weights for the regression loss of the image, so as to perform object detection or perform object detection according to the sample weights generated by the sample weight generation system.
[0146] According to an embodiment of the present invention, it is found that the predicted classification weights are unstable because the uncertainty between negative samples and positive samples is much greater than the uncertainty of regression. Therefore, the classification weights of positive samples and negative samples in each batch are averaged respectively as a smoothed version of the classification loss weight prediction. Each batch may be the number of samples defined manually during the training process.
[0147] Figures 5(a)-5(d) show the qualitative performance comparison between RetinaNet and RetinaNet+SWN on the COCO dataset. Following the common threshold of 0.5 for visualizing detected objects, the present invention only illustrates detections when their scores are higher than the threshold. As shown in Figures 5(a)-5(d), some so-called "easy" objects (such as children, sofas, baby bottles, etc.) missed by RetinaNet have been successfully detected by the enhanced RetinaNet with the Sample Weighted Network (SWN). The present invention speculates that the original RetinaNet may have focused too much on "difficult" samples. As a result, "easy" samples have received less attention and contributed less to model training. As a result, the scores of these "easy" samples have been reduced, leading to non-detection. The purpose of Figures 5(a)-5(d) is not to show the "weak points" of RetinaNet in score calibration, because "easy" samples can be detected anyway when the threshold is reduced. It actually shows that the sample weighted learning system of the present invention does not give smaller weights to "easy" samples.
[0148] There is another research method aimed at improving bounding box regression. In other words, they attempt to optimize the regression loss by using IoU as supervision or learning in combination with NMS. Based on the Faster R-CNN+ResNet-50+FPN framework, the present invention conducts a comparison on COCO val2017, and the performance comparison shows that the sample weighted learning system according to the present invention and its extension SWN+Soft-NMS are superior to IoU-Net and IoU-Net+NMS. The performance comparison further confirms the advantage of learning sample weights for both classification and regression.
[0149] Figure 6 A block diagram of an electronic device for image processing according to an embodiment of the present disclosure is schematically shown. Figure 6 The electronic device shown is merely an example and should not impose any limitation on the functions and scope of use of the embodiments of the present disclosure.
[0150] As Figure 6 shown, the electronic device 600 includes a processor 610 and a computer-readable storage medium 620. The electronic device 600 can execute the method according to the embodiment of the present disclosure.
[0151] Specifically, the processor 610 may include, for example, a general - purpose microprocessor, an instruction - set processor, and / or a related chipset, and / or a dedicated microprocessor (e.g., an application - specific integrated circuit (ASIC)), etc. The processor 610 may also include on - board memory for caching purposes. The processor 610 may be a single processing unit or multiple processing units for performing different actions of the method flow according to the embodiments of the present disclosure.
[0152] The computer - readable storage medium 620 may be, for example, a non - volatile computer - readable storage medium. Specific examples include, but are not limited to: magnetic storage devices, such as magnetic tapes or hard disk drives (HDDs); optical storage devices, such as compact discs (CD - ROMs); memories, such as random access memories (RAMs) or flash memories; etc.
[0153] The computer - readable storage medium 620 may include a computer program 621. The computer program 621 may include code / computer - executable instructions that, when executed by the processor 610, cause the processor 610 to execute the method according to the embodiments of the present disclosure or any variation thereof.
[0154] The computer program 621 may be configured to have computer program code that includes, for example, computer program modules. For example, in an exemplary embodiment, the code in the computer program 621 may include one or more program modules, such as module 621A, module 621B,.... It should be noted that the way of dividing the modules and the number of modules are not fixed. Those skilled in the art can use appropriate program modules or combinations of program modules according to the actual situation. When these combinations of program modules are executed by the processor 610, the processor 610 can execute the method according to the embodiments of the present disclosure or any variation thereof.
[0155] According to the embodiments of the present disclosure, Figure 3 at least one of the devices or apparatuses shown may be implemented as a computer program module as described with reference to Figure 6 which, when executed by the processor 610, can implement the corresponding operations described above.
[0156] The present invention also provides a computer - readable storage medium. The computer - readable storage medium may be included in the device / equipment / system described in the above - mentioned embodiments; or it may exist separately without being assembled into the device / equipment / system. The above - mentioned computer - readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present invention is implemented.
[0157] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions denoted in the blocks may occur in a different order than that denoted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0158] Those skilled in the art will appreciate that although the present invention has been shown and described with reference to specific exemplary embodiments thereof, those skilled in the art should understand that various changes in form and detail may be made therein without departing from the spirit and scope of the present invention as defined by the appended claims and their equivalents. Therefore, the scope of the present invention should not be limited to the above embodiments, but should be determined not only by the appended claims but also by the equivalents of the appended claims.
Claims
1. A sample weight generation system for image processing, the system comprising: A feature transformation device for transforming the input first feature, second feature, third feature, and fourth feature into a first dense feature, a second dense feature, a third dense feature, and a fourth dense feature respectively, where the first feature represents the classification loss for a sample, the second feature represents the regression loss for the sample, the third feature represents the intersection over union for the sample, and the fourth feature represents the classification probability for the sample; A sample feature generation device for generating a joint sample feature according to the first dense feature, the second dense feature, the third dense feature, and the fourth dense feature using a transformation function; A weight prediction device for predicting the sample weight of the classification loss and the sample weight of the regression loss for a sample according to the generated joint sample feature, where the sample weight of the classification loss and the sample weight of the regression loss are used to adjust the multi-task balance weight at the sample level; Wherein, the transformation function is adjusted based on a loss function, and the loss function is determined based on the predicted sample weight of the classification loss and the sample weight of the regression loss.
2. A sample weighted learning system for object detection in images, the system comprising: An input device for receiving the input first feature, second feature, third feature, and fourth feature for each sample, where the first feature represents the classification loss for each sample, the second feature represents the regression loss for each sample, the third feature represents the intersection over union for each sample, and the fourth feature represents the classification probability for each sample; A feature transformation device for transforming the input first feature, second feature, third feature, and fourth feature into a first dense feature, a second dense feature, a third dense feature, and a fourth dense feature respectively; A sample feature generation device for generating a joint sample feature according to the first dense feature, the second dense feature, the third dense feature, and the fourth dense feature using a transformation function; A weight prediction device for predicting the sample weight of the classification loss and the sample weight of the regression loss for each sample according to the generated joint sample feature, where the sample weight of the classification loss and the sample weight of the regression loss are used to adjust the multi-task balance weight at the sample level; A loss function calculation device for calculating a loss function according to the predicted sample weight of the classification loss and the sample weight of the regression loss; and A feature transformation adjustment device for adjusting the transformation function used by the sample feature generation device based on the calculated loss function.
3. A method for generating sample weights for image processing, the method comprising: Transform the input first feature, second feature, third feature, and fourth feature into a first dense feature, a second dense feature, a third dense feature, and a fourth dense feature respectively, where the first feature represents the classification loss for a sample, the second feature represents the regression loss for the sample, the third feature represents the intersection over union for the sample, and the fourth feature represents the classification probability for the sample; Generate a joint sample feature according to the first dense feature, the second dense feature, the third dense feature, and the fourth dense feature using a transformation function; Predict the sample weights for the classification loss and the sample weights for the regression loss for the samples according to the generated joint sample features, where the sample weights for the classification loss and the sample weights for the regression loss are used to adjust the multi-task balance weights at the sample level; Among them, the transformation function is adjusted based on the loss function, and the loss function is determined based on the predicted sample weights for the classification loss and the sample weights for the regression loss.
4. The sample weight generation method according to claim 3, wherein: The predicted sample weights for the classification loss and the sample weights for the regression loss are obtained from a first exponential function and a second exponential function respectively.
5. A method for sample weighted learning for object detection in images, the method comprising: First step: Receive the input first feature, second feature, third feature, and fourth feature for each sample, where the first feature represents the classification loss for each sample, the second feature represents the regression loss for each sample, the third feature represents the intersection over union for each sample, and the fourth feature represents the classification probability for each sample; Second step: Transform the input first feature, second feature, third feature, and fourth feature into a first dense feature, a second dense feature, a third dense feature, and a fourth dense feature respectively; Third step: Use the transformation function to generate joint sample features according to the first dense feature, the second dense feature, the third dense feature, and the fourth dense feature; Fourth step: Predict the sample weights for the classification loss and the sample weights for the regression loss for each sample according to the generated joint sample features, where the sample weights for the classification loss and the sample weights for the regression loss are used to adjust the multi-task balance weights at the sample level; Fifth step: Calculate the loss function according to the predicted sample weights for the classification loss and the sample weights for the regression loss; Sixth step: Adjust the transformation function according to the calculated loss function.
6. The sample weighted learning method according to claim 5, wherein the first step further comprises receiving a fifth feature; and the method further comprises a step of transforming the input fifth feature into a fifth dense feature.
7. The sample weighted learning method according to claim 6, wherein the fifth feature is a mask loss.
8. The sample weighted learning method according to claim 5, wherein: The predicted sample weights for the classification loss and the sample weights for the regression loss are obtained from a first exponential function and a second exponential function respectively.
9. The sample weighted learning method according to claim 5, wherein: Average the sample weights for the classification loss of a set of samples including positive and negative samples as the sample weights for the classification loss of each sample.
10. The sample weighted learning method according to claim 5, wherein the sixth step further includes: Derive the gradient according to the calculated loss function, and then adjust the transformation function according to the derived gradient.
11. A method for object detection in image processing, comprising: Receive the input image; And Use the sample weight generation method according to any one of claims 3 to 4 or the sample weighted learning method according to any one of claims 5 to 10 to obtain the sample weights for the classification loss and the sample weights for the regression loss of the image for object detection.
12. A system for object detection in image processing, comprising: An input device for receiving the input image; And The sample weight generation system according to claim 1 or the sample weighted learning system according to claim 2, which is used to obtain the sample weights for the classification loss and the sample weights for the regression loss of the input image for object detection.
13. An electronic device for image processing, the device includes a memory, a processor, and a computer program stored on the memory, wherein when the processor executes the computer program, it implements the sample weight generation method according to any one of claims 3 to 4 or the sample weighted learning method according to any one of claims 5 to 10.
14. A computer-readable storage medium storing a computer program with program code, the program code is used to execute the sample weight generation method according to any one of claims 3 to 4 or the sample weighted learning method according to any one of claims 5 to 10 when running on a computer.
Citation Information
Patent Citations
Neural network training and image processing method and device
CN110009090A
Three-dimensional object detection method and device, electronic equipment and storage medium
CN110751040A