Target detection method based on background perception and similarity matrix
By introducing background-aware classification loss and similarity matrix loss, the background recognition and category distinction capabilities of the model are optimized, and the detection error caused by complex background and inter-class feature similarity is solved, which improves the accuracy and robustness of target detection.
Patent Information
- Application Number
- CN202510521353.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-22
AI Technical Summary
When the existing object detection model deals with complex backgrounds and similar features between classes, it is easy to misjudgment the background as the target, and it is difficult to effectively distinguish similar categories, resulting in a decrease in detection accuracy.
Background-aware classification loss and similarity matrix loss are introduced. By calculating background-aware classification loss and similarity matrix loss, the background recognition and category distinction capabilities of the model are optimized. A feature extraction network is used to generate multi-scale feature maps, and the real background anchor box is screened through interleaving and comparison calculations, and the loss function is optimized in combination with the Softmax reweighting mechanism.
It significantly improves the model's ability to identify complex backgrounds and distinguish similar categories, and improves the accuracy and robustness of object detection.
Smart Images

Figure CN120355885A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition, and particularly to an optimization method for object detection in complex scenarios where foreground-background similarity and inter-class feature similarity exist. Background Art
[0002] With the continuous development of deep learning, object detection has evolved from traditional object detection algorithms that require manual feature extraction to deep learning-based object detection algorithms that can automatically learn image features. In deep learning-based object detection tasks, the detection performance of the model depends on its effective discrimination between foreground objects and background regions. Ideally, the foreground and background should have obvious distinctiveness in the feature space to ensure the accurate recognition and localization of foreground objects by the model. However, in real-world application scenarios, target images often contain a large number of complex background elements, and these background regions are often very similar to real objects in terms of visual features such as color, texture, and structure, resulting in the model easily misjudging the background as an object. In addition, in multi-class object detection, there may be a high visual similarity between different classes, with only minor differences between them. The high inter-class similarity further exacerbates the difficulty of the model in discrimination, easily leading to classification confusion and reducing the overall detection accuracy.
[0003] Although existing object detection models have introduced technologies such as feature pyramid networks, attention mechanisms, and residual networks to enhance feature expression capabilities, they still face the following main deficiencies in dealing with the two core problems of foreground-background similarity and inter-class feature overlap:
[0004] 1) Lack of a dedicated recognition mechanism for background features, which easily misjudges complex backgrounds as objects.
[0005] 2) The features between similar classes have a high degree of overlap in the high-dimensional space, and the inter-class boundaries are blurred, making it difficult for the classifier to effectively distinguish them. Summary of the Invention
[0006] To solve the above technical problems in the background, the present invention proposes a background-aware classification loss and a similarity matrix loss, aiming to address the detection difficulties brought about by the similarity between foreground object and background region features and the similarity between object category features.
[0007] To achieve the above object, the present application provides an object detection method based on background awareness and similarity matrix, and the steps include:
[0008] Obtain an input image and generate a multi-scale feature map of the image through a feature extraction network;
[0009] Calculate the intersection over union (IoU) between randomly generated anchor boxes and ground truth bounding boxes based on the multi-scale feature map to obtain true background anchor boxes;
[0010] Align the predicted background region with the true background anchor box based on the true background anchor box, and calculate the background-aware classification loss;
[0011] Construct a class similarity matrix for the predicted scores of target classes, and combine the Softmax reweighting mechanism to calculate the similarity matrix loss;
[0012] Jointly optimize the background-aware classification loss and the similarity matrix loss to improve the overall object detection accuracy and robustness.
[0013] Preferably, the method for generating the multi-scale feature map includes: first using a feature extraction network to extract features of the target to be detected; then using a multi-scale feature fusion network to perform multi-scale feature fusion on the extracted features to obtain the multi-scale feature map.
[0014] Preferably, according to a set intersection over union (IoU) threshold, filter out the anchor boxes with an overlap degree lower than the threshold with the true annotation box from the anchor boxes randomly generated based on the multi-scale feature map as the true background anchor boxes.
[0015] Preferably, according to the predicted scores of each candidate region by the model, when the maximum predicted score of all classes in the region is lower than a preset confidence threshold, determine that the region is a predicted background region.
[0016] Preferably, the binary cross-entropy loss function is used to calculate the background-aware classification loss, and the calculation formula is as follows:
[0017]
[0018] where M bg represents the true background anchor box; P bg represents the predicted background region; BCE represents the binary cross-entropy loss.
[0019] Preferably, the similarity matrix is constructed by calculating the dot product between all candidate target class scores, taking the absolute value of the dot product result, and removing the diagonal elements.
[0020] Preferably, the Softmax reweighting mechanism weights the similarity matrix, where each weight is dynamically determined by the corresponding similarity value and the exponential parameter β, and the calculation formula is as follows:
[0021] W=(softmax(M)) β ,β>1
[0022] where W is the Softmax reweighting matrix; M is the similarity matrix; β is the exponential parameter.
[0023] Preferably, the calculation formula for the similarity matrix loss is as follows:
[0024]
[0025] Among them, C is the number of categories; C(C - 1) is the number of non - diagonal elements of the similarity matrix; w jk is the re - weighted matrix value; m jk is the similarity matrix value.
[0026] Preferably, based on the background - aware classification loss and the similarity matrix loss, a total loss function is constructed; the network model is trained using the total loss function to complete object detection.
[0027] Compared with the prior art, the beneficial effects of this application are as follows:
[0028] By introducing the background - aware classification loss and the class similarity matrix loss, the present invention starts from two key aspects: foreground - background feature discrimination and inter - class feature differences. The background - aware classification loss explicitly constrains the model to identify the background region and suppress its interference with foreground objects, enhancing the saliency expression of foreground objects; at the same time, the class similarity matrix loss enables the model to focus on reducing the discriminative interference between highly similar classes during training, encouraging the model to learn the feature differences between classes, thereby improving the model's recognition ability for classes with visually similar features. The present invention effectively solves the detection errors caused by complex backgrounds or similar classes, and solves the problems in the prior art that the model is vulnerable to background interference and difficult to distinguish similar classes. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] To more clearly illustrate the technical solutions of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0030] Figure 1 is the schematic flowchart of the method of the embodiment of this application;
[0031] Figure 2 is the schematic overall network flowchart of the embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.
[0033] Embodiment 1
[0034] As shown in Figure 1 , the following are the steps of the method flow diagram of this embodiment:
[0035] S1. Obtain the input image and generate a multi-scale feature map of the image through the feature extraction network.
[0036] In order to obtain the position and category information of the target, first use the feature extraction network to extract the necessary feature information in the input image to lay the foundation for the subsequent tasks of the network; then use the multi-scale fusion network to fuse the features extracted by the feature extraction network, and the feature maps of different depths in the fusion network to enrich the feature information extracted by the multi-scale fusion network.
[0037] In this embodiment, the feature extraction network can select VGG, Resnet, Darknet, etc.; in this embodiment, the multi-scale feature fusion network can select FPN, PANet, etc.
[0038] S2. Calculate the intersection over union (IoU) between the randomly generated anchor boxes based on the multi-scale feature map and the ground truth bounding boxes to obtain the true background anchor boxes.
[0039] According to the set IoU threshold, select the anchor boxes with an overlap degree lower than the threshold with the ground truth bounding boxes from the randomly generated anchor boxes based on the multi-scale feature map as the true background anchor boxes.
[0040] S3. Align the predicted background region with the true background anchor boxes based on the true background anchor boxes, and calculate the background-aware classification loss.
[0041] According to the predicted scores of each candidate region by the model, when the maximum predicted score of all classes within the region is lower than the preset confidence threshold, determine that the region is a predicted background region.
[0042] The binary cross-entropy loss function is used to calculate the background-aware classification loss, and the calculation formula is as follows:
[0043]
[0044] where M bg represents the true background anchor boxes; P bg represents the predicted background region; BCE represents the binary cross-entropy loss.
[0045] S4. Construct a class similarity matrix for the predicted scores of the target classes, and combine the Softmax reweighting mechanism to calculate the similarity matrix loss.
[0046] The similarity matrix is constructed by calculating the dot product between all candidate target class scores, taking the absolute value of the dot product result, and removing the diagonal elements.
[0047] The Softmax reweighting mechanism weights the similarity matrix, where each weight is dynamically determined by the corresponding similarity value and the exponential parameter β, and the calculation formula is as follows:
[0048] W = (softmax(M)) β , β > 1
[0049] where W is the Softmax reweighting matrix; M is the similarity matrix; β is the exponential parameter.
[0050] The calculation formula for the similarity matrix loss is as follows:
[0051]
[0052] where C is the number of classes; C(C - 1) is the number of non - diagonal elements of the similarity matrix; w jk is the value of the reweighting matrix; m jk is the value of the similarity matrix.
[0053] S5. Jointly optimize the background - aware classification loss and the similarity matrix loss to improve the overall object detection accuracy and robustness.
[0054] As Figure 2 shown, it is the overall network flow diagram of this embodiment, and the steps include:
[0055] In the first step, obtain the input image and generate a multi - scale feature map of the image through the feature extraction network.
[0056] To obtain the location and category information of the target, first use the feature extraction network to extract the necessary feature information in the input image to lay the foundation for the subsequent tasks of the network; then use the multi - scale fusion network to fuse the features extracted by the feature extraction network, and the feature maps of different depths in the fusion network to enrich the feature information extracted by the multi - scale fusion network.
[0057] In this embodiment, the feature extraction network can be selected from VGG, Resnet, Darknet, etc.; in this embodiment, the multi - scale feature fusion network can be selected from FPN, PANet, etc.
[0058] In the second step, calculate the intersection - over - union (IoU) between the randomly generated anchor boxes based on the multi - scale feature map and the ground - truth bounding boxes to obtain the true background anchor boxes.
[0059] According to the set IoU threshold, select the anchor boxes with an overlap degree lower than the threshold with the ground - truth bounding boxes from the randomly generated anchor boxes based on the multi - scale feature map as the true background anchor boxes.
[0060] In the third step, based on the real background anchor boxes, align the predicted background regions with the real background anchor boxes, and calculate the background-aware classification loss.
[0061] According to the predicted scores of each candidate region by the model, when the maximum predicted score of all categories within the region is lower than the preset confidence threshold, the region is determined as a predicted background region.
[0062] The binary cross-entropy loss function is used to calculate the background-aware classification loss, and the calculation formula is as follows:
[0063]
[0064] where M bg represents the real background anchor box; P bg represents the predicted background region; BCE represents the binary cross-entropy loss.
[0065] In the fourth step, construct a class similarity matrix for the predicted scores of the target categories, and combine the Softmax reweighting mechanism to calculate the similarity matrix loss.
[0066] The similarity matrix is constructed by calculating the dot product between all candidate target category scores, taking the absolute value of the dot product result, and removing the diagonal elements.
[0067] The Softmax reweighting mechanism weights the similarity matrix, where each weight is dynamically determined by the corresponding similarity value and the exponential parameter β, and the calculation formula is as follows:
[0068] W = (softmax(M)) β , β > 1
[0069] where W is the Softmax reweighting matrix; M is the similarity matrix; β is the exponential parameter.
[0070] The calculation formula for the similarity matrix loss is as follows:
[0071]
[0072] where C is the number of categories; C(C - 1) is the number of non-diagonal elements of the similarity matrix; w jk is the value of the reweighting matrix; m jk is the value of the similarity matrix.
[0073] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the technical field of the present application within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A target detection method based on background awareness and similarity matrix, characterized in that, The steps are as follows: Obtain an input image and generate a multi-scale feature map of the image through a feature extraction network; Calculate the intersection over union (IoU) between randomly generated anchor boxes based on the multi-scale feature map and the ground truth bounding boxes to obtain the true background anchor boxes; Based on the true background anchor boxes, align the predicted background regions with the true background anchor boxes and calculate the background-aware classification loss; Construct a class similarity matrix for the predicted scores of target classes and calculate the similarity matrix loss in combination with the Softmax reweighting mechanism; Jointly optimize the background-aware classification loss and the similarity matrix loss to improve the overall object detection accuracy and robustness.
2. The object detection method based on background perception and similarity matrix according to claim 1, characterized in that, The method for generating the multi-scale feature map includes: first, using a feature extraction network to extract features of the object to be detected; then, using a multi-scale feature fusion network to perform multi-scale feature fusion on the extracted features to obtain the multi-scale feature map.
3. The object detection method based on background perception and similarity matrix according to claim 1, characterized in that According to the set IoU threshold, select the anchor boxes with an overlap degree lower than the threshold with the ground truth bounding boxes from the anchor boxes randomly generated based on the multi-scale feature map as the true background anchor boxes.
4. The object detection method based on background perception and similarity matrix according to claim 1, characterized in that, According to the predicted scores of each candidate region by the model, when the maximum predicted score of all classes in the region is lower than the preset confidence threshold, determine the region as a predicted background region.
5. The object detection method based on background perception and similarity matrix according to claim 1, characterized in that The binary cross-entropy loss function is used to calculate the background-aware classification loss, and the calculation formula is as follows: Among them, M bg represents the real background anchor box; P bg represents the predicted background region; BCE represents the binary cross-entropy loss.
6. The object detection method based on background perception and similarity matrix according to claim 1, characterized in that The similarity matrix is constructed by calculating the dot product between all candidate target class scores, taking the absolute value of the dot product result, and removing the diagonal elements.
7. The object detection method based on background awareness and similarity matrix according to claim 1, characterized in that The Softmax reweighting mechanism weights the similarity matrix, where each weight is dynamically determined by the corresponding similarity value and the exponential parameter β, and the calculation formula is as follows: W = (softmax(M)) β , β > 1 Where, W is the Softmax reweighting matrix; M is the similarity matrix; β is the exponential parameter.
8. The object detection method based on background awareness and similarity matrix according to claim 1, characterized in that The calculation formula for the similarity matrix loss is as follows: Among them, C is the number of categories; C(C - 1) is the number of non-diagonal elements of the similarity matrix; w jk is the reweighted matrix value; m jk is the similarity matrix value.
9. The object detection method based on background perception and similarity matrix according to claim 1, characterized in that Based on the background-aware classification loss and the similarity matrix loss, construct a total loss function; use the total loss function to train the network model to complete object detection.