A target detection method in complex backgrounds based on a convolutional attention mechanism

By introducing a convolutional attention mechanism into the object detection algorithm in complex backgrounds, building a ForegroundNet network model, extracting and fusion of foreground, background and edge features, the problem that existing algorithms are difficult to accurately detect the foreground target edge areas under complex backgrounds, and achieving higher detection accuracy and performance.

CN114882241BActive Publication Date: 2025-05-27SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210558583.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-20
Publication Date
2025-05-27
Estimated Expiration
2042-05-20

AI Technical Summary

Technical Problem

The existing object detection algorithms in complex backgrounds are difficult to accurately detect the edge areas of the foreground targets under the complex background of the picture, which often leads to misjudgment and misdetection.

Method used

A target detection method in a complex context based on convolutional attention mechanism is proposed. By constructing a ForegroundNet network model, a foreground background proposal module, a feature generation module and a convolutional attention-based feature fusion decoding module are used to extract and fuse the foreground features, background features and edge features to improve the detection ability of the foreground target edge area.

Benefits of technology

The detection accuracy of the model foreground target edge areas is improved, and the detection performance is improved, which is manifested as a significant improvement in detection indicators (such as SM, EM, WFM and MAE) in complex contexts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114882241B_ABST
    Figure CN114882241B_ABST
Patent Text Reader

Abstract

The present invention discloses an object detection method based on a convolutional attention mechanism in complex backgrounds, which can be used to accurately detect foreground objects in complex backgrounds. The invention mainly includes: obtaining an object detection data set in complex backgrounds that is publicly available, constructing a training set, a validation set, and a test set; constructing an artificial neural network model ForegroundNet model based on the convolutional attention mechanism; using the training set to perform supervised training on the ForegroundNet network model on the Pytorch deep learning platform; evaluating the detection performance of the converged ForegroundNet model on the constructed test set. Compared with the current main complex background object detection algorithms, the present invention can accurately detect the edge regions of foreground objects, thereby achieving higher detection performance. The average absolute error corresponding to the detection results of the present invention on the test set is lower, and the enhanced alignment index, structure index, and weighted F index are higher. It is a more accurate object detection algorithm in complex backgrounds.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an object detection method based on a convolutional attention mechanism in complex backgrounds, which is applicable to the technical field of object detection in complex backgrounds in computer vision. Background Art

[0002] With the development of human society, the important sources for humans to obtain information have gradually become images and videos. How to design algorithms to process and utilize these massive amounts of generated data is an urgent need in the industrial community for the development of computer vision technology. Object detection in complex backgrounds is a challenging task in the field of computer vision. The purpose is to detect foreground objects of interest under the condition of complex backgrounds and output a segmentation map of the foreground objects as the detection result. Object detection in complex backgrounds has very high research value and has extensive applications in fields such as medical, agriculture, ocean, and military.

[0003] Benefiting from the development of deep learning technology, object detection methods based on neural networks have achieved great success, and a large number of efficient general object detection algorithms have been proposed. However, under complex background conditions, pictures usually have characteristics such as chaotic colors, variable lighting conditions, and foreground objects with camouflage colors, resulting in foreground objects being easily integrated with the background of the picture and being difficult to detect. General object detection algorithms often cannot achieve good detection results in pictures with complex backgrounds. Therefore, it is necessary to optimize the network model specifically for the characteristics of complex backgrounds. Currently, most of the existing object detection algorithms in complex backgrounds are optimized from aspects such as enhancing image features and multi-scale feature fusion, and have achieved good detection performance. However, due to the interference of complex backgrounds, the colors and textures of foreground objects are usually similar to those of the background, making it difficult to judge the edge details of foreground objects. Since there is no effective optimization method for detecting the edge regions of objects from the physical meaning aspect, the current main complex background object detection algorithms can accurately detect the positions and approximate contours of the main regions of foreground objects, but cannot accurately and meticulously detect the easily confused regions at the edges of foreground objects, often resulting in misjudgment and false detection phenomena. Summary of the Invention

[0004] Aiming at the problem that the existing complex background object detection algorithms cannot accurately detect the edge regions of foreground objects, the present invention proposes an object detection method under complex backgrounds based on a convolutional attention mechanism, and constructs a ForegroundNet network model. This network model forms foreground features, background features, and edge features through a foreground-background proposal module and a feature generation module, and further aggregates the information related to foreground objects in the three types of features through convolutional attention, improving the model's detection ability for the edge regions of foreground objects, thereby improving the overall detection performance of the model. The detection performance of the ForegroundNet model converging to the optimal performance is better than the mainstream models in the field of complex background object detection. It can reduce the Mean Absolute Error (MAE) of the detection results while improving the Enhanced-alignment Measure (EM), Structure Measure (SM), and Weighted F Measure (WFM) of the detection results, indicating that the present invention can effectively improve the object detection accuracy of the model.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] An object detection method under complex backgrounds based on a convolutional attention mechanism, characterized by comprising the following steps:

[0007] Step S1: Obtain a publicly available complex background object detection dataset, and construct a training set, a validation set, and a test set;

[0008] Step S2: Construct a ForegroundNet network model based on a convolutional attention mechanism;

[0009] Step S3: Supervise and train the constructed ForegroundNet model on the constructed training set until the model converges to the optimal performance;

[0010] Step S4: Test the converged ForegroundNet model on the constructed test set to evaluate the detection performance of the model under complex backgrounds.

[0011] Further, the specific steps of step S1 include:

[0012] Step S101: Obtain a publicly available miscellaneous background object detection dataset, including the COD10K dataset, the CAMO dataset, and the CHAMELEON dataset;

[0013] Step S102: Based on the acquired complex-background object detection datasets COD10K and CAMO, construct a training set containing 4040 pairs of picture-label pairs, a validation set containing 101 pairs of image-label pairs, and a test set containing 2353 pairs of picture-label pairs.

[0014] Furthermore, the network model ForegroundNet model constructed in step S2 mainly includes: a feature extraction module, a foreground-background proposal module, a feature generation module, and a convolutional attention-based feature fusion decoding module.

[0015] Furthermore, step S3 specifically includes:

[0016] Step S301: Randomly extract training pictures from the constructed training set for preprocessing. First, use the interpolation algorithm to resize the input image and the corresponding ground truth label to H×W, where H represents the image height and W represents the image width. Subsequently, perform image data augmentation processing, and finally normalize the image and input it into the ForegroundNet network model for supervised training;

[0017] Step S302: The ForegroundNet model extracts multi-scale abstract features through the feature extraction module. The foreground-background proposal module proposes foreground-background segmentation maps based on the input features and outputs N layers of proposal results The proposal results include foreground object detection results and background object detection results Subsequently, the feature generation module generates foreground features, background features, and edge features based on the proposal results. The convolutional attention-based feature fusion decoding module fuses and decodes the foreground, background, and edge features of each layer, and outputs M layers of foreground object detection results based on the decoded features, denoted as where k = N + 1, N + 2,..., N + M;

[0018] Step S303: Based on the proposal results and detection results, calculate the loss function L of the model during the supervised training process overall , and its calculation method is shown in Equation (1).

[0019]

[0020] where and respectively represent the cross-entropy losses weighted by structural information corresponding to the k-th layer foreground object proposal result and background object proposal result ; and respectively represent the k-th layer foreground object proposal result and background object proposal result The corresponding intersection over union loss weighted by structural information. Taking the foreground object detection result as an example, the calculation methods of the cross-entropy loss and the intersection over union loss are shown in Equations (2) and (3).

[0021]

[0022] Among them, H represents the image height, W represents the image width, and mask gt (x, y) and respectively represent the true label of the foreground object and the value at the position coordinates (x, y) in the k-th foreground object proposal result. γ is a parameter related to the structural weight, and w(x, y) represents the structural weight corresponding to the position with coordinates (x, y). Its expression is as follows:

[0023]

[0024] Among them, A xy represents the set of pixels around the pixel with coordinates (x, y);

[0025] Step S304: Based on the overall loss function L overall of the model in the training stage, use the stochastic gradient descent (SGD) algorithm to iteratively update the network parameters of the ForegroundNet model;

[0026] Step S305: Fix the network parameters with the optimal convergence performance of the ForegroundNet model during training, input the image to be detected for forward calculation, and select the foreground object detection result with the largest scale among the N + M different detection results output as the final detection result mask pred of the model;

[0027] Furthermore, the specific steps of step S4 include:

[0028] Step S401: Read the images to be detected in the test set one by one, use the interpolation method to adjust their sizes to H×W, then normalize the images, and input them into the ForegroundNet model that converges to the optimal performance for forward calculation, and output the corresponding detection result mask pred ;

[0029] Step S402: According to the foreground object detection result mask pred of the model and the true label mask gt of the foreground object, calculate the objective evaluation indicators of the model on the test set, including: SM index, EM index, WFM index, and MAE index.

[0030] Finally, the target detection in complex backgrounds can be performed through the converged ForegroundNet network model. The image to be detected is input for forward calculation, and the predicted foreground target segmentation map is output as the detection result.

[0031] The beneficial effects of the present invention are as follows: The ForegroundNet model constructed by the present invention can construct foreground features, background features, and edge features through the detection results of the foreground-background proposal module, and uses the attention mechanism for layer-by-layer fusion decoding to enhance the model's detection ability for the confusing edge regions in the proposal results, thereby improving the detection accuracy of the model for foreground targets and achieving better detection performance. Specifically, the detection metrics of the ForegroundNet model on the test set are overall better than those of the mainstream models in the field of complex background target detection, and there are obvious improvements in the SM metric, EM metric, WFM metric, and MAE metric on the test set. Description of the Drawings

[0032] Figure 1 It is a flowchart of the target detection method in complex backgrounds based on the convolutional attention mechanism in Embodiment 1.

[0033] Figure 2 It is a structural diagram of the ForegroundNet network model in Embodiment 1.

[0034] Figure 3 It is an internal structural diagram of the proposal module in Embodiment 1.

[0035] Figure 4 It is an internal structural diagram of the feature mining module in Embodiment 1.

[0036] Figure 5 It is an internal structural diagram of the feature generation module in Embodiment 1.

[0037] Figure 6 It is an internal structural diagram of the convolutional attention decoding module in Embodiment 1.

[0038] Figure 7 It is an internal structural diagram of the multi-head convolutional attention decoding module in Embodiment 1.

[0039] Figure 8 It is a schematic diagram of the calculation process of convolutional attention decoding in the convolutional attention head in Embodiment 1.

[0040] Figure 9 It is a flowchart of the supervised training of the ForegroundNet model in Embodiment 1.

[0041] Figure 10 It is a comparison of the detection performance of the method of the present invention and several main methods in Embodiment 1 in terms of evaluation metrics.

[0042] Figure 11 For the comparison of the detection performance of the method of the present invention and several main methods in Example 1 in terms of visual effects. Detailed implementation mode

[0043] In order to make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0044] Example 1

[0045] See Figures 1-11 , this embodiment provides a target detection method under complex backgrounds based on a convolutional attention mechanism.

[0046] Specifically, see Figure 1 , this method specifically includes:

[0047] Step S1: Obtain a target detection data set under publicly available complex backgrounds, including: the COD10K data set, the CAMO data set, and the CHAMELEON data set, and construct a training set, a validation set, and a test set based on this;

[0048] More specifically, the constructed training set contains 3040 pairs of picture-label pairs in COD10K and 1000 pairs of image-label pairs in the CAMO data set, for a total of 4040 pieces of data; the constructed validation set contains 101 pairs of picture data pairs in the COD10K data set; the constructed test set contains 2026 pairs of picture-label pairs in COD10K, 250 pairs of picture-label pairs in the CAMO data set, and 76 pairs of picture-label pairs in the CHAMELEON data set, for a total of 2352 pieces of data.

[0049] Step S2: Construct a ForegroundNet network model based on a convolutional attention mechanism;

[0050] More specifically, the overall structure of the constructed ForegroundNet model is as Figure 2 shown. This network model is mainly composed of a feature extractor, a foreground-background proposal module, a feature generation module, and a convolutional attention-based feature fusion decoding module. Among them, the feature extractor uses the Res2Net-50 model. The feature extractor extracts 4 layers of features with different scales from the input picture and inputs them into the foreground-background proposal module. The foreground-background proposal module mainly predicts the foreground-background targets based on the 4 layers of features and outputs the proposal results. The feature generation module generates foreground features, background features, and edge features based on the proposal results, and inputs them into the feature fusion and decoding module based on convolutional attention for feature fusion and decoding, and outputs the foreground object detection results of 2 layers.

[0051] More specifically, the foreground and background proposal module is mainly composed of a receptive field enhancement module and a proposal module. After the input features pass through the receptive field enhancement module, the receptive field enhanced features are output. To the proposal module, the internal structure of the proposal module is as Figure 3 shown, mainly composed of a feature mining module and a detection module, responsible for detecting foreground objects and backgrounds based on the receptive field enhanced features.

[0052] a) The receptive field enhancement module is composed of different receptive field enhancement branches. Different receptive field branches contain two convolutional layers and one dilated convolutional layer. By setting convolutional kernels and dilation rates of different scales to simulate receptive fields of different scales, the features after multiple receptive field enhancements are concatenated in the channel dimension and then the final receptive field enhanced features are output.

[0053] b) The feature mining module is responsible for the fusion of multi-scale features, fusing the high-level feature proposal results and high-level features with the receptive field enhanced features of the current layer. Its internal structure is as Figure 4 shown. After the high-level feature input feature mining module, it first performs upsampling by a factor of two, and then passes through one convolutional layer to output the upsampled high-level features. The high-level foreground object proposal results and background object proposal results are respectively upsampled by a factor of two and then multiplied by the high-level features to obtain the weighted features F fg and F bg . The input features of the current layer are first added to F fg to enhance the foreground object information using the high-level features. After one normalization, it is then subtracted from F bg to suppress the interference brought by background irrelevant features using the high-level features. Finally, the result is output through one convolutional layer.

[0054] c) The detection module is mainly composed of a lightweight Unet network. In the encoding and decoding stages, the detection module only performs upsampling and downsampling once respectively.

[0055] More specifically, the internal structure of the feature generation module is as Figure 5 shown, responsible for generating foreground features, background features, and edge features according to the foreground object proposals background object proposals and receptive field enhanced features . First, the edge blurred region is calculated according to the method shown in the following formula

[0056]

[0057] Among them, is flipped from the background proposal result. The receptive field enhanced features are multiplied by the foreground proposal result, the background proposal result, and the edge blurred area respectively to output the foreground feature, the background feature, and the edge feature.

[0058] More specifically, the feature fusion decoding module based on convolutional attention is mainly composed of a convolutional attention decoding module, which is responsible for fusing and decoding the foreground feature, the background feature, and the edge feature of each layer. Its internal structure is as Figure 6 shown. The input feature passes through the multi-head convolutional attention decoding module to fuse and decode the foreground feature, the background feature, and the edge feature of this layer. After one layer of convolution, it forms a residual connection structure with the input feature to output the decoded feature. Among them, the multi-head convolutional attention decoding module is the key module for feature fusion decoding using the convolutional attention mechanism, and its internal structure is as Figure 7 shown. The input feature X k and the foreground feature background feature and the edge feature are split in the channel dimension in the multi-head convolutional attention module, divided into 4 components and respectively input into the convolutional attention heads for the calculation of convolutional attention fusion decoding. The calculation process is as Figure 8 shown. The input foreground, background, and edge features respectively generate key-value pair feature maps (K i , V i ) through the convolutional layer. The input feature X in generates a query feature map Q through the convolutional layer. The query feature map is respectively concatenated with the key-value feature maps K i of the three features in the depth dimension, and after one layer of convolution, the attention weight map W i is output. According to the attention weight, the value feature maps of the three features are weighted to obtain the decoding result X out of the convolutional attention head. Finally, the output results of the 4 convolutional attention heads are concatenated in the channel dimension as the final convolutional attention decoding result.

[0059] Step S3: Based on the Pytorch deep learning platform, use the constructed training set to supervise and train the ForegroundNet model, and use the SGD optimizer to optimize and update its parameters;

[0060] More specifically, the training process of the ForegroundNet model is as Figure 9 shown, including:

[0061] Step S301: Randomly and batchwise extract training images from the constructed training set for preprocessing. First, use the interpolation algorithm to resize the input image and the corresponding ground truth label to 384×384. Subsequently, perform data augmentation processing such as randomly adjusting the image attributes, randomly affine transformation, and random erasing on the image. Finally, after normalizing the image, input it into the ForegroundNet network model for supervised training;

[0062] Step S302: The ForegroundNet model extracts 4 layers of multi-scale abstract features through the feature extraction module and inputs them into the foreground-background proposal module. This module detects foreground objects and background objects based on the input features and outputs 4 layers of foreground-background proposal results The proposal results include foreground object detection results and background object detection results Subsequently, the feature generation module generates foreground features, background features, and edge features based on the proposal results. The feature fusion decoding module based on convolutional attention fuses and decodes the foreground, background, and edge features of each layer, and outputs 2 layers of foreground object detection results, denoted by ;

[0063] Step S303: Calculate the loss function L of the model during the supervised training process based on the 4 layers of proposal results and 2 layers of detection results , and its calculation method is shown in Equation (1). overall .

[0064]

[0065] Among them, and respectively represent the cross-entropy loss weighted by the structure information corresponding to the k-th layer foreground object proposal result and the background object proposal result . and respectively represent the intersection over union loss weighted by the structure information corresponding to the k-th layer foreground object proposal result and the background object proposal result . Taking the foreground object detection result as an example, the calculation methods of the cross-entropy loss and the intersection over union loss are shown in Equation (2) and Equation (3).

[0066]

[0067] Among them, H represents the image height, W represents the image width, mask gt (x,y) and respectively represent the true label of the foreground target and the value of the position coordinates (x, y) in the k-th foreground target proposal result. γ is a parameter related to the structure weight, and w(x, y) represents the structure weight corresponding to the position with coordinates (x, y), and its calculation method is as shown in Equation (6).

[0068]

[0069] Among them, A xy represents the set of pixels around the pixel with coordinates (x, y);

[0070] Step S304: Based on the loss function L of the model overall , use the stochastic gradient descent optimization algorithm to iteratively update the network parameters of the ForegroundNet model until it converges to the optimal performance;

[0071] Step S305: Fix the network parameters when the model converges to the optimal performance during training, input the image to be detected for forward calculation, and among the 6 different detection results output, select the foreground target detection result with the largest scale as the final detection result mask of the model pred .

[0072] Step S4: Perform detection on the constructed test set, and calculate detection performance indicators such as the SM index, EM index, WFM index, and MAE index based on the detection results and the true labels, so as to evaluate the detection performance of the converged ForegroundNet model.

[0073] More specifically, when testing the detection performance of the model, each image to be detected in the test set needs to be preprocessed, resized to 384×384, and then normalized and input into the converged ForegroundNet model for forward calculation to output the detection result mask pred . The detection result is readjusted to the original image size through the interpolation algorithm and compared with the true label to calculate the SM index, EM index, WFM index, and MAE index corresponding to this image. Take the average of the indicators of all images in the same dataset as the evaluation indicator of the detection performance of the ForegroundNet model on this dataset.

[0074] It should be noted that the smaller the mean absolute error MAE index, and the larger the structure index SM, enhancement-alignment index EM, and weighted F index WFM, the more accurate the output detection result. In addition, the indicators for measuring the target detection effect are not limited to the above 4 evaluation indicators, as long as they can show the similarity or discrimination degree between the predicted foreground target segmentation map and the true label.

[0075] Figure 10Shows the comparison of the detection performance evaluation metrics of the ForegroundNet network model proposed by the present invention and the main models for detecting targets in complex backgrounds. From Figure 10 As can be seen from the numerical results shown, the SM metric, EM metric, WFM metric, and MAE metric of the method of the present invention on the COD10K dataset and the CAMO dataset have greatly exceeded the main models such as PFNet, MGL, and LSR. On the CHAMELEON dataset with the least number of samples, the detection performance of the method of the present invention is slightly lower, and only the MAE metric reaches the optimal metric. Based on the above analysis, the detection performance of the method of the present invention as a whole exceeds the main models in this field.

[0076] Figure 11 This is the comparison of the detection performance in terms of visual effects between the ForegroundNet network model proposed by the present invention and the main models in the field of detecting targets in complex backgrounds. By comparing Figure 11 the first row of pictures in, it can be seen that the method of the present invention can not only completely detect the target of interest, but also can relatively accurately distinguish the foreground target area and the background occluder area. By comparing Figure 11 the second row of data in, it can be seen that the method of the present invention can overcome the interference of the camouflage color of the foreground target and restore the contour details of the target to the greatest extent. By comparing Figure 11 the fourth, fifth, and sixth rows of pictures in, it can be seen that the method of the present invention can accurately detect the main body and edge areas of the target when the texture and color of the foreground target are similar to the background due to the complex background. By comparing Figure 11 the second and ninth rows of pictures in, it can be seen that the method of the present invention can accurately detect the area where the small target is located and retain the complete contour details. Through the above comprehensive comparative analysis, the method of the present invention can improve the detection ability of the easily confused area of the target edge by introducing a convolutional attention mechanism, thereby improving the overall detection performance of the model.

[0077] Where the present invention is not described in detail are all well-known technologies to those skilled in the art.

[0078] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations according to the concept of the present invention without creative labor. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention based on the concept of the present invention through logical analysis, reasoning, or limited experiments should be within the protection scope determined by the claims.

Claims

1. A target detection method in complex backgrounds based on a convolutional attention mechanism, characterized in that, the method comprises the following steps: Step S1: Obtain a complex background target detection dataset, and construct a training set, a validation set, and a test set; Step S2: Construct a ForegroundNet network model based on a convolutional attention mechanism; Step S3: Perform supervised training on the ForegroundNet network model on the constructed training set until the model converges to the optimal performance; Step S4: Test the converged ForegroundNet network model on the constructed test set to evaluate the detection performance of the model in complex backgrounds; The ForegroundNet network model based on a convolutional attention mechanism constructed in step S2 includes: a feature extraction module, a foreground-background proposal module, a feature generation module, and a convolutional attention-based feature fusion decoding module; The specific content of step S3 includes: Step S301: Randomly extract training images from the constructed training set for preprocessing. First, use the interpolation algorithm to adjust the sizes of the input image and the corresponding ground truth label to H×W, where H represents the image height and W represents the image width; then perform image data augmentation processing, and finally normalize the image and input it into the ForegroundNet network model for supervised training; Step S302: The ForegroundNet network model extracts multi-scale abstract features through the feature extraction module. The foreground-background proposal module proposes a foreground-background segmentation map based on the input features and outputs N layers of proposal results k = 1, 2, ..., N, and the proposal results include foreground object detection results and background object detection results Subsequently, the feature generation module generates foreground features, background features, and edge features based on the proposal results. The feature fusion decoding module based on convolutional attention fuses and decodes the foreground, background, and edge features of each layer, and outputs M layers of foreground object detection results, denoted by , where k = N + 1, N + 2, ..., N + M; Step S303: Based on the proposed result and the detection result, calculate the loss function L of the model during the supervised training process overall , and its calculation method is shown in Equation (1). Among them, and respectively represent the foreground object proposal results of the k-th layer and the background object proposal results corresponding to the cross-entropy loss weighted by structural information, and respectively represent the foreground object proposal results of the k-th layer and the background object proposal results corresponding to the intersection over union loss weighted by structural information; Step S304: Based on the overall loss function L of the model overall , use the stochastic gradient descent optimization algorithm to iteratively update the network parameters of the ForegroundNet model; Step S305: The ForegroundNet network model gradually converges to the optimal performance during training. Subsequently, the network parameters of the model are solidified, and the input image to be detected is forward-computed. Among the N+M different detection results output, the foreground object detection result with the largest scale is selected as the final detection result mask of the model pred .

2. The target detection method in complex backgrounds based on a convolutional attention mechanism according to claim 1, characterized in that, the complex background target detection dataset in step S1 includes the datasets COD10K, CAMO, and CHAMELEON.

3. The target detection method in complex backgrounds based on a convolutional attention mechanism according to claim 1, characterized in that, Taking the foreground target detection result as an example, the calculation methods of the cross-entropy loss and the intersection over union loss are shown in equations (2) and (3) respectively: Among them, mask gt (x, y) and respectively represent the true label of the foreground target and the value at the position coordinates (x, y) in the foreground target proposal result of the k-th layer. γ is a parameter related to the structure weight, and w(x, y) represents the structure weight corresponding to the position with coordinates (x, y). Its expression is as follows: Among them, A xy represents the set of pixels surrounding the pixel with coordinates (x, y).

4. The target detection method in complex backgrounds based on a convolutional attention mechanism according to claim 1, characterized in that, The specific content of step S4 includes: Step S401: Read the images to be detected in the test set one by one, use the interpolation method to adjust their sizes to H×W, where H represents the image height and W represents the image width, then normalize the images, and input them into the ForegroundNet network model that converges to the optimal performance for forward calculation, and output the corresponding detection result mask pred ; Step S402: Calculate the objective evaluation metrics of the model on the test set based on the foreground object detection result mask pred of the model and the ground truth mask gt of the foreground object.

5. The target detection method in complex backgrounds based on a convolutional attention mechanism according to claim 4, characterized in that, the objective evaluation metrics include: SM metric, EM metric, WFM metric, and MAE metric.

Citation Information

Patent Citations

  • Small sample target detection method based on attention and contrast learning

    CN113392855A

  • Remote sensing image small target detection method based on improved YOLOv3

    CN113971764A