Full-convolution single-stage target detection method based on mixed domain attention mechanism
By introducing a fully convolutional single-stage object detection method with a hybrid domain attention mechanism into the object detection method, the problem of insufficient detection efficiency, accuracy and adaptability in the prior art is solved, and a more efficient and accurate object detection effect is achieved.
Patent Information
- Application Number
- CN202510687931.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-06-24
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing object detection methods have shortcomings in detection efficiency, accuracy and adaptability, especially when dealing with complex scenarios, it is difficult to accurately detect small objects and complex objects.
A full convolution single-stage object detection method based on the hybrid domain attention mechanism is adopted to extract multi-scale features from the original image through feature extraction blocks, and feature fusion is used to use the feature fusion module of the hybrid domain attention mechanism to enhance feature expression capabilities.
It improves the efficiency, accuracy and adaptability of object detection, and can better detect objects of different sizes, especially in complex scenarios.
Smart Images

Figure CN120198657A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the technical field of image recognition, and in particular, to a fully convolutional single-stage object detection method based on a hybrid-domain attention mechanism. Background Art
[0002] At present, object detection, as a very important task in computer vision, mainly aims to locate the position of the object bounding box in an image and identify the category to which the object belongs. Object detection has a wide range of applications in road detection, traffic light detection, robot vision, security monitoring, and remote sensing image analysis. Traditional candidate region-based methods require a large amount of computing resources and time, and have low detection accuracy for small objects. Regression-based methods, such as YOLO and SSD, although they improve the detection speed and the efficiency of end-to-end training, still lack a precise image preprocessing mechanism and are prone to losing small object information. Although anchor-free detection models, such as FCOS and FoveaBox, solve the problem of positive and negative sample imbalance and reduce the computational burden, there are still certain detection accuracy problems due to object box overlap or semantic ambiguity, especially when dealing with complex scenes.
[0003] It can be seen that there is an urgent need for a fully convolutional single-stage object detection method based on a hybrid-domain attention mechanism with high detection efficiency, accuracy, and adaptability. Summary of the Invention
[0004] In view of this, the embodiments of the present invention provide a fully convolutional single-stage object detection method based on a hybrid-domain attention mechanism, which at least partially solves the problems of poor detection efficiency, accuracy, and adaptability in the prior art.
[0005] The embodiments of the present invention provide a fully convolutional single-stage object detection method based on a hybrid-domain attention mechanism, including: Step 1, using a feature extraction block to extract a first-layer feature map, a second-layer feature map, a third-layer feature map, and a fourth-layer feature map with different sizes and channels from the original image; Step 2, inputting each layer of the feature map into the feature fusion module of the hybrid-domain attention mechanism from top to bottom in sequence. After obtaining a new feature map, enlarging it to the size of the upper-layer feature map and splicing the feature map with the upper-layer feature map to obtain a first-layer fusion feature map, a second-layer fusion feature map, a third-layer fusion feature map, and a fourth-layer fusion feature map; Step 3, inputting the first-layer fusion feature map, the second-layer fusion feature map, the third-layer fusion feature map, and the fourth-layer fusion feature map into a detection head to obtain prediction values and prediction probabilities, and calculating a classification loss and a location regression loss based on this; Step 4: Train the feature extraction block, feature fusion module, and detection head based on the classification loss and location regression loss to form a detection model; Step 5: Input the image to be detected into the detection model to obtain the object detection result.
[0006] According to a specific implementation manner of an embodiment of the present invention, the calculation formula of the feature extraction block is ; where represents the output of the nth residual structure, and respectively represent the mapping and input learned by the nth network.
[0007] According to a specific implementation manner of an embodiment of the present invention, the stride of the first-layer feature map, the second-layer feature map, the third-layer feature map, and the fourth-layer feature map are 4, 8, 16, and 32 respectively.
[0008] According to a specific implementation manner of an embodiment of the present invention, the feature fusion module of the hybrid domain attention mechanism includes a channel attention module and a spatial attention module connected in sequence.
[0009] According to a specific implementation manner of an embodiment of the present invention, Step 2 specifically includes: Step 2.1: Input the current-layer feature map into the channel attention module, perform average pooling operation and max pooling operation on the feature map respectively to obtain two one-dimensional tensors, representing the average pooling feature and the max pooling feature respectively, then send these two one-dimensional tensors into a multi-layer perceptron, and then after adding the two obtained features and passing through a Sigmoid activation function to obtain a weight coefficient, multiply the weight coefficient by the current feature map to obtain a scaled new feature map , ; where represents the original feature map, represents the new feature map, represents the Sigmoid activation function, represents the multi-layer perceptron, represents the average pooling operation, represents the max pooling operation; Step 2.2: Input the scaled new feature map into the spatial attention module, perform average pooling operation and max pooling operation on it respectively to obtain two channel descriptions of H×W×1, and splice these two descriptions together by channel to generate an effective feature descriptor, then pass through a 7×7 convolutional layer and a Sigmoid activation function, and after obtaining a weight coefficient, multiply the weight coefficient by the scaled new feature Multiply to obtain a new scaled feature map , ; Among them, represents the new feature map, represents the Sigmoid activation function, represents a 7×7 convolutional layer, represents the average pooling operation, represents the max pooling operation, represents the concatenation operation by channel; Step 2.3, magnify the new feature map corresponding to the current layer feature map to the size of the previous layer feature map and splice the feature map with the previous layer feature map as the new previous layer feature map; Step 2.4, repeat Steps 2.1 to 2.3 from top to bottom to obtain the new first layer feature map, second layer feature map, third layer feature map, and fourth layer feature map, and input them into the convolutional module respectively to obtain the first layer fused feature map, second layer fused feature map, third layer fused feature map, and fourth layer fused feature map.
[0010] According to a specific implementation manner of an embodiment of the present invention, the classification loss is the Focal Loss function, and the location regression loss is the IoU Loss function.
[0011] According to a specific implementation manner of an embodiment of the present invention, the expression of the Focal Loss function is ; Among them, represents the predicted probability, and are hyperparameters, represents the object detection result, 1 represents success, and 0 represents failure; The expression of the IoU Loss function is ; Among them, GT represents the ground truth, P represents the predicted value, represents the overlapping area between the predicted value and the ground truth, represents the area after merging the predicted value and the ground truth.
[0012] The full convolutional single-stage object detection scheme based on the hybrid domain attention mechanism in the embodiments of the present invention includes: Step 1, using a feature extraction block to extract the first layer feature map, the second layer feature map, the third layer feature map, and the fourth layer feature map with different sizes and channels from the original image; Step 2, inputting each layer of feature map into the feature fusion module of the hybrid domain attention mechanism from top to bottom in sequence. After obtaining a new feature map, enlarge it to the size of the upper layer feature map and splice the feature map with the upper layer feature map to obtain the first layer fusion feature map, the second layer fusion feature map, the third layer fusion feature map, and the fourth layer fusion feature map; Step 3, inputting the first layer fusion feature map, the second layer fusion feature map, the third layer fusion feature map, and the fourth layer fusion feature map into the detection head to obtain prediction values and prediction probabilities, and calculate the classification loss and the location regression loss accordingly; Step 4, training the feature extraction block, the feature fusion module, and the detection head according to the classification loss and the location regression loss to form a detection model; Step 5, inputting the image to be detected into the detection model to obtain the object detection result.
[0013] The beneficial effects of the embodiments of the present invention are as follows: Through the solution of the present invention, first, the feature extraction block is responsible for extracting the original features from the original image; then, based on the hybrid domain attention mechanism, the extraction of features is further deepened. The multi-scale fusion module not only utilizes the strong semantic features at the top layer that are beneficial for classification, but also utilizes the high-resolution information at the bottom layer for positioning, enabling different feature layers to better detect objects of different sizes; finally, the detection head is responsible for integrating the feature map and converting the feature map into a tensor for detecting category information and location information, improving the detection efficiency, accuracy, and adaptability. Description of the Drawings
[0014] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0015] Figure 1 It is a schematic flowchart of a full convolutional single-stage object detection method based on the hybrid domain attention mechanism provided by the embodiments of the present invention; Figure 2 It is a structure diagram of EfficientNet-B0 provided by the embodiments of the present invention; Figure 3 It is a structure diagram of a feature extraction block of the hybrid domain attention mechanism provided by the embodiments of the present invention; Figure 4 It is a structure diagram of a multi-scale feature fusion module based on the hybrid domain attention mechanism provided by the embodiments of the present invention; Figure 5This is the system framework diagram corresponding to a full-convolution single-stage object detection method based on a hybrid-domain attention mechanism provided by an embodiment of the present invention. Detailed implementation manners
[0016] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0017] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts belong to the scope of protection of the present invention.
[0018] It should be noted that the following describes various aspects of the embodiments within the scope of the appended claims. It should be obvious that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on the present invention, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement the device and / or practice the method. In addition, this device and / or this method can be implemented using other structures and / or functions in addition to one or more of the aspects described herein.
[0019] It should also be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in its actual implementation can be an arbitrary change, and the component layout type may also be more complex.
[0020] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.
[0021] An embodiment of the present invention provides a full-convolution single-stage object detection method based on a hybrid-domain attention mechanism. The method can be applied to the object detection process in scenarios such as road detection, traffic light detection, and robot vision.
[0022] See Figure 1 , which is a schematic flow chart of a fully convolutional single-stage object detection method based on a hybrid domain attention mechanism provided by an embodiment of the present invention. As Figure 1 shown, the method mainly includes the following steps: Step 1, use a feature extraction block to extract the first, second, third, and fourth feature maps with different sizes and channels from the original image; Specifically, the feature extraction block mainly extracts features from the original image. The method of the present invention uses the EfficientNet convolutional neural network as the feature extraction block. The main structure of EfficientNet is composed of the Mobile Inverted Bottleneck Convolution module, and at the same time, the attention idea of the Squeeze-and-Excitation Network (SENet) is introduced into this module. The Mobile Inverted Bottleneck Convolution module is obtained through neural network search. The structure of this module is similar to the Depthwise Separable Convolution (DWConv). The Mobile Inverted Bottleneck first performs a 1×1 pointwise convolution on the input and changes the output channel dimension according to the expansion ratio. For example, if the expansion ratio is 6, the channel dimension is increased by 6 times. Then a k×k depth convolution is performed. If the squeeze-and-excitation operation is required, this operation will be performed after the depth convolution, and then a 1×1 pointwise convolution is used to restore the original channel dimension. Finally, connection inactivation and skip connection of the input are performed. This operation makes the model have random depth and improves the model performance.
[0023] The method of the present invention uses the basic network EfficientNet-B0 of EfficientNet, which is composed of 16 Mobile Inverted Bottleneck Convolution modules, 2 convolutional layers, 1 global average pooling layer, and 1 classification layer. This method modifies the last convolutional layer, fully connected layer, and classification layer of the network. After the image passes through the backbone network, four feature layers C2, C3, C4, and C5 with different sizes and channels are exported to prepare for the next multi-scale feature fusion. The strides of these feature layers relative to the original image are 4, 8, 16, and 32 respectively, and their structures are as Figure 2 shown.
[0024] Use a module based on the hybrid domain attention mechanism for feature fusion, mainly to increase the semantic information representation ability of the underlying network, and enhance the expression ability of features by channel attention and spatial attention. This feature extraction block is composed of 4 convolutional layers with the same structure containing an attention mechanism. Each feature extraction block structure is as Figure 3 shown.
[0025] In this module, two 1×1 convolutions are respectively used to reduce and increase the feature dimension. The main purpose is to reduce the number of parameters, thereby reducing the computational power. The 3×3 convolution in the middle uses depthwise separable convolution, which can maintain the performance loss within a small range while significantly reducing the number of parameters. Then, a channel attention module and a spatial attention module are applied in sequence to increase the expressiveness, focus on important features and suppress unnecessary features. Finally, the entire module expresses the output as a linear superposition of the input and a non-linear transformation of the input through a residual connection. After adding the residual connection, the network can converge faster on the premise of the same number of layers. When reusing features, a shortcut for feature transmission is added between the networks, directly transmitting the feature information to the deeper layers of the neural network. This process can be represented by Equation 1.
[0026] ; denotes the output of the nth residual structure, and respectively denote the mapping learned by the nth network and the input.
[0027] Step 2: Input the feature maps of each layer into the feature fusion module of the hybrid domain attention mechanism from top to bottom in sequence. After obtaining the new feature maps, enlarge them to the size of the feature maps of the previous layer and splice the feature maps with the feature maps of the previous layer to obtain the first-layer fused feature map, the second-layer fused feature map, the third-layer fused feature map, and the fourth-layer fused feature map; Specifically, for feature fusion, the top-down process is to first reduce the number of channels of the more abstract and semantically stronger high-level feature maps by a factor of two using a 1×1 convolutional layer, and then use bilinear interpolation upsampling to enlarge the feature maps to the same size as the previous layer. Then, splice the feature maps of this layer with the feature maps of the previous layer, and finally send them into the feature extraction block containing the hybrid domain attention mechanism to obtain the fused features. The structure of the multi-scale feature fusion module based on the hybrid domain attention mechanism is as Figure 4 shown.
[0028] The attention mechanism is a data processing method in machine learning and is widely used in various types of machine learning tasks such as natural language processing, image recognition, and speech recognition. Generally speaking, the attention mechanism hopes that the network can automatically learn the areas that need attention in the picture. For example, when a person's eyes look at a painting, they do not equally distribute their attention to all the pixels in the painting, but allocate more attention to the areas that people focus on.
[0029] The attention mechanism generates a mask through the operations of a neural network, scores the values on the mask, and evaluates the scores of the points that need to be focused on currently. Therefore, the Channel Attention Module generates a mask for the channels and scores it; while the Spatial Attention Module generates a mask for the space and scores it. For example, the hybrid-domain attention mechanism used in this method evaluates and scores both the channel attention and the spatial attention simultaneously.
[0030] For the channel attention module, this method first performs average pooling and max pooling operations on the feature map respectively to obtain two one-dimensional tensors, representing the average pooling feature and the max pooling feature respectively; then these two one-dimensional tensors are fed into a shared neural network, which consists of a multi-layer perceptron (MLP) with a hidden layer; then the two features obtained are added together and passed through a Sigmoid activation function to obtain the weight coefficient; finally, the original feature is multiplied by the weight coefficient to obtain the scaled new feature, and this process can be represented by Equation 2.
[0031] ; where, F represents the original feature, represents the new feature, represents the Sigmoid activation function, represents the multi-layer perceptron, represents the average pooling operation, represents the max pooling operation.
[0032] After the channel attention module, a spatial attention module is introduced to focus on where the features are meaningful. Similar to the channel attention, first perform average pooling and max pooling operations on the feature map respectively to obtain two channel descriptions of H×W×1, and concatenate these two descriptions along the channel to generate an effective feature descriptor; then pass through a 7×7 convolutional layer and a Sigmoid activation function to obtain the weight coefficient, and finally multiply the weight system by the original feature to obtain the scaled new feature, and this process can be represented by Equation 3.
[0033] ; where, F represents the original feature, represents the new feature, represents the Sigmoid activation function, represents the 7×7 convolutional layer, represents the average pooling operation, represents the max pooling operation, Indicates the operation of splicing by channel.
[0034] Then, first use a 3×3 depthwise separable convolutional layer to double the number of channels output by the multi-scale feature fusion module, and then use a 1×1 ordinary convolution to reduce the number of channels to 24, obtaining the first fused feature map, the second fused feature map, the third fused feature map, and the fourth fused feature map.
[0035] Step 3: Input the first fused feature map, the second fused feature map, the third fused feature map, and the fourth fused feature map into the detection head to obtain prediction values and prediction probabilities, and calculate the classification loss and the location regression loss based on these; In specific implementation, the detection head mainly reorganizes the information learned by the previous convolutional layer into a structure suitable for object detection. After obtaining the first fused feature map, the second fused feature map, the third fused feature map, and the fourth fused feature map, input them into the detection head to obtain prediction values and prediction probabilities, and calculate the classification loss and the location regression loss based on these.
[0036] Step 4: Train the feature extraction block, the feature fusion module, and the detection head according to the classification loss and the location regression loss to form a detection model; In specific implementation, for any point on the i-th layer feature map , it can correspond to a point on the input image , where represents the stride, the size of the input image is picture, , that is . All the points falling within the same object bounding box are sorted according to the center-ness. If there are more than 5 points, take the top 5 points sorted by center-ness from large to small. If there are less than or equal to 5 points, take all the points. Set the selected points as positive samples, and the category of this point is , and the remaining points are marked as negative samples (background class). The location regression target of this point is is the distance from this point to the four sides of the bounding box. The target of location regression is expressed as formula 4, and the center-ness calculation formula is as formula 5.
[0037] ; ; To solve the problem of training ambiguity caused by the overlap of objects of different sizes, the features Figure 1 derived from the multi-scale feature fusion module are a total of 4, which are respectively , , , set a stride interval , where , representing the maximum value that can satisfy the upper - border regression distance of feature layer i. First, calculate , is the longest side of the ground - truth bounding box. If it satisfies then set this point as a positive sample on the i - th layer and as a negative sample on the remaining feature layers.
[0038] If there is still an overlapping situation on the same feature layer, then calculate the center - ness of this point from the centers of the two ground - truth bounding boxes, and assign this point to the bounding box with a higher center - ness for regression.
[0039] The loss function for optimizing the network is the scheduling center of the entire network model. The classification loss used in this method is Focal Loss, which is used to solve the problem of the imbalance in the ratio of positive and negative samples in single - stage object detection, as shown in Equation 6.
[0040] ; where represents the predicted probability, and are two hyperparameters, which are used to reconcile the weight ratio of positive and negative samples and reduce the loss value of simple negative samples respectively.
[0041] For the location regression loss, this method uses IoU Loss, as shown in Equation 7. Among them, IoU is the ratio of the intersection and union of the ground - truth box and the predicted box. When they completely overlap, IoU is 1. Then for the loss function, the smaller the loss value, the better, indicating a high degree of overlap between them.
[0042] ; where GT represents the ground - truth value, and P represents the predicted value. In practical applications, IoU Loss can also be simply expressed as Equation 8.
[0043] ; Step 5, input the image to be detected into the detection model to obtain the object detection result.
[0044] Specifically, when implementing, adjust the hyperparameters of the feature extraction block, feature fusion module, and detection head through training to form a detection model such as Figure 5As shown, when it is necessary to detect a certain image, the image to be detected can be input into the detection model. The feature extraction block is responsible for extracting features from the original image. The multi-scale feature fusion module based on the hybrid-domain attention mechanism not only utilizes the strong semantic features conducive to classification in the top layer, but also utilizes the high-resolution information for localization in the bottom layer, enabling different feature layers to better detect objects of different sizes. The detection head is responsible for integrating the feature map and converting the feature map into a tensor for detecting category information and location information, thereby detecting the corresponding object. The feature extraction block is responsible for extracting features from the original image. The multi-scale feature fusion module based on the hybrid-domain attention mechanism not only utilizes the strong semantic features conducive to classification in the top layer, but also utilizes the high-resolution information for localization in the bottom layer, enabling different feature layers to better detect objects of different sizes. The detection head is responsible for integrating the feature map and converting the feature map into a tensor for detecting category information and location information, thereby detecting the corresponding object and obtaining the object detection result.
[0045] The fully convolutional single-stage object detection method based on the hybrid-domain attention mechanism provided in this embodiment extracts original features from the original image through the feature extraction block. Then, based on the hybrid-domain attention mechanism, the feature extraction is further deepened. The multi-scale fusion module not only utilizes the strong semantic features conducive to classification in the top layer, but also utilizes the high-resolution information for localization in the bottom layer, enabling different feature layers to better detect objects of different sizes. Finally, the detection head is responsible for integrating the feature map and converting the feature map into a tensor for detecting category information and location information, improving the detection efficiency, accuracy, and adaptability.
[0046] The method of the present invention will be further described below in conjunction with a specific embodiment. The platform used in this embodiment is the Ubuntu 16.04.6 operating system, an eight-core Intel 2.10GHz CPU, a GPU is an Nvidia GeForce RTX 2080 Ti, 64G of memory and 2T of hard disk, and the model in this article is trained under the Pytorch 1.5 deep learning framework based on the GPU version. This article uses the pre-trained weight file provided by EfficientNet-B0 to initialize the backbone network, sets the bias to zero, and optimizes the network using SGD. The batch size is set to 32, the initial learning rate is 1e-3, and the learning rate is reduced to half of the original every 100 generations of each iteration, with a total of 500 generations of iteration.
[0047] (1) Dataset The PASCAL VOC dataset was used in the experiment. This dataset includes 20 categories, namely aero, bike, bird, boat, bottle, bus, car, cat, chair, cow, table, dog, horse, motorbike, person, plant, sheep, sofa, train, and tv. In this paper, the PASCAL VOC 2007 and PASCAL VOC 2012 datasets were used for training, with a total of 16,551 images, and the PASCAL VOC 2007 dataset was used for testing, with a total of 4,952 images.
[0048] To make full use of the training data, this paper randomly augmented the data in three ways with a 50% probability: (1) adjusting brightness, contrast, and chroma; (2) randomly cropping the images; (3) flipping the images horizontally.
[0049] (2) Experimental results and analysis Using the PASCAL VOC 2007 dataset for testing, the input image size was scaled to 320×320, and the mean average precision (mAP) and model size (MB) were used to verify the accuracy and lightweight of the model.
[0050] Table 1; ; From the data in Table 1, it can be seen that the method proposed in this paper outperforms Fast RCNN and Faster RCNN in terms of performance. Among them, there are 8 categories leading other methods, the mean average precision exceeds all comparison methods, and the number of parameters is also the lowest among all comparison methods. When the batch size is set to 1, the FPS can exceed 24 when tested on Nvidia GeForce RTX 2080 Ti, achieving the effect of real-time detection. The selected pictures are several common types in life, such as animals, people, vehicles, bottles, etc. In terms of the effect, for the first type of picture, the detection effects of this method and method c are the same, while method b detected both the correct dog and the wrong cat at the same time; for the second type of picture, which is a dining table, this method detected one more bottle than method b, while method c detected the same bottle twice; for the third type of picture, which is urban traffic, all three methods detected the bus, the taxi in the front row, and the blurred car, while this method also detected the car that was only half-exposed on the right side, the people on the non-motor vehicle on the left side, and the model image on the light board behind.
[0051] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof.
[0052] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A fully convolutional single-stage object detection method based on a hybrid-domain attention mechanism, characterized in that Including: Step 1: Use the feature extraction block to extract the first-layer feature map, second-layer feature map, third-layer feature map, and fourth-layer feature map with different sizes and channels from the original image; Step 2: Input each layer of feature map into the feature fusion module of the hybrid domain attention mechanism from top to bottom in sequence. After obtaining the new feature map, enlarge it to the size of the upper-layer feature map and splice the feature map with the upper-layer feature map to obtain the first-layer fusion feature map, second-layer fusion feature map, third-layer fusion feature map, and fourth-layer fusion feature map; Step 3: Input the first-layer fusion feature map, second-layer fusion feature map, third-layer fusion feature map, and fourth-layer fusion feature map into the detection head to obtain the predicted value and predicted probability, and calculate the classification loss and location regression loss accordingly; Step 4: Train the feature extraction block, feature fusion module, and detection head according to the classification loss and location regression loss to form a detection model; Step 5: Input the image to be detected into the detection model to obtain the object detection result.
2. The method according to claim 1, wherein The calculation formula of the feature extraction block is Among them, represents the output of the nth residual structure, and respectively represent the mapping and input learned by the nth network.
3. The method according to claim 1, characterized in that, The step sizes of the first-layer feature map, second-layer feature map, third-layer feature map, and fourth-layer feature map are 4, 8, 16, and 32 respectively.
4. The method according to claim 1, characterized in that, The feature fusion module of the hybrid domain attention mechanism includes a channel attention module and a spatial attention module connected in sequence.
5. The method according to claim 4, wherein The specific content of Step 2 includes: Step 2.1, input the current layer feature map into the channel attention module, perform average pooling operation and max pooling operation on the feature map respectively to obtain two one-dimensional tensors, representing the average pooling feature and the max pooling feature respectively. Then, send these two one-dimensional tensors into a multi-layer perceptron. After that, add the two obtained features, pass them through a Sigmoid activation function to get the weight coefficient, and multiply the weight coefficient by the current feature map to obtain the scaled new feature map , Among them, represents the original feature map, represents the new feature map, represents the Sigmoid activation function, represents the multi-layer perceptron, represents the average pooling operation, represents the max pooling operation; Step 2.2, input the scaled new feature map into the spatial attention module, perform average pooling operation and max pooling operation on it respectively to obtain two channel descriptions of H×W×1, concatenate these two descriptions by channel to generate an effective feature descriptor, then pass through a 7×7 convolutional layer and a Sigmoid activation function. After obtaining the weight coefficient, multiply the weight coefficient and the scaled new feature to get the scaled new feature map , Among them, represents the new feature map, represents the Sigmoid activation function, represents a 7×7 convolutional layer, represents the average pooling operation, represents the max pooling operation, represents the concatenation operation by channel; Step 2.3, magnify the new feature map corresponding to the current layer feature map to the size of the previous layer feature map, splice the feature map with the previous layer feature map, and use the result as the new previous layer feature map; Step 2.4: Repeat Steps 2.1 to 2.3 from top to bottom to obtain the new first-layer feature map, second-layer feature map, third-layer feature map, and fourth-layer feature map, and input them into the convolution module respectively to obtain the first-layer fusion feature map, second-layer fusion feature map, third-layer fusion feature map, and fourth-layer fusion feature map.
6. The method according to claim 1, wherein The classification loss is the Focal Loss function, and the location regression loss is the IoU Loss function.
7. The method according to claim 6, characterized in that, The expression of the Focal Loss function is Among them, represents the predicted probability, and are hyperparameters, represents the object detection result, 1 indicates success, and 0 indicates failure; The expression of the IoU Loss function is Among them, GT represents the true value, and P represents the predicted value. represents the region where the predicted value overlaps with the true value. represents the region after merging the predicted value and the true value.
Citation Information
Patent Citations
Small target detection method based on attention mechanism
CN114202672A
Unmanned aerial vehicle image small target detection method based on multiple attention fusion network
CN117853959A
Remote sensing image target detection method based on improved FCOS
CN119540758A