Haze image target detection model and algorithm
By introducing co-occurrence attention module and edge-aware loss function into the object detection model, the problem of insufficient target detection accuracy and robustness in haze weather is solved, and higher detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510740470.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
The existing target detection algorithm has significantly reduced performance under complex weather conditions such as haze, making it difficult to accurately identify and locate targets in images, and is not robust enough.
A haze image object detection model is constructed, a target co-occurrence attention module and edge-aware loss function are introduced, and local feature relationships are captured through self-attention branches. The co-occurrence relationship branch provides a global context, and the edge-aware loss function evaluates the edge overlap and geometric relationship of the bounding box.
The target detection accuracy and robustness of the model under haze weather conditions are improved, and the target detection capability in complex backgrounds is enhanced.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of object detection, and relates to a haze image object detection model and algorithm. Background Technique
[0002] Object detection is a fundamental task in the field of computer vision, aiming to identify and locate one or more target objects in an image or video frame. An object detection model not only needs to identify which category the target in the image belongs to (such as a person, a car, a cat, etc.), but also needs to determine the exact position of the target in the image, usually by drawing bounding boxes. Currently, there are mainly two categories of object detection algorithms in the field of deep learning. One is the two-stage object detection algorithm, whose processing process is to first generate a series of sample candidate boxes by the algorithm, and then use a convolutional neural network to classify these candidate sample boxes, such as algorithms in the R-CNN series, DetNe, etc.; the other is the one-stage object detection algorithm, which directly transforms the object bounding box localization problem into a regression problem for processing, such as SSD and YOLO series algorithms.
[0003] With the continuous emergence of new technologies and the continuous optimization of algorithms, object detection technology is developing towards a more intelligent and accurate direction. However, when these algorithms face complex weather conditions, such as haze, rain and snow, low light and other environments, they often show a significant decline in performance. These adverse weather conditions not only reduce the visibility of the image, but also increase the background noise, making it difficult to extract the features of the target, thus seriously affecting the accuracy and robustness of object detection. With the development of deep learning technology, although the performance of object detection has been significantly improved, its application in complex weather backgrounds is still a huge challenge.
[0004] Currently, in dealing with object detection tasks under vision degradation conditions, there are mainly two categories of methods. The first category is the object detection algorithm based on transfer learning. This type of method studies the alignment mechanism of features between the source domain and the target domain, and transfers the features learned in the source domain to the target domain for detection. The second category is the object detection algorithm based on image restoration and feature enhancement. Usually, an image enhancement module or network is cascaded with the object detection network to improve the quality of degraded images. Summary of the Invention
[0005] The technical problem solved by the present invention is to provide a haze image object detection model and algorithm, construct an edge loss function and introduce an object co-occurrence attention module, so as to achieve higher model accuracy.
[0006] The present invention is realized through the following technical solutions: A haze image target detection model includes a backbone network, a neck network, and a detection head. A relational attention module is also provided in the model, and this module contains a self-attention branch and a co-occurrence relationship branch; and the loss function inside the detection head is updated to an edge-aware loss function. The self-attention branch calculates lightweight multi-head self-attention through a local window to capture local feature relationships; the co-occurrence relationship branch generates global co-occurrence features through the detection results of the co-occurrence detection head to provide scene-level context; the output of the relational attention module weights and adds the output feature maps of the self-attention branch and the co-occurrence branch. The edge-aware loss function includes: ; ; Among them, Equation (5) is the expression of the edge intersection over union, and Equation (6) is the edge-aware loss function; is the mask of the true target bounding box, M is the mask of the predicted target bounding box, is the edge feature map obtained after depth Sobel convolution and square root; is the edge intersection over union loss, represents the Euclidean distance, b, w, and h respectively represent the center point coordinates, width, and height of the bounding box, 、 、 respectively represent the center point coordinates, width, and height of the predicted bounding box, 、 、 respectively represent the diagonal length, width, and height of the minimum bounding rectangle, is the center point distance loss between the true box and the target box, is the width loss between the true box and the target box, is the height loss between the true box and the target box.
[0007] Furthermore, the self-attention branch first adds position encoding to the feature map, and then successively performs operations including convolution, channel dimension division, window division, and window unfolding on it to obtain three elements of query, index, and content for calculating self-attention; the three obtained elements are calculated through multi-head self-attention and then the output feature map is merged, and finally a residual connection is made with the input feature map to capture local feature relationships.
[0008] Furthermore, the self-attention branch is implemented based on Equation (1) and Equation (2): ; Among them, The input feature map for co-occurrence relationship attention X The output after position encoding, represents a 1×1 convolution, represents operations of dividing by channel dimension and by window, represents the operation of unfolding by window; ; Among them, 、 、 are obtained from Equation (1), represents the number of attention heads, is the activation function, T represents transpose, represents transforming the outputs of different heads to achieve unification in shape, represents folding and restoring the calculation results of the previous operations of dividing by channel dimension and by window.
[0009] Furthermore, the co-occurrence branch first uses the co-occurrence detection head to extract the object categories existing in the scene, then maps the categories to high-dimensional embedding vectors to encode the co-occurrence relationships between the embedding vectors of different categories; then average pooling is performed to generate the global scene representation, and the global features are passed to all positions to provide scene-level context; The co-occurrence branch also learns the relationships between the embedding vectors of different categories through the co-occurrence patterns in the training data to achieve higher prediction accuracy.
[0010] Furthermore, the output of the co-occurrence branch is: ; Among them, is the co-occurrence feature obtained after average pooling and expanding to the spatial dimension, represents the i-th category, N is the total number of categories, is the feature vector after mapping the category ID extracted from the detection result D; B, C, H, W is the standard dimension format for the deep learning model to process image data, where B represents the batch size, C represents the number of channels, H and W respectively represent the height and width of the image.
[0011] Furthermore, the output of the relationship attention module is expressed as: ; is the output of the co-occurrence relationship attention module, where is a branch adjustment parameter used to balance the output of the self-attention branch and the output of the collinear branch in terms of their contribution to the output feature map.
[0012] Furthermore, the backbone network, neck network, and detection head use the baseline network yolov8, and a co-occurrence relationship attention module is added after the convolution at the end of the c2f module of the baseline network yolov8; In the constructed model, the input image first enters the backbone network, and through multiple convolutional layers alternating with the C2f module, deep semantic features are gradually extracted to generate multiple feature maps with different resolutions from P1 to P5; in the neck network, feature maps at different levels are fused through a top-down path and a bottom-up path to obtain the P3 layer, P4 layer, and P5 layer after feature fusion; finally, the detection head performs detection on them, and the loss function inside the detection head evaluates the quality of the detection.
[0013] Furthermore, the edge-aware loss function requires three inputs: The first input is the ground truth bounding box annotated in the dataset. A mask map is extracted from the ground truth bounding box through a mask extraction module and input into the edge intersection over union loss function for calculation; The second input is the original image of the dataset. After shallow convolutional processing by the model, the feature map of the intermediate layer is extracted, and it is subjected to depth Sobel convolution processing, followed by a square root operation to obtain the edge feature map, which is then input into the edge intersection over union loss function; The third input is the output prediction result of the model. A mask of the predicted target bounding box is extracted from the prediction result through a mask extraction module and input into the loss function.
[0014] The present invention also provides a detection algorithm based on the haze image target detection model, including the following operations: 1) Construct a haze image target detection model with edge-aware loss and co-occurrence relationship; and update the loss function to the edge-aware loss function; 2) Training of the network model, selecting model evaluation metrics; during the training process, the model optimizes the loss through backpropagation, making the predicted bounding box gradually approach the position and size of the ground truth bounding box; the smaller the output loss value, the higher the edge overlap degree, the closer the center points, and the more consistent the aspect ratios between the predicted bounding box and the ground truth bounding box; 3) Obtain the trained model for target detection.
[0015] Compared with the prior art, the present invention has the following beneficial technical effects: The haze image target detection model provided by the present invention uses YOLOv8 as the baseline network, introduces a target co-occurrence attention module and designs a new edge-aware loss function. The introduced target co-occurrence attention module is based on the fact that in the real world, it is easy to observe that some objects may appear in the same scene simultaneously. Therefore, the target detection model can utilize this symbiotic relationship to make up for the deficiency of visual features. Through the target co-occurrence attention module, the symbiotic relationship can be fully utilized to make up for the shortcoming that the visual features extracted from haze weather images by the target detection model are insufficient, thereby improving the accuracy of the model.
[0016] The edge-aware loss function (EAIoU) constructed in the present invention is based on the original loss function that only considers the geometric relationship of the bounding boxes, and also considers the alignment of the object edges within the bounding boxes. By combining the traditional IoU calculation and edge information, EAIoU can more accurately evaluate the similarity between the bounding boxes predicted by the model and the ground truth bounding boxes.
[0017] The present invention realizes more efficient utilization of existing information and mining of more available information, effectively improves the accuracy of the target detection model under complex weather backgrounds, has better performance and stronger robustness. Compared with the baseline network, it achieves higher model accuracy. Brief Description of the Drawings
[0018] Figure 1 It is a schematic diagram of the co-occurrence attention module of the present invention; Figure 2 It is a schematic diagram of the c2f_CRA module of the present invention; Figure 3 It is a schematic diagram of the target detection network model of the present invention; Figure 4 It is a schematic diagram of the edge loss function of the present invention; Figure 5 It is the heat map of the baseline model; Figure 6 It is the heat map of the model generated by the present invention. Detailed Description of the Invention
[0019] The following further describes the present invention in detail with reference to embodiments, which are explanations rather than limitations of the present invention.
[0020] A haze image target detection model includes a backbone network, a neck network, and a detection head. The backbone network is the main feature extraction part of the model, which consists of a deep convolutional neural network (CNN) and is used to extract high-level and semantically rich features from the input image; The neck network is an intermediate layer between the backbone network and the head. Its role is to further perform operations such as feature fusion and context enhancement based on the features extracted by the backbone network. The structure of the neck usually includes convolutional layers, pooling layers, attention mechanisms, etc. By designing the network neck, it can help the network perceive targets at different scales and provide more context information.
[0021] The detection head can also be called the head network. The head network is the output part of the model, responsible for the final task prediction, and outputs the class labels of each predicted target and the predicted values of each category.
[0022] And in the model, there is also a relational attention module, which contains two branches: self-attention and co-occurrence relationship; and the loss function inside the detection head is updated to an edge-aware loss function; The self-attention branch calculates lightweight multi-head self-attention through local windows to capture local feature relationships; the co-occurrence relationship branch generates global co-occurrence features through the detection results of the co-occurrence detection head to provide scene-level context; the output of the relational attention module weights and adds the output feature maps of the self-attention branch and the co-occurrence branch.
[0023] And the detection algorithm based on the haze image target detection model includes the following operations: a. Construct a haze image target detection model based on edge-aware loss and co-occurrence relationship; and update the loss function to an edge-aware loss function; b. Training of the network model, select the dataset and model evaluation metrics; during the training process, the model optimizes the loss through backpropagation, making the prediction box gradually approach the position and size of the true box; the smaller the output loss value, the higher the edge overlap degree between the prediction box and the true box, the closer the center points, and the more consistent the aspect ratios. c. Obtain the trained model for target detection.
[0024] The following will explain each part in detail.
[0025] 1) Construct a haze image target detection model based on edge-aware loss and co-occurrence relationship Design of the backbone network and the neck network. Currently, traditional target detection models mainly rely on visual features to achieve target recognition and localization. However, it is a challenge for target detectors to fully extract visual features from images under complex weather conditions. In real-world scenarios, certain objects tend to co-occur in the same scene. For example, the probability of a bicycle and a pedestrian appearing in the same scene is relatively high, while the probability of a bicycle and an airplane appearing simultaneously is very low. Therefore, target detectors can utilize this co-occurrence relationship to make up for the deficiency of visual features.
[0026] To achieve this goal, the present invention constructs a co-occurrence relationship attention module CRA (Co-occurrence relationship attention); the embodiment given in the present invention combines this module with c2f in the baseline network yolov8 to improve the baseline network.
[0027] This module contains two branches: self-attention and co-occurrence relationship. Among them, the self-attention branch calculates lightweight multi-head self-attention through a local window to capture local feature relationships; The co-occurrence relationship branch generates global co-occurrence features through the detection results of the co-occurrence detection head to provide scene-level context; the global co-occurrence features are the average representation of all detected categories, rather than explicitly modeling the relationships between categories; The output of the CRA module is the weighted feature fusion of the two: the output feature maps of the self-attention branch and the co-occurrence branch are weighted and added; in this way, local details and global semantics are combined through the preliminary detection results; For the self-attention branch, since traditional convolutional neural networks (CNNs) mainly rely on local receptive fields and are difficult to capture the dependencies between distant objects, such as the relationship between cars and pedestrians in a traffic scene, it is necessary to introduce a self-attention branch to calculate the dependencies between all positions in the feature map, directly model the global context information, and thus better understand the associations between objects in the scene.
[0028] See Figure 1 , the self-attention branch first adds positional encoding to the feature map, and then sequentially performs operations such as 1×1 convolution, channel-wise division, and window division on it to obtain the query , index , content Three elements. The three obtained elements are calculated through multi-head self-attention and then merged into the output feature map, which is finally connected to the input feature map by residual connection to prevent gradient disappearance in deep neural networks and at the same time help the model retain the original input information.
[0029] Specifically, the self-attention branch is implemented based on Equations (1) and (2): ; Among them, Is the input feature map of the co-occurrence relationship attention X The output after positional encoding, Represents 1×1 convolution, Represents the operations of channel-wise division and window division, Represents the operation of unfolding by window.
[0030] ; Among them, , , are derived from Equation (1). X is the input feature map obtained by transmitting the initial input image through layer-by-layer processing of the network to the co-occurrence relationship attention. represents the number of attention heads. is the activation function. T represents transpose. represents transforming the outputs of different heads to achieve unification in shape. represents folding and restoring the calculation results previously divided by channel dimension and by window.
[0031] The following gives the specific operations of the self-attention branch: Use operation to split the input feature map into multiple local windows, and the features within each window will be processed separately. Next, after generating Q, K, V, operation splits them into three parts: q, k, and v. Then, perform window unfolding on each part, restricting the query, key, and value at each position within the local window. For example, for a 7×7 window, each position can only focus on other positions within the surrounding 7×7 area, so as to capture local relationships. In the multi-head processing part, the unfolded window is converted into a multi-head structure, and each head processes different subspaces. The calculation of attention scores is carried out within each window, so that each head can focus on different local feature combinations within the window. One of the heads may focus on horizontal edges, and another head may focus on vertical edges, thus jointly capturing local feature relationships.
[0032] The co-occurrence branch dynamically adjusts the feature weights by using the preliminary prediction results of the co-occurrence detection head, enabling the model to pay more attention to the regions related to the current scene. For example, in an indoor scene, after detecting the target of the category "table", the model can enhance the feature response to categories such as "chair" or "computer".
[0033] The model learns the relationships between different category embedding vectors through the co-occurrence patterns in the training data. If "table" and "chair" frequently co-occur in the training data, the model may make their embedding vectors have high similarity in space. When "table" is detected, the co-occurrence feature is the embedding vector of "table" (single-category average). The model learns to make the embedding vector of "table" close to the embedding vector of "chair" in the vector space. After fusion, the feature response in the region related to "chair" will be indirectly enhanced. In this way, higher prediction accuracy is achieved.
[0034] The following gives the specific operations of the co-occurrence branch: First, use the co-occurrence detection head (with the same structure as the detection head, serving as the supervision signal source for the early feature layer) to extract the object categories present in the scene, then map the categories to high-dimensional embedding vectors to encode the co-occurrence relationships, specifically manifested as the relationships between the embedding vectors of different categories; finally, perform average pooling to generate the global scene representation and transfer the global features to all positions to provide scene-level context.
[0035] Specifically, the output of the co-occurrence branch is: ; Equation (3) is the output of the collinearity branch, where represents traversing each category and expanding the feature vector of this category to the [B, C, H, W] format; is the co-occurrence feature obtained after average pooling and expanding to the spatial dimension, represents the i-th category, N is the total number of categories, is the feature vector after mapping the category ID extracted from the detection result D; B, C, H, W is the standard dimension format for the deep learning model to process image data, where B represents the batch size, C represents the number of channels, H and W represent the height and width of the image respectively.
[0036] The output of the relationship attention module is expressed as: ; is the output of the co-occurrence relationship attention module, where is the branch adjustment parameter used to balance the contribution degrees of the self-attention branch and the co-linearity branch to the output feature map.
[0037] See Figure 2 , add the co-occurrence relationship attention module to the baseline network yolov8, where n represents the number of repeated bottleneck residuals; specifically, add the co-occurrence relationship attention module after the convolution at the end of the c2f module of the baseline network yolov8, and the value of n is 1 to form the c2f_CRA module.
[0038] The network model constructed based on the c2f_CRA module is as Figure 3As shown, the input image first enters the backbone network. Through the alternation of multiple convolutional layers and C2f modules, deep semantic features are gradually extracted, generating multiple feature maps with different resolutions from P1 to P5. In the neck network, feature maps at different levels are fused through a top-down path and a bottom-up path to obtain the P3, P4, and P5 layers after feature fusion. Finally, the detection head performs detection on them, and the loss function inside the detection head evaluates the quality of the detection.
[0039] 2) Design the edge loss function of the model Existing loss functions are mainly designed for datasets under normal clear conditions. When applied to datasets under complex weather conditions, the performance of the model drops significantly. Analyzing the reasons, the design of existing loss functions is based on an assumption that the detector can fully extract the visual features of the image. However, under harsh weather conditions such as haze, the visual features of the target, especially high-level features, are often masked by environmental noise, resulting in the detector being unable to extract sufficient visual features for accurate judgment. Therefore, the present invention proposes a new loss function by making full use of the information extracted by the existing detector.
[0040] As Figure 4 shown, the calculation flowchart of the edge-aware loss function (Edges Awear-IoU) proposed by the present invention details the input and processing process of this loss function. This loss function requires three inputs: The first input is the Ground Truth, that is, the real bounding boxes annotated in the dataset, and this information is usually provided by public datasets. Through the mask extraction module, the mask map is extracted from the real bounding boxes and input into the edge intersection over union loss function for calculation.
[0041] The second input is the original image of the dataset. After the shallow convolutional processing of the model, the feature map of the intermediate layer is extracted and subjected to depth Sobel convolution processing (the Sobel convolution kernel designed according to the Sobel operator, and using it for convolution operation is called depth Sobel convolution); subsequently, a square root operation is performed to obtain the edge feature map, which is then input into the edge intersection over union loss function.
[0042] The third input is the final output prediction result of the model, that is, the prediction. Through the mask extraction module, the mask of the predicted target bounding box is extracted from the prediction result and input into the loss function.
[0043] The calculation formula of the edge intersection over union loss function is as follows: ; ; Equation 5 is the expression of the edge intersection over union. is the mask for the real target bounding box, M is the mask for the predicted target bounding box, is the edge feature map obtained after depth Sobel convolution and square root. The intersection over union loss function calculates the ratio of the intersection to the union of the real target bounding box and the predicted target bounding box. The edge intersection over union loss function calculates the overlap degree of the edge information within the real bounding box and the predicted target bounding box on the edge feature map. Compared with the intersection over union, it emphasizes the role of the edge in the loss function more strongly.
[0044] Equation 6 is the equation of the edge-aware loss function, where is the edge intersection over union loss, represents the Euclidean distance, and b, w, h represent the center point coordinates, width, and height of the bounding box respectively, , , represent the center point coordinates, width, and height of the predicted bounding box respectively, , , represent the diagonal length, width, and height of the minimum bounding rectangle respectively, is the center point distance loss between the real box and the target box, is the width loss between the real box and the target box, is the height loss between the real box and the target box.
[0045] The edge-aware loss function absorbs the advantages of the loss function EIoU (Efficient Intersection over Union), considering the overlapping area, the distance between the center points, and the real differences in the length and width of the sides. By explicitly measuring the differences in the three geometric factors in the bounding box, namely the overlapping area, the center point, and the side lengths, it optimizes the bounding box regression model. On this basis, for complex weather conditions, it increases the proportion of edge information in the model decision-making, enabling the model to achieve higher performance.
[0046] The calculation result of the edge-aware loss function is a scalar value, representing the average loss value between all the predicted boxes and the real boxes in the current batch (Batch). For example, if the batch size is 64, the loss will perform a weighted average on the differences between the predicted boxes and the real boxes of these 64 samples, and finally output a numerical value. The smaller the output loss value, the higher the edge overlap degree, the closer the center points, and the more consistent the aspect ratios between the predicted box and the real box. During the training process, the model optimizes the loss through backpropagation, making the predicted box gradually approach the position and size of the real box.
[0047] 3) Training of the network model (1) Dataset selection; The dataset used by the network is the RTTS dataset, which includes 4,322 pieces of object detection data in real-world foggy scenarios, covering 5 categories: "person", "bus", "car", "motorbike", and "bike". Among them, 5,311 pictures were selected as the training set, and 3,457 pictures were selected as the validation set for the training and testing of the network.
[0048] (2) Select model evaluation metrics. In order to quantify the effect of the model and compare the advantages and disadvantages of different models, a reasonable evaluation metric is needed. mAP is one of the most important evaluation metrics in object detection algorithms. It measures the average performance of the model across different categories, obtained by calculating the average precision (AP) for each category and then taking the average. AP is the area under the Precision-Recall curve, representing the classification effect of the model for each category. mAP takes into account the performance of the model across all categories and is therefore a comprehensive evaluation metric. FPS, i.e., frames per second, is a metric used to measure the processing speed of the algorithm. It indicates how many frames of pictures the network can process (detect) per second, that is, the number of pictures that can be processed per second or the time required to process one picture. The higher the FPS, the faster the inference speed of the model, which is particularly important for application scenarios that require real-time processing.
[0049] During the training process, the model optimizes the loss through backpropagation, making the predicted bounding boxes gradually approach the position and size of the ground truth boxes; the smaller the output loss value, the higher the degree of overlap between the edges of the predicted bounding box and the ground truth box, the closer the center points, and the more consistent the aspect ratios. (3) Conduct comparative experiments. Both the baseline model and the improved model are trained on the RTTS dataset, and the experimental results are recorded.
[0050] The trained model is obtained for object detection.
[0051] Comparison of heatmaps of the improved model is as Figure 5 、 Figure 6 shown. Figure 5 is the visual saliency heatmap of the baseline model, Figure 6 is the corresponding visualization result of the improved model of the present invention. Visual saliency analysis based on color coding shows (warm colors represent high-attention regions) that although both the baseline model and the improved model can locate two types of target objects, pedestrians and electric vehicles, there are significant differences in their attention distributions. There is an obvious spatial dispersion phenomenon in the saliency region of the baseline model, and some high response values are abnormally distributed in the background regions irrelevant to the target semantics (such as ground textures and building edges).
[0052] In contrast, the present invention achieves more precise attention focusing through architecture optimization, and its significant response mainly focuses on the key feature regions of the target. For example, for electric vehicles, it focuses on the wheels and the front of the vehicle. This thermodynamic distribution feature with physical interpretability indicates that the improved model can effectively suppress background noise interference by enhancing the feature selection ability, thereby improving the credibility and robustness of the target detection task.
[0053] The present invention adds the proposed edge-aware loss function and the co-occurrence relationship attention module CRA to the baseline network YOLOv8, and conducts a comparative evaluation with the baseline network. It conducts 300 iterations of training on the RTTS dataset respectively and adopts an early stopping strategy. The map0.5, map0.5-0.95, and FPS are selected as evaluation indicators, and the comparison test results are shown in Table 1.
[0054] Table 1 Comparison table of experimental results
[0055] It can be seen from the experimental results that adding EAIoU and co-occurrence relationship attention respectively can improve the accuracy of the model, and adding EAIoU and co-occurrence attention simultaneously can also improve the accuracy of the model, which proves that both modules play a certain role. This result also proves the effectiveness and advancement of the method designed by the present invention.
[0056] The embodiments given above are the better examples for implementing the present invention, and the present invention is not limited to the above embodiments. Any non-essential addition or replacement made by those skilled in the art according to the technical features of the technical solution of the present invention shall fall within the protection scope of the present invention.
Claims
1. A haze image target detection model, comprising a backbone network, a neck network and a detection head, characterized in that, A relation attention module is also provided in the model, which includes a self-attention branch and a co-occurrence relation branch; and the loss function inside the detection head is updated to an edge-aware loss function; The self-attention branch calculates lightweight multi-head self-attention through a local window to capture local feature relationships; the co-occurrence relation branch generates global co-occurrence features through the detection results of the co-occurrence detection head to provide scene-level context; The output of the relation attention module weights and adds the output feature maps of the self-attention branch and the co-occurrence branch; The edge-aware loss function includes: ; ; Among them, Equation (5) is the expression of the edge intersection over union, and Equation (6) is the edge-aware loss function; is the mask of the real target bounding box, M is the mask of the predicted target bounding box, is the edge feature map obtained after depth Sobel convolution and square root; is the Intersection over Union (IoU) loss for the edges, represents the Euclidean distance, where b, w, and h represent the center coordinates, width, and height of the bounding box respectively, , , represent the center coordinates, width, and height of the predicted bounding box respectively, , , represent the diagonal length, width, and height of the minimum bounding rectangle respectively, is the loss of the distance between the center points of the ground truth box and the target box, is the loss of the width between the ground truth box and the target box, is the loss of the height between the ground truth box and the target box.
2. The haze image target detection model according to claim 1, wherein The self-attention branch first adds position encoding to the feature map, and then sequentially performs operations including convolution, channel-wise division, window division, and window unfolding on it to obtain three elements: query, index, and content for calculating self-attention; the three obtained elements are calculated through multi-head self-attention and then merged output feature map, and finally a residual connection is made with the input feature map to capture local feature relationships.
3. The haze image target detection model according to claim 2, wherein The self-attention branch is implemented based on Equations (1) and (2): ; Among them, is the input feature map of co-occurrence relationship attention X is the output after position encoding, represents a 1×1 convolution, represents operations of dividing by channel dimension and by window, represents the operation of unfolding by window; ; Among them, , , are derived from Equation (1), represents the number of attention heads, is the activation function, T represents transpose, represents shape conversion of the outputs of different heads to achieve unification, represents folding and restoring the calculation results previously divided by channel dimension and by window.
4. The haze image target detection model according to claim 1, wherein, The co-occurrence branch first uses the co-occurrence detection head to extract the object categories existing in the scene, and then maps the categories to high-dimensional embedding vectors to encode the co-occurrence relationships between the embedding vectors of different categories; then average pooling is used to generate a global scene representation, and the global features are passed to all positions to provide scene-level context; The co-occurrence branch also learns the relationships between the embedding vectors of different categories through the co-occurrence patterns in the training data to achieve higher prediction accuracy.
5. The haze image target detection model according to claim 4, characterized in that, The output of the co-occurrence branch is: ; Among them, is the co-occurrence feature obtained after average pooling and expanding to the spatial dimension, representing the i-th category, N being the total number of categories, is the feature vector after mapping the category ID extracted from the detection result D; B, C, H, W is the standard dimension format for the deep learning model to process image data, where B represents the batch size, C represents the number of channels, H and W represent the height and width of the image respectively.
6. The haze image target detection model according to claim 4, wherein The output of the relation attention module is expressed as: ; is the output of the co-occurrence relationship attention module, where is the branch adjustment parameter used to balance the output of the self-attention branch and the output of the collinear branch to the contribution degree of the output feature map.
7. The haze image target detection model according to claim 1, wherein The backbone network, neck network, and detection head adopt the baseline network yolov8, and a co-occurrence relation attention module is added after the convolution at the end of the c2f module of the baseline network yolov8; In the constructed model, the input image first enters the backbone network, and through the alternation of multiple convolutional layers and C2f modules, deep semantic features are gradually extracted to generate multiple feature maps with different resolutions from P1 to P5; In the neck network, the feature maps of different levels are feature fused through a top-down path and a bottom-up path to obtain the P3 layer, P4 layer, and P5 layer after feature fusion; finally, the detection head performs detection on it, and the loss function inside the detection head evaluates the quality of the detection.
8. The haze image target detection model according to claim 1, characterized in that, The edge-aware loss function requires three inputs: The first input is the ground truth bounding box annotated in the dataset. A mask map is extracted from the ground truth bounding box through a mask extraction module and input into the edge intersection over union loss function for calculation; The second input is the original image of the dataset. After being processed by the shallow convolution of the model, the feature map of the intermediate layer is extracted, and it is processed by a depth Sobel convolution, and then a square root operation is performed to obtain the edge feature map, which is then input into the edge intersection over union loss function; The third input is the output prediction result of the model. A mask of the predicted target bounding box is extracted from the prediction result through a mask extraction module and input into the loss function.
9. A detection algorithm for the haze image target detection model according to claim 1, characterized in that, Including the following operations: 1) Construct a haze image object detection model that includes edge-aware loss and co-occurrence relationships; and update the loss function to an edge-aware loss function; 2) Training of the network model, selecting a dataset and model evaluation metrics; during the training process, the model optimizes the loss through backpropagation, making the predicted bounding box gradually approach the position and size of the ground truth bounding box; the smaller the output loss value, the higher the edge overlap degree between the predicted bounding box and the ground truth bounding box, the closer the center points, and the more consistent the aspect ratios; 3) Obtain the trained model for object detection.
Citation Information
Patent Citations
Edge-guided multi-attention RGBD underwater salient target detection method
CN117095277A
Pixel-level mask assisted camouflage target detection method based on edge prior guidance
CN119251812A
Intelligent monitoring image defogging method and system based on coring attention
CN119273581A
Building height estimation method and device fusing multi-view building roof contour
CN119313721A
Apparatus for estimating under monocular infrared thermal imaging vision pose of object grasped by manipulator, and method thereof
WO2024148645A1