A Low-Light Object Detection Method Based on Multi-Dimensional Attention Mechanism

By introducing FEM-MDT and MDAM modules in the RT-DETR model, the problem of image quality reduction, noise interference and detailed information loss of target detection in low-illumination environments is solved, and higher detection accuracy and faster detection speed are achieved.

CN119832220BActive Publication Date: 2025-06-03ANHUI UNIV OF SCI & TECH GUOZHEN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510303964.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-03
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

In low-illumination environments, object detection faces problems such as image quality degradation, noise interference and loss of detail information, resulting in poor performance of traditional methods in detection accuracy and speed.

Method used

The low-illumination object detection method based on a multi-dimensional attention mechanism is adopted. The FEM-MDT module is introduced into the backbone network of the RT-DETR model for feature extraction, and the MDAM module is added to the neck structure to improve the accuracy of feature extraction and the model's understanding of features.

Benefits of technology

It significantly improves the accuracy and detection accuracy of feature extraction under low illumination conditions, enhances the model's processing ability of low illumination images, and reduces missed and missed detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832220B_ABST
    Figure CN119832220B_ABST
Patent Text Reader

Abstract

The present invention discloses a low-light target detection method based on a multi-dimensional attention mechanism. First, a low-light image dataset is collected and converted into the YOLO format. Then, a low-light target detection model is constructed for target detection. The low-light target detection model is based on the RT-DETR model as the benchmark network, including a backbone network, a neck structure, and a head structure. The backbone network is a feature extraction module based on a multi-scale dilated transformer to extract features from the input feature map. A multi-dimensional attention module is added to the neck structure. Finally, the constructed low-light target detection model is iteratively trained to obtain a trained low-light target detection model, and the trained low-light target detection model is used to detect the target to be detected in the low-light image to obtain the detection result. The low-light target detection model constructed by the present invention improves the accuracy and reliability of feature extraction under low-light conditions and enhances the model's ability to understand features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image data processing, and specifically to a low-light target detection method based on a multi-dimensional attention mechanism. Background Art

[0002] In the field of computer vision, target detection always occupies a core position, and its wide range of applications deeply affects the development directions of many industries and technologies. Target detection aims to accurately locate and identify specific targets from images or video sequences. This technology plays an indispensable cornerstone role in aspects such as autonomous driving, intelligent security monitoring, industrial inspection, and many intelligent image analysis systems. However, when the scene switches to a low-light environment, target detection faces extremely severe challenges, and its technical background becomes extremely complex. Low-light images mainly result from severely insufficient lighting conditions, such as outdoor scenes without auxiliary lighting at night, dim indoor environments, or special working areas with weak light. In this case, the image quality drops sharply, presenting many characteristics that are unfavorable for target detection.

[0003] First of all, the contrast of low-light images is significantly reduced. The weak light makes the brightness difference between the target and the background become blurred, and the original clearly distinguishable target edges and texture information are difficult to distinguish under low contrast, which poses a great obstacle to traditional methods that rely on target edge and texture features for detection. For example, in some target localization algorithms based on edge detection, low contrast will lead to inaccurate edge extraction, and thus the position and contour of the target cannot be accurately defined. Secondly, noise becomes a prominent problem in low-light images. Due to insufficient light, the image sensor needs to increase the gain to enhance the signal intensity when collecting images, but this also amplifies the noise signal. These noises exist in various forms such as salt-and-pepper noise and Gaussian noise, seriously interfering with the true feature information of the target. For traditional feature extraction methods, the existence of noise greatly reduces the stability and reliability of the extracted features, easily misjudging the noise as target features or masking the true features of the target, thus causing false detections or missed detections. Moreover, a large amount of detail information in low-light images is lost. The lack of light makes the fine structure and local features of the target difficult to clearly present. In deep learning models, this means that the effective information available for learning and identifying the target is reduced. If many deep learning models are not specially processed for low-light images during the training process, it is difficult to learn the essential features of the target from the blurred details, resulting in low accuracy in actual low-light detection tasks.

[0004] From the perspective of technology development, traditional target detection technology is mainly divided into two categories: manual feature-based and deep learning-based methods. Manual feature-based methods, such as HOG and SIFT, extract target features through manually designed feature descriptors, and then use classifiers to identify and locate targets. However, this type of method is highly dependent on professional knowledge and rich experience, and the generalization ability of manually designed features is limited, making it difficult to adapt to complex and changeable low-light scenes. At the same time, the feature extraction process is complex and the computational efficiency is not high. With the rise of deep learning, target detection technology has ushered in a new revolution. The two-stage target detection technology based on deep learning, represented by the Faster R-CNN algorithm, first generates possible candidate regions through the region proposal network, and then extracts and classifies these regions. Although it has certain advantages in detection accuracy, it has problems such as slow detection speed, complex model structure, long training time, and high hardware resource requirements. Single-stage target detection technology, such as the YOLO series, regards target detection as a regression problem and realizes real-time detection, but the early versions have relatively low detection accuracy in complex low-light scenes, and are prone to missed detection and false detection.

[0005] In recent years, attention mechanisms have been introduced into the field of target detection to improve performance. It can guide the model to focus on key areas or features in the image, and improve the perception of the target to a certain extent. However, in terms of low-light small target detection, the existing attention mechanism is still insufficient. Due to the complex background of low-light images and weak target signals, some attention modules are difficult to accurately assign attention weights, are easily disturbed by noise, or fail to highlight the focus on small targets, which ultimately affects the detection accuracy.

[0006] As an emerging target detection technology, RT-DETR (Real-Time Detection Transformer) has the advantages of fast detection speed, high accuracy, and relatively simple model structure for easy deployment and application in conventional scenarios. However, it also faces many difficulties in the field of low-light target detection. In low-light environments with complex backgrounds, such as nighttime city streets, where light reflections and shadows are intertwined, RT-DETR is difficult to accurately define the boundary between the target and the background, resulting in inaccurate positioning, misjudgment or missed targets. For small targets with weak features under low illumination and easily masked by noise, its general detection mechanism is difficult to effectively capture subtle features, resulting in a significant decline in detection accuracy, frequent missed detections and false detections. In addition, it lacks effective denoising or anti-noise strategies specifically for low-light image noise, and is easily interfered by noise during feature extraction, resulting in deviations in the extracted target features, which in turn affects the detection results. When the target changes its posture in a low-light scene, RT-DETR is also difficult to adapt quickly, and cannot accurately extract effective features of targets in different postures, which interferes with detection and classification. Summary of the invention

[0007] The technical problem to be solved by the present invention is to provide a low-light target detection method based on a multi-dimensional attention mechanism, using the RT-DETR model as the benchmark network. Its backbone network uses the FEM-MDT module to extract features from the input feature map, and the MDAM module is added to the neck structure, thereby improving the accuracy and reliability of feature extraction under low-light conditions and enhancing the model's ability to understand features.

[0008] The technical solution of the present invention is as follows:

[0009] A low-light target detection method based on a multi-dimensional attention mechanism specifically includes the following steps:

[0010] (1) Collect a low-light image dataset and convert the format of the low-light image dataset to the YOLO format;

[0011] (2) Construct a low-light target detection model. The low-light target detection model uses the RT-DETR model as the benchmark network, including a backbone network, a neck structure, and a head structure; the backbone network uses the FEM-MDT module to extract features from the input feature map; the MDAM module is added to the neck structure to fuse feature maps of different scales output by the backbone network; the head structure is a decoder that decodes the feature map output by the neck structure, decodes the target information, and outputs the final detection result;

[0012] The FEM-MDT module is a feature extraction module based on a multi-scale dilated transformer, with dilated convolutions with different dilation rates as the core, and uses sliding windows with different dilation rates applied to the image to extract features;

[0013] The MDAM module is a multi-dimensional attention module that uses a three-branch structure to capture cross-dimensional interactions of the input data, thereby calculating attention weights;

[0014] (3) Iteratively train the constructed low-light target detection model to obtain a trained low-light target detection model, and use the trained low-light target detection model to detect the target to be detected in the low-light image to obtain the detection result.

[0015] The processing process of the FEM-MDT module specifically includes the following steps:

[0016] S11. First, divide the input feature map into non-overlapping small regions, and then obtain tensors of query , key , and value through linear projection. Then, query , key , and value The tensors are respectively fed into the sliding window dilated attention with different dilation rates. Specifically: the query tensor is fed into the dilated convolution with a dilation rate of 1, and the key tensor is fed into the dilated convolution with a dilation rate of 2, and the value tensor is fed into the dilated convolution with a dilation rate of 3, obtaining feature maps with different receptive fields . The processing process is shown in the following formula (1), and the calculation formula for the receptive field convolution kernel size is shown in the following formula (2);

[0017] (1);

[0018] (2);

[0019] In formulas (1) and (2), is the sliding window dilated attention, , and respectively represent the query matrix, key matrix, and value matrix, is the dilation rate, represents the receptive field convolution kernel size, represents the original convolution kernel size, and the receptive field after query dilated convolution is , the receptive field after key dilated convolution is , and the receptive field after value dilated convolution is ;

[0020] Then, the three obtained feature maps are respectively input into the channel attention for processing. The obtained attention map and the feature map are multiplied element-wise to obtain the feature map , as shown in the following formula (3) specifically. Then, the feature map is input into the spatial attention for processing. The obtained attention map and the feature map are multiplied element-wise to obtain the feature map , as shown in the following formula (4):

[0021] (3);

[0022] (4);

[0023] In formulas (3) and (4), is the channel attention, is the spatial attention, is an element-wise multiplication operation;

[0024] Then, the three output feature maps of different scales are concatenated and linearly layered to obtain a feature map , and then are successively subjected to global average pooling, convolution, ReLU activation function, convolution, and Sigmoid function activation operations, and then element-wise multiplied with . The multiplication output is then passed through a GELU function and batch normalization, and element-wise added to the initially input feature map before output.

[0025] The backbone network includes an HGStem module, six FEM-MDT modules, and three DWConv modules. The processing process of the backbone network is shown in the following formula (5):

[0026] (5);

[0027] In formula (5), is the input of the backbone network, i.e., the collected low-light image, , , , are four feature maps of different scales extracted by the backbone network.

[0028] The neck structure includes nine Conv 1×1 modules, four Conv 3×3 modules, two MDAM modules, one AIFI module, two ConvTranspose modules, four RepC3 modules, and four Concat modules. The processing process of the neck structure is shown in the following formula (6):

[0029] (6);

[0030] In formula (6), , , , , , and are intermediate feature maps of the neck structure, , and are three feature maps output by the neck structure;

[0031] The three feature maps output by the neck structure , and are input into the decoder of the head structure to decode the target information and output the final detection result.

[0032] The processing process of the MDAM module specifically includes the following steps:

[0033] S21. Input the tensor into the MDAM module. In the first branch, call the Permute function to make the tensor be swapped to a tensor with a dimension of . The tensor successively passes through pooling, convolution, batch normalization, and the Sigmoid activation function operation to generate attention weights. The attention weights are element-wise multiplied with the tensor to obtain the tensor . Then call the Permute function again to make the tensor be swapped to a tensor with a dimension of . ; ;

[0034] In the second branch, call the Permute function to make the tensor be swapped to a tensor with a dimension of . The tensor successively passes through pooling, convolution, batch normalization, and the Sigmoid activation function operation to generate attention weights. The attention weights are element-wise multiplied with the tensor to obtain the tensor . Then call the Permute function again to make the tensor be swapped to a tensor with a dimension of . ; ;

[0035] In the third branch, the tensor successively passes through channel pooling, convolution, batch normalization, and the Sigmoid activation function operation to generate attention weights. The attention weights are element-wise multiplied with the tensor to obtain the tensor ;

[0036] Then, the tensors , and are element-wise added to obtain ;

[0037] S22. Divide into four feature maps with different scales. Specifically, divide , that is using The convolution is used for downsampling to obtain a feature map , and then for the feature map use the convolution for downsampling to obtain a feature map , and again for the feature map use the convolution for downsampling to obtain a feature map , and then the four feature maps with different scales , , and after being processed by the channel attention module, the output attention weight coefficients are respectively multiplied element-wise with the feature maps of the corresponding scales , , or , and then four feature maps with different scales are obtained through the convolution , , and . For , an upsampling operation is performed to obtain , which is directly used as . For and , downsampling operations are both performed to obtain and . Then, for , , and , after performing smooth convolution operations respectively, the four outputs of the smooth convolution operations are multiplied element-wise to obtain the finally output feature map.

[0038] The processing procedure of the pooling is shown in the following formula (7):

[0039] (7);

[0040] In formula (7), represents the input of the pooling, represents the output result after the dimensional reduction of the pooling, represents max pooling, represents average pooling, and the subscript in represents the 0th dimension for performing max pooling and average pooling operations, i.e., the batch size.

[0041] The processing process of the described channel attention module is specifically as follows: After the input feature map is processed by global average pooling and global max pooling respectively, the feature map processed by global average pooling and the feature map processed by global max pooling are added element by element, and then after being activated by the Sigmoid activation function, the attention weight coefficient is obtained.

[0042] Advantages of the present invention:

[0043] (1) The backbone network of the low-light target detection model of the present invention uses the FEM-MDT module (feature extraction module based on multi-scale dilated transformers) to extract features from the input feature map. The FEM-MDT module can obtain feature information at different scales and pay attention to the target itself and the broader context information around the target. The sliding window in the FEM-MDT module can traverse the entire image in a systematic way. By starting from the upper left corner of the image and moving the window with a fixed step size, it ensures that every part of the image is detected, thereby significantly improving the accuracy and reliability of feature extraction under low-light conditions and enhancing the feature extraction ability of the visual system.

[0044] (2) The present invention adds the MDAM module to the neck structure, and uses a three-branch structure to capture the cross-dimensional interaction of the input data, thereby calculating the attention weight, enabling the low-light target detection model to build the mutual dependence between channels and spatial positions, and improving the feature understanding ability of the low-light target detection model. Description of the Drawings

[0045] Figure 1 is the flowchart of the present invention.

[0046] Figure 2 is the network structure diagram of the low-light target detection model of the present invention.

[0047] Figure 3 is the network structure diagram of the FEM-MDT module of the present invention.

[0048] Figure 4 is the network structure diagram of the three-branch structure of the MDAM module of the present invention.

[0049] Figure 5 is the network structure diagram of the fusion of feature maps of different scales in the MDAM module of the present invention. Detailed Embodiment

[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0051] See Figure 1 , a low-light target detection method based on a multi-dimensional attention mechanism, specifically including the following steps:

[0052] (1). Collect a low-light image dataset, convert the format of the low-light image dataset into the YOLO format, and divide the low-light image dataset into a training set, a validation set, and a test set, with the proportions being 7:2:1 respectively; among them, the training set is used for model training to update the model weights; the validation set is used for model evaluation after each round of training; the test set is used to evaluate the final detection performance of the model after training ends;

[0053] (2). Build a low-light target detection model and perform target retrieval;

[0054] See Figure 2 , the low-light target detection model takes the RT-DETR model as the backbone network, including a backbone network, a neck structure, and a head structure;

[0055] The backbone network includes an HGStem module, six FEM-MDT modules, and three DWConv modules ( Figure 2 the modules numbered 0-9 in ), and the processing process of the backbone network is shown in the following formula (5):

[0056] (5);

[0057] In formula (5), is the input of the backbone network, that is, the collected low-light image, , , , are four different-scale feature maps extracted by the backbone network;

[0058] The neck structure includes nine Conv 1×1 (1×1 convolution) modules, four Conv 3×3 (3×3 convolution) modules, two MDAM modules, one AIFI module, two ConvTranspose modules, four RepC3 modules, and four Concat (concatenation) modules ( Figure 2 the modules numbered 10-35 in ), and the processing process of the neck structure is shown in the following formula (6):

[0059] (6);

[0060] In formula (6), , , , , , and are the intermediate feature maps of the neck structure, , and are the three feature maps output by the neck structure;

[0061] The three feature maps output by the neck structure , and are input into the decoder of the head structure, and the target information is decoded and the final detection result is output;

[0062] See Figure 3 , the FEM-MDT module is a feature extraction module based on a multi-scale dilated transformer, with dilated convolutions with different dilation rates as the core, and a sliding window is applied to the image by dilated convolutions with different dilation rates to extract features. The processing process of the FEM-MDT module specifically includes the following steps:

[0063] S11. First, the input feature map is divided into non-overlapping small regions, and then queries , keys , and values tensors are obtained through linear projection. Then, the query , key , and value tensors are respectively fed into the sliding window dilated attention using different dilation rates. Specifically: the query tensor is fed into the dilated convolution with a dilation rate of 1 , the key tensor is fed into the dilated convolution with a dilation rate of 2 , and the value tensor is fed into the dilated convolution with a dilation rate of 3 to obtain feature maps with different receptive fields . The processing process is shown in formula (1) below, and the calculation formula for the receptive field convolution kernel size is shown in formula (2);

[0064] (1);

[0065] (2);

[0066] In formulas (1) and (2), is the sliding window dilated attention, , and represent the query matrix, the key matrix, and the value matrix respectively, is the dilation rate, represents the convolutional kernel size of the receptive field, represents the original convolutional kernel size, and the query is calculated through Equation (2) The receptive field after dilated convolution is and the key The receptive field after dilated convolution is , and the value The receptive field after dilated convolution is ;

[0067] Then, the three obtained feature maps are respectively input into the channel attention for processing, and the obtained attention map is multiplied element-wise with the feature map to obtain the feature map , as shown in Equation (3) below. Then, the feature map is input into the spatial attention for processing, and the obtained attention map is multiplied element-wise with the feature map to obtain the feature map , as shown in Equation (4) below:

[0068] (3);

[0069] (4);

[0070] In Equations (3) and (4), is the channel attention, is the spatial attention, is the element-wise multiplication operation;

[0071] Then, the three output feature maps of different scales are concatenated (Concat) and linearly layered (Linear) to obtain the feature map , and then is successively subjected to global average pooling (Global AvgPool), convolution (Conv 1×1), ReLU activation function, convolution (Conv 1×1), and Sigmoid function activation operations, and then multiplied element-wise with . The multiplied output is then passed through the GELU function and batch normalization (BN), and then added element-wise to the initially input feature map for output;

[0072] The MDAM module is a multi-dimensional attention module that uses a three-branch structure to capture cross-dimensional interactions of the input data, thereby calculating attention weights, enabling the model to build mutual dependencies between channels and spatial positions, and improving the model's ability to understand features. The processing process of the MDAM module specifically includes the following steps:

[0073] S21, see Figure 4 , input the tensor into the MDAM module. In the first branch, call the Permute function to make the tensor switched to a tensor of dimension . The tensor goes through pooling (Z-Pool), convolution (Conv 7×7), batch normalization (BN), and Sigmoid activation function operations in sequence to generate attention weights. The attention weights are multiplied element-wise with the tensor to obtain the tensor . Then call the Permute function again to make the tensor switched to a tensor of dimension ;

[0074] In the second branch, call the Permute function to make the tensor switched to a tensor of dimension . The tensor goes through pooling (Z-Pool), convolution (Conv 7×7), batch normalization (BN), and Sigmoid activation function operations in sequence to generate attention weights. The attention weights are multiplied element-wise with the tensor to obtain the tensor . Then call the Permute function again to make the tensor switched to a tensor of dimension ;

[0075] In the third branch, the tensor goes through channel pooling (Channel Pool), convolution (Conv 7×7), batch normalization (BN), and Sigmoid activation function operations in sequence to generate attention weights. The attention weights are multiplied element-wise with the tensor to obtain the tensor ;

[0076] Then, the tensor ​​​​, and are added element by element to obtain ;

[0077] S22. See Figure 5 . Divide into four feature maps of different scales. Specifically, , that is, is used for convolution for downsampling to obtain the feature map . Then, the feature map is used for convolution for downsampling to obtain the feature map . Again, the feature map is used for convolution for downsampling to obtain the feature map . Then, the four feature maps of different scales , , and are all processed by the channel attention (CA) module. After that, the output attention weight coefficients are respectively multiplied element by element with the feature maps of the corresponding scales , , or . Then, through convolution (Conv 1×1), four feature maps of different scales , , and are obtained. An upsampling (UpSample) operation is performed on to obtain . is directly identity mapped to . Downsampling (DownSample) operations are performed on and to obtain and . Then, smooth convolution (SmooothConv) operations are respectively performed on , , and . After that, the four outputs of the smooth convolution operations are multiplied element by element to obtain the finally output feature map;

[0078] The above-mentioned processing process of pooling is shown in the following formula (7):

[0079] (7);

[0080] In formula (7), represents the input of pooling, represents the output result after dimensionality reduction of pooling, represents max pooling, represents average pooling, and the subscript in represents the 0th dimension for max pooling and average pooling operations, i.e., the batch size;

[0081] The processing process of the channel attention (CA) module is specifically as follows: After the input feature map is respectively processed by global average pooling and global max pooling, then the feature map processed by global average pooling and the feature map processed by global max pooling are added element-wise, and then after activation by the Sigmoid activation function, the attention weight coefficient is obtained;

[0082] (3) Iteratively train the constructed low-light object detection model, use the SGD optimizer to optimize the model, the batch normalization size is 16, and the number of iterations is 120 times, to obtain a trained low-light object detection model, and use the trained low-light object detection model to detect the object to be detected in the low-light image to obtain the detection result.

[0083] The low-light object detection model is iteratively trained using the ExDark dataset, and the SGD (Stochastic Gradient Descent) optimizer is used to optimize the model, the batch normalization size is 16, and the total number of training cycles is 120 times, to obtain a trained low-light object detection model.

[0084] In specific implementation, the present application provides a computer storage medium and a corresponding data processing unit. Among them, the computer storage medium can store a computer program, and the computer program can be executed by the data processing unit to run the inventive content of a low-light object detection method based on a multi-dimensional attention mechanism and some or all of the steps in each embodiment. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0085] Table 1 below shows the comparative experiments conducted using the ExDark dataset. The experimental results in Table 1 indicate that the model proposed in this embodiment (Ours) leads other existing models (TOOD, ATSS, R-CNN, Faster R-CNN, YOLOv8n, YOLOv10n, and RT-DETR) with a precision of 0.854, demonstrating its best performance in correctly identifying positive samples. In contrast, the YOLOv8n model has the lowest precision (0.756). The mAP@0.5 of the model proposed in this embodiment (Ours) is 0.817, which is also better than other existing models.

[0086] Table 1

[0087]

[0088] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A low-light target detection method based on a multi-dimensional attention mechanism, characterized in that: The specific steps include: (1) Collect low-light image datasets and convert the format of the low-light image datasets into YOLO format; (2) Construct a low-light target detection model. The low-light target detection model uses the RT-DETR model as the benchmark network, including a backbone network, a neck structure, and a head structure. The backbone network extracts features from the input feature map based on the FEM-MDT module. The MDAM module is added to the neck structure to fuse feature maps of different scales output by the backbone network. The head structure is a decoder that decodes the feature map output by the neck structure, decodes the target information, and outputs the final detection result. The processing of the FEM-MDT module includes: S11. Divide the input feature map into small non-overlapping areas, and then obtain the query ,key ,value Tensor, query ,key ,value The tensors are sent to the sliding window expansion attention with different expansion rates to obtain feature maps with different receptive fields. ; The three feature maps obtained They are input into the channel attention for processing respectively, and the attention map and feature map obtained after processing Perform element-by-element multiplication to obtain the feature map , and then the feature map Input into the spatial attention for processing, and then get the attention map and feature map Perform element-by-element multiplication to obtain the feature map ; The three feature maps of different scales are concatenated and linearized to obtain the feature map. , and then Perform global average pooling, Convolution, ReLU activation function, After convolution and Sigmoid function activation operations, Perform element-by-element multiplication, and then add the multiplication output to the original input feature map element-by-element after passing through the GELU function and batch normalization for output; The MDAM module is a multi-dimensional attention module that uses a three-branch structure to capture the cross-dimensional interaction of input data to calculate the attention weight; (3) Iteratively train the constructed low-illumination target detection model to obtain the trained low-illumination target detection model, and then detect the target to be detected in the low-illumination image to obtain the detection result.

2. The low-light target detection method based on a multi-dimensional attention mechanism according to claim 1, characterized in that: The query will ,key ,value The tensors are fed into the sliding window dilation attention with different dilation rates, specifically: query The tensor is fed into a dilation rate of 1 Dilated convolution, key The tensor is fed into a dilation rate of 2 Dilated convolution, value The tensor is fed into a dilation rate of 3 The processing process of dilated convolution is shown in the following formula (1), and the calculation formula of the convolution kernel size of the receptive field is shown in the following formula (2); (1); (2); In formula (1) and formula (2), Expanding attention for sliding windows, , and represent the query matrix, key matrix and value matrix respectively, is the expansion rate, represents the convolution kernel size of the receptive field, Represents the original convolution kernel size, and is calculated by formula (2) to obtain the query The receptive field after the dilated convolution is ,key The receptive field after the dilated convolution is ,value The receptive field after the dilated convolution is .

3. The low-light target detection method based on a multi-dimensional attention mechanism according to claim 1, characterized in that: The backbone network includes an HGStem module, six FEM-MDT modules and three DWConv modules. The processing process of the backbone network is shown in the following formula (5): (5); In formula (5), As the input of the backbone network, that is, the collected low-light image, , , , Four feature maps of different scales extracted by the backbone network.

4. The low-light target detection method based on a multi-dimensional attention mechanism according to claim 3, characterized in that: The neck structure includes nine Conv 1×1 modules, four Conv 3×3 modules, two MDAM modules, one AIFI module, two ConvTranspose modules, four RepC3 modules and four Concat modules. The processing process of the neck structure is shown in the following formula (6): (6); In formula (6), , , , , , and is the intermediate feature map of the neck structure, , and Three feature maps output for the neck structure; The three characteristic graphs output by the neck structure , and Input into the decoder of the header structure, decode the target information and output the final detection result.

5. The low-light target detection method based on a multi-dimensional attention mechanism according to claim 4, characterized in that: The processing process of the MDAM module specifically includes the following steps: S21. Tensor Input into the MDAM module, in the first branch, call the Permute function so that the tensor Convert to dimension Tensor , tensor Pass by Pooling, After convolution, batch normalization and Sigmoid activation function operations, attention weights are generated. Attention weights and tensors Perform element-by-element multiplication to obtain a tensor , and then call the Permute function again to make the tensor Convert to dimension Tensor ; In the second branch, the Permute function is called to make the tensor Convert to dimension Tensor , tensor Pass by Pooling, After convolution, batch normalization and Sigmoid activation function operations, attention weights are generated. Attention weights and tensors Perform element-by-element multiplication to obtain a tensor , and then call the Permute function again to make the tensor Convert to dimension Tensor ; In the third branch, the tensor After channel pooling, After convolution, batch normalization and Sigmoid activation function operations, attention weights are generated. Attention weights and tensors Perform element-by-element multiplication to obtain a tensor ; Then the tensor , and Add element by element, and we get ; S22, will Divided into four feature maps of different scales, specifically Right now use Convolution is performed to downsample and obtain the feature map , and then the feature map use Convolution is performed to downsample and obtain the feature map , again for the feature map use Convolution is performed to downsample and obtain the feature map , then four feature maps of different scales , , and After being processed by the channel attention module, the output attention weight coefficients are then compared with the feature maps of the corresponding scales. , , or Perform element-by-element multiplication and then pass Convolution obtains four feature maps of different scales , , and ,right Perform upsampling to obtain , Directly ,right and The downsampling operation is performed to obtain and , then , , and After performing smooth convolution operations respectively, the four outputs of the smooth convolution operation are multiplied element by element to obtain the final output feature map.

6. The low-light target detection method based on a multi-dimensional attention mechanism according to claim 5, characterized in that: The The pooling process is shown in the following formula (7): (7); In formula (7), represent Pooled input, represent The output result after pooling dimension reduction, represents maximum pooling, represents average pooling, and Subscript in Represents the 0th dimension for the maximum pooling and average pooling operations, i.e., the batch size.

7. The low-light target detection method based on a multi-dimensional attention mechanism according to claim 5, characterized in that: The processing process of the channel attention module is specifically as follows: the input feature map is processed by global average pooling and global maximum pooling respectively, and then the feature map processed by global average pooling and the feature map processed by global maximum pooling are added element by element, and then activated by Sigmoid activation function to obtain the attention weight coefficient.

Citation Information

Patent Citations

  • Light-weight crayfish quality rapid detection method based on pattern recognition

    CN119152256A

  • Medical image target detection method and device, electronic equipment and readable storage medium

    CN119600267A