SD-RTDETR personnel detection model used in low-light hazardous chemical substance storage

Through the synergy of the MSCA-SCINet image enhancement module and the DNCV-RTDETR target detection network, the problem of personnel detection in hazardous chemical warehouses under low-light conditions was solved, and efficient and accurate target feature extraction and recognition were achieved.

CN120635494APending Publication Date: 2025-09-12BEIJING INSTITUTE OF PETROCHEMICAL TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510777146.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing target detection algorithms have difficulty effectively capturing complex features under low-light conditions, resulting in poor personnel detection in hazardous chemical warehouses. In particular, under complex and low-light conditions, image quality degrades and details become blurred, making it difficult to achieve efficient and accurate personnel identification.

Method used

The MSCA-SCINet image enhancement module and DNCV-RTDETR target detection network are used to improve image quality through self-calibration illumination learning and multi-scale attention mechanism. The deformable convolution and global collaborative feature aggregation network are combined to enhance the ability to capture target features in low-light environments.

Benefits of technology

It significantly improves the accuracy and robustness of personnel detection in hazardous chemical storage scenarios, can effectively deal with the problems of image quality degradation and changing personnel postures in low-light environments, and provides an efficient and reliable detection solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635494A_ABST
    Figure CN120635494A_ABST
Patent Text Reader

Abstract

The invention discloses an SD-RTDETR personnel detection model used in a low-illumination hazardous chemical substance storage, which is used for solving the key technical problem of personnel detection in a hazardous chemical substance storage scene in a low-illumination environment, and comprises an MSCA-SCINet image enhancement module and a DNCV-RTDETR target detection network, the MSCA-SCINet image enhancement module is used for enhancing a low-illumination image, and the DNCV-RTDETR target detection network is used for detecting the DNCV-RTDETR target detection network. Then features in the enhanced image are accurately extracted through the DNCV-RTDETR target detection network, and through the synergistic effect of the MSCA-SCINet image enhancement module and the DNCV-RTDETR target detection network, the model not only can effectively deal with the image quality degradation problem in the low-light environment, but also can overcome the detection challenge brought by the changeable postures of people, and the detection accuracy is improved. And an efficient and reliable solution is provided for personnel detection in a hazardous chemical substance storage scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection models, and in particular to an SD-RTDETR personnel detection model for use in hazardous chemical warehouses under low light conditions. Background Art

[0002] With the acceleration of industrialization, the production, storage, and transportation of hazardous chemicals continues to expand, raising the issue of safety oversight. Hazardous chemicals are flammable, explosive, and toxic, and accidents can easily cause significant casualties. Research shows that a significant proportion of hazardous chemical accidents in the petrochemical industry are caused by the use of open flames, mobile phones, smoking, and other human factors. Therefore, real-time monitoring of personnel in hazardous chemical storage facilities is crucial to provide timely warnings and interventions against unsafe behaviors, minimizing the risk of accidents.

[0003] Traditional safety monitoring methods rely primarily on manual inspections and fixed cameras to monitor unsafe conditions in hazardous chemical storage facilities in real time. However, manual inspections are inefficient and prone to missed detections, while fixed cameras, limited by their viewing angles and resolution, have limited coverage and struggle to effectively monitor complex scenarios. Consequently, AI-based object detection algorithms have been put into practice in petrochemical enterprises with hazardous chemical storage facilities, providing technical support for identifying hazardous human behavior in these environments and achieving promising results. These methods automatically extract features from images, enabling efficient and accurate identification of individuals and their real-time location.

[0004] However, existing object detection algorithms are subject to practical limitations. Due to the irregular shapes of objects caused by varying human poses and occlusion by cargo, traditional convolution operations struggle to effectively capture these complex features. In particular, image quality degrades and details blur under complex and low-light conditions, resulting in poor recognition of individuals. This paper addresses this issue by providing an SD-RTDETR human detection model for low-light hazardous chemical warehouses. Summary of the Invention

[0005] The present invention provides an SD-RTDETR personnel detection model for hazardous chemical storage in low-light conditions. The model first improves image quality through an image enhancement module, and then enhances the ability to capture target features under low-light conditions through a target detection network to improve detection accuracy.

[0006] The technical solution adopted by the present invention to solve the above technical problems is:

[0007] A SD-RTDETR personnel detection model for hazardous chemicals storage in low-light environments is used to solve the key technical problems of personnel detection in hazardous chemicals storage scenarios in low-light environments, such as Figure 1As shown in the figure, it includes the MSCA-SCINet image enhancement module and the DNCV-RTDETR target detection network. The low-light image is first enhanced by the MSCA-SCINet image enhancement module, and then the features in the enhanced image are accurately extracted by the DNCV-RTDETR target detection network. Through the synergistic effect of the MSCA-SCINet image enhancement module and the DNCV-RTDETR target detection network, the model can not only effectively deal with the problem of image quality degradation in low-light environments, but also overcome the detection challenges brought about by the changeable postures of people, providing an efficient and reliable solution for personnel detection in hazardous chemical storage scenarios.

[0008] like Figure 1 As shown in the figure, the MSCA-SCINet image enhancement module is used to solve the problems of low brightness, poor image quality and blurred details of the input low-light images. The MSCA-SCINet image enhancement module is based on the self-calibrated illumination learning model SCINet (Self-Calibrated Illumination Learning Networks). In order to enable the module to better learn the features of each stage and capture multi-scale contextual information, the module introduces the multi-scale attention mechanism MSCA (multi-scale convolutional attention) in the image enhancement (Enhanced Illumination) stage, which enhances the module's perception of global and local features of the image; the MSCA-SCINet image enhancement module outputs high-quality clear images by improving image brightness, enhancing color contrast and restoring image detail information, providing reliable input for subsequent target detection tasks.

[0009] like Figure 1 As shown in the figure, the DNCV-RTDETR target detection network designs the GCFA_DCNv2 module in the deep BasicBlock basic block of the network. The GCFA_DCNv2 module combines the advantages of the global collaborative feature aggregation network GCFA (Global Collaborative Feature Aggregation Network) and the deformable convolution DCNv2 (DeformableConvNets v2). While extracting global features, it dynamically adjusts the receptive field of different deformation and scale features. At the same time, it uses deformable convolution to adaptively adjust the sampling position to more comprehensively capture image features. The DNCV-RTDETR target detection network can not only effectively adapt to changes in people's posture, but also achieve more accurate feature extraction, significantly improving detection accuracy and robustness.

[0010] The MSCA-SCINet image enhancement module includes a self-calibration illumination learning network and an illumination enhancement network. Multiple cycles are required in the MSCA-SCINet image enhancement module. In each cycle, the illumination enhancement network first enhances the low-light image, then self-calibrates the enhanced image using the self-calibration illumination learning network. The enhanced image is then fed into the next cycle for further image enhancement.

[0011] Furthermore, if Figure 2 As shown in the figure, the structure of the self-calibration illumination learning network is as follows: after the low-light image is input, it first passes through the input convolution layer in_Conv to map the input image from 3 channels (RGB) to the specified number of channels. Then, multiple self-calibration modules consisting of the convolution layer Conv, batch normalization Batch Norm and activation function ReLU are used to extract image features. At the same time, the difference between the input image and the target output is learned using the residual to alleviate the gradient vanishing problem. Finally, the residual is output by the output convolution layer out_Conv.

[0012] The self-calibration illumination learning network includes a self-calibration module and an illumination estimation module. In each cycle, the illumination enhancement network first performs image enhancement on the low-light image to obtain an enhanced image, then the illumination estimation module analyzes the enhanced image, and finally the self-calibration module performs self-calibration on the enhanced image. The self-calibrated enhanced image is input into the next cycle, that is, the output of the previous stage is used as the input of the next stage, and the cycle is repeated multiple times.

[0013] Furthermore, if Figure 3 As shown, the specific calculation process of the lighting estimation module is as follows:

[0014] The lighting estimation module is designed based on the Retinex theory. In Retinex theory, a low-light image y is regarded as the product of the illumination component x and the clear image reflection component z:

[0015] y=x⊙z (1-1),

[0016] The illumination component x refers to the part affected by external illumination, i.e., weak illumination, and the reflection component z refers to the constant part determined by the object's own properties, i.e., the clear image part.

[0017] Then introduce the mapping H θ Build a progressive lighting optimization process, the basic units of which are:

[0018]

[0019] Among them, t represents the number of stages, u t Represents the residual of the tth stage, x tRepresents the illumination component of stage t;

[0020] It should be noted that the self-calibration illumination learning network and the illumination enhancement network adopt a weight sharing mechanism. After adopting the weight sharing mechanism, the illumination estimation network at each stage maintains the structure and parameter sharing state. θ The lighting estimation networks at each stage represented are not associated with the stage number t.

[0021] The specific calculation process of the self-calibration module is as follows: the self-calibration module gradually corrects the input of each stage through the Retinex theory, thereby indirectly affecting the output image correction effect of each stage;

[0022] The network structure is designed so that the input of each stage is the output of the previous stage. In order to express the difference between the output of different stages and the original low-light image input in the first stage, a self-calibration mapping function s is introduced. The self-calibration module of the tth stage can be expressed as:

[0023]

[0024] Among them, z t is the reflection component of the tth stage, i.e. the clear image part, is the division symbol, s t is the self-calibration image at stage t, Indicates that it contains learnable parameters The introduction of parameterized operators, v t represents the input of the next stage after calibration,

[0025] Since the output of each stage is the input of the next stage, the basic unit of the illumination self-calibration process is recorded as:

[0026] F(x t )→F(g(x t )) (1-4),

[0027] Through self-calibration in each stage, convergence between stages is achieved.

[0028] Furthermore, since the detection object of the detection model of the present invention is personnel detection in hazardous chemical warehouses, its characteristics are closer to the shape of long strips. Therefore, a multi-scale attention mechanism MSCA that helps to extract strip features is introduced into the illumination enhancement network and embedded into the illumination enhancement network of the present invention.

[0029] The illumination enhancement network, which incorporates the multi-scale attention mechanism MSCA, first maps the original image into a high-dimensional feature space through an input convolutional layer. It then extracts features through multiple convolutional blocks. Each convolutional block consists of a standard convolutional layer and a multi-scale attention module MSCA. The multi-scale attention module MSCA captures the image's global and local contextual information through multi-scale convolution kernels and enhances important features through a channel blending mechanism. Finally, the network maps the features back to the image space through an output convolutional layer and adds them to the input image to generate an enhanced image. The illumination enhancement network, which incorporates the multi-scale attention mechanism MSCA, effectively improves image brightness and contrast while preserving detailed information, providing high-quality input for subsequent object detection tasks.

[0030] Furthermore, if Figure 4 As shown in the figure, the multi-scale attention mechanism MSCA includes depth-wise convolution for aggregating local information, multi-branch depth strip convolution for capturing multi-scale context, and 1×1 convolution for modeling the relationship between different channel dimensions.

[0031] When multi-branch depthwise strip convolution approximates a convolution with a kernel size of 7×7, it separates it into a pair of 7×1 and 1×7 convolutions. 11×11 convolution and 21×21 convolution are similar operations. This not only makes the network lighter, but also better captures the strip features of objects, serving as a supplement to ordinary convolution.

[0032] The output of the 1×1 convolution is directly used as the attention weight to reweight the input of MSCA. The implementation formula of the multi-scale attention mechanism MSCA is:

[0033]

[0034] Where L represents the attention map of MSCA, Branch i (i=0,1,2,3) represents the four branches output from the convolution kernel of 5×5, Fea represents the input feature, DpConv represents the depth convolution, Conv1x1 represents the 1x1 convolution,

[0035] Then the attention map is element-wise multiplied with the input features as the output, and the formula is:

[0036]

[0037] Furthermore, the DNCV-RTDETR target detection network of the present invention adopts ResNet18 as the backbone network, and uses the GCFA_DCNv2 module to replace the second convolution layer in the BasicBlock basic block to form an improved residual structure BasicBlock_GCFA_DCNv2; the GCFA_DCNv2 module includes a global collaborative feature aggregation network GCFA and a deformable convolution DCNv2.

[0038] The ResNet18 network effectively alleviates the gradient vanishing problem in deep neural network training by introducing residual connections, significantly improving the training stability of the model. The ResNet18 network is a shallow network and uses BasicBlock as its residual block structure. Figure 5 As shown in a, the original BasicBlock contains two 3×3 convolutional layers, each of which is followed by batch normalization and ReLU activation functions. At the same time, a 1×1 convolutional layer is designed in the shortcut connection to ensure the consistency of feature dimensions.

[0039] However, in the hazardous chemicals storage scenario where the present invention is applied, the non-fixed posture of personnel leads to irregular target shapes, and traditional convolution operations are difficult to effectively capture such complex features, such as Figure 5 As shown in Figure b, the present invention uses the GCFA_DCNv2 module to replace the second convolutional layer in the BasicBlock basic block to form an improved residual structure BasicBlock_GCFA_DCNv2.

[0040] Furthermore, if Figure 5 As shown in Figure 2b, the improved residual structure BasicBlock_GCFA_DCNv2 is structured as follows: the input feature map first passes through a 3×3 standard convolutional layer, then through a deformable convolutional layer DCNv2 with a global collaborative feature aggregation network (GCFA) to extract features with adaptive sampling positions. Finally, the extracted features are added to a shortcut connection of a 1×1 convolutional layer to form the improved residual structure BasicBlock_GCFA_DCNv2. This structure not only captures more complex target features but also adaptively adjusts the sampling position through deformable convolution, enhancing the model's feature extraction capabilities for irregularly shaped targets, thereby significantly improving detection accuracy.

[0041] Furthermore, the offset generation method of DCNv2_Offset_Attention in the deformable convolutional layer DCNv2 relies only on local convolution features, which makes it difficult to fully capture global context information, resulting in a relatively weak ability to generate offset-mask and an inability to reasonably allocate weights according to the features at different positions of the window. Figure 6 As shown in the figure, in order to solve the above problems, the present invention adds a global collaborative feature aggregation network GCFA to the deformable convolution layer DCNv2. The structure of the global collaborative feature aggregation network GCFA is as follows: while aggregating features in the original height and width spatial dimensions, a global pooling branch is designed, and combined with attention weight calculation, global information is integrated into the offset generation process. Figure 7 As shown in the figure, the DCNv2 module with the addition of GCFA can fully extract the features of each branch while reasonably allocating feature weights, thereby significantly improving the offset-mask generation capability.

[0042] The specific calculation process of the global collaborative feature aggregation network GCFA is as follows:

[0043] Given an input feature X, perform global average pooling on it to obtain the global information of the cth channel,

[0044]

[0045] Among them, x c (i, j) is the value of the input feature X at the cth channel and position (i, j);

[0046] Then, one-dimensional global pooling is performed using pooling kernels (H, 1) and (1, W) in the height and width dimensions respectively;

[0047] The output formula of the cth channel with height h is:

[0048]

[0049] The output formula of the cth channel with width w is:

[0050]

[0051] The two transformations (1-8) and (1-9) aggregate the features along the horizontal coordinate and vertical coordinate space respectively to generate a pair of direction-aware feature maps;

[0052] In order to effectively aggregate the features between channels, the spatial dimension features extracted by formulas (1-8) and (1-9) are first connected, and then input into the 1×1 convolution function F1 to obtain the fused features:

[0053]

[0054] Where [·,·] represents the connection operation along the spatial dimension, δ is the nonlinear activation function, r is the block size reduction ratio used to control the SE block, and C is the number of channels of the input feature map;

[0055] Split formula (1-10) into two separate tensors along the horizontal and vertical dimensions and

[0056] Similarly, after the global pooling branch performs a 1×1 convolution transformation, the intermediate feature f is obtained conv =δ[F1(z c )];

[0057] The connection feature f obtained by formula (1-10) has a dimension of 2. The mean function torch.mean is used to calculate the mean along the spatial dimension.

[0058] f mean =torch.mean(f,dim=2) (1-11),

[0059] After adjusting the feature dimension, the intermediate feature f output in the global pooling branch conv Perform splicing to obtain the global characteristics of the branch:

[0060] f a =[f conv ,f mean ] (1-12),

[0061] The horizontal dimension feature f is transformed into h , vertical dimension feature f w and global features f a Calculate the corresponding attention weight g h 、g w and g a , the formula is as follows

[0062]

[0063] Finally, output features:

[0064] y c (i,j)=x c (i,j)×g h (i)×g w (j)×g a (1-14).

[0065] The added global collaborative feature aggregation network GCFA fully considers the different attentions of different channels and global dimensions, and encodes the spatial position information, so that it can more accurately learn the attention weights of features in different positions and capture the relationship between channels, thereby effectively locating the area of ​​interest in subsequent target recognition.

[0066] The beneficial effects of the present invention are as follows:

[0067] Low-light images are first enhanced using the MSCA-SCINet image enhancement module, and then features are precisely extracted from the enhanced images using the DNCV-RTDETR object detection network. The synergistic effect of the MSCA-SCINet image enhancement module and the DNCV-RTDETR object detection network not only effectively addresses image quality degradation in low-light environments, but also overcomes the detection challenges posed by the changing postures of people, providing an efficient and reliable solution for personnel detection in hazardous chemical storage scenarios.

[0068] The MSCA-SCINet image enhancement module significantly improves the visibility of details in low-light images and solves the problem of unclear target features through a self-calibrated illumination learning network and a multi-scale contextual attention mechanism.

[0069] The GCFA_DCNv2 module introduces a deformable convolution that integrates multiple global collaborative feature aggregation networks in the BasicBlock residual module. It rationally assigns weights to features at different positions, enhances the model's adaptability to irregularly shaped targets, and can effectively distinguish irregular targets from regular targets, thereby more accurately extracting personnel features. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1 Schematic diagram of the overall architecture of the detection model of the present invention;

[0071] Figure 2 Schematic diagram of the self-calibration lighting learning network structure of the present invention;

[0072] Figure 3 A schematic diagram showing the data flow of the self-calibration lighting learning network of the present invention;

[0073] Figure 4 This is a schematic diagram of the MSCA network architecture of the multi-scale attention mechanism of the present invention;

[0074] Figure 5 Schematic diagram comparing the GCFA_DCNv2 module of the present invention with the original BasicBlock;

[0075] Figure 6 This is a schematic diagram of the network architecture of the global collaborative feature aggregation network GCFA of the present invention;

[0076] Figure 7 Schematic diagram of the GCFA_DCNv2 module model of the present invention;

[0077] Figure 8 This is an image enhancement comparison diagram in the embodiment;

[0078] Figure 9 This is the ablation experiment data table in the embodiment;

[0079] Figure 10 This is a comparison experiment diagram of the characteristic heat map in the embodiment;

[0080] Figure 11 4 is a comparison algorithm table of the embodiment. DETAILED DESCRIPTION

[0081] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0082] In the description of the present invention, it should be understood that the terms "center", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention.

[0083] The application scenario of the specific embodiment of the present invention is the hazardous chemical storage environment. When storing hazardous chemicals, most of them are stored in containers such as barrels or square boxes, which appear as regular squares or strips. After people enter, they also appear as long strips. However, for existing detection algorithms with low accuracy, it is difficult to distinguish the strip characteristics shown by people from the stored products that appear as strips. In order to accurately identify personnel in the above situation, the present invention adopts an SD-RTDETR personnel detection model for low-light hazardous chemical storage. The collected low-light image is first enhanced by the MSCA-SCINet image enhancement module to solve the problems of image quality degradation and blurred details caused by the low-light environment. Then, the DNCV-RTDETR target detection network is used to capture more complex target features. At the same time, the sampling position can be adaptively adjusted through deformable convolution, which enhances the feature extraction capability of irregular-shaped targets and can accurately distinguish the irregular shapes formed by personnel from the regular shapes formed by storage products. The accurate extraction of personnel features is achieved, which significantly improves the detection accuracy of the detection model of the present invention and enables it to be applied to low-light hazardous chemical storage environments. When the present invention is implemented, a model for personnel detection in a low-light hazardous chemical storage environment is first constructed, and then it is trained and optimized, and finally its effectiveness is verified.

[0084] This embodiment selects the public dataset ExDark and the self-built low-light hazardous chemicals warehouse personnel dataset to conduct network design effectiveness verification experiments. The ExDark dataset is an image dataset specially taken in low-light environments, with a total of 7363 images and 12 categories. The dataset is divided into a Train training set of 6622 images and a Val validation set of 741 images; the self-built low-light hazardous chemicals warehouse personnel dataset uses a real hazardous chemicals warehouse in a certain place to scale up and simulate the warehouse, and takes low-light images, finally obtaining an image dataset with a training set of 763 images and a validation set of 205 images.

[0085] This example uses precision, recall, and mean average precision (mAP@0.5) to evaluate the model.

[0086] The accuracy formula is:

[0087]

[0088] The recall formula is:

[0089]

[0090] Among them, TP means that the predicted value is the same as the true value, the predicted value is a positive sample, that is, the true value is a positive sample; FP means that the predicted value is different from the true value, the predicted value is a positive sample, that is, the true value is a negative sample; FN means that the predicted value is different from the true value, the predicted value is a negative sample, that is, the true value is a positive sample.

[0091] mAP@0.5 means that when the IoU threshold is 0.5, the precision rate is calculated for each category of samples, and then the average of these precision rates is taken. mAP@0.5:0.95 means that the average mAP value is calculated for different IoU thresholds ranging from 0.5 to 0.95 with a step size of 0.05. For each IoU threshold, this embodiment calculates the precision-recall curve for each category and calculates the average precision under the curve; then the average precision under all IoU thresholds is averaged to obtain the value of mAP@0.5:0.95. The specific formula is as follows:

[0092]

[0093] Where c represents the total number of target categories defined in the dataset;

[0094]

[0095] like Figure 8 As shown in the figure, the robustness of the detection model algorithm was verified through low-light image enhancement comparative tests. 114 test samples were randomly selected from the EaDark dataset, covering 10 typical low-light scenarios: strong nighttime light, weak nighttime light, twilight, and indoor dim light. A targeted analysis was performed based on the characteristics of hazardous chemical storage scenarios. The detection model was compared with the enhancement effects of Zero-DCE++ and the original SCI algorithm. The qualitative evaluation results are as follows:

[0096] Although the Zero-DCE++ algorithm can improve the overall brightness and performs well in improving the dark side of objects, it suffers from local overexposure in images with stronger light. Since most low-light images in hazardous chemical warehouses are indoor low-light images, it is difficult to meet the strict image fidelity requirements of industrial scenarios. The original SCI algorithm performs poorly in color reproduction, and the enhanced image exhibits an unnatural yellow tint overall, and there are significant deviations in the reconstruction of people's skin color. The detection model of the present invention overcomes the defects of traditional algorithms such as inconspicuous details and unnatural colors caused by the enhancement process, and is particularly suitable for indoor low-light personnel monitoring scenarios unique to hazardous chemical warehouses.

[0097] like Figure 9 As shown in Figure 2, ablation experiments are performed to verify the effectiveness of the detection model module. Ablation experiments are conducted on the EaDark dataset and a self-built low-light image dataset of hazardous chemical warehouses.

[0098] Ablation experiments on the ExDark dataset show that when only the M1 module is introduced, mAP@0.5 slightly increases to 58.9%. However, its attention mechanism's ability to model complex lighting conditions improves recall by 1.3%, demonstrating that this module effectively enhances robustness in object localization. When only the M2 module is introduced, mAP@0.5:0.95 increases to 36.5%, thanks to the deformable convolution's adaptive feature extraction for unusually shaped objects. However, mAP@0.5 decreases by 0.4%, due to the expansion of the local receptive field, which causes some features to be lost during detection, leading to an increase in the false detection rate for small objects. DCNv2, after integrating M3, effectively addresses this issue, with mAP@0.5 and mAP@0.5:0.95 increasing by 0.4% and 0.5%, respectively. This demonstrates that the M3 module effectively improves the original deformable convolution's feature extraction capabilities and properly allocates weights to features at different locations. Ablation experiments show that the detection model of the present invention significantly improves the detection performance through the collaborative optimization of MSCA-SCINet and GCFA_DCNv2; the improvement in mAP@0.5 indicator is 0.9%, and the improvement in mAP@0.5:0.95 indicator reaches 35.8%, which is 1.2 percentage points higher than the baseline model.

[0099] Ablation experiments on a self-built dataset of low-light images of hazardous chemical warehouses show that the introduction of the M1 module increases mAP@0.5 by 1.2%. When M2 is added alone, mAP@0.5 increases by 2.2% and recall by 3.0%. After adding the M3 module and integrating it with M2, mAP@0.5 further increases to 85.4%, and precision increases by 1.7%. When the M1, M2, and M3 modules work together, the detection model of this invention achieves optimal overall performance. The mAP@0.5 and mAP@0.5:0.95 indicators increase by 12.3% and 9.5%, respectively, and the recall rate stabilizes at 89.2%, achieving a balance between precision and recall, ensuring the system strikes a reasonable balance between missed detections and false positives.

[0100] like Figure 10 As shown in the figure, the gradient-weighted class activation mapping technique is used to visualize the model's attention area. Experiments show that the original RTDETR model lacks perception of feature texture details, and is prone to losing key semantic information, especially in low illumination. After adding the M1 module, low-light images are illuminated, and the adaptive illumination enhancement mechanism significantly optimizes the feature space distribution. By adding only the M2 and M3 modules, the dynamic receptive field adjustment strategy enhances shape adaptability and pays more attention to the target. After fusing the M1, M2, and M3 modules, attention is more focused on the target itself for irregularly shaped targets, effectively balancing the contradiction between detection accuracy and generalization performance.

[0101] like Figure 11As shown in Figure 2, the effectiveness of the detection model is verified through comparative experiments. Comparative experiments are conducted on the EaDark dataset and a self-built low-light image dataset of hazardous chemicals storage.

[0102] The comparative experimental results on the ExDark dataset show that the detection model of the present invention improves by 7.4%, 2.8%, 5.6%, and 1.7% in mAP@0.5 and mAP@0.5:0.95 compared with YOLOv5 and YOLOv8, respectively, and improves by 0.9% and 1.2% compared with RT-DETR. Compared with YOLOv10, although the mAP@0.5 is slightly lower than that of this method, the mAP@0.5:0.95 is improved by 9.8%, which verifies the detection model of the present invention's ability to optimize target positioning accuracy. At the same time, the 73.9% Precision and 53.8% Recall of the detection model of the present invention also show that it has both low false detection and high recall characteristics in a high noise background.

[0103] Comparative experiments on a self-built dataset of low-light images of hazardous chemical warehouses demonstrate that the proposed detection model achieves a mAP@0.5 of 94.6%, a 12.3% improvement over RTDETR, and 33.4%, 14.9%, and 11.5% improvements over YOLOv5, YOLOv8, and YOLOv10, respectively. The mAP@0.5:0.95 ratio is 46.0%, a 16.8%, 2.9%, and 6.3% improvement over YOLOv5, YOLOv8, and YOLOv10, respectively. Furthermore, the 93.8% Precision and 89.2% Recall demonstrate the proposed detection model's adaptability to specific scenarios.

[0104] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims rather than the foregoing description. It is intended that all variations within the meaning and range of equivalents of the claims be embraced herein, and any reference signs in the claims should not be construed as limiting the claims to which they relate.

Claims

1. An SD-RTDETR personnel detection model for hazardous chemical storage in low-light conditions, characterized by: It includes the MSCA-SCINet image enhancement module and the DNCV-RTDETR target detection network. The MSCA-SCINet image enhancement module is used to enhance low-light images, and then the DNCV-RTDETR target detection network is used to accurately extract features from the enhanced images. The MSCA-SCINet image enhancement module includes a self-calibration illumination learning network and an illumination enhancement network; In the loop, the low-light image is first enhanced by the illumination enhancement network, and then the enhanced image is self-calibrated by the self-calibration illumination learning network, and the self-calibrated enhanced image is input into the next loop for image enhancement again; The DNCV-RTDETR target detection network uses ResNet18 as the backbone network and uses the GCFA_DCNv2 module to replace the second convolutional layer in the BasicBlock basic block to form an improved residual structure BasicBlock_GCFA_DCNv2; The GCFA_DCNv2 module includes a global collaborative feature aggregation network GCFA and a deformable convolution DCNv2.

2. The SD-RTDETR personnel detection model for low-light hazardous chemical storage according to claim 1, characterized in that: The structure of the improved residual structure BasicBlock_GCFA_DCNv2 is as follows: the input feature map first passes through a 3×3 ordinary convolution layer, and then passes through a deformable convolution layer DCNv2 with a global collaborative feature aggregation network GCFA to extract features with adaptive sampling positions. Finally, the extracted features are added to the shortcut connection of the 1×1 convolution layer.

3. The SD-RTDETR personnel detection model for low-light hazardous chemical storage according to claim 2, characterized in that: The structure of the global collaborative feature aggregation network GCFA is as follows: while aggregating features in the original height and width spatial dimensions, a global pooling branch is designed, and combined with attention weight calculation, global information is integrated into the offset generation process.

4. The SD-RTDETR personnel detection model for low-light hazardous chemical storage according to claim 2, characterized in that: The specific calculation process of the global collaborative feature aggregation network GCFA is as follows: Given an input feature X, perform global average pooling on it to obtain the global information of the cth channel: Among them, x c (i, j) is the value of the input feature X at the cth channel and position (i, j), H is the height of the input feature map X in the spatial dimension, and W is the width of the input feature map X in the spatial dimension; Then, one-dimensional global pooling is performed using pooling kernels (H, 1) and (1, W) in the height and width dimensions respectively; The output formula of the cth channel with height h is: The output formula of the cth channel with width w is: The two transformations (1-8) and (1-9) aggregate the features along the horizontal coordinate and vertical coordinate space respectively to generate a pair of direction-aware feature maps; In order to effectively aggregate the features between channels, the spatial dimension features extracted by formulas (1-8) and (1-9) are first connected, and then input into the 1×1 convolution function F1 to obtain the fused features: where [·,·] represents the connection operation along the spatial dimension, δ is the nonlinear activation function, r is the block size reduction ratio used to control the SE block, and C is the number of channels of the input feature map. Split formula (1-10) into two separate tensors along the horizontal and vertical dimensions and Similarly, after the global pooling branch performs a 1×1 convolution transformation, the intermediate feature f is obtained conv =δ[F1(z c )]; The connection feature f obtained by formula (1-10) has a dimension of 2. The mean function torch.mean is used to calculate the mean along the spatial dimension. f mean =torch.mean(f,dim=2)(1-11), After adjusting the feature dimension, the intermediate feature f output in the global pooling branch conv Perform splicing to obtain the global characteristics of the branch: f a =[f conv ,f mean ](1-12), The horizontal dimension feature f is transformed into h , vertical dimension feature f w and global features f a Calculate the corresponding attention weight g h 、g w and g a , the formula is as follows Finally, output features: y c (i,j)=x c (i,j)×g h (i)×g w (j)×g a (1-14)。 5. The SD-RTDETR personnel detection model for low-light hazardous chemical storage according to claim 1, characterized in that: The self-calibration lighting learning network includes a self-calibration module and a lighting estimation module; In each cycle, the low-light image is first enhanced by the illumination enhancement network, then the enhanced image is analyzed by the illumination estimation module, and finally the enhanced image is self-calibrated by the self-calibration module, and the self-calibrated enhanced image is input into the next cycle.

6. The SD-RTDETR personnel detection model for low-light hazardous chemical storage according to claim 5, characterized in that: The specific calculation process of the lighting estimation module is as follows: The lighting estimation module is designed based on the Retinex theory. In Retinex theory, a low-light image y is regarded as the product of the illumination component x and the clear image reflection component z: y=x⊙z(1-1), The illumination component x refers to the part affected by external illumination, i.e., weak illumination; the reflection component z refers to the constant part determined by the nature of the object, i.e., the clear image part. Then introduce the mapping H θ Build a progressive lighting optimization process, the basic units of which are: Among them, t represents the number of stages, u t Represents the residual of the tth stage, x t Represents the illumination component of stage t; The specific calculation process of the self-calibration module is as follows: The network structure is designed so that the input of each stage is the output of the previous stage. In order to express the difference between the output of different stages and the original low-light image input in the first stage, a self-calibration mapping function s is introduced. The self-calibration module of the tth stage can be expressed as: Among them, z t is the reflection component of the tth stage, i.e. the clear image part, is the division symbol, s t is the self-calibration image at stage t, Indicates that it contains learnable parameters The introduction of parameterized operators, v t represents the input of the next stage after calibration, Since the output of each stage is the input of the next stage, the basic unit of the illumination self-calibration process is recorded as: F(x t )→F(g(x t ))(1-4)。 7. The SD-RTDETR personnel detection model for low-light hazardous chemical storage according to claim 1, characterized in that: The illumination enhancement network introduces a multi-scale attention mechanism MSCA, which first maps the original image to a high-dimensional feature space through the input convolution layer, and then extracts features through multiple convolution blocks. Each convolution block consists of a standard convolution layer and a multi-scale attention module MSCA. The multi-scale attention module MSCA captures the global and local contextual information of the image through multi-scale convolution kernels and enhances important features in combination with the channel mixing mechanism. Finally, the network maps the features back to the image space through the output convolution layer and adds them to the input image to generate an enhanced image.

8. The SD-RTDETR personnel detection model for low-light hazardous chemical storage according to claim 7, characterized in that: The multi-scale attention mechanism MSCA includes depth-wise convolution for aggregating local information, multi-branch depth-wise strip convolution for capturing multi-scale context, and 1×1 convolution for modeling the relationship between different channel dimensions.

9. The SD-RTDETR personnel detection model for low-light hazardous chemical storage according to claim 7, characterized in that: The implementation formula of the multi-scale attention mechanism MSCA is: Where L represents the attention map of MSCA, Branch i (i=0,1,2,3) represents the four branches output from the convolution kernel of 5×5, Fea represents the input feature, DpConv represents the depth convolution, Conv1x1 represents the 1x1 convolution, Then the attention map is element-wise multiplied with the input features as the output, and the formula is:

10. The SD-RTDETR personnel detection model for low-light hazardous chemical storage according to claim 1, characterized in that: The self-calibration lighting learning network and the lighting enhancement network adopt a weight sharing mechanism.